From 3a7385577b5f99cb88517184ebb73953e427ba67 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:41:03 +0000 Subject: [PATCH 001/206] fix(runtime): published dynamic expansions are the lock-free authority for fan-out membership A published dynamic-expansions/.json is the durable record of which generated nodes a group has, but loadOrCreateDynamicExpansion re-proved it on every workflow render and every strict lifecycle admission: it re-read and re-hashed the live source artifact and required it to match source.output_sha256, all under an O_EXCL lock that was never reclaimed. Both are self-inflicted stops: - A reset that re-runs the source (a --reset-node, or a Smithers timetravel that sweeps the source into its dependent set) empties the source's canonical artifact directory when its agent attempt starts. Every render then threw "dynamic source artifact does not exist" and, once the source wrote a different plan, DYNAMIC_EXPANSION_CHANGED, with no way back. The same check made strict admission refuse cancel, pause, why, fork and replay. - A controller killed (SIGKILL, OOM) or an observer interrupted while holding .expansion.lock left every later render and admission busy-waiting 5 s and then throwing DYNAMIC_EXPANSION_LOCKED. An existing manifest is now validated (the manifest set plus the group's static identity: run, group, source node, attempt and path, JSON and key paths, node-ID template, prompt template digest and fingerprint, limit) and returned without touching the source. The source is read only to create the manifest, and its digest is still recorded as provenance. The lock is deleted: reads need none, creation happens in the workflow render, which Smithers runs for one live driver per run, publishFileDurableExclusive never replaces a published file, and the existing re-read after publication refuses an inconsistent set. Semantic change: after a reset re-runs a source, the fan-out keeps its originally published items instead of failing. resume --retry-failed of a failed source verifier still archives the manifests (#1064) and re-expands. The tests that encoded the old behaviour (a changed source must throw, the lock must never be reclaimed, admission must fail while the source is absent) are replaced by tests that a re-run source keeps the published fan-out for renders and admission, and that a leftover lock does not block expansion. Refs #1142 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/topology-yaml.md | 13 +- packages/runtime/src/dynamic-expansion.ts | 211 +++++++----------- .../runtime/test/dynamic-expansion.test.ts | 136 ++++++----- .../runtime/test/dynamic-lifecycle.test.ts | 29 +-- 5 files changed, 192 insertions(+), 198 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..f92d1a2d4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [docs]** A published dynamic expansion manifest is now the only authority for its group's generated nodes. Workflow renders and the strict lifecycle admission behind `cancel`, `pause`, `why`, `fork`, and `replay` no longer re-read the source artifact, so a source that runs again after a reset keeps the published fan-out instead of failing every render while its artifact directory is empty and, once it writes a different plan, with `DYNAMIC_EXPANSION_CHANGED`; `resume --retry-failed` of a failed source verifier still re-expands from the new output. The expansion lock, which was never reclaimed, is removed, so a lock file left by a killed process no longer fails every later render with `DYNAMIC_EXPANSION_LOCKED` (#1142). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/topology-yaml.md b/docs/reference/topology-yaml.md index 65708ce28..a40889652 100644 --- a/docs/reference/topology-yaml.md +++ b/docs/reference/topology-yaml.md @@ -346,9 +346,16 @@ Generated children use the same global concurrency scheduler as static nodes. exceeding it fails explicitly and never truncates the source array. The first successful expansion is persisted under -`dynamic-expansions/.json`. Resume reuses that exact manifest and -rejects changes to its source bytes, prompt template, topology contract, or -dynamic-node limit instead of silently changing the graph. +`dynamic-expansions/.json`, and from then on that manifest alone +decides the group's generated nodes. Resume reuses it and rejects changes to its +prompt template, topology contract, or dynamic-node limit instead of silently +changing the graph. The source artifact is read only to create the manifest, so +a source node that runs again, for example after a reset, leaves the published +fan-out unchanged. The exception is `resume --retry-failed` for a source whose +verifier failed: it moves the published manifests to +`dynamic-expansion-history/` before the source runs again (and refuses if +another source published any of them), so the group expands again from the new +output. ## Model Fan-Out diff --git a/packages/runtime/src/dynamic-expansion.ts b/packages/runtime/src/dynamic-expansion.ts index fdb05d9ca..0ea3da7d8 100644 --- a/packages/runtime/src/dynamic-expansion.ts +++ b/packages/runtime/src/dynamic-expansion.ts @@ -227,12 +227,8 @@ export function loadOrCreateDynamicExpansion(input: { const runRoot = path.resolve(input.runRoot); assertPathInside(runRoot, input.sourceArtifactPath, "dynamic source artifact"); assertPathInside(runRoot, input.templatePath, "dynamic prompt template"); - assertNoSymlinkComponents(runRoot, input.sourceArtifactPath, "dynamic source artifact"); assertNoSymlinkComponents(runRoot, input.templatePath, "dynamic prompt template"); - assertRegularFileInside(runRoot, input.sourceArtifactPath, "dynamic source artifact"); assertRegularFileInside(runRoot, input.templatePath, "dynamic prompt template"); - const sourceBytes = fs.readFileSync(input.sourceArtifactPath); - const sourceDigest = sha256Bytes(sourceBytes); const actualTemplateDigest = sha256Bytes(fs.readFileSync(input.templatePath)); if (actualTemplateDigest !== input.templateDigest) { throw dynamicError("DYNAMIC_TEMPLATE_CHANGED", `Dynamic group ${input.groupNodeId} prompt template changed`, { @@ -248,88 +244,96 @@ export function loadOrCreateDynamicExpansion(input: { assertNoSymlinkComponents(runRoot, manifestDir, "dynamic expansion manifest directory"); const manifestPath = path.join(manifestDir, `${validateSafeId(input.groupNodeId, "dynamic group node ID")}.json`); const sourceArtifactRelativePath = path.relative(runRoot, input.sourceArtifactPath).split(path.sep).join("/"); - return withDynamicExpansionLock(manifestDir, () => { - const priorManifests = readExpansionManifests(manifestDir); - assertManifestSetMatchesInput(priorManifests, { - runId: input.runId, - maxDynamicNodes: input.maxDynamicNodes, - reservedNodeIds: input.reservedNodeIds - }); - const existing = priorManifests.find((manifest) => manifest.group_node_id === input.groupNodeId); - if (existing !== undefined) { - assertCompatibleManifest(existing, { - runId: input.runId, - groupNodeId: input.groupNodeId, - sourceNodeId: input.sourceNodeId, - sourceAttemptId: input.sourceAttemptId, - sourceArtifactPath: sourceArtifactRelativePath, - sourceDigest, - templateDigest: input.templateDigest, - templateFingerprint: input.templateFingerprint, - sourcePath: input.sourcePath, - keyPath: input.keyPath, - nodeIdTemplate: input.nodeIdTemplate, - maxDynamicNodes: input.maxDynamicNodes - }); - return existing; - } - - let sourceDocument: unknown; - try { - sourceDocument = JSON.parse(sourceBytes.toString("utf8")) as unknown; - } catch (error) { - throw dynamicError("DYNAMIC_SOURCE_JSON_INVALID", `Dynamic source artifact is not valid JSON`, { - groupNodeId: input.groupNodeId, - reason: error instanceof Error ? error.message : String(error) - }); - } - const reserved = new Set(input.reservedNodeIds ?? []); - for (const manifest of priorManifests) { - for (const item of manifest.items) reserved.add(item.node_id); - } - const manifest = planDynamicExpansion({ + // No lock. Reading a published manifest needs none, and creation happens in the workflow render, + // which Smithers runs for one live driver per run. publishFileDurableExclusive never replaces a + // published file, and the re-read after publication refuses a set that a racing second publisher + // made inconsistent. The lock this replaced was never reclaimed, so a process killed while + // holding it failed every later render and lifecycle admission (#1142). + const priorManifests = readExpansionManifests(manifestDir); + assertManifestSetMatchesInput(priorManifests, { + runId: input.runId, + maxDynamicNodes: input.maxDynamicNodes, + reservedNodeIds: input.reservedNodeIds + }); + const existing = priorManifests.find((manifest) => manifest.group_node_id === input.groupNodeId); + if (existing !== undefined) { + // A published manifest is the group's membership. `source.output_sha256` records the bytes it + // was planned from; the live source is not consulted again, because a reset re-runs the source, + // whose agent attempt wipes that artifact and then writes a new one. + assertCompatibleManifest(existing, { runId: input.runId, groupNodeId: input.groupNodeId, sourceNodeId: input.sourceNodeId, sourceAttemptId: input.sourceAttemptId, sourceArtifactPath: sourceArtifactRelativePath, - sourceDigest, - sourceDocument, + templateDigest: input.templateDigest, + templateFingerprint: input.templateFingerprint, sourcePath: input.sourcePath, keyPath: input.keyPath, nodeIdTemplate: input.nodeIdTemplate, - templateDigest: input.templateDigest, - templateFingerprint: input.templateFingerprint, - maxDynamicNodes: input.maxDynamicNodes, - sequence: priorManifests.length, - alreadyExpandedNodes: priorManifests.reduce((sum, entry) => sum + entry.items.length, 0), - reservedNodeIds: reserved - }); - // Validate the complete candidate set in memory first: a manifest that would make the set - // invalid must never reach durable storage, otherwise every automatic resume keeps failing - // until an operator removes or repairs the published file by hand. - const candidateSet = [...priorManifests, manifest]; - validateManifestSet(candidateSet, manifestDir); - assertManifestSetMatchesInput(candidateSet, { - runId: input.runId, - maxDynamicNodes: input.maxDynamicNodes, - reservedNodeIds: input.reservedNodeIds + maxDynamicNodes: input.maxDynamicNodes }); - publishFileDurableExclusive(manifestDir, `${input.groupNodeId}.json`, `${JSON.stringify(manifest, null, 2)}\n`); - const publishedManifests = readExpansionManifests(manifestDir); - assertManifestSetMatchesInput(publishedManifests, { - runId: input.runId, - maxDynamicNodes: input.maxDynamicNodes, - reservedNodeIds: input.reservedNodeIds + return existing; + } + + assertNoSymlinkComponents(runRoot, input.sourceArtifactPath, "dynamic source artifact"); + assertRegularFileInside(runRoot, input.sourceArtifactPath, "dynamic source artifact"); + const sourceBytes = fs.readFileSync(input.sourceArtifactPath); + let sourceDocument: unknown; + try { + sourceDocument = JSON.parse(sourceBytes.toString("utf8")) as unknown; + } catch (error) { + throw dynamicError("DYNAMIC_SOURCE_JSON_INVALID", `Dynamic source artifact is not valid JSON`, { + groupNodeId: input.groupNodeId, + reason: error instanceof Error ? error.message : String(error) }); - const published = publishedManifests.find((candidate) => candidate.group_node_id === input.groupNodeId); - if (published === undefined || !fs.existsSync(manifestPath)) { - throw dynamicError("DYNAMIC_MANIFEST_PUBLISH_FAILED", "Dynamic expansion manifest publication failed", { - groupNodeId: input.groupNodeId - }); - } - return published; + } + const reserved = new Set(input.reservedNodeIds ?? []); + for (const manifest of priorManifests) { + for (const item of manifest.items) reserved.add(item.node_id); + } + const manifest = planDynamicExpansion({ + runId: input.runId, + groupNodeId: input.groupNodeId, + sourceNodeId: input.sourceNodeId, + sourceAttemptId: input.sourceAttemptId, + sourceArtifactPath: sourceArtifactRelativePath, + sourceDigest: sha256Bytes(sourceBytes), + sourceDocument, + sourcePath: input.sourcePath, + keyPath: input.keyPath, + nodeIdTemplate: input.nodeIdTemplate, + templateDigest: input.templateDigest, + templateFingerprint: input.templateFingerprint, + maxDynamicNodes: input.maxDynamicNodes, + sequence: priorManifests.length, + alreadyExpandedNodes: priorManifests.reduce((sum, entry) => sum + entry.items.length, 0), + reservedNodeIds: reserved }); + // Validate the complete candidate set in memory first: a manifest that would make the set + // invalid must never reach durable storage, otherwise every automatic resume keeps failing + // until an operator removes or repairs the published file by hand. + const candidateSet = [...priorManifests, manifest]; + validateManifestSet(candidateSet, manifestDir); + assertManifestSetMatchesInput(candidateSet, { + runId: input.runId, + maxDynamicNodes: input.maxDynamicNodes, + reservedNodeIds: input.reservedNodeIds + }); + publishFileDurableExclusive(manifestDir, `${input.groupNodeId}.json`, `${JSON.stringify(manifest, null, 2)}\n`); + const publishedManifests = readExpansionManifests(manifestDir); + assertManifestSetMatchesInput(publishedManifests, { + runId: input.runId, + maxDynamicNodes: input.maxDynamicNodes, + reservedNodeIds: input.reservedNodeIds + }); + const published = publishedManifests.find((candidate) => candidate.group_node_id === input.groupNodeId); + if (published === undefined || !fs.existsSync(manifestPath)) { + throw dynamicError("DYNAMIC_MANIFEST_PUBLISH_FAILED", "Dynamic expansion manifest publication failed", { + groupNodeId: input.groupNodeId + }); + } + return published; } export function dynamicStorageId(groupNodeId: string, generatedNodeId: string): string { @@ -778,61 +782,6 @@ function assertManifestSetMatchesInput( } } -function withDynamicExpansionLock(manifestDir: string, operation: () => T): T { - const lockPath = path.join(manifestDir, ".expansion.lock"); - const deadline = Date.now() + 5_000; - const token = `${process.pid}:${crypto.randomBytes(16).toString("hex")}`; - let descriptor: number | undefined; - while (descriptor === undefined) { - try { - descriptor = fs.openSync(lockPath, "wx", 0o600); - fs.writeFileSync(descriptor, `${token}\n`, "utf8"); - } catch (error) { - if (!isAlreadyExistsError(error)) throw error; - let stat: fs.Stats; - try { - stat = fs.lstatSync(lockPath); - } catch (statError) { - // The holder may release between our exclusive-create failure and the - // inspection. That is ordinary lock contention, not a run failure. - if (isNoEntryError(statError)) continue; - throw statError; - } - if (stat.isSymbolicLink() || !stat.isFile()) { - throw dynamicError("DYNAMIC_EXPANSION_LOCK_INVALID", "Dynamic expansion lock is unsafe", { lockPath }); - } - // Never steal a lock based on age. Between an age check and unlink, the - // observed inode can disappear and a new owner can publish a fresh lock - // at the same path. Only the token-owning holder releases the lock; - // contenders fail after the bounded wait and leave recovery explicit. - if (Date.now() >= deadline) { - throw dynamicError("DYNAMIC_EXPANSION_LOCKED", "Dynamic expansion is already being materialized", { - lockPath - }); - } - Atomics.wait(new Int32Array(new SharedArrayBuffer(4)), 0, 0, 10); - } - } - try { - return operation(); - } finally { - fs.closeSync(descriptor); - try { - if (fs.readFileSync(lockPath, "utf8").trim() === token) fs.unlinkSync(lockPath); - } catch { - // A missing lock after the operation cannot weaken manifest validation. - } - } -} - -function isAlreadyExistsError(error: unknown): boolean { - return error instanceof Error && "code" in error && error.code === "EEXIST"; -} - -function isNoEntryError(error: unknown): boolean { - return error instanceof Error && "code" in error && error.code === "ENOENT"; -} - type ManifestFailure = (reason: string, details?: Record) => never; function assertExactKeys( @@ -933,7 +882,6 @@ function assertCompatibleManifest( sourceNodeId: string; sourceAttemptId: string; sourceArtifactPath: string; - sourceDigest: string; templateDigest: string; templateFingerprint: string; sourcePath: string; @@ -948,7 +896,6 @@ function assertCompatibleManifest( ["source node ID", manifest.source.node_id, expected.sourceNodeId], ["source attempt ID", manifest.source.attempt_id, expected.sourceAttemptId], ["source artifact", manifest.source.artifact_path, expected.sourceArtifactPath], - ["source output", manifest.source.output_sha256, expected.sourceDigest], ["source path", manifest.source.json_path, expected.sourcePath], ["key path", manifest.template.key_path, expected.keyPath], ["node ID template", manifest.template.node_id, expected.nodeIdTemplate], diff --git a/packages/runtime/test/dynamic-expansion.test.ts b/packages/runtime/test/dynamic-expansion.test.ts index 0d9190a11..6c5c00996 100644 --- a/packages/runtime/test/dynamic-expansion.test.ts +++ b/packages/runtime/test/dynamic-expansion.test.ts @@ -212,7 +212,7 @@ test("dynamic expansion rejects duplicate keys, reserved IDs, and malformed huma assert.equal(dotted.items[0]?.variables["oracle.v2:stale-price"], "stale prices"); }); -test("persisted expansion rejects tampering, transplantation, source changes, and symlink manifests", () => { +test("persisted expansion rejects tampering, transplantation, and symlink manifests", () => { const tampered = expansionFixture({ runId: "tampered", items: [item(0)] }); tampered.invoke(); const manifestPath = path.join(tampered.runRoot, "dynamic-expansions", "fanout.json"); @@ -226,14 +226,6 @@ test("persisted expansion rejects tampering, transplantation, source changes, an (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_MANIFEST_INVALID" ); - const changed = expansionFixture({ runId: "changed", items: [item(0)] }); - changed.invoke(); - fs.writeFileSync(changed.sourcePath, `${JSON.stringify({ goals: [item(1)] })}\n`, "utf8"); - assert.throws( - () => changed.invoke(), - (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_EXPANSION_CHANGED" - ); - const transplanted = expansionFixture({ runId: "origin-run", items: [item(0)] }); transplanted.invoke(); assert.throws( @@ -504,49 +496,16 @@ test("explicit source retry re-derives the base runtime controls after archiving ); }); -test("dynamic expansion retries when a contended lock disappears before inspection", () => { - const fixture = expansionFixture({ runId: "lock-release-race", items: [item(0)] }); - const originalOpenSync = fs.openSync; - let simulated = false; - fs.openSync = ((...args: Parameters) => { - if (!simulated && String(args[0]).endsWith(".expansion.lock")) { - simulated = true; - throw Object.assign(new Error("simulated released contender"), { code: "EEXIST" }); - } - return originalOpenSync(...args); - }) as typeof fs.openSync; - - try { - assert.equal(fixture.invoke().items.length, 1); - assert.equal(simulated, true); - } finally { - fs.openSync = originalOpenSync; - } -}); - -test("dynamic expansion never steals an old lock owned by another materializer", () => { - const fixture = expansionFixture({ runId: "old-lock", items: [item(0)] }); +test("a lock file left by a killed materializer does not block expansion", () => { + // Earlier builds serialized every expansion read behind this file and never reclaimed it, so a + // process killed while holding it failed every later render and admission check (#1142). + const fixture = expansionFixture({ runId: "stale-lock", items: [item(0)] }); const manifestDirectory = path.join(fixture.runRoot, "dynamic-expansions"); - const lockPath = path.join(manifestDirectory, ".expansion.lock"); fs.mkdirSync(manifestDirectory); - fs.writeFileSync(lockPath, "other-owner\n", "utf8"); - const old = new Date("2020-01-01T00:00:00.000Z"); - fs.utimesSync(lockPath, old, old); - const originalNow = Date.now; - const base = originalNow(); - let reads = 0; - Date.now = () => base + reads++ * 6_000; - - try { - assert.throws( - () => fixture.invoke(), - (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_EXPANSION_LOCKED" - ); - } finally { - Date.now = originalNow; - } - assert.equal(fs.readFileSync(lockPath, "utf8"), "other-owner\n"); - assert.equal(fs.statSync(lockPath).mtimeMs, old.getTime()); + fs.writeFileSync(path.join(manifestDirectory, ".expansion.lock"), "999999:deadbeef\n", "utf8"); + const created = fixture.invoke(); + assert.equal(created.items.length, 1); + assert.deepEqual(fixture.invoke(), created); }); test("runtime materialization preserves required inputs and allows partial review joins", () => { @@ -645,6 +604,83 @@ test("runtime materialization preserves required inputs and allows partial revie } }); +test("a re-run dynamic source keeps the published fan-out for renders and admission", () => { + const runId = "source-rerun"; + const projectRoot = tempDirectory(); + const runRoot = path.join(projectRoot, "runs", runId); + const sourceArtifactPath = path.join(runRoot, "artifacts", "planner", "plan.json"); + const templatePath = path.join(runRoot, "templates", "worker.md"); + const graphPath = path.join(runRoot, "graph.json"); + const tasksPath = path.join(runRoot, "smithers", "tasks.json"); + const baseGraphPath = path.join(runRoot, "smithers", "runtime-base-graph.json"); + const baseTasksPath = path.join(runRoot, "smithers", "runtime-base-tasks.json"); + fs.mkdirSync(path.dirname(sourceArtifactPath), { recursive: true }); + fs.mkdirSync(path.dirname(templatePath), { recursive: true }); + fs.mkdirSync(path.dirname(tasksPath), { recursive: true }); + fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals: [item(0), item(1)] })}\n`, "utf8"); + fs.writeFileSync(templatePath, "Investigate {{item.goal_prompt}}.\n", "utf8"); + const joinTask = compiledTask(projectRoot, runRoot, "join", "join", undefined, ["fanout"]); + const group: CompiledSmithersDynamicGroup = { + groupNodeId: "fanout", + logicalNodeId: "fanout", + source: { + concreteNodeId: "planner", + attemptId: "planner", + verifierSmithersNodeId: "verify:planner", + artifactPath: sourceArtifactPath + }, + sourcePath: "$.goals", + keyPath: "id", + nodeIdTemplate: "dynamic:item:{{ item.id }}", + templatePath, + templateDigest: digest(fs.readFileSync(templatePath)), + templateFingerprint: digest("template-fingerprint"), + continueOnFail: true, + maxDynamicNodes: 100, + reservedNodeIds: ["planner", "fanout", "join"], + taskTemplates: [compiledTask(projectRoot, runRoot, "fanout", "fanout", templatePath)], + promptContext: promptContext(projectRoot, runRoot) + }; + fs.writeFileSync(graphPath, `${JSON.stringify(plannedGraph(runId))}\n`, "utf8"); + fs.writeFileSync( + tasksPath, + `${JSON.stringify({ schema_version: "1.0", run_id: runId, tasks: [joinTask], dynamic_groups: [group] })}\n`, + "utf8" + ); + fs.copyFileSync(graphPath, baseGraphPath); + fs.copyFileSync(tasksPath, baseTasksPath); + const controls = { + runId, + projectRoot, + runRoot, + graphPath, + tasksPath, + baseGraphPath, + baseTasksPath, + baseTasks: [joinTask], + groups: [group] + }; + const render = (readyGroupIds: string[]) => + dynamicRuntimeFingerprint(materializeDynamicRuntime({ ...controls, readyGroupIds })); + + const published = render(["fanout"]); + // A reset re-runs the planner: its agent attempt first wipes the canonical artifact directory, + // then writes a new plan, and the verifier only accepts that plan later. Every render in between + // still sees the published manifest, and so do the lifecycle admission checks. + fs.rmSync(sourceArtifactPath); + assert.equal(render([]), published); + assert.equal(dynamicRuntimeFingerprint(verifyDynamicRuntimeMaterialization(controls)), published); + fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals: [item(2)] })}\n`, "utf8"); + assert.equal(render([]), published); + assert.equal(render(["fanout"]), published); + assert.equal(dynamicRuntimeFingerprint(verifyDynamicRuntimeMaterialization(controls)), published); + const tasks = JSON.parse(fs.readFileSync(tasksPath, "utf8")) as { tasks: CompiledSmithersTask[] }; + assert.deepEqual( + tasks.tasks.flatMap((task) => task.metadata.node.dynamic?.expansionKey ?? []), + ["goal-0", "goal-1"] + ); +}); + test("100 generated attempts remain queued under the ordinary concurrency projection", () => { const createdAt = "2026-08-04T00:00:00.000Z"; const attempts = Array.from({ length: 100 }, (_, index) => `dynamic-item-${index}`); diff --git a/packages/runtime/test/dynamic-lifecycle.test.ts b/packages/runtime/test/dynamic-lifecycle.test.ts index 86e7679f8..8878bcca7 100644 --- a/packages/runtime/test/dynamic-lifecycle.test.ts +++ b/packages/runtime/test/dynamic-lifecycle.test.ts @@ -734,20 +734,21 @@ test("a half-published dynamic expansion stays readable while execution stays cl assert.equal(fs.existsSync(generated.renderedPromptPath!), false); }); -test("event observers stay readable while a retried dynamic source is rematerialized", async () => { +test("a re-running dynamic source keeps published controls admissible and observable", async () => { const fixture = await createDynamicFixture({ runId: "dynamic-source-retry-events" }); const [firstTask] = fixture.generatedTasks; assert.ok(firstTask); const eventNodeId = firstTask.smithersNodeId; setLifecycle(fixture, [], [{ type: "NodeStarted", nodeId: eventNodeId, attempt: 2 }]); + // A reset re-runs the planner, whose agent attempt first wipes its canonical artifact directory. const sourcePath = path.join(fixture.runRoot, "artifacts", "planner", "plan.json"); + const publishedPlan = fs.readFileSync(sourcePath, "utf8"); fs.rmSync(sourcePath); + // The published expansion manifest decides the fan-out, so the strict admission that cancel, + // pause, fork, and replay require keeps re-deriving the same controls without the source. const strict = await readLinkedWorkflowEvidence(fixture.project, fixture.runId); - assert.equal(strict.ok, false); - if (!strict.ok) { - assert.match(strict.diagnostics[0]?.message ?? "", /dynamic source artifact does not exist/u); - } + assert.equal(strict.ok, true, "diagnostics" in strict ? JSON.stringify(strict.diagnostics) : ""); const queried = await queryWorkflowEvents({ projectRoot: fixture.project, @@ -757,12 +758,9 @@ test("event observers stay readable while a retried dynamic source is rematerial assert.equal(queried.ok, true, JSON.stringify(queried.diagnostics)); assert.equal(queried.value?.events[0]?.category, "NodeStarted"); assert.equal(queried.value?.events[0]?.node_id, eventNodeId); - assert.ok( - queried.diagnostics.some( - (diagnostic) => - diagnostic.code === "WORKFLOW_CONTROL_EVIDENCE_DIVERGED" && - /dynamic source artifact does not exist/u.test(diagnostic.message) - ), + assert.equal( + queried.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_CONTROL_EVIDENCE_DIVERGED"), + false, JSON.stringify(queried.diagnostics) ); @@ -775,13 +773,18 @@ test("event observers stay readable while a retried dynamic source is rematerial }); assert.equal(watched.ok, true, JSON.stringify(watched.diagnostics)); assert.equal(streamed[0]?.category, "NodeStarted"); - assert.ok( + assert.equal( watched.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_CONTROL_EVIDENCE_DIVERGED"), + false, JSON.stringify(watched.diagnostics) ); - // Observation must not recreate or otherwise repair the dynamic source on the controller's behalf. assert.equal(fs.existsSync(sourcePath), false); + + // The re-run then writes a different plan; admission still follows the published manifest. + fs.writeFileSync(sourcePath, `${publishedPlan}\n`, "utf8"); + const rewritten = await readLinkedWorkflowEvidence(fixture.project, fixture.runId); + assert.equal(rewritten.ok, true, "diagnostics" in rewritten ? JSON.stringify(rewritten.diagnostics) : ""); }); test("controller refresh preserves the sealed dynamic base after runtime materialization", async () => { From d3ba05cef7ebceb6c80a531d0b257aa8d1ccab37 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:41:35 +0000 Subject: [PATCH 002/206] fix(runtime): a leftover expansion lock no longer refuses a dynamic source retry planDynamicExpansionRetryArchive refuses a manifest directory that holds anything but the published manifests, so the .expansion.lock an older build left behind after a crash (the failure the previous commit removes) made `resume --retry-failed` of the dynamic source fail with DYNAMIC_RETRY_EXPANSION_INVALID until an operator deleted it by hand. The temporary file of an interrupted publishFileDurableExclusive has the same effect. Dot entries are never manifests (readExpansionManifests already skips them) and the directory rename archives them along with everything else, so the refusal now ignores them. Any other unrecognized entry is still refused; the existing test for that case now uses a non-dot file. Refs #1142 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- .../runtime/src/dynamic-expansion-retry.ts | 6 ++++- .../runtime/test/dynamic-expansion.test.ts | 27 ++++++++++++------- 3 files changed, 24 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f92d1a2d4..be52d50db 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime] [docs]** A published dynamic expansion manifest is now the only authority for its group's generated nodes. Workflow renders and the strict lifecycle admission behind `cancel`, `pause`, `why`, `fork`, and `replay` no longer re-read the source artifact, so a source that runs again after a reset keeps the published fan-out instead of failing every render while its artifact directory is empty and, once it writes a different plan, with `DYNAMIC_EXPANSION_CHANGED`; `resume --retry-failed` of a failed source verifier still re-expands from the new output. The expansion lock, which was never reclaimed, is removed, so a lock file left by a killed process no longer fails every later render with `DYNAMIC_EXPANSION_LOCKED` (#1142). +- **[runtime] [docs]** A published dynamic expansion manifest is now the only authority for its group's generated nodes. Workflow renders and the strict lifecycle admission behind `cancel`, `pause`, `why`, `fork`, and `replay` no longer re-read the source artifact, so a source that runs again after a reset keeps the published fan-out instead of failing every render while its artifact directory is empty and, once it writes a different plan, with `DYNAMIC_EXPANSION_CHANGED`; `resume --retry-failed` of a failed source verifier still re-expands from the new output. The expansion lock, which was never reclaimed, is removed, so a lock file left by a killed process no longer fails every later render with `DYNAMIC_EXPANSION_LOCKED` or refuses a source retry (#1142). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/packages/runtime/src/dynamic-expansion-retry.ts b/packages/runtime/src/dynamic-expansion-retry.ts index 0f14c86c3..a5b570b95 100644 --- a/packages/runtime/src/dynamic-expansion-retry.ts +++ b/packages/runtime/src/dynamic-expansion-retry.ts @@ -113,7 +113,11 @@ export function planDynamicExpansionRetryArchive(input: { ); } const expectedEntries = new Set(manifests.map((manifest) => `${manifest.group_node_id}.json`)); - const unexpectedEntries = fs.readdirSync(manifestDir).filter((entry) => !expectedEntries.has(entry)); + // A dot entry is never a manifest (readExpansionManifests skips it): an interrupted publication's + // temporary file, or the `.expansion.lock` older builds left behind. It moves with the directory. + const unexpectedEntries = fs + .readdirSync(manifestDir) + .filter((entry) => !entry.startsWith(".") && !expectedEntries.has(entry)); if (unexpectedEntries.length > 0) { throw dynamicError( "DYNAMIC_RETRY_EXPANSION_INVALID", diff --git a/packages/runtime/test/dynamic-expansion.test.ts b/packages/runtime/test/dynamic-expansion.test.ts index 6c5c00996..edf19c8b2 100644 --- a/packages/runtime/test/dynamic-expansion.test.ts +++ b/packages/runtime/test/dynamic-expansion.test.ts @@ -363,21 +363,21 @@ test("explicit source retry archives a complete expansion generation and rejects ]); // Unrecognized manifest state is refused before anything is reset or moved. - const locked = expansionFixture({ runId: "retry-archive-locked", items: [item(0)] }); - locked.invoke(); - fs.writeFileSync(path.join(locked.runRoot, "dynamic-expansions", ".expansion.lock"), "other-owner\n", "utf8"); + const unrecognized = expansionFixture({ runId: "retry-archive-unrecognized", items: [item(0)] }); + unrecognized.invoke(); + fs.writeFileSync(path.join(unrecognized.runRoot, "dynamic-expansions", "notes.txt"), "operator note\n", "utf8"); assert.throws( () => planDynamicExpansionRetryArchive({ - projectRoot: locked.runRoot, - runRoot: locked.runRoot, + projectRoot: unrecognized.runRoot, + runRoot: unrecognized.runRoot, sourceNodeIds: ["node:planner"] }), (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_RETRY_EXPANSION_INVALID" ); - assert.deepEqual(fs.readdirSync(path.join(locked.runRoot, "dynamic-expansions")).sort(), [ - ".expansion.lock", - "fanout.json" + assert.deepEqual(fs.readdirSync(path.join(unrecognized.runRoot, "dynamic-expansions")).sort(), [ + "fanout.json", + "notes.txt" ]); }); @@ -496,7 +496,7 @@ test("explicit source retry re-derives the base runtime controls after archiving ); }); -test("a lock file left by a killed materializer does not block expansion", () => { +test("a lock file left by a killed materializer blocks neither expansion nor a source retry", () => { // Earlier builds serialized every expansion read behind this file and never reclaimed it, so a // process killed while holding it failed every later render and admission check (#1142). const fixture = expansionFixture({ runId: "stale-lock", items: [item(0)] }); @@ -506,6 +506,15 @@ test("a lock file left by a killed materializer does not block expansion", () => const created = fixture.invoke(); assert.equal(created.items.length, 1); assert.deepEqual(fixture.invoke(), created); + const plan = planDynamicExpansionRetryArchive({ + projectRoot: fixture.runRoot, + runRoot: fixture.runRoot, + sourceNodeIds: ["node:planner"] + }); + assert.deepEqual( + plan?.manifests.map((manifest) => manifest.group_node_id), + ["fanout"] + ); }); test("runtime materialization preserves required inputs and allows partial review joins", () => { From eb5487f63e1d87c9b8763e8fb9e65e6be021aca0 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:51:33 +0000 Subject: [PATCH 003/206] fix(security): require eyJ header and payload segments in the JWT rule The publication secret gate scans agent output in positive-only mode and fails the node on any hit. Its supplemental JWT rule matched any three dotted runs of base64url characters (20/10/10 minimum lengths), which also describes ordinary qualified identifiers: a report citing `ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck`, or a JSON path such as `auditProfileResolution.settingOrigins.property_priority_threshold`. The gate rejected such a file after the agent had finished its work. A JWT's header and payload are base64url-encoded JSON objects. In the compact form that JWT libraries emit, each opens with '{"' and a letter, which encodes to "eyJ". The rule now adds a (?=eyJ) lookahead to the header and payload segments and keeps the old length minimums, so it matches exactly the old matches whose first two segments begin with "eyJ". The existing k07 fixture and an HS256 token are still rejected. Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/test/artifacts.test.ts | 17 +++++++++++++++++ packages/security/src/sensitive-redaction.ts | 8 +++++--- .../security/test/sensitive-redaction.test.ts | 12 ++++++++++++ 3 files changed, 34 insertions(+), 3 deletions(-) diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 390b021cc..0f0f571c8 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -1118,6 +1118,23 @@ test("canonical publication secret gate fails closed without rewriting bytes", ( ) ); + // The positive rules must not fire on ordinary contract output either: a + // qualified identifier with three long dotted segments was JWT-shaped until + // the rule required the eyJ header and payload. + assert.doesNotThrow(() => + assertArtifactPublicationsContainNoSecrets( + new Map([ + [ + "setup/call-graph.md", + Buffer.from( + "`ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck` reverts on re-entry\n", + "utf8" + ) + ] + ]) + ) + ); + // Every credential format the previous hand-rolled patterns could name is // still rejected — via a secretlint library finding or a documented // supplemental pattern. diff --git a/packages/security/src/sensitive-redaction.ts b/packages/security/src/sensitive-redaction.ts index 5cd358d51..2954f6cc0 100644 --- a/packages/security/src/sensitive-redaction.ts +++ b/packages/security/src/sensitive-redaction.ts @@ -49,9 +49,11 @@ const SUPPLEMENTAL_SECRET_PATTERNS: readonly RegExp[] = [ /\b(?:ak|as)-[A-Za-z0-9_]{16,}\b/gu, // No Google OAuth access-token rule in the recommended preset. /\bya29\.[A-Za-z0-9._-]{20,}\b/gu, - // No JWT rule in the recommended preset; a three-part base64url token is a - // positive format identification. - /\b[A-Za-z0-9_-]{20,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\b/gu, + // No JWT rule in the recommended preset. A JWT's header and payload are + // base64url JSON objects, which encode to "eyJ" when they open with '{"' and + // a letter; that prefix is the identification. Three dotted segments alone + // also match qualified identifiers such as Contract.function.check. + /\b(?=eyJ)[A-Za-z0-9_-]{20,}\.(?=eyJ)[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\b/gu, // Provider-keyed RPC URLs embed the credential in the path; no secretlint // rule covers Alchemy/Infura project keys. /\b(?:https?|wss?):\/\/[^\s"'`]*(?:alchemy\.com\/v2\/|infura\.io\/v3\/)[A-Za-z0-9_-]{16,}\b/giu diff --git a/packages/security/test/sensitive-redaction.test.ts b/packages/security/test/sensitive-redaction.test.ts index dfe094049..59d4d20df 100644 --- a/packages/security/test/sensitive-redaction.test.ts +++ b/packages/security/test/sensitive-redaction.test.ts @@ -222,6 +222,18 @@ test("the English word bearer does not redact following prose while Bearer token ); }); +test("the JWT rule requires eyJ header and payload segments, so dotted identifiers publish", () => { + const jwt = "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dBjftJeZ4CVPmB92K27uhbUJU1p1r_wW1gFWFOEjXk"; // gitleaks:allow -- fake credential fixture for the redaction tests + assert.equal(redactSecretsInText(`session=${jwt};`, undefined, [], "positive-only"), "session=;"); + for (const fixture of [ + "ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck", + "IExampleLendingPoolCore.liquidationCall.healthFactorBefore", + "auditProfileResolution.settingOrigins.property_priority_threshold" + ]) { + assert.equal(containsSensitiveSecrets(fixture, [], "positive-only"), false, fixture); + } +}); + test("in-content secretlint-disable comments cannot suppress detection", () => { // Scanned content is agent-controlled; the filter-comments rule is disabled // so contaminated output cannot exempt itself from the fail-closed gate. From 247c75ad530c3e84b29899af66fd6751fceafe0f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:52:18 +0000 Subject: [PATCH 004/206] fix(security): let the publication gate pass Anvil's public dev mnemonic and keys Anvil prints the "test test test test test test test test test test test junk" mnemonic and the 10 keys it derives at m/44'/60'/0'/0/0-9 on every startup. Foundry tests and scripts sign with them: forge-std's own test_DeriveRememberKey quotes the mnemonic and key (0). The fail-closed publication gate rejected both. The mnemonic has a valid BIP39 checksum, and `uint256 privateKey = 0xac09...ff80;` satisfies the context-labeled private-key rule. So a generated test that derived or signed with a default Anvil account failed its node's verification. The labeled-key rule now skips those exact 10 keys (compared case-insensitively, with or without 0x), and the mnemonic rule skips a window that is exactly that phrase. The keys and the mnemonic were checked against `anvil` 1.8.3's startup output and `cast wallet private-key --mnemonic-index 0..9`. The keys are public, so exempting them hides nothing. Any other labeled 64-hex key, including key (0) with one digit changed, and any other valid mnemonic, including one that shares the first 11 words, is still rejected. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + packages/artifacts/test/artifacts.test.ts | 14 +++++- packages/security/src/sensitive-redaction.ts | 24 +++++++++- .../security/test/sensitive-redaction.test.ts | 45 +++++++++++++++++++ 4 files changed, 81 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..b44e86fbb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[security]** The fail-closed publication secret gate no longer rejects Anvil's public development accounts or long dotted identifiers in agent output. The JWT rule now requires the `eyJ` prefix that a JSON header and payload encode to, so a qualified name such as `ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck` no longer reads as a token; the labeled-private-key and mnemonic rules skip the published `test test … junk` mnemonic and the 10 keys `anvil` prints, which Foundry tests sign with. Real JWTs, other labeled 64-hex keys, and other valid mnemonics are still rejected. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 0f0f571c8..7acc0cee3 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -1119,11 +1119,21 @@ test("canonical publication secret gate fails closed without rewriting bytes", ( ); // The positive rules must not fire on ordinary contract output either: a - // qualified identifier with three long dotted segments was JWT-shaped until - // the rule required the eyJ header and payload. + // Foundry test deriving and signing with Anvil's published mnemonic and + // account (0) key, and a qualified identifier with three long dotted + // segments (JWT-shaped until the rule required the eyJ header and payload). assert.doesNotThrow(() => assertArtifactPublicationsContainNoSecrets( new Map([ + [ + "generated-tests/AnvilSigner.t.sol", + Buffer.from( + 'string memory mnemonic = "test test test test test test test test test test test junk";\n' + + "uint256 privateKey = 0xac0974bec39a17e36ba4a6b4d238ff944bacb478cbed5efcae784d7bf4f2ff80;\n" + // gitleaks:allow -- public Anvil dev key fixture for the redaction tests + "assertEq(vm.deriveKey(mnemonic, 0), privateKey);\n", + "utf8" + ) + ], [ "setup/call-graph.md", Buffer.from( diff --git a/packages/security/src/sensitive-redaction.ts b/packages/security/src/sensitive-redaction.ts index 2954f6cc0..73d7f6261 100644 --- a/packages/security/src/sensitive-redaction.ts +++ b/packages/security/src/sensitive-redaction.ts @@ -338,6 +338,26 @@ function redactExactSecretValues(value: string, placeholder: string, forbiddenSe return redacted; } +/** + * Anvil's and Hardhat's default mnemonic, and the keys of the 10 dev accounts + * Anvil derives from it and prints on startup (m/44'/60'/0'/0/0-9). They are + * published, not secret, and Foundry tests and scripts sign with them: + * forge-std's own test_DeriveRememberKey quotes the mnemonic and key (0). + */ +const PUBLIC_DEVELOPMENT_MNEMONIC = "test test test test test test test test test test test junk"; +const PUBLIC_DEVELOPMENT_PRIVATE_KEYS: ReadonlySet = new Set([ + "ac0974bec39a17e36ba4a6b4d238ff944bacb478cbed5efcae784d7bf4f2ff80", + "59c6995e998f97a5a0044966f0945389dc9e86dae88c7a8412f4603b6b78690d", + "5de4111afa1a4b94908f83103eb1f1706367c2e68ca870fc3fb9a804cdab365a", + "7c852118294e51e653712a81e05800f419141751be58f605c371e15141b007a6", + "47e179ec197488593b187f80a00eb0da91f1b9d0b13f8733639f19c30a34926a", + "8b3a350cf5c34c9194ca85829a2df0ec3153be0318b5e2d3348e872092edffba", + "92db14e403b83dfe3df233f83dfa3a0d7096f21ca9b0d6d6b8d88b2b4ec1564e", + "4bbbf85ce3377467afe5d46f804f221813b2bb87f24d81f60f1fcdbf7cbf4356", + "dbda1821b80551c9d65939329250298aa3472ba22feea921c0cf5d620ea67b97", + "2a871d0798f97d79848a013d4936a73bf4cc922c825d33c1cf7073dff6d409c6" +]); + function redactBip39Mnemonics(value: string, placeholder: string): string { type WordToken = { word: string; start: number; end: number }; let run: WordToken[] = []; @@ -359,7 +379,8 @@ function redactBip39Mnemonics(value: string, placeholder: string): string { for (const wordCount of BIP39_WORD_COUNTS) { if (run.length < wordCount) continue; const window = run.slice(-wordCount); - if (!validateMnemonic(window.map((token) => token.word).join(" "), englishWordlist)) continue; + const phrase = window.map((token) => token.word).join(" "); + if (phrase === PUBLIC_DEVELOPMENT_MNEMONIC || !validateMnemonic(phrase, englishWordlist)) continue; ranges.push({ start: window[0]!.start, end: window.at(-1)!.end }); run = []; break; @@ -382,6 +403,7 @@ function redactUnlabeledFortyHexSecrets(value: string, placeholder: string): str function redactContextLabeledPrivateKeys(value: string, placeholder: string): string { THIRTY_TWO_BYTE_HEX_PATTERN.lastIndex = 0; return value.replace(THIRTY_TWO_BYTE_HEX_PATTERN, (candidate, offset: number) => + !PUBLIC_DEVELOPMENT_PRIVATE_KEYS.has(candidate.replace(/^0x/iu, "").toLowerCase()) && /(?:private|secret|signing|wallet|ethereum|evm)[-_\s]*(?:key|scalar)\s*(?:(?:[:=]|\bis\b)\s*)?["'`]?$/iu.test( value.slice(Math.max(0, offset - 96), offset) ) diff --git a/packages/security/test/sensitive-redaction.test.ts b/packages/security/test/sensitive-redaction.test.ts index 59d4d20df..7216e8d36 100644 --- a/packages/security/test/sensitive-redaction.test.ts +++ b/packages/security/test/sensitive-redaction.test.ts @@ -222,6 +222,51 @@ test("the English word bearer does not redact following prose while Bearer token ); }); +test("positive-only scans publish Anvil's dev mnemonic and keys while other mnemonics and keys stay detected", () => { + // The mnemonic and keys `anvil` prints on startup. Foundry tests and scripts + // sign with them, so generated tests and reports quote them under a label. + const anvilMnemonic = "test test test test test test test test test test test junk"; + assert.equal(containsSensitiveSecrets(`string memory mnemonic = "${anvilMnemonic}";`, [], "positive-only"), false); + // A different valid mnemonic sharing its first 11 words is still a secret. + assert.equal( + redactSecretsInText(`mnemonic: ${anvilMnemonic.replace(/junk$/u, "absent")}`, undefined, [], "positive-only"), + "mnemonic: " + ); + + const anvilKeys = [ + "0xac0974bec39a17e36ba4a6b4d238ff944bacb478cbed5efcae784d7bf4f2ff80", + "0x59c6995e998f97a5a0044966f0945389dc9e86dae88c7a8412f4603b6b78690d", + "0x5de4111afa1a4b94908f83103eb1f1706367c2e68ca870fc3fb9a804cdab365a", + "0x7c852118294e51e653712a81e05800f419141751be58f605c371e15141b007a6", + "0x47e179ec197488593b187f80a00eb0da91f1b9d0b13f8733639f19c30a34926a", + "0x8b3a350cf5c34c9194ca85829a2df0ec3153be0318b5e2d3348e872092edffba", + "0x92db14e403b83dfe3df233f83dfa3a0d7096f21ca9b0d6d6b8d88b2b4ec1564e", + "0x4bbbf85ce3377467afe5d46f804f221813b2bb87f24d81f60f1fcdbf7cbf4356", + "0xdbda1821b80551c9d65939329250298aa3472ba22feea921c0cf5d620ea67b97", + "0x2a871d0798f97d79848a013d4936a73bf4cc922c825d33c1cf7073dff6d409c6" + ]; + for (const key of anvilKeys) { + for (const fixture of [ + `uint256 privateKey = ${key};`, + `PRIVATE_KEY=${key}`, + `signing key: ${key.slice(2).toUpperCase()}` + ]) { + assert.equal(containsSensitiveSecrets(fixture, [], "positive-only"), false, fixture); + } + } + // Only the exact published keys are exempt: key (0) with its last digit + // changed is an unknown key. + assert.equal( + redactSecretsInText( + "uint256 privateKey = 0xac0974bec39a17e36ba4a6b4d238ff944bacb478cbed5efcae784d7bf4f2ff81;", // gitleaks:allow -- fake credential fixture for the redaction tests + undefined, + [], + "positive-only" + ), + "uint256 privateKey = ;" + ); +}); + test("the JWT rule requires eyJ header and payload segments, so dotted identifiers publish", () => { const jwt = "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dBjftJeZ4CVPmB92K27uhbUJU1p1r_wW1gFWFOEjXk"; // gitleaks:allow -- fake credential fixture for the redaction tests assert.equal(redactSecretsInText(`session=${jwt};`, undefined, [], "positive-only"), "session=;"); From 6fffcf40c7a05877c10b68c583d2e2e763620640 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:58:29 +0000 Subject: [PATCH 005/206] fix(runtime): restate whole-run accounting in runtime report presentations The final-report agent copies a run summary that the host captures when the report task starts, and both runtime presentations (the verified terminal publication and the unchecked fallback report) republished that copy unchanged. Elapsed time, models, tokens and spend therefore excluded the report task itself and everything that finished after it started. On a local smoke run the published report showed 2,363,772 tokens, $5.64 and 20m 17s, while run.json recorded 4,591,556 tokens and $10.17 and the run finished after 36m 23s. projectTerminalReport and captureUnverifiedReport already read the final run.json and state.json, so withWholeRunSummary() restates elapsed time (run.json created_at to state.json finished_at) and models, tokens, spend and partial_pricing from accounting.cumulative. A missing, malformed or "unavailable" value keeps the agent's copy, and the helper never throws. The agent's own report.json and report.md are untouched. Unchecked reports also never received the run-root goal-search census, so they always said "Goal search coverage is unknown". They now read goal-search-coverage.json the same way the verified path does; an unreadable census still renders as unknown coverage instead of failing the report. Refs #1151 Co-Authored-By: Claude Opus 5.5 --- .../runtime/src/terminal-report-projection.ts | 51 +++++++ .../runtime/src/unverified-report-inputs.ts | 12 ++ packages/runtime/src/unverified-report.ts | 19 ++- packages/runtime/test/runtime.test.ts | 9 +- .../test/terminal-report-projection.test.ts | 131 +++++++++++++++++- .../runtime/test/unverified-report.test.ts | 56 ++++++++ packages/runtime/test/verified-output.test.ts | 9 +- 7 files changed, 278 insertions(+), 9 deletions(-) diff --git a/packages/runtime/src/terminal-report-projection.ts b/packages/runtime/src/terminal-report-projection.ts index 5f70cc42e..ebc6094cd 100644 --- a/packages/runtime/src/terminal-report-projection.ts +++ b/packages/runtime/src/terminal-report-projection.ts @@ -51,8 +51,59 @@ export function projectTerminalReport(input: TerminalReportProjectionInput): Can if (sourceRunId !== undefined && Reflect.get(reportMetadata, "source_run_id") !== sourceRunId) { throw new Error("Verified final report has a different source run identity"); } + report.run_metadata = withWholeRunSummary(reportMetadata as Record, metadata, state.finished_at); report.completion = completion; return projectCanonicalFinalReport(report, context); } throw new ReportUnavailableError("no successful report-agent output is available"); } + +/** + * The report agent copies a run summary that the host captured when the report task started, so its + * elapsed time and accounting miss the report task itself and anything that finished later. Runtime + * presentations restate them from run.json and the recorded finish time. A value those records do + * not provide keeps the agent's copy; malformed records are ignored, never thrown. + */ +export function withWholeRunSummary( + runMetadata: Record, + metadata: unknown, + finishedAt: unknown +): Record { + const summary = { ...runMetadata }; + const elapsed = elapsedTime(field(metadata, "created_at"), finishedAt); + if (elapsed !== undefined) summary.elapsed_time = elapsed; + const cumulative = field(field(metadata, "accounting"), "cumulative"); + const models = field(cumulative, "models"); + if (isNonEmptyStringList(models)) summary.models_used = [...models]; + const tokens = field(cumulative, "tokens_used"); + if (availableLabel(tokens)) summary.tokens_used = tokens; + const spend = field(cumulative, "estimated_spend"); + if (availableLabel(spend)) { + summary.estimated_spend = spend; + summary.partial_pricing = field(cumulative, "partial_pricing") === true; + } + return summary; +} + +function field(value: unknown, key: string): unknown { + return typeof value === "object" && value !== null && !Array.isArray(value) ? Reflect.get(value, key) : undefined; +} + +function isNonEmptyStringList(value: unknown): value is string[] { + return Array.isArray(value) && value.length > 0 && value.every((entry) => typeof entry === "string" && entry !== ""); +} + +function availableLabel(value: unknown): value is string { + return typeof value === "string" && value.trim() !== "" && value.trim().toLowerCase() !== "unavailable"; +} + +/** Same format as the host's report-start projection: `42.0s`, `5m 07s`, or `6h 02m`. */ +function elapsedTime(createdAt: unknown, finishedAt: unknown): string | undefined { + if (typeof createdAt !== "string" || typeof finishedAt !== "string") return undefined; + const totalSeconds = (Date.parse(finishedAt) - Date.parse(createdAt)) / 1_000; + if (!Number.isFinite(totalSeconds) || totalSeconds < 0) return undefined; + if (totalSeconds < 60) return `${totalSeconds.toFixed(1)}s`; + const totalMinutes = Math.floor(totalSeconds / 60); + if (totalMinutes < 60) return `${String(totalMinutes)}m ${String(Math.floor(totalSeconds % 60)).padStart(2, "0")}s`; + return `${String(Math.floor(totalMinutes / 60))}h ${String(totalMinutes % 60).padStart(2, "0")}m`; +} diff --git a/packages/runtime/src/unverified-report-inputs.ts b/packages/runtime/src/unverified-report-inputs.ts index ff9e12e2c..32c1850fa 100644 --- a/packages/runtime/src/unverified-report-inputs.ts +++ b/packages/runtime/src/unverified-report-inputs.ts @@ -13,6 +13,7 @@ import { reportSchema, type ReportVerification } from "@ultrafuzz/artifacts"; +import { loadGoalSearchCoverageSnapshot } from "./final-report-markdown.js"; import { ReportUnavailableError } from "./report-unavailable.js"; type JsonRecord = Record; @@ -31,6 +32,7 @@ export interface UnverifiedReportInputs { observed: ObservedReportCompletion; verification: ReportVerification; agentReport: JsonRecord; + goalSearchCoverage?: unknown; sources_sha256: string; } @@ -53,10 +55,20 @@ export function readUnverifiedReportInputs(root: string): UnverifiedReportInputs observed, verification: { status: "not-checked", reason_codes: [...reader.reasons].sort() }, agentReport, + goalSearchCoverage: readGoalSearchCoverage(root), sources_sha256: sha256Bytes(Buffer.from(JSON.stringify(reader.sources))) }; } +/** An unreadable census renders as unknown goal-search coverage instead of hiding the report. */ +function readGoalSearchCoverage(root: string): unknown { + try { + return loadGoalSearchCoverageSnapshot(root); + } catch { + return undefined; + } +} + class ReportInputReader { readonly reasons = new Set(["verification-unavailable"]); readonly sources: [string, string][] = []; diff --git a/packages/runtime/src/unverified-report.ts b/packages/runtime/src/unverified-report.ts index 466d93414..8489c9b32 100644 --- a/packages/runtime/src/unverified-report.ts +++ b/packages/runtime/src/unverified-report.ts @@ -14,6 +14,7 @@ import { } from "@ultrafuzz/artifacts"; import { projectCanonicalFinalReport } from "./final-report-markdown.js"; +import { withWholeRunSummary } from "./terminal-report-projection.js"; import { loadCurrentFinalReportSnapshot, publishTerminalReport, @@ -111,11 +112,19 @@ function publishUncheckedPresentation(root: string, file: string, bytes: Buffer) function captureUnverifiedReport(inputs: UnverifiedReportInputs): ReportSnapshot { validateSafeId(inputs.runId, "run ID"); - const projection = projectCanonicalFinalReport({ - ...inputs.agentReport, - verification: inputs.verification, - observed_completion: inputs.observed - }); + const projection = projectCanonicalFinalReport( + { + ...inputs.agentReport, + run_metadata: withWholeRunSummary( + inputs.agentReport.run_metadata as Record, + inputs.metadata, + inputs.state?.finished_at + ), + verification: inputs.verification, + observed_completion: inputs.observed + }, + { goalSearchCoverage: inputs.goalSearchCoverage } + ); const jsonBytes = Buffer.from(`${JSON.stringify(projection.report, null, 2)}\n`, "utf8"); const markdownBytes = Buffer.from(projection.markdown, "utf8"); const generation = sha256Bytes(Buffer.concat([jsonBytes, markdownBytes])); diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..786d1298f 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -19526,7 +19526,14 @@ test("syncRun publishes a report after recovery of a failed report agent", async assert.equal(recoveredReport.completion?.counts.succeeded, 1); assert.equal(recoveredReport.completion?.counts.failed, 0); assert.doesNotMatch(recoveredReport.markdown, /^# Ultrafuzz report — PARTIAL/u); - assert.deepEqual(recoveredReport.json, { ...finalReport.report, completion: recoveredReport.completion }); + // The run summary restates elapsed time from run.json and state.json; all review content is the agent's. + const elapsed = (recoveredReport.json as { run_metadata: { elapsed_time: string } }).run_metadata.elapsed_time; + assert.match(elapsed, /^(?:\d+\.\ds|\d+m \d{2}s)$/u); + assert.deepEqual(recoveredReport.json, { + ...finalReport.report, + run_metadata: { ...(finalReport.report.run_metadata as Record), elapsed_time: elapsed }, + completion: recoveredReport.completion + }); }); for (const variant of ["failed-verifier", "changed-output", "exhausted-loop"] as const) { diff --git a/packages/runtime/test/terminal-report-projection.test.ts b/packages/runtime/test/terminal-report-projection.test.ts index 9f1dabcd7..43c7e5b3e 100644 --- a/packages/runtime/test/terminal-report-projection.test.ts +++ b/packages/runtime/test/terminal-report-projection.test.ts @@ -133,8 +133,14 @@ test("terminal projection preserves all verified review data and warnings for co const input = { completion: census, state, metadata: metadata(), agentReport: agent }; const original = structuredClone(input); const result = projectTerminalReport(input); - assert.deepEqual(result.report, { ...before, completion: census }); - assert.deepEqual(result, projectCanonicalFinalReport({ ...before, completion: census })); + // run.json carries no accounting here, so only the whole-run elapsed time replaces the agent's copy. + const expected = { + ...before, + run_metadata: { ...(before.run_metadata as Record), elapsed_time: "2m 00s" }, + completion: census + }; + assert.deepEqual(result.report, expected); + assert.deepEqual(result, projectCanonicalFinalReport(expected)); assert.deepEqual(input, original); assert.equal(result.markdown.startsWith("# Ultrafuzz report — PARTIAL"), partial); } @@ -190,3 +196,124 @@ test("terminal projection enforces bounded and internally consistent completion assert.deepEqual(result.report.completion, bounded); assert.match(result.markdown, /identities omitted from this bounded census: `1`/u); }); + +/** A valid run.json whose cumulative accounting already includes the report task's own usage. */ +function metadataWithAccounting(cumulative: { estimated_spend?: string; models?: string[] } = {}): RunMetadataDocument { + const summary = { + uncached_input_tokens: 9_000_000, + input_tokens: 9_000_000, + output_tokens: 3_345_678, + cache_read_tokens: 0, + cache_write_tokens: 0, + reasoning_tokens: 0, + inclusive_token_total: 12_345_678, + billable_token_total: 12_345_678, + total_tokens: 12_345_678, + tokens_used: "12,345,678", + estimated_spend: "$41.20+", + estimated_spend_usd: 41.2, + component_costs_usd: { uncached_input: 30, cache_read: 0, cache_write: 0, output: 11.2, reasoning: 0 }, + usage_complete: true, + usage_incomplete_reasons: [], + pricing_complete: false, + pricing_incomplete_reasons: [{ code: "model-pricing-unavailable" as const, model: "model-b" }], + partial_pricing: true, + cache_read_pricing_estimated: false, + event_count: 2, + priced_event_count: 1, + unpriced_event_count: 1, + models: ["model-a", "model-b"], + agents: ["agent-a"] + }; + const segment = { + ...summary, + control_generation: DIGEST, + workflow_run_id: "workflow-1", + source_event_sequences: [1, 2], + attempts: [ + { node_id: "review", iteration: 0, attempt: 0 }, + { node_id: "final-report", iteration: 0, attempt: 0 } + ] + }; + return { + ...metadata(), + workflow_ids: ["workflow-1"], + workflow: { + run_id: "workflow-1", + compiled_run_id: "compiled-1", + name: "workflow", + path: "workflow.tsx", + evidence_path: "evidence.json", + expanded_graph_path: "expanded-graph.json", + config_path: "config.json", + input_path: "input.json", + tasks_path: "tasks.json", + control_integrity_path: "control-integrity.json", + control_generation: DIGEST, + workflow_link_id: "123e4567-e89b-42d3-a456-426614174000", + execution_snapshot_path: "execution-snapshot.json", + task_node_ids: ["review", "final-report"] + }, + accounting: { + schema_version: "ultrafuzz.accounting.v4", + source: "usage-ledger", + workflow_run_id: "workflow-1", + current: structuredClone(segment), + segments: [structuredClone(segment)], + cumulative: { ...summary, source_run_ids: [], ...cumulative }, + checkpoint: { + schema_version: "ultrafuzz.accounting-checkpoint.v1", + ledger_event_count: 2, + last_source_event_sequence: 2, + control_generation: DIGEST, + workflow_run_id: "workflow-1" + }, + pricing_catalog: { + source: "configured-catalog", + status: "available", + fetched_at: CREATED_AT, + resolved_models: ["model-a"], + unresolved_models: ["model-b"], + model_prices: { "model-a": { inputUsdPerMillion: 1, outputUsdPerMillion: 2 } } + }, + updated_at: FINISHED_AT + } + }; +} + +function runSummaryLines(markdown: string): string[] { + return markdown + .split("\n") + .filter((line) => /^- (?:Elapsed time|Models used|Tokens used|Estimated spend):/u.test(line)); +} + +test("terminal projection restates whole-run accounting instead of the report-start snapshot", () => { + const input = { + completion: completion(), + state: { ...terminalState(), finished_at: "2026-09-01T06:02:00.000Z" }, + metadata: metadataWithAccounting(), + agentReport: agentReport() + }; + const result = projectTerminalReport(input); + assert.deepEqual(runSummaryLines(result.markdown), [ + "- Elapsed time: `6h 02m`", + "- Models used: `model-a, model-b`", + "- Tokens used: `12,345,678`", + "- Estimated spend: `$41.20+`" + ]); + assert.equal((result.report.run_metadata as Record).partial_pricing, true); + assert.deepEqual(result.report.issues, agentReport().issues); + + // Values the run records do not have keep the agent's copy. + const sparse = projectTerminalReport({ + ...input, + metadata: metadataWithAccounting({ estimated_spend: "unavailable", models: [] }) + }); + assert.deepEqual(runSummaryLines(sparse.markdown), [ + "- Elapsed time: `6h 02m`", + "- Models used: `example-model`", + "- Tokens used: `12,345,678`", + "- Estimated spend: `$0.01`" + ]); + assert.equal((sparse.report.run_metadata as Record).partial_pricing, false); +}); diff --git a/packages/runtime/test/unverified-report.test.ts b/packages/runtime/test/unverified-report.test.ts index 6e7456ee3..80fbb773d 100644 --- a/packages/runtime/test/unverified-report.test.ts +++ b/packages/runtime/test/unverified-report.test.ts @@ -231,3 +231,59 @@ test("failure to save the status cache does not suppress explicit report access" assert.equal(loadReportSnapshot(root).terminal, true); assert.equal(readReportPublicationStatus(root, state).status, "unknown"); }); + +test("unchecked reports restate whole-run accounting from run.json", () => { + const root = reportRun("whole-run-accounting"); + writeAgentReport(root); + const runId = path.basename(root); + const state = JSON.parse(fs.readFileSync(path.join(root, "state.json"), "utf8")); + fs.writeFileSync( + path.join(root, "state.json"), + JSON.stringify({ ...state, finished_at: "2026-09-01T03:30:00.000Z" }) + ); + fs.writeFileSync( + path.join(root, "run.json"), + JSON.stringify({ + run_id: runId, + created_at: "2026-09-01T00:00:00.000Z", + accounting: { + cumulative: { models: ["model-a"], tokens_used: "4,321", estimated_spend: "$3.00", partial_pricing: false } + } + }) + ); + const report = loadReportSnapshot(root); + assert.equal(report.verification, "not-checked"); + assert.match(report.markdown, /^- Elapsed time: `3h 30m`$/mu); + assert.match(report.markdown, /^- Models used: `model-a`$/mu); + assert.match(report.markdown, /^- Tokens used: `4,321`$/mu); + assert.match(report.markdown, /^- Estimated spend: `\$3\.00`$/mu); + assert.equal(reportSchema.parse(report.json).run_metadata.partial_pricing, false); + assertReportSnapshotRemainedCurrent(report); +}); + +test("unchecked reports render the run's goal-search census", () => { + const root = reportRun("goal-search-census"); + writeAgentReport(root); + const censusPath = path.join(root, "goal-search-coverage.json"); + fs.writeFileSync( + censusPath, + JSON.stringify({ + schema_version: "ultrafuzz.goal-search-coverage.v1", + run_id: path.basename(root), + totals: { planned: 2 }, + goals: [ + { node_id: "dynamic:class:1", logical_node_id: "class-goals", status: "completed-no-findings" }, + { node_id: "dynamic:class:2", logical_node_id: "class-goals", status: "stopped-early" } + ] + }) + ); + const report = loadReportSnapshot(root); + assert.match(report.markdown, /^- Targeted goal search lanes: `2`\n- Completed with a verified result: `1`$/mu); + assert.doesNotMatch(report.markdown, /Goal search coverage is unknown/u); + + // A census the reader refuses still leaves the report readable, with coverage stated as unknown. + const outside = path.join(temporaryRoot("ultrafuzz-outside-"), "goal-search-coverage.json"); + fs.renameSync(censusPath, outside); + fs.symlinkSync(outside, censusPath); + assert.match(loadReportSnapshot(root).markdown, /Goal search coverage is unknown/u); +}); diff --git a/packages/runtime/test/verified-output.test.ts b/packages/runtime/test/verified-output.test.ts index 89a23a7cc..8046cb947 100644 --- a/packages/runtime/test/verified-output.test.ts +++ b/packages/runtime/test/verified-output.test.ts @@ -245,7 +245,14 @@ test("terminal presentation discloses tolerated failures and preserves verified assert.equal(published.completion?.counts.failed, 1); assert.equal(published.completion?.outcome, "partial"); const expected = JSON.parse(fixture.reportBytes.toString("utf8")) as Record; - assert.deepEqual(published.json, { ...expected, completion: published.completion }); + // The run summary restates elapsed time from run.json and state.json; all review content is the agent's. + const elapsed = (published.json as { run_metadata: { elapsed_time: string } }).run_metadata.elapsed_time; + assert.match(elapsed, /^\d+\.\ds$/u); + assert.deepEqual(published.json, { + ...expected, + run_metadata: { ...(expected.run_metadata as Record), elapsed_time: elapsed }, + completion: published.completion + }); assert.match(published.markdown, /producer/u); }); From b6a2a6805ace3bda88677648de71f51badabcb29 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:58:29 +0000 Subject: [PATCH 006/206] fix(runtime): stop the canonical report renderer rejecting its own output projectCanonicalFinalReport re-validates the Markdown it just rendered with regex rules written for hand-authored reports. publicProse did not escape `](`, so upstream prose that the review gates require byte-for-byte (a spec link, an image, `handlers[id](payload)`, a relative link whose `_` gets escaped) tripped the link or image rule. The link rule also rejected the renderer's own index anchors for any issue title with a non-ASCII letter, and the public projection threw when a redacted path was followed by `(` (`[redacted-path](line 12)`). The agent cannot repair byte-preserved fields, so every retry failed the same way and the run ended without a report. publicProse now escapes `(` after every `]`, so prose cannot form an inline link or image, and the image and link rules are deleted; the raw-HTML rule stays. Rendering of prose without `](` is unchanged, so existing reports keep verifying. The audit-context renderer (appendAuditContext, safeReportLink, SAFE_REPORT_RELATIVE_LINK_PATTERN) is deleted: report@3 has no audit_context field and rejects unknown keys, so it could never run. The prompt's `## Audit context` instructions, which asked the agent to write a section that the byte-exact canonical render never contains, are removed. Refs #1151 Co-Authored-By: Claude Opus 5.5 --- .ultrafuzz/prompts/review/final-report.md | 19 ---- packages/runtime/src/final-report-markdown.ts | 41 ++------- .../test/final-report-markdown.test.ts | 87 +++++++++++++++++++ 3 files changed, 93 insertions(+), 54 deletions(-) diff --git a/.ultrafuzz/prompts/review/final-report.md b/.ultrafuzz/prompts/review/final-report.md index 70008c7ae..ab9a4965c 100644 --- a/.ultrafuzz/prompts/review/final-report.md +++ b/.ultrafuzz/prompts/review/final-report.md @@ -385,25 +385,8 @@ Ultrafuzz is an automated smart-contract fuzzing campaign assistant. Issues belo - Tokens used: `` - Estimated spend: `` - Audit profile: `` - -## Audit context - -- Threat model: [THREAT_MODEL.md](); [threat-model.json]() -- Goal plan: [goal-plan.json]() ``` -Render `## Audit context` with exactly this heading, bullet order, and link -text, immediately after `## Run summary`. Use repository-relative or -report-relative paths to the run's own `threat-model` and `goal-plan` artifacts; -never absolute paths or external URLs. Omit an individual link whose artifact -the run did not produce, omit the `Goal plan` bullet when there is no goal plan, -and omit the whole section when the run produced none of them. Do not invent a -different heading, ordering, or link text: `ultrafuzz report` regenerates this -exact section deterministically from the run's own artifacts and overwrites -anything else. -Keep detailed threat content in those dedicated artifacts; do not duplicate it -in `report.md`. - Each production issue entry must use exactly this Markdown section order. The following example is structural only; replace the title, actor names, actions, outcomes, explanations, code, variants, and strategy IDs with issue-specific @@ -825,8 +808,6 @@ Before finishing, verify that: - `report.md` contains `## Property provenance`, including every property-derived finding and no invented property IDs for non-property findings. -- `report.md` renders the fixed `## Audit context` section for every artifact - the run produced, without copying their detailed analysis. - `report.md` contains `## Property implementation coverage` rendered from the exact runtime-authoritative coverage object. - `report.md` contains `## Goal search coverage` with counts recomputed from the diff --git a/packages/runtime/src/final-report-markdown.ts b/packages/runtime/src/final-report-markdown.ts index e60e04716..03c17f2b2 100644 --- a/packages/runtime/src/final-report-markdown.ts +++ b/packages/runtime/src/final-report-markdown.ts @@ -21,7 +21,6 @@ import { redactSecretsInText, type SecretScanMode } from "@ultrafuzz/security"; export const MAX_FINAL_REPORT_JSON_BYTES = 64 * 1024 * 1024; export const MAX_FINAL_REPORT_MARKDOWN_BYTES = 16 * 1024 * 1024; -const SAFE_REPORT_RELATIVE_LINK_PATTERN = /^\.\.\/(?:(?!\.\.?\/)[A-Za-z0-9._-]+\/)+(?!\.\.?$)[A-Za-z0-9._-]+$/u; /** * Secret placeholder for the public projection only. Two constraints pick it: * @@ -29,10 +28,9 @@ const SAFE_REPORT_RELATIVE_LINK_PATTERN = /^\.\.\/(?:(?!\.\.?\/)[A-Za-z0-9._-]+\ * the placeholder must be a fixed point of the redaction pass. The key-name assignment rule's * unquoted value class stops at whitespace, `,`, `;`, `]`, and `}`, so a placeholder containing * any of those is re-redacted on the next pass (`token=[redacted]` becomes `token=[redacted]]`). - * - The final-review Markdown gate rejects raw HTML, images, and links outside fenced code, and - * redacted values land unescaped in inline code (run summary values, coverage paths, source - * nodes), so the placeholder must not read as HTML (``), a link (`[redacted](`), - * an image, or emphasis (`*`, `_`). + * - The final-review Markdown gate rejects raw HTML outside fenced code, including inside inline + * code, where redacted values land unescaped (run summary values, coverage paths, source nodes), + * so the placeholder must not read as HTML (``). * * A bare uppercase word satisfies both. Every other redaction keeps the security package's default. */ @@ -362,12 +360,7 @@ function finalReportProseDirectiveViolation(prose: string): string | undefined { /(?:^|\n)- (?:Strategy loops|Audit profile catalog digest|Topology digest|Prompt digest|Expanded graph fingerprint):/imu, "contains legacy run metadata" ], - [/<[A-Za-z][^>]*>/u, "contains raw HTML outside fenced code"], - [/!\[[^\]]*\]\(/u, "contains an embedded image outside fenced code"], - [ - /(?]*>/u, "contains raw HTML outside fenced code"] ]; return forbiddenPatterns.find(([pattern]) => pattern.test(prose))?.[1]; } @@ -695,7 +688,6 @@ function renderCanonicalReport(report: JsonRecord, goalSearchCoverage: unknown): lines, isRecord(report.run_metadata) ? report.run_metadata.artifact_validation_warnings : undefined ); - appendAuditContext(lines, report.audit_context); const campaignDidNotRun = appendCampaignOutcome(lines, report.campaign_outcome); appendCoverageEvidence(lines, report.coverage_evidence); const goalCoverage = summarizeGoalSearchCoverage(goalSearchCoverage); @@ -961,29 +953,6 @@ function appendRunSummary(lines: string[], metadata: JsonRecord): void { } } -function appendAuditContext(lines: string[], value: unknown): void { - if (!isRecord(value)) return; - const threat = recordField(value, "threat_model"); - const goalPlan = recordField(value, "goal_plan"); - const threatMarkdown = safeReportLink(threat?.markdown); - const threatJson = safeReportLink(threat?.json); - const goalPlanJson = safeReportLink(goalPlan?.json); - if (threatMarkdown === undefined && threatJson === undefined && goalPlanJson === undefined) return; - lines.push("", "## Audit context", ""); - if (threatMarkdown !== undefined || threatJson !== undefined) { - const links = [ - threatMarkdown === undefined ? undefined : `[THREAT_MODEL.md](${threatMarkdown})`, - threatJson === undefined ? undefined : `[threat-model.json](${threatJson})` - ].filter((entry): entry is string => entry !== undefined); - lines.push(`- Threat model: ${links.join("; ")}`); - } - if (goalPlanJson !== undefined) lines.push(`- Goal plan: [goal-plan.json](${goalPlanJson})`); -} - -function safeReportLink(value: unknown): string | undefined { - return typeof value === "string" && SAFE_REPORT_RELATIVE_LINK_PATTERN.test(value) ? value : undefined; -} - function appendProductionIssue(lines: string[], rendered: RenderedIssue): void { const { issue } = rendered; lines.push("", renderedIssueHeading(rendered), "", publicProse(issueDescription(issue)), "", "### Severity", ""); @@ -1522,6 +1491,7 @@ function recordTitle(record: JsonRecord, fallback: string): string { return typeof record.title === "string" && record.title.trim().length > 0 ? record.title.trim() : fallback; } +/** Escaping `(` after every `]` keeps byte-preserved prose from forming an inline link or image. */ function publicProse(value: string): string { return value .replace(/\s+/gu, " ") @@ -1533,6 +1503,7 @@ function publicProse(value: string): string { .replaceAll("!", "\\!") .replaceAll("#", "\\#") .replaceAll("~", "\\~") + .replaceAll("](", "]\\(") .replaceAll("<", "<") .replaceAll(">", ">"); } diff --git a/packages/runtime/test/final-report-markdown.test.ts b/packages/runtime/test/final-report-markdown.test.ts index 9c1dd9a5e..ff8555706 100644 --- a/packages/runtime/test/final-report-markdown.test.ts +++ b/packages/runtime/test/final-report-markdown.test.ts @@ -27,6 +27,7 @@ test("final reports retain artifact warnings and their context without changing import { validateSafeId, type ReportCompletion } from "@ultrafuzz/artifacts"; import { redactSecretsInText } from "@ultrafuzz/security"; +import { fromMarkdown } from "mdast-util-from-markdown"; import { isDirectiveConformingFinalReportMarkdown, @@ -1406,3 +1407,89 @@ test("goal search coverage is rendered into the Markdown report instead of only ); assert.match(roamingOnly.markdown, /^No issues were reported, but no targeted goal search lane ran, /mu); }); + +interface MarkdownNode { + type: string; + url?: string; + value?: string; + children?: MarkdownNode[]; +} + +/** Flatten the CommonMark tree so assertions describe what a reader sees, not escape bytes. */ +function markdownNodes(markdown: string): Array<{ type: string; url?: string; text: string }> { + const text = (node: MarkdownNode): string => node.value ?? (node.children ?? []).map(text).join(""); + const nodes: Array<{ type: string; url?: string; text: string }> = []; + const walk = (node: MarkdownNode): void => { + nodes.push({ type: node.type, ...(node.url === undefined ? {} : { url: node.url }), text: text(node) }); + for (const child of node.children ?? []) walk(child); + }; + walk(fromMarkdown(markdown) as unknown as MarkdownNode); + return nodes; +} + +test("upstream prose with link or image syntax renders as literal text", () => { + const report = renderableReport(); + const [issue] = report.issues as Array>; + if (issue === undefined) throw new Error("missing issue fixture"); + const description = + "See [the spec](https://example.com/spec), ![flow](https://example.com/flow.png), " + + "[THREAT_MODEL.md](../threat-model/THREAT_MODEL.md), and handlers[id](payload)."; + issue.description = description; + (issue.proof_of_concept as Record).scenario = ["Call handlers[id](payload).", "Observe it."]; + const before = structuredClone(report); + + const projection = projectCanonicalFinalReport(report); + assert.deepEqual(projection.report, before); + const nodes = markdownNodes(projection.markdown); + assert.deepEqual( + nodes.filter((node) => node.type === "image" || (node.type === "link" && !node.url?.startsWith("#"))), + [] + ); + assert.ok(nodes.some((node) => node.type === "paragraph" && node.text === description)); + assert.ok(nodes.some((node) => node.type === "paragraph" && node.text === "Call handlers[id](payload).")); +}); + +test("issue titles with non-ASCII letters render with index anchors that resolve to their headings", () => { + // GitHub heading slugs: lowercase, keep letters, marks, digits, spaces, "-" and "_", then spaces become "-". + const slug = (text: string): string => + text + .trim() + .toLowerCase() + .replace(/[^\p{L}\p{M}\p{N}\s_-]/gu, "") + .replace(/\s/gu, "-"); + for (const title of ["Δ-neutral rebalance drifts", "Naïve [share] math"]) { + const report = renderableReport(); + const [issue] = report.issues as Array>; + if (issue === undefined) throw new Error("missing issue fixture"); + issue.title = `[L-01] - ${title}`; + for (const entry of report.property_provenance as Array>) entry.title = issue.title; + + const nodes = markdownNodes(projectCanonicalFinalReport(report).markdown); + const headings = new Set(nodes.filter((node) => node.type === "heading").map((node) => slug(node.text))); + const anchors = nodes.filter((node) => node.type === "link" && node.url?.startsWith("#")); + assert.ok(anchors.length > 0, title); + for (const anchor of anchors) assert.ok(headings.has(anchor.url?.slice(1) ?? ""), `${title}: ${anchor.url ?? ""}`); + } +}); + +test("public projection keeps a redacted path followed by a parenthesis as literal text", () => { + const report = renderableReport(); + const [issue] = report.issues as Array>; + if (issue === undefined) throw new Error("missing issue fixture"); + issue.description = "The reproducer at /srv/customer/private/Repro.t.sol(line 12) fails."; + + const published = projectPublicCanonicalFinalReport(report); + assert.equal( + (published.report.issues as Array>)[0]?.description, + "The reproducer at [redacted-path](line 12) fails." + ); + const nodes = markdownNodes(published.markdown); + assert.equal( + nodes.some((node) => node.type === "link" && !node.url?.startsWith("#")), + false + ); + assert.ok( + nodes.some((node) => node.type === "paragraph" && node.text === "The reproducer at [redacted-path](line 12) fails.") + ); + assertPublicProjectionFixedPoint(published); +}); From c51b310a4a9cd3c473c5fa89504ebe0a8be6fb59 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:58:29 +0000 Subject: [PATCH 007/206] docs: describe whole-run report summaries and literal prose links The artifacts reference claimed report.md links to THREAT_MODEL.md, threat-model.json and goal-plan.json (only dead code ever did) and that terminal presentations keep the report-start accounting snapshot. Document what runtime presentations now restate, that prose link syntax renders as literal text, and the upgrade effect on existing terminal publications. Refs #1151 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/artifacts-reports.md | 35 +++++++++++++++++------------ 2 files changed, 22 insertions(+), 14 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..732c71a9d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [prompts] [docs]** Runtime report presentations (the verified terminal publication and unchecked reports) now restate elapsed time, models, tokens, and estimated spend from `run.json` and the recorded finish time instead of the snapshot the report agent received when its task started, and unchecked reports render the run's goal-search census instead of always stating unknown coverage. The canonical renderer escapes `](` in prose, so it no longer rejects its own output when byte-preserved upstream prose contains link or image syntax, when an issue title contains a non-ASCII letter, or when the public projection redacts a path followed by `(`; the dead audit-context renderer and the prompt's contradictory `## Audit context` instructions are removed. Markdown for reports without `](` in prose is byte-identical, but terminal publications written by earlier versions read back as `not-checked` once the restated summary differs from their stored bytes (usually at least the elapsed time), and an agent report whose prose held a previously accepted `](` form (a `#anchor` or `../` link, or `\](`) no longer re-verifies (#1151). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index 2683bb70e..dff578046 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -537,8 +537,9 @@ artifacts/final-report/report.json `report.json` must satisfy `ultrafuzz/report@3` with the exact `ultrafuzz.report.v3` version literal. After a run stops, the runtime can format -that report and attach whole-run completion information without changing the -agent's files. Verified publications use: +that report, attach whole-run completion information, and restate the run +summary's elapsed time and accounting without changing the agent's files. +Verified publications use: ```text review/runtime-report//report.json @@ -741,11 +742,12 @@ this section with the typed handoff and rejects missing, duplicated, reordered, or bare coverage scores. Raw `covg-eval` output is for iteration only and defines neither published declaration-completeness view. -Current-run `report.md` contains concise links to `THREAT_MODEL.md`, -`threat-model.json`, and `goal-plan.json`, plus source-node provenance for each -production issue. Detailed threat analysis stays in the dedicated threat-model -artifacts and is not duplicated into the report. `report.json` preserves the -same `source_nodes` arrays. +Current-run `report.md` contains source-node provenance for each production +issue and does not link to other run files. Detailed threat analysis stays in +the dedicated threat-model artifacts and is not duplicated into the report. +`report.json` preserves the same `source_nodes` arrays. Inline link and image +syntax inside report prose, including prose preserved byte-for-byte from +upstream findings, renders as literal text. When workflow usage data is available, run metadata includes `accounting.cumulative.tokens_used` and @@ -755,13 +757,18 @@ available cumulative values into the markdown run summary and into persisted estimate is partial because some token usage did not have pricing data. -If cumulative metadata has not synchronized when the final-report producer -starts, its live Smithers fallback is a snapshot through that producer's start. -It includes earlier attempts but cannot include the producer's own eventual -duration, model fallback, tokens, or cost. A terminal presentation of an existing -verified agent report preserves those accounting values. Report v3 has no -metric-scope field, so use -`ultrafuzz stats` after terminal synchronization for closed-run accounting. +The final-report producer receives its run summary when its task starts: from +cumulative metadata when it has synchronized, otherwise from a live Smithers +fallback. Either way it is a snapshot through that producer's start. It includes +earlier attempts but cannot include the producer's own eventual duration, model +fallback, tokens, or cost, and the agent's `report.json` and `report.md` keep +that snapshot. Runtime presentations (the verified terminal publication and +unchecked reports) restate the run summary instead: elapsed time from +`run.json#created_at` to `state.json#finished_at`, and models, tokens, +estimated spend, and `partial_pricing` from the current +`accounting.cumulative`. A value those records lack, or record as +`unavailable`, keeps the agent's copy. Use `ultrafuzz stats` for the full +accounting breakdown. `accounting.segments` publishes one rollup per checkpoint generation, and `accounting.current` identifies the latest segment. Each segment retains every From 87c57d95d422fd24eb24d1d769ba9d60b86d8811 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:01:07 +0000 Subject: [PATCH 008/206] fix(runtime): resume keeps the run's agent auth config Native continuation set ULTRAFUZZ_CONFIG_PATH to /smithers/resolved-config.json. Every stock agent adapter reads that path with its TOML reader, which finds no [agents.*] tables in JSON and returns an empty config, so each adapter fell back to its defaults on resume: auth, api_key_env and config_dir were silently dropped. With the default ultrafuzz.toml (CodexAgent auth = "api-key"), every resumed Codex task ran with subscription auth and a cleared OPENAI_API_KEY. Launch hands adapters a sealed copy of /smithers/execution-config.toml, the TOML rendering of the same resolved config. Resume now points them at that file. It is written at compile time since #635, and resume only sets the variable when resolved-config.json parses as the current v4 schema (#1120), so every run that reaches this branch was compiled with it. The new regression test resumes a run, rebuilds the generated CodexAgent from the config path the runner received, and checks it still carries the API key. It fails on main (CODEX_API_KEY is empty) and passes with this change. "native resume delegates the persisted workflow after mutable project sources are replaced" asserted the JSON path, i.e. the bug; it now asserts the TOML path. Co-Authored-By: Claude Opus 5.5 --- docs/reference/cli.md | 3 +- packages/runtime/src/start-run.ts | 6 +++- packages/runtime/test/runtime.test.ts | 46 ++++++++++++++++++++++++++- 3 files changed, 52 insertions(+), 3 deletions(-) diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..68bd3bd33 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -323,7 +323,8 @@ schema bindings, and metadata projections remain provenance for inspection; they are not resume authorization. Smithers decides which finished rows can be reused and which newly rendered or unfinished tasks run. Ultrafuzz does not rewrite historical artifacts or automatically reset, replay, timetravel, or -fork completed work. +fork completed work. Agent adapters in the continued workflow read the same +agent config as at launch, from the run's `smithers/execution-config.toml`. `resume --refresh-controller` first renders the currently installed Ultrafuzz controller and stock adapters beside the historical source, then delegates to diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index ba32e842b..2e8f44da9 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -526,6 +526,10 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { const smithersRoot = safeResolveInside(layout.root, "smithers", "Smithers evidence"); const tasksPath = safeResolveInside(smithersRoot, "tasks.json", "workflow task manifest"); const configPath = safeResolveInside(smithersRoot, "resolved-config.json", "workflow config"); + // Agent adapters parse ULTRAFUZZ_CONFIG_PATH as TOML; given the JSON above + // they find no agent tables and fall back to default auth. Launch writes + // the same config as TOML beside it and hands adapters a copy of that file. + const agentConfigPath = safeResolveInside(smithersRoot, "execution-config.toml", "workflow agent config"); let taskDocument: SmithersTaskManifestDocument | undefined; let config: ResolvedConfig | undefined; if (fs.existsSync(tasksPath)) { @@ -571,7 +575,7 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { ...forgeGuard.env, ULTRAFUZZ_ARTIFACTS_MODULE: import.meta.resolve("@ultrafuzz/artifacts"), ULTRAFUZZ_RUNTIME_MODULE: import.meta.resolve("@ultrafuzz/runtime"), - ...(config === undefined ? {} : { ULTRAFUZZ_CONFIG_PATH: configPath }), + ...(config === undefined ? {} : { ULTRAFUZZ_CONFIG_PATH: agentConfigPath }), ULTRAFUZZ_WORKFLOW_PERSISTED_PATH: workflowPath }; let trustedCli: TrustedCliEnvironment = { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..5fd3a3f9f 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -2503,6 +2503,16 @@ function controllerRefreshTerminalEnv( }); } +/** Make a fake lifecycle runner append `$` to `logPath` on every `up`. */ +function logFakeRunnerUpVariable(env: Record, name: string, logPath: string): void { + const shim = env.SMITHERS_BIN; + assert.ok(shim); + fs.writeFileSync( + shim, + fs.readFileSync(shim, "utf8").replace(" up)\n", ` up)\n printf '%s\\n' "$${name}" >> ${shellQuote(logPath)}\n`) + ); +} + function workflowEvents( workflowRunId: string, events: Array<{ @@ -15383,7 +15393,7 @@ test("native resume delegates the persisted workflow after mutable project sourc assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); const consumed = fs.readFileSync(snapshotBytesLog, "utf8"); assert.equal(consumed.includes(`workflow=${mutableWorkflow}\n`), true); - assert.match(consumed, /^config=.*\/smithers\/resolved-config\.json$/mu); + assert.match(consumed, /^config=.*\/smithers\/execution-config\.toml$/mu); assert.match(consumed, /^agent=.*\/\.smithers\/agents\/codex\.ts$/mu); assert.match(consumed, /HostileReplacement/u); assert.match(consumed, /export const hostile/u); @@ -25060,6 +25070,40 @@ test("native continuation does not use historical trusted CLI identity as an aut ); }); +test("native continuation hands generated agents the run's TOML config, so CodexAgent keeps API-key auth", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "continuation-agent-config"; + const env = controllerRefreshTerminalEnv(project, runId); + const configPathLog = path.join(project, "fake-smithers-up-config-path.log"); + logFakeRunnerUpVariable(env, "ULTRAFUZZ_CONFIG_PATH", configPathLog); + const launched = await startRun({ projectRoot: project, runId, env }); + assert.equal(launched.ok, true, JSON.stringify(launched.diagnostics)); + fs.writeFileSync(configPathLog, "", "utf8"); + const resumed = await resumeRun({ projectRoot: project, runId, env }); + assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); + assert.equal(resumed.value?.submitted, true); + + // Build the stock adapter from the config path the resumed runner received. + // The init config selects `auth = "api-key"` for CodexAgent; an adapter that + // cannot read it falls back to subscription auth and clears the key. + const { createCodexAgent } = await loadGeneratedCodexAgent(project); + const previous = { config: process.env.ULTRAFUZZ_CONFIG_PATH, key: process.env.OPENAI_API_KEY }; + process.env.ULTRAFUZZ_CONFIG_PATH = fs.readFileSync(configPathLog, "utf8").trim(); + process.env.OPENAI_API_KEY = "continuation-codex-key"; + try { + const agent = createCodexAgent() as { opts: { env: Record } }; + assert.equal(agent.opts.env.CODEX_API_KEY, "continuation-codex-key"); + assert.equal(agent.opts.env.OPENAI_API_KEY, "continuation-codex-key"); + } finally { + if (previous.config === undefined) delete process.env.ULTRAFUZZ_CONFIG_PATH; + else process.env.ULTRAFUZZ_CONFIG_PATH = previous.config; + if (previous.key === undefined) delete process.env.OPENAI_API_KEY; + else process.env.OPENAI_API_KEY = previous.key; + } +}); + test("controller refresh authenticates newly required sealed runner patches and rejects source drift", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From 628f4c436922b4601a62f57dea7d46f6711ec50e Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:10:39 +0000 Subject: [PATCH 009/206] test(runtime): pin why prose escapes the parenthesis rather than the bracket Escaping `](` as `\](` would break the public projection on two shapes: after publicProse doubles a prose backslash, `\\\](` matches the UNC private-path pattern, and `token=REDACTED\](` extends the redacted assignment value so the fixed-point re-scan redacts it again. Both public projections succeed with the `]\(` escape; with `\](` this test fails. Refs #1151 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/test/final-report-markdown.test.ts | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/packages/runtime/test/final-report-markdown.test.ts b/packages/runtime/test/final-report-markdown.test.ts index ff8555706..215bc4cc7 100644 --- a/packages/runtime/test/final-report-markdown.test.ts +++ b/packages/runtime/test/final-report-markdown.test.ts @@ -1492,4 +1492,11 @@ test("public projection keeps a redacted path followed by a parenthesis as liter nodes.some((node) => node.type === "paragraph" && node.text === "The reproducer at [redacted-path](line 12) fails.") ); assertPublicProjectionFixedPoint(published); + + // An escape inserted before `]` would read as a UNC path after a doubled backslash, and would + // extend a redacted assignment value so the fixed-point re-scan redacts it again. + for (const description of ["Match a literal \\](x) in the parser.", "Set token=synthetic-escape-secret](x)."]) { + issue.description = description; + assertPublicProjectionFixedPoint(projectPublicCanonicalFinalReport(report)); + } }); From 336a2f8e57cf6685e28c69a3497a2baa60c8ed7f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:10:46 +0000 Subject: [PATCH 010/206] fix(runtime): literal braces in prompt text no longer fail the workflow render renderAgentPrompt substituted the operator and task prompts into the trusted agent-prompt template and then regex-scanned the whole result for leftover `{{word}}` placeholders. Inserted text that merely contained such a sequence threw inside the Smithers render: an escaped `\{{word}}` example in a prompt template (the prompt renderer emits it as literal `{{word}}`), the same escape inside a dynamic item value (kept verbatim), or an operator `--prompt` note. Every render builds every task's prompt, so the run failed, and failed again on every resume. Check only the template's own placeholders, inside the replace callback, as renderAgentPreambleTemplate already does. A template placeholder with no value still throws, and now names the placeholder. The dynamic prompt renderer also resolves `{{...}}` inside goal-plan replacement values, and an unbound name there throws in the same render. The goal-plan contract now rejects `{{` in replacement values, which are plain-text labels, so that failure lands on goal-plan's own verify instead of on every render. The check is a zod refinement; goal-plan.schema.json is unchanged. Co-Authored-By: Claude Opus 5.5 --- docs/reference/prompt-variables.md | 4 ++ packages/artifacts/src/goal-plan.ts | 6 +++ .../test/threat-goal-artifacts.test.ts | 28 ++++++++++++++ .../templates/smithers/workflows/workflow.tsx | 14 +++---- .../test/generated-workflow-verifier.test.ts | 38 +++++++++++++++++++ 5 files changed, 83 insertions(+), 7 deletions(-) diff --git a/docs/reference/prompt-variables.md b/docs/reference/prompt-variables.md index 32d78ba75..fab6c9744 100644 --- a/docs/reference/prompt-variables.md +++ b/docs/reference/prompt-variables.md @@ -174,6 +174,10 @@ bounded, item-scoped `replacements` map. Namespaced keys such as `{{liquidation:overdue}}` resolve only from that item. A value such as `{{item.goal_prompt}}` may retain those placeholders for the bounded nested replacement pass; unresolved, cyclic, non-scalar, or non-item references fail. +That failure happens while the workflow renders, so it stops the whole run +rather than one node. The shipped `ultrafuzz/goal-plan@1` contract therefore +rejects `{{` anywhere in a replacement value, and a bad label fails +`goal-plan`'s own verification instead. ## Output Contract diff --git a/packages/artifacts/src/goal-plan.ts b/packages/artifacts/src/goal-plan.ts index b45e679f2..7cadda1ad 100644 --- a/packages/artifacts/src/goal-plan.ts +++ b/packages/artifacts/src/goal-plan.ts @@ -123,6 +123,12 @@ const replacementValue = nonEmptyString message: "Replacement values must be a human-readable title, not a serialized JSON record: keep the record in the " + "artifact directory and reference it by path so the goal sentence stays one sentence" + }) + // The dynamic-node renderer resolves `{{...}}` inside replacement values, and a reference it cannot + // bind throws inside the workflow render, which fails the whole run on every resume. A label has no + // reason to carry template syntax, so reject it here, where the failure is goal-plan's own verify. + .refine((value) => !value.includes("{{"), { + message: "Replacement values must be plain-text labels and must not contain template braces '{{'" }); const replacementsSchema = z diff --git a/packages/artifacts/test/threat-goal-artifacts.test.ts b/packages/artifacts/test/threat-goal-artifacts.test.ts index 1c0df73cb..03f89ade0 100644 --- a/packages/artifacts/test/threat-goal-artifacts.test.ts +++ b/packages/artifacts/test/threat-goal-artifacts.test.ts @@ -1348,6 +1348,34 @@ test("goal-plan replacements are bounded titles, not inlined JSON records", () = assert.equal(validateGoalPlan(withReplacement(true)).ok, true); }); +test("goal-plan replacement values cannot carry template placeholders into the dynamic render", () => { + const plan = goalPlanFixture(); + const threatId = "liquidation:overdue"; + const withReplacement = (value: string): Record => { + const next = structuredClone(plan); + const [goal] = next.threat_goals as Array>; + assert.ok(goal); + goal.replacements = { [threatId]: value }; + return next; + }; + + // The renderer resolves `{{...}}` inside a replacement value, and an unbound name such as `amount` + // throws inside the workflow render. Item references and escaped braces are template syntax too, + // which a label never needs, so the verifier rejects every form rather than re-deriving the renderer. + for (const label of [ + "Late tick lets {{amount}} round down", + "Overdue liquidation of {{item.id}}", + "Late tick lets \\{{amount}} round down" + ]) { + const rejected = validateGoalPlan(withReplacement(label)); + assert.equal(rejected.ok, false, label); + assert.match(rejected.issues.map((issue) => issue.message).join("; "), /must not contain template braces/u); + } + + // Single braces are prose, as in the bounded-titles test above. + assert.equal(validateGoalPlan(withReplacement("{withdraw} settles before the late tick")).ok, true); +}); + test("escaped required goal placeholders are rejected instead of rendering as literals", () => { const plan = goalPlanFixture(); const goal = (plan.threat_goals as Array>)[0]!; diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index ceab70669..d22794bc4 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -1678,13 +1678,13 @@ function renderAgentPrompt(values: { runtimeContext: string; operatorPrompt: str ["operator_prompt", values.operatorPrompt], ["task_prompt", values.taskPrompt] ]); - const rendered = agentPromptTemplate.replace(/\{\{\s*([A-Za-z0-9_]+)\s*\}\}/gu, (match: string, key: string) => - replacements.has(key) ? replacements.get(key)! : match - ); - if (/\{\{\s*[A-Za-z0-9_]+\s*\}\}/u.test(rendered)) { - throw new Error("agent prompt template contains an unresolved variable"); - } - return rendered; + // Check only the trusted template's own placeholders. The inserted prompts may legitimately + // contain literal `{{word}}` text, and rejecting it here would fail every render of the run. + return agentPromptTemplate.replace(/\{\{\s*([A-Za-z0-9_]+)\s*\}\}/gu, (_match: string, key: string) => { + const value = replacements.get(key); + if (value === undefined) throw new Error(`agent prompt template contains an unresolved variable: ${key}`); + return value; + }); } function sourceUsesPinnedBranch(): boolean { diff --git a/packages/runtime/test/generated-workflow-verifier.test.ts b/packages/runtime/test/generated-workflow-verifier.test.ts index 839f7f256..b3e714775 100644 --- a/packages/runtime/test/generated-workflow-verifier.test.ts +++ b/packages/runtime/test/generated-workflow-verifier.test.ts @@ -10,6 +10,7 @@ import test from "node:test"; import { isDeepStrictEqual } from "node:util"; import * as ts from "typescript"; import { z } from "zod/v4"; +import { loadAgentPreambleTemplate } from "@ultrafuzz/prompts"; import { artifactContractDefinition, @@ -508,6 +509,24 @@ function loadJsonValidatorPreflight(options: { failure?: unknown; stdout?: strin }; } +type AgentPromptRenderer = (values: { runtimeContext: string; operatorPrompt: string; taskPrompt: string }) => string; + +function loadAgentPromptRenderer(template: string): AgentPromptRenderer { + const source = fs.readFileSync(workflowTemplatePath, "utf8"); + const helperStart = source.indexOf("function renderAgentPrompt"); + const helperEnd = source.indexOf("\nfunction sourceUsesPinnedBranch", helperStart); + assert.ok(helperStart >= 0 && helperEnd > helperStart, source); + const helper = ts.transpileModule(source.slice(helperStart, helperEnd), { + compilerOptions: { module: ts.ModuleKind.None, target: ts.ScriptTarget.ES2022 } + }).outputText; + return new Function( + "agentPromptTemplate", + "authorizedDefensiveSecurityContext", + "untrustedContentBoundary", + `${helper}; return renderAgentPrompt;` + )(template, "SECURITY CONTEXT", "UNTRUSTED BOUNDARY") as AgentPromptRenderer; +} + function loadFinalReportPromptAuthorityHarness(maxAuthorityBytes = 128 * 1024 * 1024): { materialize(task: unknown, coverage: unknown, execution: unknown): void; assertUnchanged(task: unknown): void; @@ -6742,6 +6761,25 @@ test("generated validator preflight budgets a contended CLI start and reports th ); }); +test("generated agent prompt inserts literal braces from task and operator prompts verbatim", () => { + const render = loadAgentPromptRenderer(loadAgentPreambleTemplate("agent-prompt")); + // Model-authored goal text, an escaped prompt example and an operator note are all inserted text, + // not placeholders of the trusted template. Any one of them used to throw inside every render. + const taskPrompt = "Find where a late tick lets {{amount}} round down.\nThe plan escaped \\{{window}} on purpose.\n"; + const operatorPrompt = "Prefer {{ reentrancy }} leads.\n\n"; + assert.equal( + render({ runtimeContext: "## Topology Runtime Context", operatorPrompt, taskPrompt }), + `SECURITY CONTEXT\n\nUNTRUSTED BOUNDARY\n\n## Topology Runtime Context\n\n${operatorPrompt}${taskPrompt}` + ); + + // A placeholder the template itself cannot bind is still a controller defect, reported by name. + const unbound = loadAgentPromptRenderer("{{runtime_context}}\n{{retry_failure}}\n{{task_prompt}}"); + assert.throws( + () => unbound({ runtimeContext: "context", operatorPrompt: "", taskPrompt: "task" }), + /agent prompt template contains an unresolved variable: retry_failure$/u + ); +}); + test("generated Smithers workflow does not precreate runtime-owned workspace patch outputs", () => { const source = fs.readFileSync(workflowTemplatePath, "utf8"); const helperStart = source.indexOf("function prepareArtifactMirror"); From c6f0060ea071a977f892ef1c55b7c2754b16334d Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:10:46 +0000 Subject: [PATCH 011/206] fix(runtime): run the JSON validator preflight once per engine process prepareArtifactMirror spawned `ultrafuzz json validate` against the smoke fixture in every prepare, in every agent-attempt reset, when an agent is built after a controller restart, and again in the zero-retry verify task. The trusted launcher's cold start was measured at ~35 s under contention (#1026), so each spawn was another chance to fail an attempt, including one whose agent work had already finished. What the spawn proves -- that this process can launch the agent-facing validator -- does not depend on the task: materializePromptSchemas has just digest-checked the workspace's schema copy, and start-run already runs the trusted-launcher preflight at launch and resume. Remember the first success in the engine process. A failure is not remembered, so the next caller spawns again. Co-Authored-By: Claude Opus 5.5 --- .../templates/smithers/workflows/workflow.tsx | 11 ++++++++-- .../test/generated-workflow-verifier.test.ts | 21 +++++++++++++++++++ 2 files changed, 30 insertions(+), 2 deletions(-) diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index d22794bc4..44d9634f8 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -4084,10 +4084,10 @@ function assertTaskOutputSchemaBindings(task: (typeof taskSpecs)[number]): void } /** - * How long the per-node validator preflight may spend inside the Ultrafuzz CLI. + * How long the validator preflight may spend inside the Ultrafuzz CLI. * * `"ultrafuzz"` resolves to the run-owned trusted launcher, which `composeSmithersCommandPath` puts - * first on PATH, so every node preparation pays the CLI's own cold start: ~2.5 s on an idle box, + * first on PATH, so the preflight pays the CLI's own cold start: ~2.5 s on an idle box, * ~35 s once a dozen agents are building against the same cores. Below that the step reports a bare * `spawnSync ultrafuzz ETIMEDOUT`, which names neither the contention nor a schema, and which cost * the smoke lane two of its three targets in #1026. Roughly five times that measured worst case, @@ -4095,8 +4095,14 @@ function assertTaskOutputSchemaBindings(task: (typeof taskSpecs)[number]): void * inside the attempt. */ const JSON_VALIDATOR_PREFLIGHT_TIMEOUT_MS = 180_000; +// The preflight proves that this process can launch the agent-facing validator, which does not vary +// by task: `materializePromptSchemas` has already digest-checked each workspace's schema copy. One +// success per engine process is enough. Re-spawning it in every prepare, attempt reset and +// zero-retry verify only added CLI cold starts that could fail a finished attempt. +let jsonValidatorPreflightPassed = false; function preflightJsonValidator(schemaDirectory: string): void { + if (jsonValidatorPreflightPassed) return; const findings = artifactSchemaRegistry().find( (entry: { filename: string }) => entry.filename === "findings.schema.json" ); @@ -4130,6 +4136,7 @@ function preflightJsonValidator(schemaDirectory: string): void { cause: error }); } + jsonValidatorPreflightPassed = true; } function taskPublishesWorkspacePatch(task: (typeof taskSpecs)[number]): boolean { diff --git a/packages/runtime/test/generated-workflow-verifier.test.ts b/packages/runtime/test/generated-workflow-verifier.test.ts index b3e714775..c02e5c441 100644 --- a/packages/runtime/test/generated-workflow-verifier.test.ts +++ b/packages/runtime/test/generated-workflow-verifier.test.ts @@ -471,6 +471,7 @@ function loadFinalReportAgentExecutionAuthority( function loadJsonValidatorPreflight(options: { failure?: unknown; stdout?: string } = {}): { preflight(): void; observedTimeoutMs(): number | undefined; + spawnCount(): number; budgetMs: number; } { const source = fs.readFileSync(workflowTemplatePath, "utf8"); @@ -481,6 +482,7 @@ function loadJsonValidatorPreflight(options: { failure?: unknown; stdout?: strin compilerOptions: { module: ts.ModuleKind.None, target: ts.ScriptTarget.ES2022 } }).outputText; let observedTimeout: number | undefined; + let spawns = 0; const loaded = new Function( "artifactSchemaRegistry", "artifactValidatorSmokeFixturePath", @@ -495,6 +497,7 @@ function loadJsonValidatorPreflight(options: { failure?: unknown; stdout?: strin () => [{ filename: "findings.schema.json" }], () => path.join(path.sep, "fixture", "findings.json"), (_file: string, _args: readonly string[], spawnOptions: { timeout?: number }) => { + spawns += 1; observedTimeout = spawnOptions.timeout; if (options.failure !== undefined) throw options.failure; return options.stdout ?? "{}"; @@ -505,6 +508,7 @@ function loadJsonValidatorPreflight(options: { failure?: unknown; stdout?: strin return { preflight: () => loaded.preflight(path.join(path.sep, "fixture", "schemas")), observedTimeoutMs: () => observedTimeout, + spawnCount: () => spawns, budgetMs: loaded.budgetMs }; } @@ -6761,6 +6765,23 @@ test("generated validator preflight budgets a contended CLI start and reports th ); }); +test("generated validator preflight spawns the CLI once per engine process and never remembers a failure", () => { + // One engine process runs every prepare, agent-attempt reset and zero-retry verify. The CLI answer + // does not depend on the task, so only the first success spawns it. + const harness = loadJsonValidatorPreflight(); + harness.preflight(); + harness.preflight(); + harness.preflight(); + assert.equal(harness.spawnCount(), 1); + + const failing = loadJsonValidatorPreflight({ + failure: Object.assign(new Error("spawnSync ultrafuzz ETIMEDOUT"), { code: "ETIMEDOUT" }) + }); + assert.throws(() => failing.preflight(), /JSON validator preflight failed/u); + assert.throws(() => failing.preflight(), /JSON validator preflight failed/u); + assert.equal(failing.spawnCount(), 2, "a failed preflight must be retried by the next caller"); +}); + test("generated agent prompt inserts literal braces from task and operator prompts verbatim", () => { const render = loadAgentPromptRenderer(loadAgentPreambleTemplate("agent-prompt")); // Model-authored goal text, an escaped prompt example and an operator note are all inserted text, From 9a92ecd551b760f766b66c2857adb582f501f6ac Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:12:56 +0000 Subject: [PATCH 012/206] refactor(runtime): delete the unused smithersSnapshotUnverifiedDependencies helper The helper regex-scanned a Smithers inspect snapshot for the generated workflow's "artifact dependency has not passed verification" message so a resume could name the dependency behind a failed `prepare:` wrapper (the R43 shape behind #272). Nothing in src calls it: prepare-wrapper failures are now attributed to their durable node (#288), and the lens sanitizer that caused R43 was deleted (#558). Its only caller was its own unit test, which goes too. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/smithers.ts | 26 -------------------------- packages/runtime/test/runtime.test.ts | 27 --------------------------- 2 files changed, 53 deletions(-) diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..6ed094d30 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -6164,32 +6164,6 @@ function requiredCurrentInspectEnum( return value as Values[number]; } -/** - * Dependency attempt ids whose verified artifacts no longer match their verification marker. - * The message is emitted by the generated workflow's own `assertVerifiedDependency`, so the shape - * is stable, and it is the only signal that reaches the resume side: the failure lands on the - * dependent's `prepare:` task and leaves no failed node behind, so the run row's `error_json` is - * where it surfaces. Recovery on top of this is tracked separately in #288. - */ -export function smithersSnapshotUnverifiedDependencies(snapshot: SmithersCommandSnapshot): string[] { - const evidence = [ - snapshot.stdout, - snapshot.stderr, - snapshot.error ?? "", - snapshot.json === undefined ? "" : JSON.stringify(snapshot.json) - ].join("\n"); - const dependencies = new Set(); - for (const match of evidence.matchAll( - /artifact dependency has not passed verification ([A-Za-z0-9._-]+) for [A-Za-z0-9._-]+/gu - )) { - const dependency = match[1]; - if (dependency !== undefined && dependency.trim() !== "" && !dependency.includes("..")) { - dependencies.add(dependency); - } - } - return [...dependencies].sort(); -} - function isCompatibleSmithersRunId(value: string): boolean { return /^[a-z0-9_-]{1,64}$/u.test(value); } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..f09dc47d6 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -26682,33 +26682,6 @@ test("resume continues a run-level render failure in place without a no-op rewin ); }); -test("unverified dependency detection reads a dependent prepare failure off the run row", async () => { - const { smithersSnapshotUnverifiedDependencies } = await import("../src/smithers.js"); - const runError = { - name: "SmithersError", - code: "SESSION_ERROR", - message: "Task failed: prepare:property-specification-fanin", - cause: { - message: - "artifact-contract failure: artifact dependency has not passed verification " + - "property-specification-crytic for property-specification-fanin" - } - }; - const snapshot = { - command: ["inspect", "ultrafuzz-r43", "--format", "json"], - ok: true, - stdout: "", - stderr: "", - json: { ok: true, data: { run: { id: "ultrafuzz-r43", status: "failed", error: runError } } } - }; - - assert.deepEqual(smithersSnapshotUnverifiedDependencies(snapshot), ["property-specification-crytic"]); - assert.deepEqual( - smithersSnapshotUnverifiedDependencies({ ...snapshot, json: undefined, stderr: "unrelated failure" }), - [] - ); -}); - test("resume --reset-node does not repeat a committed reset after a failed continuation", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From da3653e9acb913ef3f5f6bc0823e8bb7808a3560 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:13:16 +0000 Subject: [PATCH 013/206] fix(runtime): name the runner error when a workflow fails with no failed node, and pin same-id recovery A Smithers run can end `failed` without failing any task. That is a run-level error such as WORKFLOW_RENDER_FAILED, which the runner raises when the generated workflow throws while rendering, and it leaves durable work pending. WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE could then only say "failing workflow task(s): unreported", because parseCurrentSmithersInspect accepted data.run.error and dropped it, and that error is the only record of why the run stopped. parseCurrentSmithersInspect now returns an optional runError read loosely from data.run.error. It takes the cause's message (the runner records what was thrown as `cause` under its own summary and `smithers up` recovery advice), falling back to the message, redacted and scrubbed like other runner text and capped at 1,000 characters. An unexpected shape yields nothing and never fails the parse. The WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE message appends "; workflow run error : ". Its details, and so the durable workflow-failure-unattributed event payload, are unchanged: that record stays ids-only and needs no schema change. A new integration test drives the pinned runner and Ultrafuzz's own resumeRun (the `resume --force --retry-failed` entry) through this shape. While the render-time cause persists, the resume is refused with WORKFLOW_LIFECYCLE_FAILED and the run is left exactly as it was. Once the cause is removed, the same run ID resumes and runs only the pending task, and the finished producer is not run again. The test also checks the new reader against the runner's real error JSON. No rewind, fork or replacement lineage is added: each would re-render the same program and inputs and hit the same cause (#272). Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/cli.md | 8 + packages/runtime/src/smithers.ts | 27 ++- packages/runtime/src/workflow-sync.ts | 23 ++- packages/runtime/test/runtime.test.ts | 46 ++++- ...ithers-terminal-resume.integration.test.ts | 175 ++++++++++++++++++ 6 files changed, 271 insertions(+), 9 deletions(-) create mode 100644 packages/runtime/test/smithers-terminal-resume.integration.test.ts diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..88a607f50 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime]** A run that ends `failed` with no failed durable node now carries the runner's own run-level error in its `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE` diagnostic message (for example `workflow run error WORKFLOW_RENDER_FAILED: `), redacted and capped at 1,000 characters; the durable `workflow-failure-unattributed` event stays ids-only. A new pinned-runner integration test covers recovering such a run through `resume`: while its render-time cause persists the resume is refused and the run is left unchanged, and once the cause is removed the same run ID resumes and runs only its pending task. Removes the unused `smithersSnapshotUnverifiedDependencies` helper (#272). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..43eb6323b 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -325,6 +325,14 @@ reused and which newly rendered or unfinished tasks run. Ultrafuzz does not rewrite historical artifacts or automatically reset, replay, timetravel, or fork completed work. +A run that ends `failed` with no failed durable node was stopped by something +no task owns, for example an exception thrown while rendering the workflow. +`status` reports it as `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE` and appends the +workflow runner's own error to the message. Fix that cause, then `resume` the +same run; it continues from the tasks that already finished. While the cause +persists, the resume either fails with `WORKFLOW_LIFECYCLE_FAILED` or the run +fails again. + `resume --refresh-controller` first renders the currently installed Ultrafuzz controller and stock adapters beside the historical source, then delegates to that same Smithers run. It does not publish or authenticate a historical diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 6ed094d30..f037ba5d2 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -3149,6 +3149,8 @@ export interface CurrentSmithersInspect { nodes: CurrentSmithersInspectNode[]; failedChildKeys: string[]; exhaustedLoops: CurrentSmithersExhaustedLoop[]; + /** The runner's run-level error (for example WORKFLOW_RENDER_FAILED), redacted and capped. Advisory text only. */ + runError?: { code?: string; message: string }; } /** @@ -5965,7 +5967,30 @@ export function parseCurrentSmithersInspect( if (exhaustedLoops.length > 0 && parsedRunState !== "succeeded" && parsedRunState !== "succeeded-with-failures") { throw new Error("Smithers inspect data.exhaustedLoops is only valid for a succeeded workflow state"); } - return { runStatus, runState: parsedRunState, nodes, failedChildKeys, exhaustedLoops }; + const runError = currentSmithersRunError(run.error); + return { + runStatus, + runState: parsedRunState, + nodes, + failedChildKeys, + exhaustedLoops, + ...(runError === undefined ? {} : { runError }) + }; +} + +// A run-level failure (for example a render exception) names no task, so the +// run row's error is the only record of why the run stopped. Read it loosely: +// it explains a stop and never gates one, so an unexpected shape yields nothing. +function currentSmithersRunError(value: unknown): CurrentSmithersInspect["runError"] { + if (!isObjectRecord(value)) return undefined; + // The runner records the thrown error as `cause` under its own summary. + const cause = isObjectRecord(value.cause) ? value.cause.message : undefined; + const message = [cause, value.message].find((text): text is string => typeof text === "string" && text.trim() !== ""); + if (message === undefined) return undefined; + return { + ...(typeof value.code === "string" ? { code: value.code } : {}), + message: scrubWorkflowRunnerText(redactSecretsInText(message)).slice(0, 1_000) + }; } function parseCurrentSmithersExhaustedLoops(value: unknown): CurrentSmithersExhaustedLoop[] { diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index b8810c84e..a8ae6f54a 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -145,6 +145,7 @@ interface WorkflowInspect { steps: WorkflowStep[]; failedWorkflowTaskIds: string[]; exhaustedLoops: CurrentSmithersInspect["exhaustedLoops"]; + runError?: CurrentSmithersInspect["runError"]; } type SmithersRunStatus = CurrentSmithersInspect["runStatus"]; @@ -6092,11 +6093,12 @@ function nonBlockingRuntimeNodeIds(graph: PlannedGraph): ReadonlySet { return ids; } -// A workflow that ends terminally failed while every durable node is still -// non-terminal is unrecoverable by node-level retry: there is nothing to reset -// and the next resume re-finalizes identically. That is a defect in failure -// attribution, so name the workflow tasks the failure was charged to instead of -// leaving the run indistinguishable from an idle one. +// A workflow that ends terminally failed while no durable node failed stopped on +// something no durable node owns: a run-level runner error (for example +// WORKFLOW_RENDER_FAILED) or a failed workflow task outside the durable graph. +// Name the failing workflow tasks and the runner's own error so the run is not +// indistinguishable from an idle one. Recovery is a same-id resume once that +// cause is removed (smithers-terminal-resume.integration.test.ts). function unattributedTerminalWorkflowFailure( inspect: WorkflowInspect, nodeStatuses: Map, @@ -6114,11 +6116,17 @@ function unattributedTerminalWorkflowFailure( return undefined; } const failedWorkflowTasks = inspect.failedWorkflowTaskIds; + const runError = inspect.runError; + // Message only: the durable event payload built from `details` stays ids-only. + const runErrorText = + runError === undefined + ? "" + : `; workflow run error${runError.code === undefined ? "" : ` ${runError.code}`}: ${runError.message}`; return { code: "WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE", message: `workflow run ended ${workflowState} with no failed durable node; failing workflow task(s): ${ failedWorkflowTasks.length === 0 ? "unreported" : failedWorkflowTasks.join(", ") - }`, + }${runErrorText}`, severity: "error", source: "workflow", details: { @@ -6321,7 +6329,8 @@ function parseInspectSnapshot(snapshot: SmithersCommandSnapshot, expectedWorkflo runState: current.runState, steps: current.nodes.map((node) => ({ id: node.nodeId, state: node.state, attempt: node.attempt })), failedWorkflowTaskIds: [...failedWorkflowTaskIds].sort(), - exhaustedLoops: current.exhaustedLoops + exhaustedLoops: current.exhaustedLoops, + ...(current.runError === undefined ? {} : { runError: current.runError }) }; } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index f09dc47d6..578c4b987 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -18333,6 +18333,48 @@ test("parseCurrentSmithersInspect admits the 0.35.0 envelope and bounds its new ); }); +test("parseCurrentSmithersInspect reads the run-level error loosely and never fails on its shape", async () => { + const { parseCurrentSmithersInspect } = await import("../src/smithers.js"); + const workflowRunId = "ultrafuzz-run-error"; + const runError = (error: unknown) => + parseCurrentSmithersInspect( + { + command: ["inspect"], + ok: true, + stdout: "", + stderr: "", + json: workflowInspect({ + workflowRunId, + status: "failed", + state: "failed", + error, + steps: [{ id: "node:a", state: "pending" }] + }) + }, + workflowRunId + ).runError; + + // The runner records what was thrown as `cause` under its own summary and recovery advice. + assert.deepEqual( + runError({ + code: "WORKFLOW_RENDER_FAILED", + message: 'Rendering workflow "w.tsx" threw: boom. Resume with: smithers up w.tsx --resume true', + cause: { name: "Error", message: "boom" } + }), + { code: "WORKFLOW_RENDER_FAILED", message: "boom" } + ); + assert.deepEqual(runError({ message: "Task failed: prepare:x", cause: "opaque" }), { + message: "Task failed: prepare:x" + }); + // The text reaches a printed diagnostic, so it is redacted and bounded like other runner output. + assert.deepEqual(runError({ message: "token=sk-private-secret" }), { message: "token=" }); + assert.equal(runError({ message: "x".repeat(5_000) })?.message.length, 1_000); + // Advisory only: an unexpected shape yields nothing rather than a new parse failure. + for (const opaque of ["opaque", 7, null, [], { code: 7 }, { code: "X", message: " " }]) { + assert.equal(runError(opaque), undefined, JSON.stringify(opaque)); + } +}); + // Acceptance criterion 4: a stopped 0.34.0 run's durable store must survive the // four migrations 0.35.0 adds (0041-0044) with its history intact, and the // upgrade must be re-runnable and recoverable. The migrations are forward-only @@ -19621,7 +19663,7 @@ test("syncRun reports a typed diagnostic when a terminal workflow failure has no workflowRunId, status: "failed", state: "failed", - error: { message: "Task failed: ultrafuzz-agent-tasks" }, + error: { name: "SmithersError", code: "SESSION_ERROR", message: "Task failed: ultrafuzz-agent-tasks" }, failedChildKeys: ["ultrafuzz-agent-tasks::0"], steps: [ { id: "ultrafuzz-agent-tasks", state: "failed" }, @@ -19645,6 +19687,8 @@ test("syncRun reports a typed diagnostic when a terminal workflow failure has no assert.equal(diagnostic?.severity, "error"); assert.deepEqual(diagnostic?.details?.failed_workflow_tasks, ["ultrafuzz-agent-tasks"]); assert.equal(diagnostic?.details?.workflow_state, "failed"); + // A failure no durable node owns is otherwise unexplained, so the message carries the runner's own error. + assert.match(diagnostic?.message ?? "", /; workflow run error SESSION_ERROR: Task failed: ultrafuzz-agent-tasks$/u); }); test("syncRun records an unattributed terminal workflow failure durably and only once", async () => { diff --git a/packages/runtime/test/smithers-terminal-resume.integration.test.ts b/packages/runtime/test/smithers-terminal-resume.integration.test.ts new file mode 100644 index 000000000..461189426 --- /dev/null +++ b/packages/runtime/test/smithers-terminal-resume.integration.test.ts @@ -0,0 +1,175 @@ +import assert from "node:assert/strict"; +import { execFileSync } from "node:child_process"; +import fs from "node:fs"; +import path from "node:path"; +import test from "node:test"; +import { fileURLToPath } from "node:url"; + +import { parseCurrentSmithersInspect, type CurrentSmithersInspect } from "../src/smithers.js"; +import { resumeRun } from "../src/start-run.js"; +import { temporaryRoot } from "./temporary-root.js"; + +// #272: the pinned runner can end a run `failed` with no failed task (a run-level error such as +// WORKFLOW_RENDER_FAILED) while dependent work is still pending. Recovery is a same-id +// `ultrafuzz resume`, not a rewind or a replacement lineage, so this drives that command's runtime +// entry against a real run: while the render-time cause persists the resume is refused and the run +// is left as it was; once the cause is gone the same run runs only its pending task. +test("a run-level render failure resumes under the same id once its cause is gone, without re-running finished work", async () => { + const root = temporaryRoot("ufz-terminal-resume-"); + const runtimeRoot = runtimePackageRoot(); + const runId = `terminal-resume-${process.pid}-${Date.now()}`; + const workflowPath = path.join(root, ".smithers", "workflows", "terminal-resume.tsx"); + const executionLog = path.join(root, "execution.log"); + const poison = path.join(root, "render-poison"); + const runRoot = path.join(root, ".ultrafuzz", "runs", runId); + fs.mkdirSync(path.dirname(workflowPath), { recursive: true }); + fs.mkdirSync(path.join(runRoot, "smithers"), { recursive: true }); + execFileSync("git", ["init", "--quiet", "--initial-branch=main"], { cwd: root }); + const fixtureModules = path.join(root, ".smithers", "node_modules"); + fs.symlinkSync(path.dirname(fs.realpathSync(path.join(runtimeRoot, "node_modules", "smthrs"))), fixtureModules); + fs.writeFileSync(workflowPath, workflowSource({ executionLog, poison })); + fs.writeFileSync(poison, ""); + execFileSync( + smithersBinary(runtimeRoot), + ["up", workflowPath, "--detach", "--run-id", runId, "--root", root, "--input", "{}", "--format", "json"], + { cwd: root, encoding: "utf8", env: { ...process.env, SMITHERS_POST_FAILURE: "0" } } + ); + + const failed = await stoppedRun(runtimeRoot, root, runId); + assert.equal(failed.runState, "failed"); + assert.deepEqual(nodeStates(failed), { dependent: "pending", producer: "finished@1" }); + // No task owns this failure, so the run row's error is the only record of why it stopped. + assert.deepEqual(failed.runError, { code: "WORKFLOW_RENDER_FAILED", message: "synthetic run-level render failure" }); + + // Continuation runs on Ultrafuzz's own attested controller, never on packages in the target + // (#973), so the fixture's direct `smithers up` link goes before the first resume. + fs.unlinkSync(fixtureModules); + fs.writeFileSync( + path.join(runRoot, "run.json"), + `${JSON.stringify({ + run_id: runId, + workflow_ids: [runId], + workflow: { run_id: runId, path: path.relative(root, workflowPath) } + })}\n` + ); + + const refused = await resume(root, runId); + assert.equal(refused.ok, false); + assert.equal(refused.diagnostics?.[0]?.code, "WORKFLOW_LIFECYCLE_FAILED"); + assert.match(refused.diagnostics?.[0]?.message ?? "", /synthetic run-level render failure/u); + assert.deepEqual(await stoppedRun(runtimeRoot, root, runId), failed); + assert.deepEqual(executed(executionLog), ["producer"]); + + fs.rmSync(poison); + const resumed = await resume(root, runId); + assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); + assert.equal(resumed.value?.workflow_run_id, runId); + assert.equal(resumed.value?.submitted, true); + const finished = await stoppedRun(runtimeRoot, root, runId); + assert.equal(finished.runState, "succeeded"); + assert.equal(finished.runError, undefined); + assert.deepEqual(nodeStates(finished), { dependent: "finished@1", producer: "finished@1" }); + assert.deepEqual(executed(executionLog), ["producer", "dependent"]); +}); + +function workflowSource(input: { executionLog: string; poison: string }): string { + return `/** @jsxImportSource smthrs */ +import fs from "node:fs"; +import { createSmithers } from "smthrs"; +import { z } from "zod/v4"; + +const { Workflow, Parallel, Task, smithers, outputs } = createSmithers({ + input: z.object({}), + result: z.object({ value: z.string() }) +}); +const record = (value: string) => { + fs.appendFileSync(${JSON.stringify(input.executionLog)}, value + "\\n"); + return { value }; +}; + +export default smithers((ctx) => { + // Render-time I/O that starts once the producer has output, like the generated workflow's reads + // of a published dynamic expansion and of the prompts of the tasks it adds. Every later render + // hits it, including the preflight render a detached resume runs; the first launch's does not. + if (ctx.outputMaybe(outputs.result, { nodeId: "producer" }) !== undefined && fs.existsSync(${JSON.stringify(input.poison)})) { + throw new Error("synthetic run-level render failure"); + } + return ( + + + {() => record("producer")} + {() => record("dependent")} + + + ); +}); +`; +} + +// `ultrafuzz resume --force --retry-failed`, through the runtime entry the CLI calls. +async function resume(root: string, runId: string) { + return resumeRun({ + projectRoot: root, + runId, + force: true, + retryFailed: true, + env: { PATH: process.env.PATH, SMITHERS_POST_FAILURE: "0" } + }); +} + +// A detached submission returns only after the runner has re-activated the run, so the first +// stopped state observed after an accepted resume is the continuation's own. +async function stoppedRun(runtimeRoot: string, root: string, runId: string): Promise { + const deadline = Date.now() + 60_000; + let last: unknown; + while (Date.now() < deadline) { + let stdout: string; + let json: { data?: { run?: { status?: unknown } } }; + try { + stdout = execFileSync(smithersBinary(runtimeRoot), ["inspect", runId, "--format", "json", "--full-output"], { + cwd: root, + encoding: "utf8" + }); + json = JSON.parse(stdout) as typeof json; + } catch (error) { + // Retry a transient inspect failure until the deadline, as the sibling integration tests do. + last = error; + await new Promise((resolve) => setTimeout(resolve, 100)); + continue; + } + last = json.data?.run?.status; + if (last === "finished" || last === "failed" || last === "cancelled") { + // The production reader, on the runner's real output. + return parseCurrentSmithersInspect({ command: ["inspect", runId], ok: true, stdout, stderr: "", json }, runId); + } + await new Promise((resolve) => setTimeout(resolve, 100)); + } + assert.fail(`run ${runId} did not stop; last observation ${String(last)}`); +} + +function nodeStates(inspect: CurrentSmithersInspect): Record { + return Object.fromEntries( + inspect.nodes.map((node) => [node.nodeId, node.attempt === 0 ? node.state : `${node.state}@${node.attempt}`]) + ); +} + +function executed(executionLog: string): string[] { + return fs.readFileSync(executionLog, "utf8").trim().split("\n"); +} + +function smithersBinary(runtimeRoot: string): string { + return path.join(runtimeRoot, "node_modules", ".bin", process.platform === "win32" ? "smithers.cmd" : "smithers"); +} + +function runtimePackageRoot(): string { + let directory = path.dirname(fileURLToPath(import.meta.url)); + while (directory !== path.dirname(directory)) { + const candidate = path.join(directory, "package.json"); + if (fs.existsSync(candidate)) { + const metadata = JSON.parse(fs.readFileSync(candidate, "utf8")) as { name?: string }; + if (metadata.name === "@ultrafuzz/runtime") return directory; + } + directory = path.dirname(directory); + } + throw new Error("runtime package root not found"); +} From 0da2113fb74cbb5836c831167777977fe54a7523 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:15:50 +0000 Subject: [PATCH 014/206] fix(runtime): give agent retries a real wait and the whole planned chain Agent retries used an exponential backoff from 1s, so a three-attempt budget was spent in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can need to clear (#1084). The agent Task also inherited Smithers' default identical-failure stall verdict (3), which ended a chain before later same-agent attempts (exhaustive plans five) or any [retry].agents fallback profile ran. Against real Smithers 0.35.0, a [fail, fail, fail, fallback] chain ended `stalled` after attempt 3 and never ran the fallback; with maxIdenticalFailures: 0 the fallback ran at attempt 4 and the run finished. - compileTask: initialDelayMs 60_000, so retries wait 60s, 120s, 240s, then Smithers' 300s cap. - Agent Task: maxIdenticalFailures: 0, so the planned chain is the budget. It is set in the template only; the compiled manifest, cloud handoff schema and sealed task documents keep their exact shape. - The real `smithers graph` smoke test never ran: it required a workspace-root .smithers install that no checkout has, and could not resolve the sealed module paths. It now uses the runtime package's own Smithers and asserts the retryPolicy Smithers receives. Refs #1084 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/config.md | 12 +- packages/runtime/src/smithers.ts | 5 +- .../templates/smithers/workflows/workflow.tsx | 6 +- packages/runtime/test/runtime.test.ts | 153 +++++++++--------- 5 files changed, 89 insertions(+), 88 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..0b2766f60 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [docs]** Agent retries now wait one minute, doubling up to Smithers' five-minute cap, instead of 1s and 2s, which spent a three-attempt budget in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can take; Smithers' three-identical-failures stall verdict no longer ends the planned chain early, so later same-agent attempts and `[retry].agents` fallback profiles run (#1084). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/config.md b/docs/config.md index 4ccc55e3e..fbe9f9df0 100644 --- a/docs/config.md +++ b/docs/config.md @@ -50,11 +50,13 @@ project primary count. The shipped `default` profile uses three attempts; five, and `invariant-only` inherits three. An explicit project `[retry]` value overrides the profile. The complete primary-plus-fallback chain may contain at most 100 attempts. Omitting `agents`, or leaving it empty, keeps model fallback disabled. -Retries use bounded exponential backoff, a fresh session, and the same effective -task prompt, including Smithers' safety contracts; Ultrafuzz does not inspect -provider error text. The -planned chain and actual producer are recorded in the task manifest, attempt -ledger, and final report. +A retry waits one minute, then two, then four, and at most five minutes (the +Smithers cap), and uses a fresh session and the same effective task prompt, +including Smithers' safety contracts; Ultrafuzz does not inspect provider error +text. The planned chain is the whole budget: repeated identical failures do not +end it before later attempts or fallback profiles run. The planned chain and +actual producer are recorded in the task manifest, attempt ledger, and final +report. Retry chains currently require local execution. Cloud planning accepts one effective attempt, and local fallback across different agent implementations diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..ac22464d5 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -7914,7 +7914,10 @@ function compileTask(input: { timeoutMs, heartbeatTimeoutMs, retries, - retryPolicy: { backoff: "exponential", initialDelayMs: 1_000 }, + // Retries wait 60s, 120s, 240s, then Smithers' 5-minute cap. A 1s base + // spent a three-attempt budget in about 25s, inside the minute a + // contended Claude Code OAuth refresh can take to clear (#1084). + retryPolicy: { backoff: "exponential", initialDelayMs: 60_000 }, ...(input.sourceRevision === undefined ? {} : { sourceRevision: input.sourceRevision, sourceRef: input.sourceRef! }), diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index ceab70669..f24ef08b8 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -10110,7 +10110,11 @@ export default smithers((ctx) => { timeoutMs={task.timeoutMs} heartbeatTimeoutMs={task.heartbeatTimeoutMs} retries={task.retries} - retryPolicy={task.retryPolicy} + // The planned chain is the whole retry budget. Smithers' + // default stall verdict would end it after three identical + // failures, before later same-agent attempts or fallback + // profiles run (#1084). + retryPolicy={{ ...task.retryPolicy, maxIdenticalFailures: 0 }} metadata={task.metadata} > {fullTaskPrompt} diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..eceaa6746 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -11782,7 +11782,7 @@ test("compileSmithersWorkflow exhausts same-profile retries before ordered fallb ] ); assert.equal(task.retries, 3); - assert.deepEqual(task.retryPolicy, { backoff: "exponential", initialDelayMs: 1_000 }); + assert.deepEqual(task.retryPolicy, { backoff: "exponential", initialDelayMs: 60_000 }); assert.deepEqual(task.metadata.retryPolicy, { maxAttempts: 4, sameAgentAttempts: 3, @@ -26870,92 +26870,83 @@ test("ordinary resume restores missing static presentation prompts", async () => ); }); -testWhen(realSmithersGraphUnavailable() === false)( - "compiled Smithers workflow passes a real non-executing graph smoke", - async () => { - const project = tempProject(); - initProject({ projectRoot: project, force: true }); - writeSmallTopology(project); - fs.symlinkSync( - path.join(workspaceRoot(), ".smithers", "node_modules"), - path.join(project, ".smithers", "node_modules"), - "dir" - ); +test("compiled Smithers workflow passes a real non-executing graph smoke", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + fs.symlinkSync( + path.dirname(fs.realpathSync(path.join(process.cwd(), "node_modules", "smthrs"))), + path.join(project, ".smithers", "node_modules"), + "dir" + ); - const plan = await planRun({ projectRoot: project, runId: "graph-smoke", env: {} }); - assert.equal(plan.ok, true, JSON.stringify(plan.diagnostics)); - const { compileSmithersWorkflow } = await import("../src/smithers.js"); - const compiled = compileSmithersWorkflow({ - projectRoot: project, - config: plan.value!.resolved_config, - graph: plan.value!.expanded_graph, - runLayout: plan.value!.layout, - workflowName: "ultrafuzz-graph-smoke", - renderedPrompts: plan.value!.rendered_prompts, - operatorPrompt: "graph smoke" - }); + const plan = await planRun({ projectRoot: project, runId: "graph-smoke", env: {} }); + assert.equal(plan.ok, true, JSON.stringify(plan.diagnostics)); + assert.ok(plan.value); + const { compileSmithersWorkflow } = await import("../src/smithers.js"); + const compiled = compileSmithersWorkflow({ + projectRoot: project, + config: plan.value.resolved_config, + graph: plan.value.expanded_graph, + runLayout: plan.value.layout, + workflowName: "ultrafuzz-graph-smoke", + renderedPrompts: plan.value.rendered_prompts, + operatorPrompt: "graph smoke" + }); - const graphProcess = spawnSync( - "smithers", - [ - "graph", - compiled.evidenceWorkflowPath, - "--run-id", - compiled.smithersRunId, - "--root", - project, - "--input", - fs.readFileSync(compiled.inputPath, "utf8"), - "--compact", - "--format", - "json" - ], - { - cwd: project, - encoding: "utf8", - maxBuffer: 1024 * 1024 * 16, - env: { ...process.env, OPENAI_API_KEY: "test-openai-api-key" } + const graphProcess = spawnSync( + path.join(process.cwd(), "node_modules", ".bin", "smithers"), + [ + "graph", + compiled.evidenceWorkflowPath, + "--run-id", + compiled.smithersRunId, + "--root", + project, + "--input", + fs.readFileSync(compiled.inputPath, "utf8"), + "--compact", + "--format", + "json" + ], + { + cwd: project, + encoding: "utf8", + maxBuffer: 1024 * 1024 * 16, + env: { + ...process.env, + OPENAI_API_KEY: "test-openai-api-key", + // No launch staged the sealed module copies the workflow imports by + // default; render against this checkout's builds instead. + ULTRAFUZZ_ARTIFACTS_MODULE: pathToFileURL(path.join(workspaceRoot(), "packages/artifacts/dist/index.js")).href, + ULTRAFUZZ_RUNTIME_MODULE: pathToFileURL(path.join(workspaceRoot(), "packages/runtime/dist/index.js")).href } - ); - if (graphProcess.status !== 0 || graphProcess.stdout.trim() === "") { - throw ( - graphProcess.error ?? - new Error( - [ - `smithers graph exited ${String(graphProcess.status)}`, - graphProcess.stderr.trim(), - graphProcess.stdout.trim() - ] - .filter(Boolean) - .join("\n") - ) - ); } - const graphJson = graphProcess.stdout; - const graph = JSON.parse(graphJson) as { tasks?: Array<{ nodeId?: string }> }; - assert.equal(graph.tasks?.[0]?.nodeId, "prepare:project-discovery"); - assert.equal( - graph.tasks?.some((task) => task.nodeId === "node:project-discovery"), - true - ); - assert.equal( - graph.tasks?.some((task) => task.nodeId === "verify:project-discovery"), - true + ); + if (graphProcess.status !== 0 || graphProcess.stdout.trim() === "") { + throw ( + graphProcess.error ?? + new Error( + [`smithers graph exited ${String(graphProcess.status)}`, graphProcess.stderr.trim(), graphProcess.stdout.trim()] + .filter(Boolean) + .join("\n") + ) ); } -); - -function realSmithersGraphUnavailable(): string | false { - try { - execFileSync("smithers", ["graph", "--help"], { stdio: "ignore" }); - } catch { - return "smithers CLI is not installed"; - } - if (!fs.existsSync(path.join(workspaceRoot(), ".smithers", "node_modules", "smthrs"))) { - return ".smithers Smithers dependencies are not installed"; - } - return false; -} + const graph = JSON.parse(graphProcess.stdout) as { tasks?: Array<{ nodeId?: string; retryPolicy?: unknown }> }; + assert.equal(graph.tasks?.[0]?.nodeId, "prepare:project-discovery"); + // Smithers receives the planned chain as the whole retry budget: a real wait + // between attempts and no identical-failure stall verdict (#1084). + assert.deepEqual(graph.tasks?.find((task) => task.nodeId === "node:project-discovery")?.retryPolicy, { + backoff: "exponential", + initialDelayMs: 60_000, + maxIdenticalFailures: 0 + }); + assert.equal( + graph.tasks?.some((task) => task.nodeId === "verify:project-discovery"), + true + ); +}); function workspaceRoot(): string { let current = process.cwd(); From 0a71fe3640938c38596357db6e7d3f0b9703472c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:24:42 +0000 Subject: [PATCH 015/206] fix(runtime): cancel and pause work even when run control evidence has diverged `cancel` and `pause` are how an operator stops a run, but both read the run through the strict, execution-grade mode of `readLinkedWorkflowEvidence`. That mode exists to authorize executing workflow code: it re-derives every sealed control document and binding against the current build and throws on the first divergence. So after a rebuild changed the validator build identity, or after an operator hot-patched the published workflow, both commands refused with WORKFLOW_CONTROL_EVIDENCE_INVALID and the run kept executing, while `status` read the same run and `resume` handed it to Smithers. Neither command executes workflow code; each only asks the workflow runner to cancel or pause the linked run. Both now read evidence the way `status` and `events` already do (observe-only, divergence tolerant) and invoke the runner through the same published-snapshot environment as before. A confirmed cancellation still persists `canceled`. A run whose published execution snapshot files changed is still refused, as `status` refuses it. The two tests that asserted `cancelRun` refuses a diverged run now assert that it cancels (and the contract-binding one also pauses). The hand-patched-workflow test also pauses and cancels, and the execution-file test pins that cancel still refuses there without invoking the runner. The shared fake runner answers `cancel` with the engine's confirmed-cancellation contract. Refs #674, #921 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/cli.md | 7 ++++ packages/runtime/src/lifecycle-inspection.ts | 8 +++- packages/runtime/src/start-run.ts | 13 ++++-- packages/runtime/test/runtime.test.ts | 42 +++++++++++++++++--- 5 files changed, 60 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..c9f6685fc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [docs]** `pause` and `cancel` can stop a run whose sealed control evidence diverged. Both read evidence observe-only and divergence-tolerant, as `status` does, instead of requiring execution-grade evidence, which refused with `WORKFLOW_CONTROL_EVIDENCE_INVALID` after, for example, a rebuild changed the validator build identity or an operator hot-patched the published workflow, leaving the run executing with no way to stop it through Ultrafuzz. A confirmed cancellation still persists `canceled`, and a run whose published execution snapshot files changed is still refused (#674, #921). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..5efd717db 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -480,6 +480,13 @@ failing. Both outcomes append distinct product events. Failures use the stable `WORKFLOW_CANCEL_FAILED` diagnostic. +`pause` and `cancel` read run evidence the way `status` does, without the +workflow control lock. A run whose sealed control documents diverged, for +example a hand-patched published workflow or a planned graph that no longer +matches the current build's artifact contracts after a rebuild, can therefore +still be paused or cancelled; `status` reports the divergence. Like `status`, +they still refuse a run whose published execution snapshot files changed. + `why` returns a deterministic diagnosis: a summary, the current node, and typed blockers with `kind`, `node_id`, `iteration`, `reason`, `unblocker`, `waiting_since`, `attempt`, and `max_attempts`. Blocker kinds are diff --git a/packages/runtime/src/lifecycle-inspection.ts b/packages/runtime/src/lifecycle-inspection.ts index 3aa38ac59..cb81f78a6 100644 --- a/packages/runtime/src/lifecycle-inspection.ts +++ b/packages/runtime/src/lifecycle-inspection.ts @@ -237,7 +237,13 @@ function readRunStatusIfPresent(layout: RunLayout): RunStatus | undefined { export async function cancelRun(input: CancelRunInput) { const projectRoot = path.resolve(input.projectRoot); - const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId); + // Cancelling runs none of the run's workflow code: it only asks the workflow runner to stop the + // linked run. So it reads evidence the way `status` does, and a run whose sealed control documents + // diverged can still be stopped instead of refusing with WORKFLOW_CONTROL_EVIDENCE_INVALID. + const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId, { + tolerateControlDivergence: true, + observeOnly: true + }); if (!evidence.ok) { return runtimeFailure(evidence.diagnostics); } diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index ba32e842b..f8cf4d22c 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -789,7 +789,12 @@ export async function forkRun(input: WorkflowLifecycleInput) { export async function pauseRun(input: PauseRunInput) { const projectRoot = path.resolve(input.projectRoot); - const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId); + // Like `cancel`, pausing only asks the runner to park the linked run, so diverged control + // documents must not block it. + const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId, { + tolerateControlDivergence: true, + observeOnly: true + }); if (!evidence.ok) { return runtimeFailure(evidence.diagnostics); } @@ -1323,9 +1328,9 @@ export async function readLinkedWorkflowEvidence( throw new Error("run metadata workflow IDs do not exactly match the active workflow run"); } - // Observers pass `tolerateControlDivergence` so a divergent control file downgrades to a reported - // warning instead of hiding a live run entirely (issue #674). Execution callers omit it and keep - // failing closed. + // Observers, `pause` and `cancel` pass `tolerateControlDivergence` so a divergent control file + // downgrades to a reported warning instead of hiding a live run entirely, or leaving it unstoppable + // (issue #674). Execution callers omit it and keep failing closed. const tolerateDivergence = options.tolerateControlDivergence === true; const verifiedControl = verifyWorkflowControlSnapshot(resolvedProjectRoot, layout, { tolerateDivergence }); // A live document replaced under the completeness re-derivation is a transient race, not a diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..18a4995cd 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -1916,6 +1916,10 @@ function fakeSmithersEnv(project: string): Record { " exit 2", " fi", " ;;", + " cancel)", + ' printf \'%s\\n\' \'{"ok":true,"data":{"status":"cancelled"}}\'', + " exit 2", + " ;;", " inspect)", ` if [ -f ${shellQuote(inspectStateOverride)} ]; then`, ` inspect_state=$(cat ${shellQuote(inspectStateOverride)})`, @@ -13302,6 +13306,20 @@ test("a divergent published control file leaves status readable while native res assert.equal(skipped.length, 1, JSON.stringify(health.diagnostics)); assert.equal(skipped[0]?.severity, "warning"); + // The operator can still stop the run: pause and cancel reach the runner, and a confirmed + // cancellation is still persisted as the terminal state. + const paused = await pauseRun({ projectRoot: project, runId, env }); + assert.equal(paused.ok, true, JSON.stringify(paused.diagnostics)); + assert.equal(paused.value?.status, "pause-requested"); + const cancelled = await cancelRun({ projectRoot: project, runId, env }); + assert.equal(cancelled.ok, true, JSON.stringify(cancelled.diagnostics)); + assert.equal(cancelled.value?.run_status, "canceled"); + assert.ok(run.value); + assert.equal(readRunState(layoutForRunRoot(run.value.run_root)).status, "canceled"); + const commands = fs.readFileSync(path.join(project, "smithers-commands.log"), "utf8"); + assert.match(commands, new RegExp(`^pause ultrafuzz-${runId} --format json`, "mu")); + assert.match(commands, new RegExp(`^cancel ultrafuzz-${runId} --format json`, "mu")); + // The divergence is never silently repaired. assert.equal(fs.readFileSync(snapshotWorkflowPath, "utf8"), `${pristine}\n// diverged\n`); }); @@ -13338,6 +13356,11 @@ test("an observer still refuses a run whose sealed execution files diverged", as const health = await getRunHealth({ projectRoot: project, runId, env }); assert.equal(health.ok, false); assert.equal(health.diagnostics[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); + // Cancel reaches the runner through the same snapshot environment, so it refuses before invoking it. + const cancelled = await cancelRun({ projectRoot: project, runId, env }); + assert.equal(cancelled.ok, false); + assert.equal(cancelled.diagnostics[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); + assert.doesNotMatch(fs.readFileSync(path.join(project, "smithers-commands.log"), "utf8"), /^cancel /mu); }); test("a sealed manifest that stops re-deriving leaves status readable while native resume delegates", async () => { @@ -13378,9 +13401,6 @@ test("a sealed manifest that stops re-deriving leaves status readable while nati } const resumed = await resumeRun({ projectRoot: project, runId, env }); assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - const cancelled = await cancelRun({ projectRoot: project, runId, env }); - assert.equal(cancelled.ok, false); - assert.equal(cancelled.diagnostics[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); // An observer reads the same run and is told which document stopped agreeing, naming the task the // re-derivation tripped on so the divergence is diagnosable without reproducing it by hand. @@ -13414,6 +13434,11 @@ test("a sealed manifest that stops re-deriving leaves status readable while nati assert.equal(skipped.length, 1, JSON.stringify(health.diagnostics)); assert.equal(skipped[0]?.severity, "warning"); + // The divergence does not keep the operator from stopping the run either. + const cancelled = await cancelRun({ projectRoot: project, runId, env }); + assert.equal(cancelled.ok, true, JSON.stringify(cancelled.diagnostics)); + assert.equal(cancelled.value?.run_status, "canceled"); + // Reporting the divergence must never repair it. assert.equal( (JSON.parse(fs.readFileSync(tasksPath, "utf8")) as typeof manifest).tasks @@ -13453,9 +13478,6 @@ test("a planned graph that stops matching this build's contracts leaves status r assert.equal(strict.ok, false); const resumed = await resumeRun({ projectRoot: project, runId, env }); assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - const cancelled = await cancelRun({ projectRoot: project, runId, env }); - assert.equal(cancelled.ok, false); - assert.equal(cancelled.diagnostics[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); const observed = await readLinkedWorkflowEvidence(project, runId, { tolerateControlDivergence: true }); assert.equal(observed.ok, true, JSON.stringify(observed.ok ? [] : observed.diagnostics)); @@ -13489,6 +13511,14 @@ test("a planned graph that stops matching this build's contracts leaves status r 1, JSON.stringify(health.diagnostics) ); + + // A checkout that moved on must not cost the operator the ability to stop the run. + const paused = await pauseRun({ projectRoot: project, runId, env }); + assert.equal(paused.ok, true, JSON.stringify(paused.diagnostics)); + assert.equal(paused.value?.status, "pause-requested"); + const cancelled = await cancelRun({ projectRoot: project, runId, env }); + assert.equal(cancelled.ok, true, JSON.stringify(cancelled.diagnostics)); + assert.equal(cancelled.value?.run_status, "canceled"); }); test("sealed Bun startup controls from another build are reported, not equated", () => { From 6c2020c6fe573eb409938282463c510964c069fb Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:25:14 +0000 Subject: [PATCH 016/206] fix(runtime): leave paused runs out of the workflow deadline and warn on a failed deadline cancel The workflow deadline is checked only when something synchronizes the run (#1110). An operator-paused run executes nothing, and resuming it records a new deadline, so cancelling it at the first observation past its deadline bounded nothing. The control projection now leaves paused runs out of deadlineExceeded. A failed deadline cancel was an error diagnostic, so `status` returned ok:false and `status --watch` stopped polling, and nothing asked again. It is now a warning: the run stays active and the next synchronization requests cancellation again. Refs #1110 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/workflow-control.ts | 5 +- packages/runtime/src/workflow-sync.ts | 3 +- packages/runtime/test/runtime.test.ts | 62 +++++++++++++++++++ .../runtime/test/workflow-control.test.ts | 18 ++++++ 4 files changed, 86 insertions(+), 2 deletions(-) diff --git a/packages/runtime/src/workflow-control.ts b/packages/runtime/src/workflow-control.ts index b6e38439c..2626a8d53 100644 --- a/packages/runtime/src/workflow-control.ts +++ b/packages/runtime/src/workflow-control.ts @@ -225,10 +225,13 @@ export function projectWorkflowControlState(input: WorkflowControlProjectionInpu const controlTransition = controlStateChanged(previous, state); state.last_transition_at = controlTransition ? now : previous.last_transition_at; + // A paused run executes nothing and resuming it resets the deadline, so + // cancelling it at the next observation would bound nothing. const deadlineExceeded = state.workflow_deadline_at !== undefined && input.nowMs >= timestampMs(state.workflow_deadline_at, Number.POSITIVE_INFINITY, "workflow deadline") && - !isTerminalRunStatus(state.status); + !isTerminalRunStatus(state.status) && + state.status !== "paused"; return { state, diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index b8810c84e..a70ad8a2e 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -1642,7 +1642,8 @@ export async function synchronizeLinkedWorkflowRun( workflowControl.state.last_transition_at = new Date(observedAtMs).toISOString(); deadlineApplied = true; } catch (error) { - diagnostics.push(smithersDiagnostic(error, "WORKFLOW_DEADLINE_CANCEL_FAILED")); + // The run stays active and the next synchronization requests cancellation again. + diagnostics.push({ ...smithersDiagnostic(error, "WORKFLOW_DEADLINE_CANCEL_FAILED"), severity: "warning" }); } } const preControlMutationBudgetDiagnostic = synchronizationBudgetDiagnostic(control, synchronizationClock(control)); diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..c3f0ce2b5 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -13111,6 +13111,68 @@ test("getRunHealth returns live runner health when observation synchronization e assert.deepEqual(fs.readFileSync(statePath), stateBefore); }); +/** The run documents a synchronization can write, so a test can prove that a pass wrote none of them. */ +function runStateDocumentBytes(runRoot: string): Record { + return Object.fromEntries( + ["state.json", "events.jsonl", "run.json", "attempts.jsonl", "usage.jsonl"].map((file) => { + const filePath = path.join(runRoot, file); + return [file, fs.existsSync(filePath) ? fs.readFileSync(filePath).toString("base64") : null]; + }) + ); +} + +/** Makes one fake runner subcommand run `script` first, e.g. to stall or fail it. */ +function prependFakeSmithersCase(env: Record, subcommand: string, script: string): void { + const smithers = env.SMITHERS_BIN; + assert.ok(smithers !== undefined); + const source = fs.readFileSync(smithers, "utf8"); + assert.ok(source.includes(` ${subcommand})\n`), `fake runner has no ${subcommand} case`); + fs.writeFileSync(smithers, source.replace(` ${subcommand})\n`, ` ${subcommand})\n${script}`)); +} + +test("a failed deadline cancel is a warning and the next status requests it again", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "deadline-cancel-retry"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "pending", attempt: 0 }] + }), + events: workflowEvents(workflowRunId, [{ type: "NodePending", nodeId: "node:project-discovery", attempt: 0 }]), + status: currentStatusEnvelope(workflowRunId) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const layout = layoutForRunRoot(run.value.run_root, runId); + const state = JSON.parse(fs.readFileSync(layout.statePath, "utf8")) as RunState; + state.workflow_deadline_at = "2000-01-01T00:00:00.000Z"; + fs.writeFileSync(layout.statePath, `${JSON.stringify(state, null, 2)}\n`, "utf8"); + prependFakeSmithersCase(env, "cancel", " printf '%s\\n' 'runner busy' >&2\n exit 1\n"); + const commandLog = env.SMITHERS_FAKE_LOG; + assert.ok(commandLog !== undefined); + fs.writeFileSync(commandLog, "", "utf8"); + + for (const poll of [1, 2]) { + const health = await getRunHealth({ projectRoot: project, runId, env }); + + assert.equal(health.ok, true, JSON.stringify(health.diagnostics)); + assert.ok( + health.diagnostics.some( + (diagnostic) => diagnostic.code === "WORKFLOW_DEADLINE_CANCEL_FAILED" && diagnostic.severity === "warning" + ), + JSON.stringify(health.diagnostics) + ); + assert.equal(readRunState(layout).status, "running"); + assert.equal((fs.readFileSync(commandLog, "utf8").match(/^cancel /gmu) ?? []).length, poll); + } +}); + test("getRunHealth stays readable while execution holds the workflow control lock", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); diff --git a/packages/runtime/test/workflow-control.test.ts b/packages/runtime/test/workflow-control.test.ts index a6f76aeb7..232844a87 100644 --- a/packages/runtime/test/workflow-control.test.ts +++ b/packages/runtime/test/workflow-control.test.ts @@ -227,6 +227,24 @@ test("workflow deadline decisions are deterministic at the fake-clock boundary", assert.equal(deadlineProjection.deadlineExceeded, true); }); +test("a paused run is not cancelled by an observation past its workflow deadline", () => { + const graph = syntheticGraph([node("pending")]); + const paused = initialState(graph, 1, 10); + paused.status = "paused"; + + const projection = projectWorkflowControlState({ + previousState: structuredClone(paused), + state: paused, + graph, + tasks: tasksFor(graph), + workflowStates: new Map(), + workflowState: "paused", + nowMs: BASE_MS + 10_000 + }); + + assert.equal(projection.deadlineExceeded, false); +}); + test("controller recovery rejects malformed present timestamps instead of substituting the clock", () => { const graph = syntheticGraph([node("pending")]); const state = initialState(graph, 1, 60, 45); From abf81468ae01b25927a7d4a8f47f286750322c9b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:25:14 +0000 Subject: [PATCH 017/206] fix(runtime): stop rewriting unchanged runs on observation and bound runner queries Every control projection renews the controller lease and advances the concurrency clock, and any difference counted as a change, so each `status` poll rewrote state.json and a finished run kept an "active" lease renewed by whoever looked at it (#1145). The projection now also reports whether those clocks are its only change; an observe-only pass (status) and any pass over a terminal run skip writing such a projection. Explicit passes over a live run (syncRun, inspect, why, stats) still renew its lease, which the eval and Modal pumps read as liveness. Read-only runner queries had no timeout unless the caller passed one, so a wedged runner process blocked status and the synchronization pumps indefinitely. runSmithersInspectionCommand now bounds the runner process (never executable preparation) at 120 s by default; ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS overrides it, capped at 600 s. A timeout becomes the existing failed-query diagnostic, with a message naming the bound. The runner's orphan reason recommends `smithers supervise -r `, which the public rewrite turned into `workflow runner supervise -r `, a command nobody can run. It now recommends `ultrafuzz resume `. Refs #1145 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/smithers.ts | 36 ++++- packages/runtime/src/state-export.ts | 17 +- packages/runtime/src/workflow-control.ts | 20 ++- packages/runtime/src/workflow-sync.ts | 6 +- packages/runtime/test/runtime.test.ts | 153 +++++++++++++++++- .../runtime/test/workflow-control.test.ts | 28 ++++ 6 files changed, 247 insertions(+), 13 deletions(-) diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..4be9f06f5 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -5635,6 +5635,20 @@ export async function assertSmithersControllerRefreshable(input: { return { status: "present", snapshot: inspection, inspect }; } +const DEFAULT_RUNNER_QUERY_TIMEOUT_MS = 120_000; +const MAX_RUNNER_QUERY_TIMEOUT_MS = 600_000; + +/** + * Bounds one read-only runner query, so a wedged runner becomes a failed snapshot instead of blocking + * `status`, `inspect`, `stats`, or a synchronization pump forever. `ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS` + * overrides the default with a positive number of milliseconds, capped at ten minutes. + */ +function runnerQueryTimeoutMs(env: Record | undefined): number { + const configured = (env ?? process.env).ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS; + const parsed = configured !== undefined && /^[1-9]\d*$/u.test(configured) ? Number(configured) : Number.NaN; + return Number.isSafeInteger(parsed) ? Math.min(parsed, MAX_RUNNER_QUERY_TIMEOUT_MS) : DEFAULT_RUNNER_QUERY_TIMEOUT_MS; +} + export async function runSmithersInspectionCommand(input: { args: readonly string[]; projectRoot: string; @@ -5644,6 +5658,7 @@ export async function runSmithersInspectionCommand(input: { timeoutMs?: number; }): Promise { const command = [...input.args]; + const commandTimeoutMs = runnerQueryTimeoutMs(input.env); try { const result = await execSmithersCli({ args: command, @@ -5651,7 +5666,8 @@ export async function runSmithersInspectionCommand(input: { env: input.env, environmentVariableNames: input.environmentVariableNames, signal: input.signal, - timeoutMs: input.timeoutMs + timeoutMs: input.timeoutMs, + commandTimeoutMs }); return { command: result.command, @@ -5662,16 +5678,24 @@ export async function runSmithersInspectionCommand(input: { }; } catch (error) { const record = - error && typeof error === "object" ? (error as { stdout?: unknown; stderr?: unknown; message?: unknown }) : {}; + error && typeof error === "object" + ? (error as { stdout?: unknown; stderr?: unknown; message?: unknown; killed?: unknown; signal?: unknown }) + : {}; const stdout = typeof record.stdout === "string" ? record.stdout : ""; const stderr = typeof record.stderr === "string" ? record.stderr : ""; + // Node reports the kill of its own execution timeout as `killed` with SIGTERM. + const timedOut = record.killed === true && record.signal === "SIGTERM"; return { command: smithersDisplayCommand(command), ok: false, stdout, stderr, ...jsonField(stdout), - error: error instanceof Error ? error.message : String(error) + error: timedOut + ? `workflow runner query exceeded its time limit and was stopped (ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS=${String(commandTimeoutMs)})` + : error instanceof Error + ? error.message + : String(error) }; } } @@ -6236,6 +6260,8 @@ async function execSmithersCli(input: { acceptedExitCodes?: readonly number[]; signal?: AbortSignal; timeoutMs?: number; + /** Bounds only the runner process, never executable preparation. */ + commandTimeoutMs?: number; }): Promise<{ stdout: string; stderr: string; command: string[]; exitCode: number }> { const command = [...input.args]; const executionDeadline = input.timeoutMs === undefined ? undefined : Date.now() + input.timeoutMs; @@ -6243,8 +6269,10 @@ async function execSmithersCli(input: { signal: input.signal, timeoutMs: input.timeoutMs }); - const commandTimeoutMs = + const remainingMs = executionDeadline === undefined ? undefined : Math.max(1, Math.ceil(executionDeadline - Date.now())); + const commandTimeoutMs = + remainingMs === undefined ? input.commandTimeoutMs : Math.min(remainingMs, input.commandTimeoutMs ?? remainingMs); const { anchored, snapshotAnchor, executableAnchor } = acquireAnchoredSmithersController(command, commandEnvironment); try { const { stdout, stderr } = await execFileAsync( diff --git a/packages/runtime/src/state-export.ts b/packages/runtime/src/state-export.ts index 157b8564c..393006c95 100644 --- a/packages/runtime/src/state-export.ts +++ b/packages/runtime/src/state-export.ts @@ -307,7 +307,7 @@ export async function getRunHealth(input: { })) ]); } - const health = parseRunHealth(snapshot.json, evidence.smithersRunId); + const health = parseRunHealth(snapshot.json, evidence.smithersRunId, input.runId); if (health === undefined) { return runtimeFailure([ ...syncDiagnostics, @@ -830,7 +830,8 @@ export const CURRENT_SMITHERS_STATUS_KEY_CONTRACT = { function parseRunHealth( value: unknown, - expectedWorkflowRunId: string + expectedWorkflowRunId: string, + runId: string ): | Omit | undefined { @@ -978,7 +979,7 @@ function parseRunHealth( return { workflow_status: workflowStatus, verdict, - reason: publicHealthReason(reason), + reason: publicHealthReason(reason, runId), counts: typedCounts, model_mix: modelMix, throughput: { @@ -1165,9 +1166,13 @@ function parseRunHealthOneshotControl(value: unknown): RunHealthValue["oneshot_c }; } -function publicHealthReason(value: string): string { - // `ultrafuzz why` now wraps the engine diagnosis, so recommend it directly. - return value.replace(/`?smithers\s+why`?/giu, "`ultrafuzz why`").replace(/smithers/giu, "workflow runner"); +function publicHealthReason(value: string, runId: string): string { + // Point at the Ultrafuzz commands that wrap the runner's own: `ultrafuzz why` for its diagnosis and + // `ultrafuzz resume` to continue an orphaned run. + return value + .replace(/`?smithers\s+why`?/giu, "`ultrafuzz why`") + .replace(/`?smithers\s+supervise\s+-r\s+[^\s`;,]+`?/giu, `\`ultrafuzz resume ${runId}\``) + .replace(/smithers/giu, "workflow runner"); } function objectRecord(value: unknown): Record | undefined { diff --git a/packages/runtime/src/workflow-control.ts b/packages/runtime/src/workflow-control.ts index 2626a8d53..80e147858 100644 --- a/packages/runtime/src/workflow-control.ts +++ b/packages/runtime/src/workflow-control.ts @@ -32,6 +32,8 @@ export interface WorkflowControlProjectionInput { export interface WorkflowControlProjection { state: RunState; changed: boolean; + /** The only change is the lease renewal and concurrency clock that every projection advances. */ + observationOnly: boolean; transitioned: boolean; deadlineExceeded: boolean; recoveryDue: boolean; @@ -233,15 +235,31 @@ export function projectWorkflowControlState(input: WorkflowControlProjectionInpu !isTerminalRunStatus(state.status) && state.status !== "paused"; + const changed = JSON.stringify(state) !== JSON.stringify(input.state); return { state, - changed: JSON.stringify(state) !== JSON.stringify(input.state), + changed, + observationOnly: + changed && + JSON.stringify(withoutObservationClock(state)) === JSON.stringify(withoutObservationClock(input.state)), transitioned: controlTransition, deadlineExceeded, recoveryDue }; } +function withoutObservationClock(state: RunState): unknown { + const { renewed_at: _renewedAt, expires_at: _expiresAt, ...lease } = state.controller_lease; + const { + observed_at: _observedAt, + queued_duration_ms: _queuedMs, + active_duration_ms: _activeMs, + idle_duration_ms: _idleMs, + ...concurrency + } = state.concurrency; + return { ...state, controller_lease: lease, concurrency }; +} + function assertExactWorkflowStates( nodeStates: ReadonlyMap, runState: WorkflowRunState | undefined diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index a70ad8a2e..d994ddc63 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -22,6 +22,7 @@ import { createNodeAttemptLedgerEntry, createUsageLedgerEntry, getNodeArtifactDir, + isTerminalRunStatus, layoutForRunRoot, manifestDigest, nodeAttemptLedgerIdentity, @@ -1650,7 +1651,10 @@ export async function synchronizeLinkedWorkflowRun( if (preControlMutationBudgetDiagnostic !== undefined) { return { ok: false, diagnostics: [preControlMutationBudgetDiagnostic] }; } - if (workflowControl.changed || deadlineApplied) { + // Only an explicit synchronization of a live run persists a lease renewal on its own. Otherwise every + // status poll, and every observation of a finished run, would rewrite state.json with a new clock. + const persistObservation = control.observeOnly !== true && !isTerminalRunStatus(workflowControl.state.status); + if (deadlineApplied || (workflowControl.changed && (persistObservation || !workflowControl.observationOnly))) { writeRunState(layout, workflowControl.state, { forbiddenSecretValues }); } if (deadlineApplied) { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index c3f0ce2b5..3cbcc5048 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -13130,6 +13130,146 @@ function prependFakeSmithersCase(env: Record, subcom fs.writeFileSync(smithers, source.replace(` ${subcommand})\n`, ` ${subcommand})\n${script}`)); } +test("repeated observations leave an unchanged finished or orphaned run byte-identical", async () => { + for (const scenario of ["finished", "orphaned"] as const) { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = `idempotent-observation-${scenario}`; + const workflowRunId = `ultrafuzz-${runId}`; + const status = currentStatusEnvelope(workflowRunId); + Object.assign( + status.data as Record, + scenario === "finished" + ? { + status: "finished", + verdict: "done", + reason: "run finished", + liveness: { state: "succeeded" }, + finishedAtMs: 2_000 + } + : { verdict: "orphaned", reason: "engine heartbeat is stale", liveness: { state: "orphaned" } } + ); + const env = fakeLifecycleSmithersEnv(project, { + inspect: + scenario === "finished" + ? workflowInspect({ workflowRunId, steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] }) + : workflowInspect({ + workflowRunId, + status: "running", + state: "orphaned", + steps: [{ id: "node:project-discovery", state: "in-progress", attempt: 1 }] + }), + events: workflowEvents( + workflowRunId, + scenario === "finished" + ? [ + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ] + : [{ type: "RunStarted" }, { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }] + ), + status + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + if (scenario === "finished") { + writeRequiredArtifactSet(runRoot, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + } + const first = await getRunHealth({ projectRoot: project, runId, env }); + assert.equal(first.ok, true, `${scenario}: ${JSON.stringify(first.diagnostics)}`); + const settled = readRunState(layoutForRunRoot(runRoot, runId)); + assert.equal(settled.status, scenario === "finished" ? "succeeded" : "running"); + assert.equal(settled.controller_lease.status, scenario === "finished" ? "active" : "expired"); + const before = runStateDocumentBytes(runRoot); + await new Promise((resolve) => setTimeout(resolve, 20)); + + const second = await getRunHealth({ projectRoot: project, runId, env }); + + assert.equal(second.ok, true, `${scenario}: ${JSON.stringify(second.diagnostics)}`); + assert.equal(second.value?.verdict, scenario === "finished" ? "done" : "orphaned"); + assert.deepEqual(runStateDocumentBytes(runRoot), before, `${scenario}: a repeated status rewrote run state`); + if (scenario === "finished") { + const explicit = await syncRun({ projectRoot: project, runId, env }); + assert.equal(explicit.ok, true, JSON.stringify(explicit.diagnostics)); + assert.deepEqual(runStateDocumentBytes(runRoot), before, "an explicit sync renewed a finished run's lease"); + } + } +}); + +test("an explicit synchronization keeps renewing a live run's controller lease", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "renewed-live-lease"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "in-progress", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const layout = layoutForRunRoot(run.value.run_root, runId); + assert.equal((await syncRun({ projectRoot: project, runId, env })).ok, true); + const renewedAt = readRunState(layout).controller_lease.renewed_at; + await new Promise((resolve) => setTimeout(resolve, 20)); + + assert.equal((await syncRun({ projectRoot: project, runId, env })).ok, true); + + // The eval and Modal pumps rely on this renewal to see that a run is still owned. + assert.ok(Date.parse(readRunState(layout).controller_lease.renewed_at) > Date.parse(renewedAt)); +}); + +test("status reports a runner query that stops answering instead of waiting on it", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "stalled-runner-query"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "in-progress", attempt: 1 }] + }), + status: currentStatusEnvelope(workflowRunId) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + prependFakeSmithersCase(env, "status", " sleep 60\n"); + let timer: NodeJS.Timeout | undefined; + + const outcome = await Promise.race([ + getRunHealth({ projectRoot: project, runId, env: { ...env, ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS: "3000" } }), + new Promise<"waiting">((resolve) => { + timer = setTimeout(() => resolve("waiting"), 40_000); + }) + ]); + + clearTimeout(timer); + if (outcome === "waiting") assert.fail("status waited on a runner query that stopped answering"); + assert.equal(outcome.ok, false); + const failure = outcome.diagnostics.find((diagnostic) => diagnostic.code === "WORKFLOW_STATUS_FAILED"); + assert.match( + failure?.message ?? "", + /exceeded its time limit.*ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS=3000/u, + JSON.stringify(outcome.diagnostics) + ); +}); + test("a failed deadline cancel is a warning and the next status requests it again", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); @@ -13688,7 +13828,11 @@ test("getRunHealth accepts strict 0.35 orphan, cancel-pending, quota, and operat data: { ...base, verdict, - reason: `run is ${verdict}`, + // The orphaned reason is the pinned runner's own wording, remediation included. + reason: + verdict === "orphaned" + ? "engine heartbeat is stale and its process is gone (last heartbeat 2026-08-13T00:00:00.000Z); the run is orphaned — resume it with `smithers supervise -r ultrafuzz-health-034-shapes`" + : `run is ${verdict}`, liveness: { state: verdict, unhealthy: { kind: "engine-heartbeat-stale", lastHeartbeatAt: "2026-08-13T00:00:00.000Z" } @@ -13712,6 +13856,13 @@ test("getRunHealth accepts strict 0.35 orphan, cancel-pending, quota, and operat const health = await getRunHealth({ projectRoot: project, runId: "health-034-shapes", env }); assert.equal(health.ok, true, `${verdict}: ${JSON.stringify(health.diagnostics)}`); assert.equal(health.value?.verdict, verdict); + if (verdict === "orphaned") { + // `workflow runner supervise -r ...` is not a command an operator can run; resume is. + assert.match( + health.value?.reason ?? "", + /the run is orphaned — resume it with `ultrafuzz resume health-034-shapes`$/u + ); + } assert.equal(health.value?.started_by?.session_id, "session-1"); assert.equal(health.value?.attention?.crossed_count, 2); assert.equal(health.value?.oneshot_control?.message_id, "message-1"); diff --git a/packages/runtime/test/workflow-control.test.ts b/packages/runtime/test/workflow-control.test.ts index 232844a87..b1c7898e2 100644 --- a/packages/runtime/test/workflow-control.test.ts +++ b/packages/runtime/test/workflow-control.test.ts @@ -245,6 +245,34 @@ test("a paused run is not cancelled by an observation past its workflow deadline assert.equal(projection.deadlineExceeded, false); }); +test("re-projecting an unchanged run only advances its observation clock", () => { + const graph = syntheticGraph([node("active")]); + const state = initialState(graph, 1); + const active = state.nodes.active; + assert.ok(active); + active.status = "running"; + const project = (previous: RunState, workflowState: "running" | "orphaned", nowMs: number) => + projectWorkflowControlState({ + previousState: structuredClone(previous), + state: previous, + graph, + tasks: tasksFor(graph), + workflowStates: new Map([["active", "in-progress"]]), + workflowState, + nowMs + }); + const first = project(state, "running", BASE_MS + 1_000); + + const renewed = project(first.state, "running", BASE_MS + 20_000); + assert.equal(renewed.changed, true); + assert.equal(renewed.observationOnly, true); + assert.equal(renewed.state.controller_lease.renewed_at, new Date(BASE_MS + 20_000).toISOString()); + + const orphaned = project(first.state, "orphaned", BASE_MS + 40_000); + assert.equal(orphaned.observationOnly, false); + assert.equal(orphaned.state.controller_lease.status, "expired"); +}); + test("controller recovery rejects malformed present timestamps instead of substituting the clock", () => { const graph = syntheticGraph([node("pending")]); const state = initialState(graph, 1, 60, 45); From 4653c4b707c60ebfdef492bfb2ff202f71203d94 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:25:14 +0000 Subject: [PATCH 018/206] fix(runtime): run one synchronization pass per run at a time and drop redundant artifact events status, inspect, why, stats, the dashboard and the eval runner all run the full mutating synchronization, and nothing serialized them (#1152). Two passes over one newly finished node both finalize it and race on its manifest temp file, state.json and the journals; the issue analysis reproduced duplicate node-synced events, a false failure that the immutable attempt ledger then kept, and passes that died midway through a write group. synchronizeLinkedWorkflowRun now runs each pass under a non-blocking proper-lockfile lock on /.workflow-sync. A pass that finds it held returns ok with an info WORKFLOW_SYNC_IN_PROGRESS diagnostic and writes nothing, so status never waits or fails on contention. Acquisition never waits, so it cannot deadlock with the control lock, and its target is not the run root because proper-lockfile keys its in-process registry by target and the control lock already locks that path. The lock lives in a small wrapper that re-enters the unchanged pass body. node-artifacts-verified and node-artifacts-missing duplicated node provenance, had no reader, and reported `missing: []` for every schema, semantic or authority failure. They are no longer written; their schema variants stay so existing journals replay. A failed node's durable last_error now prefixes each error with its run-relative path and JSON pointer, never a host path. The eventProvenanceForTask plumbing is deleted: createEventRecord never persisted it. Refs #1152 Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/src/events.ts | 1 + packages/runtime/src/workflow-sync.ts | 151 ++++++++++++++++---------- packages/runtime/test/runtime.test.ts | 120 ++++++++++++++++++++ 3 files changed, 212 insertions(+), 60 deletions(-) diff --git a/packages/artifacts/src/events.ts b/packages/artifacts/src/events.ts index 194a22787..108a7e95d 100644 --- a/packages/artifacts/src/events.ts +++ b/packages/artifacts/src/events.ts @@ -57,6 +57,7 @@ export const EVENT_RECORD_TYPES = [ "workflow-failure-unattributed", "run-recovered", "node-synced", + // No longer emitted (node-synced and node state carry the outcome); kept so existing journals replay. "node-artifacts-verified", "node-artifacts-missing", "node-controller-refinalization-intent", diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index d994ddc63..852bdc704 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -3,6 +3,7 @@ import fs from "node:fs"; import path from "node:path"; import { isDeepStrictEqual } from "node:util"; import { parseResolvedConfigJsonBytes } from "@ultrafuzz/config"; +import lockfile from "proper-lockfile"; import { ARTIFACT_MANIFEST_FILE, @@ -351,10 +352,7 @@ interface NodeFinalization { type PendingNodeAppendEvent = Extract< AppendEventInput, - { - eventType: - "node-artifacts-verified" | "node-artifacts-missing" | "findings-validated" | "artifact-manifest-written"; - } + { eventType: "findings-validated" | "artifact-manifest-written" } >; type PendingNodeEvent = PendingNodeAppendEvent extends infer Event ? Event extends PendingNodeAppendEvent @@ -1286,6 +1284,75 @@ async function assertStoppedResetAuthorityUnchanged(input: { } } +/** Controls of passes that hold their run's synchronization lock. */ +const exclusiveSynchronizations = new WeakSet(); + +/** + * Two passes over one newly finished node would both finalize it and race on its manifest, state + * and journals. A pass therefore runs only under its run's synchronization lock, and a pass that + * finds the lock held leaves the run to the pass in progress instead of waiting for it. + */ +async function synchronizeExclusively( + input: SyncRunInput, + control: WorkflowSynchronizationControl +): Promise< + { ok: true; value: SyncRunValue; diagnostics: RuntimeDiagnostic[] } | { ok: false; diagnostics: RuntimeDiagnostic[] } +> { + const exclusive = { ...control }; + exclusiveSynchronizations.add(exclusive); + const layoutResult = await checkedRunLayout(path.resolve(input.projectRoot), input.runId); + if (!layoutResult.ok || !fs.existsSync(layoutResult.layout.root)) { + // A missing or invalid run has nothing to lock; the pass reports it. + return synchronizeLinkedWorkflowRun(input, exclusive); + } + const layout = layoutResult.layout; + const release = await tryAcquireSynchronizationLock(layout); + if (release === undefined) { + return { + ok: true, + diagnostics: [ + { + code: "WORKFLOW_SYNC_IN_PROGRESS", + message: + "another synchronization of this run is in progress, so this one was skipped; local run state may lag until it finishes", + severity: "info", + source: "workflow", + path: layout.root + } + ], + value: { run_id: layout.runId, run_root: layout.root, status: readRunState(layout).status, synced_nodes: 0 } + }; + } + try { + return await synchronizeLinkedWorkflowRun(input, exclusive); + } finally { + // Failing to remove the lock never fails a finished pass; a lock left behind goes stale. + await release().catch(() => undefined); + } +} + +/** + * Acquires the lock without waiting. Its target is not the run root: the control lock locks that + * path, and proper-lockfile keys its in-process registry by target path, so a second lock on it + * would release the first. + */ +async function tryAcquireSynchronizationLock(layout: RunLayout): Promise<(() => Promise) | undefined> { + const target = path.join(layout.root, ".workflow-sync"); + try { + return await lockfile.lock(target, { + lockfilePath: `${target}.lock`, + realpath: false, + stale: 300_000, + update: 60_000, + // The default handler throws from a timer and would kill an observer or the eval runner. + onCompromised: () => undefined + }); + } catch (error) { + if (error instanceof Error && "code" in error && error.code === "ELOCKED") return undefined; + throw error; + } +} + export async function syncRun(input: SyncRunInput, control: WorkflowSynchronizationControl = {}) { const result = await synchronizeLinkedWorkflowRun(input, control); if (!result.ok) { @@ -1300,6 +1367,7 @@ export async function synchronizeLinkedWorkflowRun( ): Promise< { ok: true; value: SyncRunValue; diagnostics: RuntimeDiagnostic[] } | { ok: false; diagnostics: RuntimeDiagnostic[] } > { + if (!exclusiveSynchronizations.has(control)) return synchronizeExclusively(input, control); let synchronizationNowMs = synchronizationClock(control); const budgetDiagnostic = synchronizationBudgetDiagnostic(control, synchronizationNowMs); if (budgetDiagnostic !== undefined) { @@ -4097,12 +4165,10 @@ async function synchronizeTasks(input: { } syncedNodes += 1; if (previous?.status !== patchStatus) { - const eventProvenance = eventProvenanceForTask(task); appendEvent(input.layout, { eventType: "node-synced", nodeId: task.attemptId, status: patchStatus, - ...(eventProvenance === undefined ? {} : { provenance: eventProvenance }), payload: { workflow_run_id: input.workflowRunId, workflow_task_id: attemptEvidence.taskId, @@ -4305,19 +4371,7 @@ async function finalizeTerminalTask(input: { } : {}) }, - events: - verifierOutputGate !== undefined - ? [ - { - eventType: "node-artifacts-missing", - status: "failed", - payload: { - output_contracts: input.node.outputs, - missing: verifierOutputGate.missing - } - } - ] - : [] + events: [] }; } @@ -4363,25 +4417,6 @@ async function finalizeTerminalTask(input: { authenticatedGateSnapshots(input.node, verifierAuthority) ); diagnostics.push(...gate.diagnostics); - if (gate.ok) { - events.push({ - eventType: "node-artifacts-verified", - status: "succeeded", - payload: { - output_contracts: input.node.outputs, - missing: gate.missing - } - }); - } else { - events.push({ - eventType: "node-artifacts-missing", - status: "failed", - payload: { - output_contracts: input.node.outputs, - missing: gate.missing - } - }); - } let findingsCount: number | undefined; let artifactManifestSha256: string | undefined; @@ -4547,7 +4582,7 @@ async function finalizeTerminalTask(input: { return { status: "failed", diagnostics, - lastError: errorDiagnostics.map((diagnostic) => diagnostic.message).join("; "), + lastError: errorDiagnostics.map((diagnostic) => durableDiagnosticText(input.layout, diagnostic)).join("; "), provenance: { // A controller publication failure is never an output-contract // success. In particular, the run-state schema forbids publishing @@ -4590,6 +4625,22 @@ async function finalizeTerminalTask(input: { }; } +/** + * Durable failure text names where each error is, relative to the run, so state.json and the attempt + * ledger say which artifact failed without carrying a host path. + */ +function durableDiagnosticText(layout: RunLayout, diagnostic: RuntimeDiagnostic): string { + if (diagnostic.path === undefined) return diagnostic.message; + const pointerStart = diagnostic.path.indexOf("#"); + const filePath = pointerStart === -1 ? diagnostic.path : diagnostic.path.slice(0, pointerStart); + const pointer = pointerStart === -1 ? "" : diagnostic.path.slice(pointerStart); + const relative = path.isAbsolute(filePath) ? path.relative(layout.root, filePath) : filePath; + if (relative === "" || relative === ".." || relative.startsWith(`..${path.sep}`) || path.isAbsolute(relative)) { + return diagnostic.message; + } + return `${relative.split(path.sep).join("/")}${pointer}: ${diagnostic.message}`; +} + function admittedManifestPrerequisiteAttemptIds( task: StoredWorkflowTask, admittedDependencyAttemptIds: readonly string[] @@ -5008,31 +5059,11 @@ function appendNodeEvents( events: PendingNodeEvent[], forbiddenSecretValues: readonly string[] ): void { - const provenance = eventProvenanceForTask(task); for (const event of events) { - appendEvent(layout, { - ...event, - nodeId: task.attemptId, - ...(provenance === undefined ? {} : { provenance }), - forbiddenSecretValues - }); + appendEvent(layout, { ...event, nodeId: task.attemptId, forbiddenSecretValues }); } } -function eventProvenanceForTask(task: StoredWorkflowTask): Record | undefined { - const producerNodeId = task.metadata?.node?.producerNodeId; - const storageId = task.metadata?.node?.storageId; - const dynamic = task.metadata?.node?.dynamic; - if (producerNodeId === undefined && storageId === undefined && dynamic === undefined) return undefined; - return { - producer_node_id: producerNodeId ?? task.concreteNodeId, - concrete_node_id: task.concreteNodeId, - strategy_attempt_id: task.attemptId, - ...(storageId === undefined ? {} : { storage_id: storageId }), - ...(dynamic === undefined ? {} : { dynamic }) - }; -} - function taskAttemptInputManifestDigest(layout: RunLayout, task: StoredWorkflowTask): string { const state = readRunState(layout); return manifestDigest( diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 3cbcc5048..65db19c72 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -13270,6 +13270,126 @@ test("status reports a runner query that stops answering instead of waiting on i ); }); +test("a synchronization that finds another in progress skips it without waiting or writing", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = "single-sync-writer"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ]), + status: currentStatusEnvelope(workflowRunId) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + writeRequiredArtifactSet(runRoot, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + // Park the first synchronization inside its runner inspection until the test releases it. + const entered = path.join(project, "inspect-entered"); + const gate = path.join(project, "inspect-gate"); + prependFakeSmithersCase( + env, + "inspect", + ` : > ${shellQuote(entered)}\n while [ ! -f ${shellQuote(gate)} ]; do sleep 0.05; done\n` + ); + const inProgress = syncRun({ projectRoot: project, runId, env }); + try { + const waitStartedAt = Date.now(); + while (!fs.existsSync(entered)) { + assert.ok(Date.now() - waitStartedAt < 120_000, "the first synchronization never reached its inspection"); + await new Promise((resolve) => setTimeout(resolve, 20)); + } + const before = runStateDocumentBytes(runRoot); + let timer: NodeJS.Timeout | undefined; + + const concurrent = await Promise.race([ + getRunHealth({ projectRoot: project, runId, env }), + new Promise<"waiting">((resolve) => { + timer = setTimeout(() => resolve("waiting"), 30_000); + }) + ]); + + clearTimeout(timer); + if (concurrent === "waiting") assert.fail("status waited on the synchronization in progress"); + assert.equal(concurrent.ok, true, JSON.stringify(concurrent.diagnostics)); + assert.ok( + concurrent.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_SYNC_IN_PROGRESS"), + JSON.stringify(concurrent.diagnostics) + ); + assert.deepEqual(runStateDocumentBytes(runRoot), before, "the skipped synchronization wrote run state"); + } finally { + fs.writeFileSync(gate, ""); + } + const completed = await inProgress; + assert.equal(completed.ok, true, JSON.stringify(completed.diagnostics)); + assert.equal(completed.value?.status, "succeeded"); + const finalizations = replayEvents(layoutForRunRoot(runRoot, runId), Number.MAX_SAFE_INTEGER).records.filter( + (event) => event.event_type === "node-synced" && event.node_id === "project-discovery" + ); + assert.equal(finalizations.length, 1); +}); + +test("a node whose outputs fail validation names each failing artifact in its last error", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = "located-output-failure"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + writeRequiredArtifactSet(runRoot, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + // The verification marker binds these bytes, so only the host's schema check rejects them. + fs.writeFileSync( + path.join(runRoot, "artifacts", "project-discovery", "findings.json"), + JSON.stringify([{ schema_version: "ultrafuzz.finding.v2", id: "finding-without-required-fields" }]) + ); + writeCurrentArtifactVerificationMarker(runRoot, "project-discovery"); + + const sync = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + const layout = layoutForRunRoot(runRoot, runId); + const node = readRunState(layout).nodes["project-discovery"]; + assert.equal(node?.status, "failed"); + assert.match( + node?.last_error ?? "", + /(?:^|; )artifacts\/project-discovery\/findings\.json#\/0: must have required property '/u, + node?.last_error + ); + assert.equal(node?.last_error?.includes(runRoot), false, "last_error carries a host path"); + // node-synced and the node's state carry the outcome; an empty "missing artifacts" event only misled. + assert.deepEqual( + replayEvents(layout, Number.MAX_SAFE_INTEGER) + .records.map((event) => event.event_type) + .filter((type) => type.startsWith("node-artifacts-")), + [] + ); +}); + test("a failed deadline cancel is a warning and the next status requests it again", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From aa3574ac243f54088fd40b700839efb56e053a4c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:25:14 +0000 Subject: [PATCH 019/206] fix(runtime): status reports the runner's health when its state refresh fails status ran its observe-only synchronization strictly. A failed or malformed runner events query made it return ok:false, exiting 1 and stopping --watch, although the runner's own health was available, and any error thrown by the refresh (for example a torn usage.jsonl append) escaped getRunHealth. stats already tolerated the transient class. The observe-only pass now sets tolerateInvalidEventStreams and downgrades the same transient code set as stats; workflow-sync exports it once as TRANSIENT_SYNC_DIAGNOSTIC_CODES for both. An error thrown by the refresh becomes a WORKFLOW_STATE_SYNC_FAILED warning (the snapshot race keeps WORKFLOW_STATE_SYNC_RACED), and status goes on to query the runner. getRunHealth also passes the evidence it already verified into the pass instead of reading and verifying the sealed snapshot a second time. Refs #1145 Co-Authored-By: Claude Opus 5.5 --- packages/cli/src/commands/stats.ts | 15 +++----- packages/runtime/src/state-export.ts | 46 ++++++++++++++----------- packages/runtime/src/workflow-sync.ts | 26 ++++++++++++-- packages/runtime/test/runtime.test.ts | 49 +++++++++++++++++++++++++++ 4 files changed, 103 insertions(+), 33 deletions(-) diff --git a/packages/cli/src/commands/stats.ts b/packages/cli/src/commands/stats.ts index fcbbe3f8c..a5ed23685 100644 --- a/packages/cli/src/commands/stats.ts +++ b/packages/cli/src/commands/stats.ts @@ -17,6 +17,7 @@ import { validateSafeId } from "@ultrafuzz/artifacts"; import { + TRANSIENT_SYNC_DIAGNOSTIC_CODES, describeObservationSynchronizationDeadline, observationSynchronizationDeadline, runsRootForProject, @@ -162,17 +163,9 @@ async function loadLocalEvidence( const snapshot = readCoherentLocalEvidenceSnapshot(layout); const runMetadata = assertRunMetadataDocument(parseLocalJson(snapshot.runMetadata, layout.runMetadataPath), runId); if (!synchronized.ok && runMetadata.workflow !== undefined) { - const transientCodes = new Set([ - "WORKFLOW_INSPECT_FAILED", - "WORKFLOW_INSPECT_INVALID", - "WORKFLOW_EVENTS_FAILED", - "WORKFLOW_EVENTS_INVALID", - "WORKFLOW_TOKEN_EVENTS_FAILED", - "WORKFLOW_TOKEN_EVENTS_INVALID", - "WORKFLOW_SYNC_CANCELLED", - "WORKFLOW_SYNC_DEADLINE_EXCEEDED" - ]); - const authorityFailure = synchronized.diagnostics.find((diagnostic) => !transientCodes.has(diagnostic.code)); + const authorityFailure = synchronized.diagnostics.find( + (diagnostic) => !TRANSIENT_SYNC_DIAGNOSTIC_CODES.has(diagnostic.code) + ); if (authorityFailure !== undefined) { throw new Error(`linked workflow authority is invalid: ${authorityFailure.message}`); } diff --git a/packages/runtime/src/state-export.ts b/packages/runtime/src/state-export.ts index 393006c95..991b0221a 100644 --- a/packages/runtime/src/state-export.ts +++ b/packages/runtime/src/state-export.ts @@ -44,8 +44,13 @@ import { summarizeRunProgress } from "./run-progress.js"; import { diagnosticFromError, runtimeFailure, runtimeResult } from "./utils.js"; import { parseCurrentSmithersInspect, runSmithersInspectionCommand, type SmithersCommandSnapshot } from "./smithers.js"; import { workflowControlDivergenceDiagnostics } from "./control-divergence-diagnostics.js"; -import { linkedWorkflowExecutionEnvironment, readLinkedWorkflowEvidence } from "./start-run.js"; import { + linkedWorkflowExecutionEnvironment, + readLinkedWorkflowEvidence, + type LinkedWorkflowEvidence +} from "./start-run.js"; +import { + TRANSIENT_SYNC_DIAGNOSTIC_CODES, describeObservationSynchronizationDeadline, observationSynchronizationDeadline, synchronizeLinkedWorkflowRun @@ -269,10 +274,7 @@ export async function getRunHealth(input: { const syncDiagnostics: RuntimeDiagnostic[] = [...controlDiagnostics]; if (controlDiagnostics.length === 0) { syncDiagnostics.push( - ...(await synchronizeObservedWorkflowRun( - { projectRoot, runId: input.runId, env: input.env }, - evidence.layout.root - )) + ...(await synchronizeObservedWorkflowRun({ projectRoot, runId: input.runId, env: input.env }, evidence)) ); } else { syncDiagnostics.push({ @@ -358,35 +360,41 @@ export async function getRunHealth(input: { * Observe-only synchronization reads state.json and run.json strictly while the run's controller, or * another concurrent `status`, keeps replacing them by atomic rename, so it can hit the same transient * snapshot race the direct reads in `getRunHealth` retry. It is retried within the same bounded - * budget. Once that budget is spent the run is still reported: health comes from the workflow runner - * and local run state is simply the last coherent snapshot, which is said in a warning, the way a - * skipped synchronization is reported. Every other failure propagates unchanged. + * budget. * - * The refresh is also bounded by the opt-in observation deadline when one is configured. Health below - * comes from the direct runner query, so an exceeded deadline is a warning about possibly stale local - * state, not a failure. One absolute deadline spans every retry, so racing reads cannot extend the - * observer's wall-clock budget. + * Health below comes from the direct runner query, so a refresh that fails never hides it: a failed + * or malformed runner query, an exceeded opt-in observation deadline, an exhausted race budget, and any + * other synchronization error are warnings that local run state may be stale. One absolute deadline + * spans every retry, so racing reads cannot extend the observer's wall-clock budget. */ -async function synchronizeObservedWorkflowRun(input: SyncRunInput, runRoot: string): Promise { +async function synchronizeObservedWorkflowRun( + input: SyncRunInput, + evidence: LinkedWorkflowEvidence +): Promise { const deadlineMs = observationSynchronizationDeadline(input.env); try { const sync = await retryTransientSnapshotObservation(() => - synchronizeLinkedWorkflowRun(input, { observeOnly: true, deadlineMs }) + synchronizeLinkedWorkflowRun(input, { + observeOnly: true, + tolerateInvalidEventStreams: true, + deadlineMs, + evidence + }) ); return sync.diagnostics.map((diagnostic) => - diagnostic.code === "WORKFLOW_SYNC_DEADLINE_EXCEEDED" + TRANSIENT_SYNC_DIAGNOSTIC_CODES.has(diagnostic.code) ? describeObservationSynchronizationDeadline({ ...diagnostic, severity: "warning" as const }) : diagnostic ); } catch (error) { - if (!isTransientSnapshotRace(error)) throw error; + const reason = error instanceof Error ? error.message : String(error); return [ { - code: "WORKFLOW_STATE_SYNC_RACED", - message: `run state synchronization was skipped because ${error.message}; reported counts come from the workflow runner and local run state may be stale`, + code: isTransientSnapshotRace(error) ? "WORKFLOW_STATE_SYNC_RACED" : "WORKFLOW_STATE_SYNC_FAILED", + message: `run state synchronization was skipped because ${reason}; reported counts come from the workflow runner and local run state may be stale`, severity: "warning", source: "runtime", - path: runRoot + path: evidence.layout.root } ]; } diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index 852bdc704..8a739571e 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -386,8 +386,26 @@ export interface WorkflowSynchronizationControl { allowMissingWorkflowRun?: boolean; /** Status synchronization authenticates published evidence without taking or repairing control state. */ observeOnly?: boolean; + /** Evidence an observer already authenticated for this run, so the pass does not verify it again. */ + evidence?: LinkedWorkflowEvidence; } +/** + * Synchronization failures an observer reports as warnings: a runner query failed or returned + * unusable output, or the observation budget ran out. Local state stays the last coherent snapshot + * and the next poll retries. + */ +export const TRANSIENT_SYNC_DIAGNOSTIC_CODES: ReadonlySet = new Set([ + "WORKFLOW_INSPECT_FAILED", + "WORKFLOW_INSPECT_INVALID", + "WORKFLOW_EVENTS_FAILED", + "WORKFLOW_EVENTS_INVALID", + "WORKFLOW_TOKEN_EVENTS_FAILED", + "WORKFLOW_TOKEN_EVENTS_INVALID", + "WORKFLOW_SYNC_CANCELLED", + "WORKFLOW_SYNC_DEADLINE_EXCEEDED" +]); + const MAX_OBSERVATION_SYNC_TIMEOUT_MS = 60_000; /** @@ -1394,9 +1412,11 @@ export async function synchronizeLinkedWorkflowRun( }; } - const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId, { - ...(control.observeOnly === true ? { observeOnly: true } : {}) - }); + const evidence = + control.evidence ?? + (await readLinkedWorkflowEvidence(projectRoot, input.runId, { + ...(control.observeOnly === true ? { observeOnly: true } : {}) + })); if (!evidence.ok) { return { ok: false, diagnostics: evidence.diagnostics }; } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 65db19c72..463b492a2 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -13390,6 +13390,55 @@ test("a node whose outputs fail validation names each failing artifact in its la ); }); +test("status keeps reporting runner health when run-state synchronization fails", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "observer-survives-sync-failure"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "in-progress", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 } + ]), + status: currentStatusEnvelope(workflowRunId) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + const failEvents = path.join(project, "fail-events"); + prependFakeSmithersCase( + env, + "events", + ` if [ -f ${shellQuote(failEvents)} ]; then printf '%s\\n' 'event store busy' >&2; exit 1; fi\n` + ); + const reportedHealth = async (label: string) => { + const health = await getRunHealth({ projectRoot: project, runId, env }); + assert.equal(health.ok, true, `${label}: ${JSON.stringify(health.diagnostics)}`); + assert.equal(health.value?.verdict, "running-healthy", label); + return health.diagnostics; + }; + + fs.writeFileSync(failEvents, ""); + const failedQuery = await reportedHealth("failed events query"); + assert.ok( + failedQuery.some((diagnostic) => diagnostic.code === "WORKFLOW_EVENTS_FAILED" && diagnostic.severity === "warning"), + JSON.stringify(failedQuery) + ); + fs.rmSync(failEvents); + + // A torn ledger append makes the synchronization itself throw. + fs.appendFileSync(path.join(runRoot, "usage.jsonl"), '{"torn":'); + await reportedHealth("torn usage ledger"); +}); + test("a failed deadline cancel is a warning and the next status requests it again", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From 689e92e08b5f018c043fd6904f973b19c2912d71 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:25:14 +0000 Subject: [PATCH 020/206] docs: document the runner query timeout, synchronization contention and deadline changes configuration.md gains the ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS row and a paragraph on skipped passes, status warnings and when state.json is left untouched. The workflow_deadline_seconds paragraph from #1157 now says paused runs are not cancelled and a failed cancel is a retried warning, and the supervisor paragraph no longer claims that reporting checks the deadline. Refs #1110, #1145, #1152 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/configuration.md | 29 ++++++++++++++++++++--------- 2 files changed, 21 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..428e503b3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [cli] [artifacts] [docs]** Run-state synchronizations of one run no longer overlap: a pass that finds another in progress skips with `WORKFLOW_SYNC_IN_PROGRESS` instead of finalizing the same nodes a second time, and `status` or any observation of a finished run no longer rewrites `state.json` only to advance the controller-lease and concurrency clocks. Read-only runner queries time out after 120 s by default (`ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS`, capped at 600 s), and `status` reports failed runner `inspect`/`events` queries and unexpected synchronization errors as warnings beside the runner's health instead of failing. The unread `node-artifacts-verified`/`node-artifacts-missing` events are no longer written (old journals still replay), a failed node's `last_error` names each error's run-relative path, an orphaned run's status recommends `ultrafuzz resume `, a paused run is no longer cancelled at its workflow deadline, and a failed deadline cancel is a warning that the next synchronization retries (#1145, #1152, #1110). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/configuration.md b/docs/reference/configuration.md index 1095c7155..0744c0dcb 100644 --- a/docs/reference/configuration.md +++ b/docs/reference/configuration.md @@ -144,13 +144,15 @@ unattended run keeps executing, and incurring provider cost, past its deadline. The deadline is checked only when a command synchronizes the run: `ultrafuzz status` (including each `--watch` poll), `inspect`, `why`, and `stats`, plus the dashboard's inspect action and the eval runner's poll loop. `ultrafuzz run` -does not check it. If the run is not yet terminal (running, pending, or paused) -at the first synchronization after the deadline, that synchronization requests -cancellation and, when the request succeeds, marks the run `timed-out` and -appends a `workflow-deadline-exceeded` event. A failed request is reported as a -`WORKFLOW_DEADLINE_CANCEL_FAILED` error and leaves the run active; a -synchronization that fails or is skipped (for example -`WORKFLOW_STATE_SYNC_SKIPPED`) does not check the deadline at all. A run that +does not check it. If the run is running or pending at the first +synchronization after the deadline, that synchronization requests cancellation +and, when the request succeeds, marks the run `timed-out` and appends a +`workflow-deadline-exceeded` event. A paused run is not cancelled: it executes +nothing, and resuming it records a new deadline. A failed request is reported as +a `WORKFLOW_DEADLINE_CANCEL_FAILED` warning and leaves the run active, and the +next synchronization requests cancellation again; a synchronization that fails +or is skipped (for example `WORKFLOW_STATE_SYNC_SKIPPED` or +`WORKFLOW_SYNC_IN_PROGRESS`) does not check the deadline at all. A run that finished first keeps its terminal outcome, with no timeout record. To bound an unattended run, run `ultrafuzz status ` periodically (for example from cron) and act on its warnings, or cancel it with `ultrafuzz cancel `. @@ -174,8 +176,7 @@ Every submitted workflow starts a run-scoped recovery supervisor. The supervisor renews controller ownership through runner heartbeats and uses an atomic claim before taking over expired ownership, so completed work is not resubmitted. The workflow deadline is separate from per-node timeouts and is -checked whenever run state is synchronized by status, inspect, reporting, or -eval watchers. +checked only when run state is synchronized, as described above. ## Execution @@ -387,6 +388,7 @@ is local and remains the default. | `ULTRAFUZZ_PRICING_CATALOG_URL` | Live model-pricing catalog URL, or `disabled`, `none`, or `off`. | | `ULTRAFUZZ_PRICING_TIMEOUT_MS` | Positive catalog request timeout in milliseconds, capped at 60 seconds. | | `ULTRAFUZZ_OBSERVATION_SYNC_TIMEOUT_MS` | Optional deadline in milliseconds for the run-state refresh before `status`, `inspect`, `why`, and `stats`. Unset means no deadline (observers wait for full synchronization); a positive value bounds it, capped at 60000; `0` or `off` is the same as unset. | +| `ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS` | Timeout in milliseconds for each read-only workflow runner query (`inspect`, `events`, `node`, `status`, `why`, `ps`, `timeline`, `snapshots`). Defaults to 120000, capped at 600000. A query that exceeds it is reported as a failed query. | Boolean values accept `1`, `true`, `yes`, `on`, `0`, `false`, `no`, and `off`. @@ -400,6 +402,15 @@ the same as unset. When a configured deadline passes, the command reports a `WORKFLOW_SYNC_DEADLINE_EXCEEDED` warning and continues with the local run state, whose node states, attempt ledgers, and usage counts may then be stale. +A command or eval poll that finds another synchronization of the same run in +progress skips its own pass instead of running alongside it, reports +`WORKFLOW_SYNC_IN_PROGRESS`, and continues with the local run state. During its +refresh, `status` reports a failed or malformed runner `inspect` or `events` +query, and an unexpected synchronization error, as warnings, so it still shows +the workflow runner's health. When nothing but the observation time changed, +`status` and any synchronization of a finished run leave `state.json` untouched; +any other synchronization of a live run still renews its controller lease. + Custom pricing catalogs must use HTTPS without credentials, query parameters, or fragments and must resolve entirely to public addresses. The validated DNS address is pinned for the request, redirects are rejected, and response bodies From 7db436fe6a78fcd5929acf9f938ec915066338e0 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:29:46 +0000 Subject: [PATCH 021/206] perf(runtime): write execution-snapshot files with their final mode and flush each once Snapshot publication wrote every file with mode 0400 and fsynced it, then the permission seal reopened every file, fchmodded it (to 0500 for executables, otherwise to the 0400 it already had) and fsynced it again. A launch of the test fixture's 12,552-file snapshot therefore issued 27,876 fsyncs, two per file plus two per directory, and launch time is fsync-bound. writeSnapshotFile now creates each file with its final mode and sets that mode through the creating descriptor (the umask can clear creation bits) before its single flush, so the bytes and the mode are durable together, as the second flush previously made them. The permission seal now only makes directories read-only; the publication verification that already follows it still checks every entry's type, mode, link count and bytes. The same fixture now issues 15,324 fsyncs (one per file plus two per directory). The new test republishes a launched run's snapshot and bounds its fsync calls at one per file plus two per directory; origin/main fails it with 27,876 flushes for 12,552 files in 1,386 directories. Refs #921 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/workflow-integrity.ts | 117 ++++++++------------- packages/runtime/test/runtime.test.ts | 44 ++++++++ 2 files changed, 89 insertions(+), 72 deletions(-) diff --git a/packages/runtime/src/workflow-integrity.ts b/packages/runtime/src/workflow-integrity.ts index 4dcae3d23..315f1d487 100644 --- a/packages/runtime/src/workflow-integrity.ts +++ b/packages/runtime/src/workflow-integrity.ts @@ -1304,14 +1304,17 @@ function publishWorkflowExecutionSnapshot( inode: temporaryStat.ino }; assertSnapshotPublicationBoundary(boundary, "workflow execution snapshot creation"); - for (const [relative, contents] of expectedFiles) writeSnapshotFile(accessRoot, relative, contents, boundary); + const executables = new Set(executablePaths); + for (const [relative, contents] of expectedFiles) { + writeSnapshotFile(accessRoot, relative, contents, executables.has(relative) ? 0o500 : 0o400, boundary); + } for (const [relative, target] of expectedLinks) writeSnapshotLink(accessRoot, relative, target, boundary); - sealSnapshotPermissions(accessRoot, new Set(executablePaths), boundary); + sealSnapshotPermissions(accessRoot, boundary); verifyPublishedWorkflowExecutionSnapshot( accessRoot, expectedFiles, expectedLinks, - new Set(executablePaths), + executables, temporaryLexicalPath ); assertSnapshotPublicationBoundary(boundary, "workflow execution snapshot publication"); @@ -1449,6 +1452,7 @@ function writeSnapshotFile( root: string, relativePath: string, contents: Buffer, + mode: number, boundary: SnapshotPublicationBoundary ): void { snapshotPath(root, relativePath, "workflow execution snapshot file"); @@ -1459,7 +1463,7 @@ function writeSnapshotFile( const descriptor = fs.openSync( destination, fs.constants.O_RDWR | fs.constants.O_CREAT | fs.constants.O_EXCL | (fs.constants.O_NOFOLLOW ?? 0), - 0o400 + mode ); try { let offset = 0; @@ -1468,6 +1472,9 @@ function writeSnapshotFile( if (written <= 0) throw new Error(`workflow execution snapshot write made no progress: ${relativePath}`); offset += written; } + // The umask can clear bits of the creation mode. Setting the final mode through the + // creating descriptor lets this file's single flush make its bytes and mode durable. + fs.fchmodSync(descriptor, mode); fs.fsyncSync(descriptor); const observed = Buffer.alloc(contents.byteLength); let readOffset = 0; @@ -1890,11 +1897,7 @@ function workflowSnapshotEnvironment( }); } -function sealSnapshotPermissions( - root: string, - executablePaths: ReadonlySet, - boundary: SnapshotPublicationBoundary -): void { +function sealSnapshotPermissions(root: string, boundary: SnapshotPublicationBoundary): void { if (root !== boundary.accessRoot || boundary.descriptor === undefined) { throw new Error("workflow execution snapshot permission sealing requires its root descriptor"); } @@ -1907,17 +1910,21 @@ function sealSnapshotPermissions( device: boundary.device, inode: boundary.inode }, - "", - executablePaths + "" ); assertSnapshotPublicationBoundary(boundary, "workflow execution snapshot durability flush"); } +/** + * Make every snapshot directory read-only, deepest first. Files already carry + * their final mode from `writeSnapshotFile` and links have none, so neither is + * reopened here; publication verification then checks every entry's type, + * mode, link count, and bytes. + */ function sealSnapshotDirectoryPermissions( boundary: SnapshotPublicationBoundary, directory: OpenedSnapshotPublicationDirectory, - relativeDirectory: string, - executablePaths: ReadonlySet + relativeDirectory: string ): void { assertSnapshotPublicationBoundary(boundary, "workflow execution snapshot permission seal"); assertOpenedPublicationDirectoryCurrent(directory, "workflow execution snapshot permission seal"); @@ -1938,70 +1945,36 @@ function sealSnapshotDirectoryPermissions( ) { throw new Error(`workflow execution snapshot entry changed during permission sealing: ${relative}`); } - if (accessed.isSymbolicLink()) continue; - if (accessed.isDirectory()) { - const descriptor = openSnapshotDirectory(candidate); - if (descriptor === undefined) { - throw new Error("workflow execution snapshot directory has no descriptor-rooted permission support"); - } - try { - const opened = fs.fstatSync(descriptor); - if (!opened.isDirectory() || opened.dev !== accessed.dev || opened.ino !== accessed.ino) { - throw new Error(`workflow execution snapshot directory changed while sealing: ${relative}`); - } - const descriptorPath = verifiedSnapshotDescriptorPath( - descriptor, - opened.dev, - opened.ino, - "workflow execution snapshot directory" - ); - if (descriptorPath === undefined) { - throw new Error("workflow execution snapshot directory lost descriptor-rooted permission support"); - } - sealSnapshotDirectoryPermissions( - boundary, - { - accessPath: descriptorPath, - lexicalPath: lexicalCandidate, - descriptor, - device: opened.dev, - inode: opened.ino - }, - relative, - executablePaths - ); - } finally { - fs.closeSync(descriptor); - } - continue; - } - if (!accessed.isFile()) { - throw new Error(`workflow execution snapshot contains a non-regular entry while sealing: ${relative}`); + if (!accessed.isDirectory()) continue; + const descriptor = openSnapshotDirectory(candidate); + if (descriptor === undefined) { + throw new Error("workflow execution snapshot directory has no descriptor-rooted permission support"); } - const descriptor = fs.openSync(candidate, fs.constants.O_RDONLY | (fs.constants.O_NOFOLLOW ?? 0)); try { const opened = fs.fstatSync(descriptor); - if (!opened.isFile() || opened.dev !== accessed.dev || opened.ino !== accessed.ino || opened.nlink !== 1) { - throw new Error(`workflow execution snapshot file changed while sealing: ${relative}`); + if (!opened.isDirectory() || opened.dev !== accessed.dev || opened.ino !== accessed.ino) { + throw new Error(`workflow execution snapshot directory changed while sealing: ${relative}`); } - fs.fchmodSync(descriptor, executablePaths.has(relative) ? 0o500 : 0o400); - fs.fsyncSync(descriptor); - const completed = fs.fstatSync(descriptor); - const completedAccess = fs.lstatSync(candidate); - const completedLexical = fs.lstatSync(lexicalCandidate); - for (const observed of [completedAccess, completedLexical]) { - if ( - !observed.isFile() || - observed.isSymbolicLink() || - observed.dev !== completed.dev || - observed.ino !== completed.ino || - observed.mode !== completed.mode || - observed.nlink !== completed.nlink || - observed.size !== completed.size - ) { - throw new Error(`workflow execution snapshot file changed after sealing: ${relative}`); - } + const descriptorPath = verifiedSnapshotDescriptorPath( + descriptor, + opened.dev, + opened.ino, + "workflow execution snapshot directory" + ); + if (descriptorPath === undefined) { + throw new Error("workflow execution snapshot directory lost descriptor-rooted permission support"); } + sealSnapshotDirectoryPermissions( + boundary, + { + accessPath: descriptorPath, + lexicalPath: lexicalCandidate, + descriptor, + device: opened.dev, + inode: opened.ino + }, + relative + ); } finally { fs.closeSync(descriptor); } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..641daa3af 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -14794,6 +14794,50 @@ test("snapshot recovery removes a nested read-only stale current-generation publ assert.equal(fs.existsSync(savedSnapshot), true); }); +// Publication is fsync-bound: a snapshot holds thousands of dependency files. +test("publishing an execution snapshot flushes each file once", async (context) => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "snapshot-single-flush"; + const run = await startRun({ projectRoot: project, runId, env: fakeSmithersEnv(project) }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + const evidence = await readLinkedWorkflowEvidence(project, runId); + assert.equal(evidence.ok, true, "diagnostics" in evidence ? JSON.stringify(evidence.diagnostics) : ""); + if (!evidence.ok) return; + const published = evidence.executionSnapshot.root; + fs.chmodSync(published, 0o700); + fs.renameSync(published, `${path.dirname(published)}.saved-${evidence.verifiedControl.generation}`); + + const flushes = context.mock.method(fs, "fsyncSync"); + const started = performance.now(); + const republished = materializeWorkflowExecutionSnapshot({ + projectRoot: project, + layout: evidence.layout, + snapshot: evidence.verifiedControl + }); + const elapsed = Math.round(performance.now() - started); + flushes.mock.restore(); + + // Recursive readdir would follow the snapshot's dependency links, which form cycles. + let files = 0; + let directories = 0; + const pending = [republished.root]; + for (let directory = pending.pop(); directory !== undefined; directory = pending.pop()) { + directories += 1; + for (const entry of fs.readdirSync(directory, { withFileTypes: true })) { + if (entry.isDirectory()) pending.push(path.join(directory, entry.name)); + else if (entry.isFile()) files += 1; + } + } + assert.ok(files > 100, `expected a dependency-sized snapshot, got ${files} files`); + const observed = `${flushes.mock.callCount()} flushes for ${files} files in ${directories} directories (${elapsed} ms)`; + context.diagnostic(observed); + // One flush per file. Each directory adds one when it is created (its parent) and one when it is + // sealed, and the publishing rename adds one. + assert.ok(flushes.mock.callCount() <= files + 2 * directories + 1, observed); +}); + test("snapshot recovery rejects matching symlink and non-directory publications without escaping", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From c0f9625160303f52e62df2e02fba6064555dbf60 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:29:56 +0000 Subject: [PATCH 022/206] fix(runtime): doctor requires only the agent CLIs a run can dispatch to doctor required the executable of every configured model profile, so a fresh `ultrafuzz init` on a Codex-only host reported "needs attention" because the scaffold's opt-in Claude, Kimi and Pi profiles were not installed, although no node would run them. It now requires an agent's CLI only when the selected topology or its [retry] agents fallback chain can dispatch to that agent (the same selection the OpenRouter credential check already used). Other configured profiles' CLIs are still probed and listed, as not required, and the human output marks them "(not required)". Agents that share a CLI (Codex and OpenRouter, Claude and DeepSeek) require it if any of them is selected. doctor also built DOCTOR_WORKFLOW_ENGINE_* diagnostics and check statuses for the project-local engine and then discarded them, reporting both checks as "unknown". The builders are deleted; layout_status and the per-patch posture are still reported, and the reference docs no longer list the five codes that were never emitted. Launch and resume install the workflow engine controller under the OS temporary directory, and a native resume keeps its install there for the detached engine (#921 cost 1: a tmpfs /tmp exhausted a host's memory). A new temporary-directory check warns, without failing the verdict, when that directory is a tmpfs or has under 2 GiB free, and reports how many ultrafuzz-controller-* directories it holds and their total size. It never removes them, because a live engine may still use them. Refs #921 Co-Authored-By: Claude Opus 5.5 --- docs/config.md | 7 +- docs/reference/cli.md | 32 ++- docs/tutorials/first-campaign.md | 7 +- packages/cli/src/commands/doctor.ts | 2 +- packages/cli/test/lifecycle-commands.test.ts | 4 +- packages/runtime/src/doctor.ts | 264 +++++++----------- .../runtime/test/lifecycle-inspection.test.ts | 176 +++++++++++- 7 files changed, 306 insertions(+), 186 deletions(-) diff --git a/docs/config.md b/docs/config.md index 4ccc55e3e..e2ca9002b 100644 --- a/docs/config.md +++ b/docs/config.md @@ -321,10 +321,9 @@ else. `config_dir` is the XDG parent, not OpenCode's own directory: `XDG_DATA_HOME` becomes `/data`, so pointing it at `~/.config/opencode` or `~/.local/share/opencode` picks nothing up. -`ultrafuzz doctor` requires the `opencode` executable whenever any configured -profile uses `OpenCodeAgent` — every profile in `[models.*]` is checked, not -only the one a run selects, so keeping the shipped `[models.opencode]` profile -means every contributor needs the CLI installed. +`ultrafuzz doctor` requires the `opencode` executable only when a node of the +selected topology, or its `[retry] agents` fallback chain, uses an +`OpenCodeAgent` profile; otherwise the executable is listed as not required. Default triage requires quorum `3` from a panel size of `4`: diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..58165d053 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -533,31 +533,33 @@ non-launching configuration contract unchanged. Doctor reports: - config, topology, prompt, and reference validation posture; - required topology backends, toolchain, and configured agent executable availability in the configured execution environment (the local `PATH` for - local runs or a transient probe of the provider image for cloud runs); + local runs or a transient probe of the provider image for cloud runs). Only + the executables of agents the selected topology can dispatch to, including + its `[retry] agents` fallbacks, are required; other configured profiles' + executables are listed as not required; - the bundled workflow engine version, the version the generated project requires, and the installed project-local version and bin target; - npm's latest published stable engine version when the registry check is available; -- whether the installed dependency layout passes Ultrafuzz's exact - manifest and path validation; -- whether required compatibility patches, or their upstream replacements, are - present. A source carrying neither the patch nor the shape Ultrafuzz patches - is reported as modified or incompatible, because the next run fails in that - state. +- whether the project-local dependency layout passes Ultrafuzz's exact + manifest and path validation, and the posture of each compatibility patch in + it. Both are informational: a launch installs, patches, and seals its own + operator-owned controller; +- the OS temporary directory, where launch and resume install that controller: + its free space and how many `ultrafuzz-controller-*` directories it holds, + with their total size. Doctor warns when the directory is a RAM-backed tmpfs + or has less than 2 GiB free, and never removes those directories, because a + native resume keeps its controller there for the detached engine. Doctor does not create project run state or install, upgrade, or repair local dependencies. For cloud execution, checking required commands may create the configured provider app on first use and uses a transient sandbox so the probe runs inside the same image as workflow nodes. -Diagnostics are stable: `DOCTOR_TOOLCHAIN_MISSING`, -`DOCTOR_TOOLCHAIN_PROBE_FAILED`, -`DOCTOR_WORKFLOW_ENGINE_MISSING`, -`DOCTOR_WORKFLOW_ENGINE_VERSION_MISMATCH`, -`DOCTOR_WORKFLOW_ENGINE_LAYOUT_INVALID`, -`DOCTOR_WORKFLOW_ENGINE_PATCHES_PENDING`, -`DOCTOR_WORKFLOW_ENGINE_PATCHES_INCOMPATIBLE`, -`DOCTOR_WORKFLOW_ENGINE_OUTDATED`, and `DOCTOR_REGISTRY_UNAVAILABLE`. A registry or network failure produces a +Diagnostics are stable: `DOCTOR_AGENT_CREDENTIAL_MISSING`, +`DOCTOR_TOOLCHAIN_MISSING`, `DOCTOR_TOOLCHAIN_PROBE_FAILED`, +`DOCTOR_WORKFLOW_ENGINE_OUTDATED`, `DOCTOR_REGISTRY_UNAVAILABLE`, and +`DOCTOR_TEMPORARY_DIRECTORY_CONSTRAINED`. A registry or network failure produces a warning and an `unknown` latest version instead of failing an otherwise valid offline project. Doctor never installs, mutates, or upgrades dependencies. diff --git a/docs/tutorials/first-campaign.md b/docs/tutorials/first-campaign.md index eb23371fd..3b4863e9a 100644 --- a/docs/tutorials/first-campaign.md +++ b/docs/tutorials/first-campaign.md @@ -12,9 +12,10 @@ You need: - Node.js `22.19.0` or newer and the repository-pinned `pnpm` `11.1.1` for the TypeScript workspace. - Bun `1.3` or newer for the Smithers workflow executable, plus Git. -- The CLI executables for the configured agent profiles and their credentials. - `doctor` checks executables for every configured model profile, including - profiles that are not selected for a run. +- The CLI executables for the agent profiles your topology selects, and their + credentials. `doctor` requires the executables of the agents the selected + topology and its retry fallbacks use, and lists the other configured + profiles' executables as not required. - Foundry (`forge`) and the target project's test dependencies and layout. See [Development Commands](../reference/development.md) for host runtime and diff --git a/packages/cli/src/commands/doctor.ts b/packages/cli/src/commands/doctor.ts index 21192c00a..d0eb9315c 100644 --- a/packages/cli/src/commands/doctor.ts +++ b/packages/cli/src/commands/doctor.ts @@ -34,7 +34,7 @@ function renderDoctor(value: DoctorValue): string { `- ${entry.name}: ${ entry.available ? `${entry.path ?? "available"}${entry.version == null ? "" : ` (${entry.version})`}` - : "missing from execution environment" + : `missing from execution environment${entry.required ? "" : " (not required)"}` }` ), "Workflow engine:", diff --git a/packages/cli/test/lifecycle-commands.test.ts b/packages/cli/test/lifecycle-commands.test.ts index 66bc223b5..c5b8cabac 100644 --- a/packages/cli/test/lifecycle-commands.test.ts +++ b/packages/cli/test/lifecycle-commands.test.ts @@ -601,7 +601,9 @@ test("doctor reports install posture in human and JSON output", async () => { assert.match(human.stdout + human.stderr, /- compatibility patches: .*supervisor descriptor /u); assert.match(human.stdout + human.stderr, /- compatibility patches: .*terminal state restore /u); assert.match(human.stdout + human.stderr, /- compatibility patches: .*resume hydration /u); - assert.match(human.stdout + human.stderr, /- recon: missing from execution environment/u); + assert.match(human.stdout + human.stderr, /^- recon: missing from execution environment$/mu); + // The scaffold's opt-in Kimi profile is configured but not selected by the topology. + assert.match(human.stdout + human.stderr, /^- kimi: missing from execution environment \(not required\)$/mu); assert.doesNotMatch(human.stdout + human.stderr, /smthrs/u); assert.doesNotMatch(human.stdout + human.stderr, /smthrs/iu); diff --git a/packages/runtime/src/doctor.ts b/packages/runtime/src/doctor.ts index 0353a8151..2a724bb41 100644 --- a/packages/runtime/src/doctor.ts +++ b/packages/runtime/src/doctor.ts @@ -1,9 +1,11 @@ import { execFile } from "node:child_process"; +import fs from "node:fs"; +import os from "node:os"; import path from "node:path"; import { promisify } from "node:util"; import { referencesStatus } from "./references.js"; -import { inspectSmithersInstallation, type SmithersInstallationPosture } from "./smithers.js"; +import { inspectSmithersInstallation } from "./smithers.js"; import { SMITHERS_PACKAGE_NAME, SMITHERS_VERSION } from "./smithers-package.js"; import type { DoctorCheck, @@ -21,6 +23,11 @@ const execFileAsync = promisify(execFile); const REGISTRY_LOOKUP_TIMEOUT_MS = 10_000; +/** Linux `statfs` filesystem type of a tmpfs mount (`TMPFS_MAGIC`). */ +const TMPFS_MAGIC = 0x01021994; + +const MIN_TEMPORARY_DIRECTORY_FREE_BYTES = 2 * 1024 ** 3; + /** Local commands every run needs regardless of which agent backend is selected. */ const REQUIRED_TOOLCHAIN_COMMANDS = ["git", "node", "forge"] as const; @@ -64,9 +71,10 @@ export async function diagnoseProject(input: DoctorInput) { }); diagnostics.push(...validation.diagnostics); - const openRouterSelected = - resolved.config !== undefined && - activeTopologyAgentRefs(projectRoot, resolved.config, input.topologyPath).includes("OpenRouterAgent"); + // The agents the selected topology can dispatch to, including retry fallbacks. + const selectedAgentRefs = + resolved.config === undefined ? [] : activeTopologyAgentRefs(projectRoot, resolved.config, input.topologyPath); + const openRouterSelected = selectedAgentRefs.includes("OpenRouterAgent"); const openRouterCredential = openRouterSelected ? resolved.config?.agents.OpenRouterAgent?.apiKeyEnv : undefined; const openRouterCredentialReady = !openRouterSelected || (openRouterCredential !== undefined && (env[openRouterCredential] ?? "").trim() !== ""); @@ -96,18 +104,18 @@ export async function diagnoseProject(input: DoctorInput) { }); diagnostics.push(...references.diagnostics); - const agentRefs = configuredAgentRefs(resolved.config?.models.profiles); - const topologyCommands = validation.value?.topology?.required_commands ?? []; - const commandRequirements = [ - ...REQUIRED_TOOLCHAIN_COMMANDS.map((name) => ({ name, required: true })), - ...topologyCommands.map((name) => ({ name, required: true })), - ...agentRefs.map((agentRef) => ({ - name: AGENT_EXECUTABLES[agentRef] ?? agentRef, - // Only demand a CLI for agents whose executable Ultrafuzz actually - // knows; an unrecognised ref is reported without being required. - required: AGENT_EXECUTABLES[agentRef] !== undefined - })) - ].filter((entry, index, entries) => entries.findIndex((candidate) => candidate.name === entry.name) === index); + const requiredByName = new Map(); + for (const name of [...REQUIRED_TOOLCHAIN_COMMANDS, ...(validation.value?.topology?.required_commands ?? [])]) { + requiredByName.set(name, true); + } + for (const agentRef of configuredAgentRefs(resolved.config?.models.profiles)) { + // Every configured agent's CLI is reported, but only a selected agent's + // known CLI is required. Agents can share a CLI, so any requirer wins. + const name = AGENT_EXECUTABLES[agentRef] ?? agentRef; + const required = AGENT_EXECUTABLES[agentRef] !== undefined && selectedAgentRefs.includes(agentRef); + requiredByName.set(name, requiredByName.get(name) === true || required); + } + const commandRequirements = [...requiredByName].map(([name, required]) => ({ name, required })); let probeFailure: string | undefined; const probes = resolved.config === undefined @@ -157,27 +165,29 @@ export async function diagnoseProject(input: DoctorInput) { }); } - const engineCheck = workflowEngineCheck(installation); - const observedEngineStatus = engineCheck.check.status; - checks.push({ - ...engineCheck.check, - status: "unknown" as const, - summary: - "project-local workflow engine posture is informational and ignored; the pinned operator-owned controller is installed, patched, and sealed at launch" - }); - - const patchCheck = compatibilityPatchCheck(installation); - checks.push({ - ...patchCheck.check, - status: "unknown" as const, - summary: - "project-local compatibility-patch posture is informational and ignored; operator-owned controller patches are sealed at launch" - }); + checks.push( + { + name: "workflow-engine-install", + status: "unknown", + summary: + "project-local workflow engine posture is informational and ignored; the pinned operator-owned controller is installed, patched, and sealed at launch" + }, + { + name: "workflow-engine-patches", + status: "unknown", + summary: + "project-local compatibility-patch posture is informational and ignored; operator-owned controller patches are sealed at launch" + } + ); const latestCheck = registryCheck(latest); checks.push(latestCheck.check); diagnostics.push(...latestCheck.diagnostics); + const temporaryCheck = temporaryDirectoryCheck(os.tmpdir()); + checks.push(temporaryCheck.check); + diagnostics.push(...temporaryCheck.diagnostics); + const value: DoctorValue = { project_root: projectRoot, ok: checks.every((check) => check.status !== "error"), @@ -194,7 +204,10 @@ export async function diagnoseProject(input: DoctorInput) { installed_bin_target: installation.installed_bin_target, bin_path: installation.bin_path, latest_published_version: latest !== undefined && "version" in latest ? latest.version : "unknown", - layout_status: observedEngineStatus, + layout_status: + installation.installed_version === installation.required_version && installation.layout_error === null + ? "ok" + : "error", layout_detail: installation.layout_error, compatibility_patches: installation.compatibility_patches } @@ -202,136 +215,71 @@ export async function diagnoseProject(input: DoctorInput) { return runtimeResult(value.ok, value, diagnostics); } -function workflowEngineCheck(installation: SmithersInstallationPosture): { - check: DoctorCheck; - diagnostics: RuntimeDiagnostic[]; -} { - if (installation.installed_version === null) { - return { - check: { - name: "workflow-engine-install", - status: "error", - summary: `pinned workflow engine ${installation.required_version} is not installed for this project` - }, - diagnostics: [ - { - code: "DOCTOR_WORKFLOW_ENGINE_MISSING", - message: `pinned workflow engine ${installation.required_version} is not installed; it installs automatically on the next run`, - severity: "error", - source: "doctor" - } - ] - }; - } - if (installation.installed_version !== installation.required_version) { - return { - check: { - name: "workflow-engine-install", - status: "error", - summary: `installed workflow engine ${installation.installed_version} does not match the required ${installation.required_version}` - }, - diagnostics: [ - { - code: "DOCTOR_WORKFLOW_ENGINE_VERSION_MISMATCH", - message: `installed workflow engine is ${installation.installed_version} but ${installation.required_version} is required`, - severity: "error", - source: "doctor" - } - ] - }; - } - if (installation.layout_error !== null) { - return { - check: { - name: "workflow-engine-install", - status: "error", - summary: `installed workflow engine layout is not usable: ${installation.layout_error}` - }, - diagnostics: [ - { - code: "DOCTOR_WORKFLOW_ENGINE_LAYOUT_INVALID", - message: `installed workflow engine layout failed validation: ${installation.layout_error}`, - severity: "error", - source: "doctor" - } - ] - }; +/** + * Launch and resume install the workflow engine controller under the OS + * temporary directory, and a native resume keeps its install there for the + * detached engine. Warn when that directory is RAM-backed or nearly full. The + * controller roots are only reported: a live engine may still be using them. + */ +function temporaryDirectoryCheck(directory: string): { check: DoctorCheck; diagnostics: RuntimeDiagnostic[] } { + const warning = (message: string) => ({ + check: { name: "temporary-directory", status: "warning" as const, summary: message }, + diagnostics: [ + { code: "DOCTOR_TEMPORARY_DIRECTORY_CONSTRAINED", message, severity: "warning" as const, source: "doctor" } + ] + }); + let stats: fs.StatsFs; + let roots: string[]; + try { + stats = fs.statfsSync(directory); + roots = fs + .readdirSync(directory, { withFileTypes: true }) + .filter((entry) => entry.isDirectory() && entry.name.startsWith("ultrafuzz-controller-")) + .map((entry) => path.join(directory, entry.name)); + } catch (error) { + return warning( + `could not inspect the temporary directory ${directory}: ${error instanceof Error ? error.message : String(error)}` + ); } - return { - check: { - name: "workflow-engine-install", - status: "ok", - summary: `workflow engine ${installation.installed_version} installed and passing manifest and path validation` - }, - diagnostics: [] - }; + const free = stats.bavail * stats.bsize; + const rootBytes = roots.reduce((total, root) => total + regularFileBytes(root), 0); + const usage = `${formatBytes(free)} free; ${String(roots.length)} ultrafuzz-controller-* ${roots.length === 1 ? "directory holds" : "directories hold"} ${formatBytes(rootBytes)}`; + const problems = [ + ...(stats.type === TMPFS_MAGIC ? ["is a RAM-backed tmpfs"] : []), + ...(free < MIN_TEMPORARY_DIRECTORY_FREE_BYTES ? ["has less than 2 GiB free"] : []) + ]; + return problems.length === 0 + ? { check: { name: "temporary-directory", status: "ok", summary: `${directory}: ${usage}` }, diagnostics: [] } + : warning( + `temporary directory ${directory} ${problems.join(" and ")}; launch and resume install the workflow engine controller there (${usage}). Set TMPDIR to a disk-backed directory with more free space.` + ); } -function compatibilityPatchCheck(installation: SmithersInstallationPosture): { - check: DoctorCheck; - diagnostics: RuntimeDiagnostic[]; -} { - const entries = Object.entries(installation.compatibility_patches); - const incompatible = entries.filter(([, posture]) => posture === "incompatible").map(([name]) => name); - const missing = entries.filter(([, posture]) => posture === "missing").map(([name]) => name); - const unknown = entries.filter(([, posture]) => posture === "unknown").map(([name]) => name); - if (incompatible.length > 0) { - // The next run hard-fails in this state, so doctor must not call it healthy. - return { - check: { - name: "workflow-engine-patches", - status: "error", - summary: `installed engine source is modified or incompatible for: ${incompatible.join(", ")}` - }, - diagnostics: [ - { - code: "DOCTOR_WORKFLOW_ENGINE_PATCHES_INCOMPATIBLE", - message: `installed workflow engine source no longer matches the shape Ultrafuzz patches for: ${incompatible.join(", ")}; reinstall the pinned engine`, - severity: "error", - source: "doctor" - } - ] - }; - } - if (missing.length > 0) { - return { - check: { - name: "workflow-engine-patches", - status: "warning", - summary: `compatibility patches not applied yet: ${missing.join(", ")}; they apply on the next run` - }, - diagnostics: [ - { - code: "DOCTOR_WORKFLOW_ENGINE_PATCHES_PENDING", - message: `required workflow engine compatibility patches are not applied: ${missing.join(", ")}`, - severity: "warning", - source: "doctor" - } - ] - }; +/** Total size of the regular files under a directory, skipping entries that vanish or cannot be read. */ +function regularFileBytes(directory: string): number { + let entries: fs.Dirent[]; + try { + entries = fs.readdirSync(directory, { withFileTypes: true }); + } catch { + return 0; } - if (unknown.length > 0) { - return { - check: { - name: "workflow-engine-patches", - status: "unknown", - summary: `compatibility patch posture unavailable for: ${unknown.join(", ")}` - }, - diagnostics: [] - }; + let total = 0; + for (const entry of entries) { + const entryPath = path.join(directory, entry.name); + if (entry.isDirectory()) total += regularFileBytes(entryPath); + else if (entry.isFile()) { + try { + total += fs.lstatSync(entryPath).size; + } catch { + // Removed while scanning. + } + } } - const upstream = entries.filter(([, posture]) => posture === "upstream").map(([name]) => name); - return { - check: { - name: "workflow-engine-patches", - status: "ok", - summary: - upstream.length === entries.length - ? "the installed engine already provides every patched behavior upstream" - : `compatibility patches applied${upstream.length > 0 ? `; upstream now covers ${upstream.join(", ")}` : ""}` - }, - diagnostics: [] - }; + return total; +} + +function formatBytes(bytes: number): string { + return bytes >= 1024 ** 3 ? `${(bytes / 1024 ** 3).toFixed(1)} GiB` : `${String(Math.round(bytes / 1024 ** 2))} MiB`; } function registryCheck(latest: LatestPublishedEngine | { error: string } | undefined): { diff --git a/packages/runtime/test/lifecycle-inspection.test.ts b/packages/runtime/test/lifecycle-inspection.test.ts index d56320c5e..c560aa504 100644 --- a/packages/runtime/test/lifecycle-inspection.test.ts +++ b/packages/runtime/test/lifecycle-inspection.test.ts @@ -1332,6 +1332,174 @@ test("diagnoseProject reports commands required by the active topology", async ( ); }); +test("diagnoseProject requires only the CLIs of agents the selected topology can dispatch to", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + // The scaffold also configures Claude, DeepSeek, Kimi, and Pi profiles; only Codex is installed. + const installed = new Set(["git", "node", "forge", "codex"]); + const diagnose = () => + diagnoseProject({ + projectRoot: project, + env: { PATH: "/usr/bin" }, + offline: true, + requiredCommandProbe: async (names) => + names.map((name) => ({ + name, + available: installed.has(name), + path: installed.has(name) ? `/usr/bin/${name}` : null, + version: null + })) + }); + + const codexOnly = await diagnose(); + assert.equal(codexOnly.value?.checks.find((check) => check.name === "toolchain")?.status, "ok"); + assert.deepEqual( + codexOnly.value?.toolchain.map((entry) => [entry.name, entry.required]), + [ + ["git", true], + ["node", true], + ["forge", true], + ["claude", false], + ["codex", true], + ["kimi", false], + ["pi", false] + ] + ); + + // A retry fallback is dispatched to as well. + const configPath = path.join(project, "ultrafuzz.toml"); + const config = fs.readFileSync(configPath, "utf8"); + fs.writeFileSync(configPath, `${config}\n[retry]\nagents = ["default", "kimi"]\n`, "utf8"); + const withFallback = await diagnose(); + assert.equal(withFallback.value?.toolchain.find((entry) => entry.name === "kimi")?.required, true); + assert.ok( + withFallback.diagnostics.some( + (entry) => entry.code === "DOCTOR_TOOLCHAIN_MISSING" && entry.message.endsWith(": kimi") + ) + ); + + // Codex also runs OpenRouterAgent, so an unselected Codex profile cannot waive that requirement. + fs.writeFileSync(configPath, config, "utf8"); + addOpenRouterProfile(project); + const topologyPath = path.join(project, ".ultrafuzz", "topology.yml"); + fs.writeFileSync( + topologyPath, + fs + .readFileSync(topologyPath, "utf8") + .replace( + " prompt: setup/project-discovery.md\n", + " prompt: setup/project-discovery.md\n model_profiles:\n - openrouter\n" + ), + "utf8" + ); + const openRouterOnly = await diagnose(); + assert.equal(openRouterOnly.value?.toolchain.find((entry) => entry.name === "codex")?.required, true); +}); + +test("diagnoseProject reports controller roots in the temporary directory and leaves them in place", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const temporary = temporaryRoot("ufz-doctor-tmpdir-"); + for (const [name, mebibytes] of [ + ["ultrafuzz-controller-first", 1], + ["ultrafuzz-controller-second", 2], + ["unrelated", 4] + ] as const) { + fs.mkdirSync(path.join(temporary, name, ".smithers"), { recursive: true }); + fs.writeFileSync(path.join(temporary, name, ".smithers", "engine.js"), Buffer.alloc(mebibytes * 1024 * 1024)); + } + + const doctor = await withTemporaryDirectory(temporary, () => + diagnoseProject({ + projectRoot: project, + env: { PATH: "/usr/bin" }, + offline: true, + requiredCommandProbe: allAvailable + }) + ); + + assert.match( + doctor.value?.checks.find((check) => check.name === "temporary-directory")?.summary ?? "", + /; 2 ultrafuzz-controller-\* directories hold 3 MiB/u + ); + assert.ok(fs.existsSync(path.join(temporary, "ultrafuzz-controller-first", ".smithers", "engine.js"))); + assert.ok(fs.existsSync(path.join(temporary, "ultrafuzz-controller-second", ".smithers", "engine.js"))); +}); + +const devShmIsTmpfs = (() => { + try { + return fs.statfsSync("/dev/shm").type === 0x01021994; + } catch { + return false; + } +})(); + +test( + "diagnoseProject warns, without failing, when the temporary directory is RAM-backed", + { skip: devShmIsTmpfs ? false : "/dev/shm is not a tmpfs mount on this host" }, + async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const temporary = registerTemporaryPath(fs.mkdtempSync("/dev/shm/ufz-doctor-")); + + const doctor = await withTemporaryDirectory(temporary, () => + diagnoseProject({ + projectRoot: project, + env: { PATH: "/usr/bin" }, + offline: true, + requiredCommandProbe: allAvailable + }) + ); + + assert.equal(doctor.value?.checks.find((check) => check.name === "temporary-directory")?.status, "warning"); + const warning = doctor.diagnostics.find((entry) => entry.code === "DOCTOR_TEMPORARY_DIRECTORY_CONSTRAINED"); + assert.equal(warning?.severity, "warning"); + assert.match(warning?.message ?? "", /is a RAM-backed tmpfs/u); + } +); + +test("diagnoseProject warns when the temporary directory has little free space", async (context) => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const temporary = temporaryRoot("ufz-doctor-full-"); + const statfsSync = fs.statfsSync; + context.mock.method(fs, "statfsSync", (target: fs.PathLike) => ({ ...statfsSync(target), bavail: 1 })); + + const doctor = await withTemporaryDirectory(temporary, () => + diagnoseProject({ + projectRoot: project, + env: { PATH: "/usr/bin" }, + offline: true, + requiredCommandProbe: allAvailable + }) + ); + + assert.equal(doctor.value?.checks.find((check) => check.name === "temporary-directory")?.status, "warning"); + assert.match( + doctor.diagnostics.find((entry) => entry.code === "DOCTOR_TEMPORARY_DIRECTORY_CONSTRAINED")?.message ?? "", + /has less than 2 GiB free/u + ); +}); + +async function allAvailable(names: readonly string[]) { + return names.map((name) => ({ name, available: true, path: `/usr/bin/${name}`, version: null })); +} + +async function withTemporaryDirectory(directory: string, operation: () => Promise): Promise { + const previous = process.env.TMPDIR; + process.env.TMPDIR = directory; + try { + return await operation(); + } finally { + if (previous === undefined) delete process.env.TMPDIR; + else process.env.TMPDIR = previous; + } +} + test("diagnoseProject rejects cwd-dependent PATH entries that are unavailable in task worktrees", async () => { for (const searchPath of ["bin", ""]) { const project = tempProject(); @@ -1600,7 +1768,6 @@ test("diagnoseProject reports a posture for every tracked compatibility patch", assert.ok(Object.hasOwn(reported, "resume_hydration")); // Target-owned dependencies never become controller authority. assert.equal(doctor.value?.checks.find((check) => check.name === "workflow-engine-patches")?.status, "unknown"); - assert.ok(!doctor.diagnostics.some((entry) => entry.code === "DOCTOR_WORKFLOW_ENGINE_PATCHES_INCOMPATIBLE")); }); test("diagnoseProject reports a missing install and a version mismatch", async () => { @@ -1608,17 +1775,18 @@ test("diagnoseProject reports a missing install and a version mismatch", async ( const missing = await diagnoseProject({ projectRoot: project, env, offline: true }); assert.equal(missing.value?.workflow_engine.installed_version, null); - assert.ok(!missing.diagnostics.some((entry) => entry.code === "DOCTOR_WORKFLOW_ENGINE_MISSING")); + assert.equal(missing.value?.workflow_engine.layout_status, "error"); writeFakeInstalledEngine(project, { version: "0.29.0" }); const mismatched = await diagnoseProject({ projectRoot: project, env, offline: true }); assert.equal(mismatched.value?.workflow_engine.installed_version, "0.29.0"); - assert.ok(!mismatched.diagnostics.some((entry) => entry.code === "DOCTOR_WORKFLOW_ENGINE_VERSION_MISMATCH")); + assert.equal(mismatched.value?.workflow_engine.layout_status, "error"); writeFakeInstalledEngine(project, { version: SMITHERS_VERSION, binTarget: "dist/other.js" }); const doctor = await diagnoseProject({ projectRoot: project, env, offline: true }); assert.equal(doctor.value?.workflow_engine.installed_bin_target, "dist/other.js"); assert.equal(doctor.value?.workflow_engine.layout_status, "error"); - assert.ok(!doctor.diagnostics.some((entry) => entry.code === "DOCTOR_WORKFLOW_ENGINE_LAYOUT_INVALID")); + // The project-local engine is informational: launch installs its own controller. + assert.equal(doctor.value?.checks.find((check) => check.name === "workflow-engine-install")?.status, "unknown"); }); test("diagnoseProject keeps an offline registry lookup non-fatal", async () => { From 332bdab50b834ca899340f52878e92806475c67a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:30:06 +0000 Subject: [PATCH 023/206] fix(runtime): keep refreshed controllers out of the governed target identity `resume --refresh-controller` renders its current controller into /.smithers/continuations//, but that directory was missing from controllerOwnedGovernancePaths. Unless the project ignores .smithers, the rendered files are untracked target files: the next launch in that project records the target as dirty, which a private campaign (the default policy) and every cloud launch reject, and the changed worktree digest invalidates disclosure acknowledgements computed for the target. The directory is now controller-owned, like .smithers/workflows. The new test renders files in that layout and checks the target identity is unchanged; origin/main reports the target dirty. Refs #921 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/data-governance.ts | 2 ++ packages/runtime/test/data-governance.test.ts | 15 +++++++++++++++ 2 files changed, 17 insertions(+) diff --git a/packages/runtime/src/data-governance.ts b/packages/runtime/src/data-governance.ts index 72f6f1d2a..888959361 100644 --- a/packages/runtime/src/data-governance.ts +++ b/packages/runtime/src/data-governance.ts @@ -465,6 +465,8 @@ export function controllerOwnedGovernancePaths(projectRoot: string, runRoot: str path.join(projectRoot, ".ultrafuzz", "runs"), path.join(projectRoot, ".smithers", "node_modules"), path.join(projectRoot, ".smithers", "workflows"), + // `resume --refresh-controller` renders each refreshed controller here. + path.join(projectRoot, ".smithers", "continuations"), // The workflow engine opens its SQLite database in the target root, so a // launched run leaves engine state in the governed worktree. path.join(projectRoot, "smithers.db"), diff --git a/packages/runtime/test/data-governance.test.ts b/packages/runtime/test/data-governance.test.ts index 63213c6f0..f74e8c563 100644 --- a/packages/runtime/test/data-governance.test.ts +++ b/packages/runtime/test/data-governance.test.ts @@ -7,6 +7,7 @@ import test from "node:test"; import { sha256Bytes, type RunDataGovernanceReference } from "@ultrafuzz/artifacts"; import type { ResolvedConfig } from "@ultrafuzz/config"; import { + controllerOwnedGovernancePaths, DATA_DISCLOSURE_ACKNOWLEDGEMENTS_ENV, DATA_DISCLOSURE_ACKNOWLEDGEMENT_SCHEMA_VERSION, DATA_GOVERNANCE_POLICY_ENV, @@ -437,6 +438,20 @@ test("target identity excludes only exact controller-owned untracked outputs", ( fs.writeFileSync(path.join(root, "unrelated"), "target\n"); assert.equal(targetIdentity(root, owned).dirty, true); }); +test("a refreshed controller rendered into the target leaves its governed identity unchanged", () => { + const root = repository(), + owned = controllerOwnedGovernancePaths(root, path.join(root, ".ultrafuzz", "runs", "refresh-run")), + initial = targetIdentity(root, owned); + // The layout `resume --refresh-controller` renders into (renderCurrentSmithersController). + const generation = path.join(root, ".smithers", "continuations", "9d4c1f3e-2b7a-4c55-8e1d-6f0a3b2c4d5e"); + for (const file of ["agents/codex.ts", "workflows/ultrafuzz-refresh-run.tsx"]) { + fs.mkdirSync(path.dirname(path.join(generation, file)), { recursive: true }); + fs.writeFileSync(path.join(generation, file), "controller\n"); + } + + assert.equal(initial.dirty, false); + assert.deepEqual(targetIdentity(root, owned), initial); +}); test("governance precedes preflight and rejects a policy mutated during it", async () => { const project = temporaryRoot("ufz-governance-plan-"); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); From e7a5e500c8de8975a69489c9d6c5441659f29496 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:30:30 +0000 Subject: [PATCH 024/206] fix(runtime): parse the control seal with an item limit that covers its schema Runtime documents are parsed with a 100,000-item limit that the strict JSON parser counts across the whole document. The control seal's schema admits 100,000 execution files plus three identity arrays of up to 100,000 entries each, and launch writes the seal after schema validation alone. A closure near the execution-file bound plus the run's task bindings would therefore produce a seal that every later command fails to parse. (An earlier measurement put the production closure at about 51,000 files; it was not re-measured for this change.) The seal is now parsed with a 400,000-item limit, the schema's total; its property total (four per execution file plus 35 fixed) already fits the default limit. Byte limits are unchanged. The new test parses a seal with the schema's maximum execution files and the fixture's bindings; origin/main rejects it with "JSON exceeds the item limit of 100000". Refs #921 Co-Authored-By: Claude Opus 5.5 --- .../runtime/src/runtime-document-codec.ts | 32 +++++++++++++------ .../test/runtime-document-contracts.test.ts | 22 +++++++++++++ 2 files changed, 45 insertions(+), 9 deletions(-) diff --git a/packages/runtime/src/runtime-document-codec.ts b/packages/runtime/src/runtime-document-codec.ts index 29e116c2c..fb81d01e5 100644 --- a/packages/runtime/src/runtime-document-codec.ts +++ b/packages/runtime/src/runtime-document-codec.ts @@ -1,10 +1,26 @@ -import { parseStrictJsonBytes, writeJsonDurable } from "@ultrafuzz/artifacts"; +import { parseStrictJsonBytes, writeJsonDurable, type StrictJsonLimits } from "@ultrafuzz/artifacts"; -import { type RuntimeDocumentForSchemaId, type RuntimeDocumentSchemaId } from "./runtime-contracts.js"; +import { + WORKFLOW_CONTROL_INTEGRITY_JSON_SCHEMA_ID, + type RuntimeDocumentForSchemaId, + type RuntimeDocumentSchemaId +} from "./runtime-contracts.js"; import { assertRuntimeJsonSchema } from "./schema-registry.js"; import { assertRuntimeDocumentSemantics } from "./runtime-semantic-gates.js"; -const MAX_RUNTIME_DOCUMENT_BYTES = 128 * 1024 * 1024; +const RUNTIME_DOCUMENT_LIMITS: StrictJsonLimits = { + maxBytes: 128 * 1024 * 1024, + maxDepth: 64, + maxItems: 100_000, + maxProperties: 500_000 +}; + +// The parser counts items across the whole document. The control seal's schema +// admits 100_000 execution files plus three identity arrays of up to 100_000 +// entries each, so a seal whose every array is within its schema bound can hold +// 400_000 items. Its property total (four per execution file plus 35 fixed) stays +// under the default limit. +const CONTROL_SEAL_LIMITS: StrictJsonLimits = { ...RUNTIME_DOCUMENT_LIMITS, maxItems: 400_000 }; export function assertRuntimeDocument( schemaId: SchemaId, @@ -22,12 +38,10 @@ export function parseRuntimeDocumentBytes { - const value = parseStrictJsonBytes(bytes, { - maxBytes: MAX_RUNTIME_DOCUMENT_BYTES, - maxDepth: 64, - maxItems: 100_000, - maxProperties: 500_000 - }); + const value = parseStrictJsonBytes( + bytes, + schemaId === WORKFLOW_CONTROL_INTEGRITY_JSON_SCHEMA_ID ? CONTROL_SEAL_LIMITS : RUNTIME_DOCUMENT_LIMITS + ); return assertRuntimeDocument(schemaId, value, label); } diff --git a/packages/runtime/test/runtime-document-contracts.test.ts b/packages/runtime/test/runtime-document-contracts.test.ts index fd4142dc4..138c1371f 100644 --- a/packages/runtime/test/runtime-document-contracts.test.ts +++ b/packages/runtime/test/runtime-document-contracts.test.ts @@ -230,6 +230,28 @@ test("workflow control execution-file uniqueness stays in the linear semantic ga assert.equal(assertRuntimeDocument(WORKFLOW_CONTROL_INTEGRITY_JSON_SCHEMA_ID, seal, "seal"), seal); }); +// Launch writes the seal after schema validation alone, so every seal the schema +// accepts has to parse back, bindings included. +test("a control seal at its schema's execution-file bound parses back", () => { + const properties = workflowControlIntegrityJsonSchema.properties as Record; + const maxExecutionFiles = properties.execution_files?.maxItems; + assert.ok(typeof maxExecutionFiles === "number"); + const seal = fixture("workflow-control-integrity.schema.json"); + const executionFile = (seal.execution_files as Array>)[0]; + assert.ok(executionFile); + seal.execution_files = Array.from({ length: maxExecutionFiles }, (_, index) => { + const identity = index.toString().padStart(6, "0"); + return { + ...executionFile, + source_path: `/generic/source/${identity}`, + snapshot_path: `generic/${identity}.json` + }; + }); + const bytes = Buffer.from(JSON.stringify(seal), "utf8"); + + assert.deepEqual(parseRuntimeDocumentBytes(WORKFLOW_CONTROL_INTEGRITY_JSON_SCHEMA_ID, bytes, "seal"), seal); +}); + test("strict runtime document parsing rejects duplicate keys before schema validation", () => { assert.throws( () => From 3d4a26e00d510c4c43c508b6a1d4cf050c297343 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:30:31 +0000 Subject: [PATCH 025/206] docs: changelog entry for launch I/O and doctor fixes Refs #921 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..7419964e2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [cli] [docs]** Execution-snapshot publication now flushes each file once instead of twice: files are created with their final mode and permission sealing only touches directories. `doctor` requires only the executables of agents the selected topology or its retry chain can dispatch to (other configured profiles are listed as not required), drops engine diagnostics it built and never reported, and warns when the temporary directory is a tmpfs or has under 2 GiB free, reporting how much its `ultrafuzz-controller-*` directories hold. A controller refresh no longer dirties the governed target through `.smithers/continuations`, and the control seal is parsed with an item limit that covers everything its schema admits (#921). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 87413054e92702c7ca02913cba3ba4685499e854 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:47:21 +0000 Subject: [PATCH 026/206] test(runtime): stop pinning adapter template bytes and sizes agent-adapter-boundaries.test.ts pinned every adapter template's SHA-256 twice (a structural policy and a separate responsibility policy) plus line and syntax-node ceilings. Those assertions check no behaviour: a comment edit or the one-line missing `import path` fix in deepseek.tsx failed the gate with "changed from its reviewed source fingerprint", so every adapter fix needed two hash edits. Merge the two tables into one policy per source (purpose, declared responsibilities, upstream links) and delete the fingerprints, the ceilings, and the tests that only exercised them. The real boundary checks stay: every source needs a policy, only registered adapters may own responsibilities, non-adapter helpers may carry no orchestration signals, and a statically detected responsibility still fails until it is declared. Co-Authored-By: Claude Opus 5.5 --- docs/reference/agent-adapter-boundaries.md | 60 ++- .../test/agent-adapter-boundaries.test.ts | 345 ++---------------- 2 files changed, 60 insertions(+), 345 deletions(-) diff --git a/docs/reference/agent-adapter-boundaries.md b/docs/reference/agent-adapter-boundaries.md index 179ae3fcf..902d4f73d 100644 --- a/docs/reference/agent-adapter-boundaries.md +++ b/docs/reference/agent-adapter-boundaries.md @@ -16,22 +16,20 @@ and the gate requires an explicit policy for every adapter actually registered in `agentFactories`, including a factory imported under an alias or registered without a matching re-export. -| Adapter | Baseline | Classification | Existing option or missing surface | Upstream dependency | -| ---------------- | -------: | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | -| `claude.tsx` | 86 | Thin mapping. Its override applies Ultrafuzz's child-environment policy but does not rebuild an orchestrator responsibility. | Uses `model`, `extraArgs`, `addDir`, `permissionMode`, `settingSources`, `apiKey`, `configDir`, and `env`. | None. | -| `codex.tsx` | 162 | Partly avoidable. The local argv rewrite and resume awareness exist because a working constructor option is serialized incorrectly upstream. Provider-home inspection and child-environment filtering are local policy. | `addDir` exists, but multiple values become one `--add-dir` occurrence. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622) | -| `deepseek.tsx` | 341 | Inherent after removing its avoidable reasoning-effort argv mapping. Route, auth, and effort are now thin mappings; result parsing and token normalization have no typed upstream surface. | Uses `model`, first-class `effort`, `addDir`, `permissionMode`, `settingSources`, `env`, and `configDir`; a custom-provider usage normalizer is missing. | [smithers#1624](https://github.com/smithersai/smithers/issues/1624) | -| `kimi.tsx` | 1,520 | Inherent with the current dependency except for Ultrafuzz-specific bounded-I/O and credential-governance checks. Usage discovery, actual-session recovery, argv compatibility, and runtime-home isolation cannot be expressed by constructor options. | `model`, `extraArgs`, `env`, `configDir`, and `session` exist; invocation-local usage, actual-session resolution, separate credential/runtime homes, and a Kimi Code 0.29.x command dialect are missing. | [smithers#1623](https://github.com/smithersai/smithers/issues/1623), [smithers#1626](https://github.com/smithersai/smithers/issues/1626) | -| `openrouter.tsx` | 1,234 | Inherent with the current dependency except for local credential/config materialization. Provider-output quarantine and exact-session retry are orchestration responsibilities; the inherited Codex argv workaround is separately avoidable. | `config`, `configDir`, `env`, `model`, and `addDir` cover the route; a bounded provider-recovery policy is missing. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622), [smithers#1625](https://github.com/smithersai/smithers/issues/1625) | +| Adapter | Classification | Existing option or missing surface | Upstream dependency | +| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | +| `claude.tsx` | Thin mapping. Its override applies Ultrafuzz's child-environment policy but does not rebuild an orchestrator responsibility. | Uses `model`, `extraArgs`, `addDir`, `permissionMode`, `settingSources`, `apiKey`, `configDir`, and `env`. | None. | +| `codex.tsx` | Partly avoidable. The local argv rewrite and resume awareness exist because a working constructor option is serialized incorrectly upstream. Provider-home inspection and child-environment filtering are local policy. | `addDir` exists, but multiple values become one `--add-dir` occurrence. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622) | +| `deepseek.tsx` | Inherent after removing its avoidable reasoning-effort argv mapping. Route, auth, and effort are now thin mappings; result parsing and token normalization have no typed upstream surface. | Uses `model`, first-class `effort`, `addDir`, `permissionMode`, `settingSources`, `env`, and `configDir`; a custom-provider usage normalizer is missing. | [smithers#1624](https://github.com/smithersai/smithers/issues/1624) | +| `kimi.tsx` | Inherent with the current dependency except for Ultrafuzz-specific bounded-I/O and credential-governance checks. Usage discovery, actual-session recovery, argv compatibility, and runtime-home isolation cannot be expressed by constructor options. | `model`, `extraArgs`, `env`, `configDir`, and `session` exist; invocation-local usage, actual-session resolution, separate credential/runtime homes, and a Kimi Code 0.29.x command dialect are missing. | [smithers#1623](https://github.com/smithersai/smithers/issues/1623), [smithers#1626](https://github.com/smithersai/smithers/issues/1626) | +| `openrouter.tsx` | Inherent with the current dependency except for local credential/config materialization. Provider-output quarantine and exact-session retry are orchestration responsibilities; the inherited Codex argv workaround is separately avoidable. | `config`, `configDir`, `env`, `model`, and `addDir` cover the route; a bounded provider-recovery policy is missing. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622), [smithers#1625](https://github.com/smithersai/smithers/issues/1625) | -The line ceilings deliberately allow only a small formatting margin: Claude -100, Codex 175, DeepSeek 350, Kimi 1,525, and OpenRouter 1,250 lines. The listed -responsibilities are explicit review declarations, not conclusions inferred -from comments or prose. The gate supplements them with a conservative AST lower -bound over executable source: static or dynamic filesystem-walking imports, -command `args` access and CLI-looking argument arrays, output-interpreter hooks -and JSON parsing of line-oriented output, common session-continuation fields, -and common token-usage fields. Every detected category must be declared; +The listed responsibilities are explicit review declarations, not conclusions +inferred from comments or prose. The gate supplements them with a conservative +AST lower bound over executable source: static or dynamic filesystem-walking +imports, command `args` access and CLI-looking argument arrays, +output-interpreter hooks and JSON parsing of line-oriented output, common +session-continuation fields, and common token-usage fields. Every detected category must be declared; declarations may include additional inherited or semantically reviewed responsibilities that the detector cannot infer. Type-only declarations, comments, strings outside argument arrays, and simple values or literal option @@ -39,27 +37,21 @@ arrays passed by a `create*Agent` factory into its returned `*Agent` constructor, named config-file reads, and empty diagnostic argv are deliberately excluded. -Every `.ts` and `.tsx` source under the adapter tree has two independent review -records. Its structural policy records a reviewed purpose, TypeScript -syntax-node ceiling, and exact SHA-256 source fingerprint. A separate central -responsibility policy repeats the exact fingerprint beside the declared -responsibilities and upstream links. Any byte change therefore invalidates both -records: updating the ordinary source fingerprint and ceiling cannot reuse a -stale responsibility review, even when the AST lower bound does not recognize -the new semantic form. Refreshing the second literal attests that the central -classification was reviewed; leaving the responsibility set unchanged is an -explicit reviewed-unchanged decision. The syntax count uses the -workspace-pinned TypeScript parser. This is an auditable two-stage source -freeze, not a claim that CI can infer every semantic behavior. +Every `.ts` and `.tsx` source under the adapter tree has one policy record: its +reviewed purpose, its declared responsibilities, and the upstream links that +justify them. The gate does not pin source bytes, line counts, or syntax-node +counts, so a comment edit or a one-line fix needs no policy change; a change +that adds a detected responsibility still fails until the policy declares it. +This is a conservative static lower bound, not a claim that CI can infer every +semantic behavior. The inventory walk is recursive. A new helper, either `.ts` or `.tsx`, fails -until it receives both an explicit structural policy and an independent -responsibility review. Non-adapter helpers must attest an empty responsibility -set and may not contain any of the lower-bound orchestration signals, so moving -such code out of a registered adapter cannot bypass either review layer. A -future shared inherent workaround must add explicit adapter ownership rather -than weakening that default. Any policy update must classify a changed adapter -responsibility in the same reviewed pull request. +until it receives an explicit policy. Non-adapter helpers must attest an empty +responsibility set and may not contain any of the lower-bound orchestration +signals, so moving such code out of a registered adapter cannot bypass the +review. A future shared inherent workaround must add explicit adapter ownership +rather than weakening that default. Any policy update must classify a changed +adapter responsibility in the same reviewed pull request. The gate is `packages/runtime/test/agent-adapter-boundaries.test.ts`. It scans every TypeScript source file in the adapter directory, derives shipped adapters diff --git a/packages/runtime/test/agent-adapter-boundaries.test.ts b/packages/runtime/test/agent-adapter-boundaries.test.ts index c8899c0f3..daf8c2d7a 100644 --- a/packages/runtime/test/agent-adapter-boundaries.test.ts +++ b/packages/runtime/test/agent-adapter-boundaries.test.ts @@ -1,6 +1,5 @@ import assert from "node:assert/strict"; import { temporaryRoot } from "./temporary-root.js"; -import crypto from "node:crypto"; import { existsSync, mkdirSync, readFileSync, readdirSync, rmSync, writeFileSync } from "node:fs"; import path from "node:path"; import test from "node:test"; @@ -10,19 +9,12 @@ import * as ts from "typescript"; type OrchestratorResponsibility = "argv-construction" | "filesystem-walking" | "output-interpretation" | "session-handling" | "token-accounting"; -type ResponsibilityPolicy = { - classifiedSourceSha256: string; +type AdapterPolicy = { + purpose: "adapter" | "data-governance" | "provider-home" | "registry" | "strict-input" | "toml"; responsibilities: readonly OrchestratorResponsibility[]; upstreamIssues: readonly string[]; }; -type SourcePolicy = { - maxLines: number; - maxSyntaxNodes: number; - purpose: "adapter" | "data-governance" | "provider-home" | "registry" | "strict-input" | "toml"; - sourceSha256: string; -}; - type SourceUnit = { ast: ts.SourceFile; source: string; @@ -82,38 +74,25 @@ const TOKEN_ACCOUNTING_SIGNALS = new Set([ "total_tokens" ]); -// This deliberately duplicates each exact source fingerprint in a separate, -// classification-owned policy. A source edit must update both the structural -// policy below and this responsibility review, even when the reviewer decides -// that the declared responsibility set remains unchanged. -const responsibilityPolicies: Record = { - "claude.tsx": { - classifiedSourceSha256: "6d8346e882f9ed1e5274d19d6e7cacf27e2743e489a3b931b9ee0488d62ffe76", - responsibilities: [], - upstreamIssues: [] - }, +// Every .ts/.tsx source under the adapter tree declares what it is for and +// which orchestrator responsibilities it owns. Only registered adapters may own +// responsibilities, and each one must link the upstream gap that forces it. +const adapterPolicies: Record = { + "claude.tsx": { purpose: "adapter", responsibilities: [], upstreamIssues: [] }, "codex.tsx": { - classifiedSourceSha256: "a44eb5c47e6374457476a86420eca0c5ff23616fc8201637f0139d37e13b92f5", + purpose: "adapter", responsibilities: ["argv-construction", "session-handling"], upstreamIssues: ["https://github.com/smithersai/smithers/issues/1622"] }, "deepseek.tsx": { - classifiedSourceSha256: "ea7c6eec70883126e3ee9588b6d3349169652688756b519f1d49c3b5b794e887", + purpose: "adapter", responsibilities: ["output-interpretation", "token-accounting"], upstreamIssues: ["https://github.com/smithersai/smithers/issues/1624"] }, - "environment.tsx": { - classifiedSourceSha256: "72d93b360e5b969e00646cfa39a24cac8a3720ce2aef1c7469c51038ef0206cd", - responsibilities: [], - upstreamIssues: [] - }, - "index.tsx": { - classifiedSourceSha256: "ce5f94b3bf12ae40c5b59ebd587a77d1e80e532d92c785f3353d272e627d79e4", - responsibilities: [], - upstreamIssues: [] - }, + "environment.tsx": { purpose: "data-governance", responsibilities: [], upstreamIssues: [] }, + "index.tsx": { purpose: "registry", responsibilities: [], upstreamIssues: [] }, "kimi.tsx": { - classifiedSourceSha256: "1551f53080570f026aa8f6f60d533ef51abe8bac453eb0610e0ef11a9f974530", + purpose: "adapter", responsibilities: [ "argv-construction", "filesystem-walking", @@ -126,13 +105,9 @@ const responsibilityPolicies: Record = { "https://github.com/smithersai/smithers/issues/1626" ] }, - "opencode.tsx": { - classifiedSourceSha256: "7d4e22b674e06b00cd537c531c0e7e95d800ffc07b50d02b566a5fd029d0b489", - responsibilities: [], - upstreamIssues: [] - }, + "opencode.tsx": { purpose: "adapter", responsibilities: [], upstreamIssues: [] }, "openrouter.tsx": { - classifiedSourceSha256: "834b8d1893f1a6b7da9667cd5f55ecd7fd5e7dc9179ab2b307562a9913f410b4", + purpose: "adapter", responsibilities: ["argv-construction", "output-interpretation", "session-handling", "token-accounting"], upstreamIssues: [ "https://github.com/monad-developers/ultrafuzz/issues/1006", @@ -142,7 +117,7 @@ const responsibilityPolicies: Record = { ] }, "pi.tsx": { - classifiedSourceSha256: "1b9e81f7d79a7c7f551c1b6d7d1f794e2d49ee698b328eee4a872b914bfc8acc", + purpose: "adapter", responsibilities: ["argv-construction", "output-interpretation", "session-handling", "token-accounting"], upstreamIssues: [ "https://github.com/monad-developers/ultrafuzz/issues/1006", @@ -151,156 +126,11 @@ const responsibilityPolicies: Record = { "https://github.com/smithersai/smithers/issues/1629" ] }, - "provider-home.tsx": { - classifiedSourceSha256: "31085a2bad1d6d82b3709946464df332fe1d22e13236708dfb840c8fbd7d5744", - responsibilities: [], - upstreamIssues: [] - }, - "strict-json.tsx": { - classifiedSourceSha256: "16c909eb1f01c82e1174db61877a30028b58a49466714865f1243293f10b186b", - responsibilities: [], - upstreamIssues: [] - }, - "toml.tsx": { - classifiedSourceSha256: "51b15d0f75a09b49a53a33709cf9127c74b2a35814ad8770ce529f0638a4f6ae", - responsibilities: [], - upstreamIssues: [] - } + "provider-home.tsx": { purpose: "provider-home", responsibilities: [], upstreamIssues: [] }, + "strict-json.tsx": { purpose: "strict-input", responsibilities: [], upstreamIssues: [] }, + "toml.tsx": { purpose: "toml", responsibilities: [], upstreamIssues: [] } }; -// Syntax-node ceilings and exact source fingerprints are the reviewed PR shape -// rooted at main@fe0922ea. The fingerprint makes every replacement visible -// even when it preserves or reduces aggregate structure; line ceilings retain -// a small formatting/documentation margin. -const sourcePolicies: Record = { - "claude.tsx": { - maxLines: 100, - maxSyntaxNodes: 452, - purpose: "adapter", - sourceSha256: "6d8346e882f9ed1e5274d19d6e7cacf27e2743e489a3b931b9ee0488d62ffe76" - }, - "codex.tsx": { - maxLines: 250, - maxSyntaxNodes: 1_300, - purpose: "adapter", - sourceSha256: "a44eb5c47e6374457476a86420eca0c5ff23616fc8201637f0139d37e13b92f5" - }, - "deepseek.tsx": { - // Raised with the 0.35.0 pin bump: the pinned BaseCliAgent now rejects - // unknown constructor options, so the adapter carries a thin constructor - // that splits the Ultrafuzz-only credential off `this.opts`. - maxLines: 360, - maxSyntaxNodes: 1_675, - purpose: "adapter", - sourceSha256: "ea7c6eec70883126e3ee9588b6d3349169652688756b519f1d49c3b5b794e887" - }, - "environment.tsx": { - // Shared native-continuation PATH and Pi home filtering (#1035). - maxLines: 425, - maxSyntaxNodes: 2_325, - purpose: "data-governance", - sourceSha256: "72d93b360e5b969e00646cfa39a24cac8a3720ce2aef1c7469c51038ef0206cd" - }, - "index.tsx": { - maxLines: 30, - maxSyntaxNodes: 125, - purpose: "registry", - sourceSha256: "ce5f94b3bf12ae40c5b59ebd587a77d1e80e532d92c785f3353d272e627d79e4" - }, - "kimi.tsx": { - // Raised with the 0.35.0 pin bump, for the same reason as deepseek.tsx. - maxLines: 1_600, - maxSyntaxNodes: 9_250, - purpose: "adapter", - sourceSha256: "1551f53080570f026aa8f6f60d533ef51abe8bac453eb0610e0ef11a9f974530" - }, - "opencode.tsx": { - maxLines: 150, - maxSyntaxNodes: 650, - purpose: "adapter", - sourceSha256: "7d4e22b674e06b00cd537c531c0e7e95d800ffc07b50d02b566a5fd029d0b489" - }, - "openrouter.tsx": { - // Raised after the accounting review for cumulative response aggregation, - // retry/failure usage, and adapter-recorded cost preservation (#1006). - maxLines: 1_750, - maxSyntaxNodes: 9_075, - purpose: "adapter", - sourceSha256: "834b8d1893f1a6b7da9667cd5f55ecd7fd5e7dc9179ab2b307562a9913f410b4" - }, - "pi.tsx": { - // Raised for cumulative per-response usage, session-aware progress, and - // adapter-recorded cost preservation in addition to terminal mapping. - maxLines: 475, - maxSyntaxNodes: 2_625, - purpose: "adapter", - sourceSha256: "1b9e81f7d79a7c7f551c1b6d7d1f794e2d49ee698b328eee4a872b914bfc8acc" - }, - "provider-home.tsx": { - maxLines: 75, - maxSyntaxNodes: 556, - purpose: "provider-home", - sourceSha256: "31085a2bad1d6d82b3709946464df332fe1d22e13236708dfb840c8fbd7d5744" - }, - "strict-json.tsx": { - maxLines: 350, - maxSyntaxNodes: 1_965, - purpose: "strict-input", - sourceSha256: "16c909eb1f01c82e1174db61877a30028b58a49466714865f1243293f10b186b" - }, - "toml.tsx": { - maxLines: 120, - maxSyntaxNodes: 526, - purpose: "toml", - sourceSha256: "51b15d0f75a09b49a53a33709cf9127c74b2a35814ad8770ce529f0638a4f6ae" - } -}; - -function lineCount(source: string): number { - return source.replace(/\n$/u, "").split("\n").length; -} - -function syntaxNodeCount(sourceFile: ts.SourceFile): number { - let count = 0; - const visit = (node: ts.Node): void => { - if (node !== sourceFile) count += 1; - ts.forEachChild(node, visit); - }; - visit(sourceFile); - return count; -} - -function sourceFingerprint(source: string): string { - return crypto.createHash("sha256").update(source, "utf8").digest("hex"); -} - -function assertSourceMatchesPolicy(relativePath: string, source: SourceUnit, policy: SourcePolicy): void { - const lines = lineCount(source.source); - const syntaxNodes = syntaxNodeCount(source.ast); - assert.ok(lines <= policy.maxLines, `${relativePath} grew past its ${policy.maxLines}-line review ceiling`); - assert.ok( - syntaxNodes <= policy.maxSyntaxNodes, - `${relativePath} grew past its ${policy.maxSyntaxNodes}-node structural ceiling; classify the change before accepting it` - ); - assert.equal( - sourceFingerprint(source.source), - policy.sourceSha256, - `${relativePath} changed from its reviewed source fingerprint; audit responsibilities and update the policy explicitly` - ); -} - -function assertResponsibilityReviewMatchesSource( - relativePath: string, - source: SourceUnit, - policy: ResponsibilityPolicy -): void { - assert.equal( - sourceFingerprint(source.source), - policy.classifiedSourceSha256, - `${relativePath} changed from its responsibility-reviewed source fingerprint; audit its responsibility declaration and refresh the independent classification policy` - ); -} - function detectedOrchestratorResponsibilities(source: SourceUnit): ReadonlySet { const detected = new Set(); const filesystemWalkingBindings = new Set(); @@ -443,7 +273,7 @@ function detectedOrchestratorResponsibilities(source: SourceUnit): ReadonlySet ): readonly OrchestratorResponsibility[] { const detected = [...detectedOrchestratorResponsibilities(source)].sort(); const undeclared = detected.filter((responsibility) => !policy.responsibilities.includes(responsibility)); @@ -881,79 +711,7 @@ test("source inventory recurses through both TypeScript source extensions", () = } }); -test("syntax budgets ignore names and comments but catch structural orchestration additions", () => { - const baseline = syntaxNodeCount( - parseSource("fixture.ts", "export async function run(task: string) { return task; }\n").ast - ); - const renamed = syntaxNodeCount( - parseSource( - "fixture.ts", - "// resumeSession and prompt_tokens are documentation, not classification markers.\n" + - "export async function invoke(prompt: string) { return prompt; }\n" - ).ast - ); - assert.equal(renamed, baseline); - assert.notEqual( - sourceFingerprint("export async function run(task: string) { return task; }\n"), - sourceFingerprint( - "// resumeSession and prompt_tokens are documentation, not classification markers.\n" + - "export async function invoke(prompt: string) { return prompt; }\n" - ), - "an equal-size replacement must still require an explicit source-policy review" - ); - - for (const addition of [ - 'import { readdir } from "node:fs/promises"; export async function run() { return readdir("."); }\n', - 'export function run(prompt: string) { const launchArguments = ["--resume", prompt]; return launchArguments; }\n', - "export function run(usage: { prompt_tokens: number; completion_tokens: number }) { return usage.prompt_tokens + usage.completion_tokens; }\n" - ]) { - assert.ok(syntaxNodeCount(parseSource("fixture.ts", addition).ast) > baseline); - } -}); - -test("source-policy acknowledgment cannot reuse a stale responsibility review for opaque behavior", () => { - const baselineSource = "export function passThrough(payload: Uint8Array) { return payload; }\n"; - const changedSource = - 'import { decodeProviderEvent } from "./provider-parser";\n' + - "export function passThrough(payload: Uint8Array) { return decodeProviderEvent(payload); }\n"; - const changed = parseSource("fixture.tsx", changedSource); - const acknowledgedSourcePolicy: SourcePolicy = { - maxLines: lineCount(changedSource), - maxSyntaxNodes: syntaxNodeCount(changed.ast), - purpose: "adapter", - sourceSha256: sourceFingerprint(changedSource) - }; - const staleResponsibilityPolicy: ResponsibilityPolicy = { - classifiedSourceSha256: sourceFingerprint(baselineSource), - responsibilities: [], - upstreamIssues: [] - }; - - assert.deepEqual( - [...detectedOrchestratorResponsibilities(changed)], - [], - "the fixture must exercise a semantic form outside the conservative static lower bound" - ); - assert.doesNotThrow(() => assertSourceMatchesPolicy("fixture.tsx", changed, acknowledgedSourcePolicy)); - assert.throws( - () => assertResponsibilityReviewMatchesSource("fixture.tsx", changed, staleResponsibilityPolicy), - /responsibility-reviewed source fingerprint/u - ); - - const refreshedResponsibilityPolicy: ResponsibilityPolicy = { - classifiedSourceSha256: sourceFingerprint(changedSource), - responsibilities: ["output-interpretation"], - upstreamIssues: ["https://example.invalid/upstream"] - }; - assert.doesNotThrow(() => - assertResponsibilityReviewMatchesSource("fixture.tsx", changed, refreshedResponsibilityPolicy) - ); - assert.doesNotThrow(() => - assertDetectedResponsibilitiesDeclared("fixture.tsx", changed, refreshedResponsibilityPolicy) - ); -}); - -test("fingerprint and ceiling acknowledgment cannot retain stale responsibility classifications", () => { +test("a detected responsibility fails until the source declares it", () => { const changedSources: Array<[OrchestratorResponsibility, string]> = [ [ "filesystem-walking", @@ -980,33 +738,13 @@ test("fingerprint and ceiling acknowledgment cannot retain stale responsibility for (const [responsibility, changedSource] of changedSources) { const parsed = parseSource("fixture.tsx", changedSource); - const acknowledgedSourcePolicy: SourcePolicy = { - maxLines: lineCount(changedSource), - maxSyntaxNodes: syntaxNodeCount(parsed.ast), - purpose: "adapter", - sourceSha256: sourceFingerprint(changedSource) - }; - assert.doesNotThrow( - () => assertSourceMatchesPolicy("fixture.tsx", parsed, acknowledgedSourcePolicy), - `${responsibility} fixture must model an independently acknowledged fingerprint and ceiling` - ); - const staleCentralPolicy: ResponsibilityPolicy = { - classifiedSourceSha256: sourceFingerprint(changedSource), - responsibilities: [], - upstreamIssues: [] - }; - assert.doesNotThrow(() => assertResponsibilityReviewMatchesSource("fixture.tsx", parsed, staleCentralPolicy)); assert.throws( - () => assertDetectedResponsibilitiesDeclared("fixture.tsx", parsed, staleCentralPolicy), + () => assertDetectedResponsibilitiesDeclared("fixture.tsx", parsed, { responsibilities: [] }), /static signals for undeclared orchestrator responsibilities/u, responsibility ); assert.doesNotThrow(() => - assertDetectedResponsibilitiesDeclared("fixture.tsx", parsed, { - classifiedSourceSha256: sourceFingerprint(changedSource), - responsibilities: [responsibility], - upstreamIssues: ["https://example.invalid/upstream"] - }) + assertDetectedResponsibilitiesDeclared("fixture.tsx", parsed, { responsibilities: [responsibility] }) ); } }); @@ -1095,7 +833,7 @@ test("non-adapter helpers cannot hide orchestrator responsibilities", () => { }); test("OpenRouter retains the manually reviewed argv responsibility inherited from Codex", () => { - assert.equal(responsibilityPolicies["openrouter.tsx"]!.responsibilities.includes("argv-construction"), true); + assert.equal(adapterPolicies["openrouter.tsx"]?.responsibilities.includes("argv-construction"), true); }); test("main agent registry and recursive sources stay inside reviewed adapter boundaries", (context) => { @@ -1106,20 +844,15 @@ test("main agent registry and recursive sources stay inside reviewed adapter bou const sources = readSourceTree(path.join(packageRoot, "src/templates/smithers/agents")); assert.deepEqual( - Object.keys(sourcePolicies).sort(), + Object.keys(adapterPolicies).sort(), [...sources.keys()].sort(), - "every recursive .ts/.tsx adapter source must have an explicit structural policy" - ); - assert.deepEqual( - Object.keys(responsibilityPolicies).sort(), - [...sources.keys()].sort(), - "every recursive .ts/.tsx source must have an independent responsibility-review policy" + "every recursive .ts/.tsx adapter source must declare its purpose and responsibilities" ); const registered = registeredAdapterSources(sources); const registeredSourcePaths = [...new Set(registered.values())].sort(); assert.deepEqual( - Object.entries(sourcePolicies) + Object.entries(adapterPolicies) .filter(([, policy]) => policy.purpose === "adapter") .map(([relativePath]) => relativePath) .sort(), @@ -1127,23 +860,16 @@ test("main agent registry and recursive sources stay inside reviewed adapter bou "only adapter sources registered in agentFactories may carry the adapter purpose" ); for (const [relativePath, source] of sources) { - const policy = sourcePolicies[relativePath]!; - const responsibilityPolicy = responsibilityPolicies[relativePath]!; - const lines = lineCount(source.source); - const syntaxNodes = syntaxNodeCount(source.ast); - context.diagnostic( - `${relativePath}: ${lines} lines; ${syntaxNodes} syntax nodes; reviewed purpose: ${policy.purpose}` - ); - assertSourceMatchesPolicy(relativePath, source, policy); - assertResponsibilityReviewMatchesSource(relativePath, source, responsibilityPolicy); + const policy = adapterPolicies[relativePath]; + assert.ok(policy, `${relativePath} has no adapter policy`); if (policy.purpose !== "adapter") { assert.deepEqual( - responsibilityPolicy.responsibilities, + policy.responsibilities, [], `${relativePath} is not a registered adapter and cannot own orchestrator responsibilities` ); assert.deepEqual( - responsibilityPolicy.upstreamIssues, + policy.upstreamIssues, [], `${relativePath} is not a registered adapter and cannot own adapter debt` ); @@ -1152,12 +878,10 @@ test("main agent registry and recursive sources stay inside reviewed adapter bou } for (const adapterSource of registeredSourcePaths) { - const policy = responsibilityPolicies[adapterSource]!; - const detectedResponsibilities = assertDetectedResponsibilitiesDeclared( - adapterSource, - sources.get(adapterSource)!, - policy - ); + const policy = adapterPolicies[adapterSource]; + const source = sources.get(adapterSource); + assert.ok(policy && source, `${adapterSource} has no adapter policy`); + const detectedResponsibilities = assertDetectedResponsibilitiesDeclared(adapterSource, source, policy); const registrations = [...registered] .filter(([, sourcePath]) => sourcePath === adapterSource) .map(([agentRef]) => agentRef) @@ -1167,7 +891,6 @@ test("main agent registry and recursive sources stay inside reviewed adapter bou policy.responsibilities.join(", ") || "none" }; statically detected lower bound: ${detectedResponsibilities.join(", ") || "none"}` ); - assert.equal(sourcePolicies[adapterSource]?.purpose, "adapter"); if (policy.responsibilities.length > 0) { assert.ok(policy.upstreamIssues.length > 0, `${adapterSource} debt must link an upstream issue`); } From 95ed8b3f19dbf52b575309fd67233e7d0a244f1e Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:50:26 +0000 Subject: [PATCH 027/206] fix(runtime): import path in the DeepSeek adapter and lint templates for undefined names deepseek.tsx called path.join without importing path, so a render without ULTRAFUZZ_CONFIG_PATH threw "path is not defined" from the factory. Nothing caught it: runtime templates are copied into projects and run by Bun, tsconfig only includes src/**/*.ts, and ESLint turns no-undef off for every .ts/.tsx file. Enable no-undef for packages/runtime/src/templates/**/*.tsx and declare the __ULTRAFUZZ_*__ placeholders the compiler substitutes as readonly globals. Against main the rule reports exactly this bug (deepseek.tsx 200:59 'path' is not defined) and nothing else. Co-Authored-By: Claude Opus 5.5 --- eslint.config.js | 27 +++++++++++++++++++ .../templates/smithers/agents/deepseek.tsx | 1 + 2 files changed, 28 insertions(+) diff --git a/eslint.config.js b/eslint.config.js index 615493e8d..78564f911 100644 --- a/eslint.config.js +++ b/eslint.config.js @@ -116,6 +116,33 @@ export default tseslint.config( ] } }, + { + // Runtime templates are copied into projects and executed by Bun without a + // type check, so an undefined name would first fail when a run renders it. + // The declared globals are the placeholders the compiler substitutes. + files: ["packages/runtime/src/templates/**/*.tsx"], + languageOptions: { + globals: Object.fromEntries( + [ + "AGENT_PROMPT_TEMPLATE", + "AUTHORIZED_DEFENSIVE_SECURITY_CONTEXT", + "COMPILED_TASKS", + "DYNAMIC_GROUPS", + "MAX_DYNAMIC_NODES", + "REPLACE_PROMPT_SCHEMAS", + "RETRY_FAILURE_TEMPLATE", + "RUN_ID_LITERAL", + "RUN_ROOT_RELATIVE", + "SOURCE_PROJECT_ROOT", + "TASK_SPECS", + "UNTRUSTED_CONTENT_BOUNDARY", + "WORKFLOW_NAME", + "WORKFLOW_PATH_RELATIVE" + ].map((name) => [`__ULTRAFUZZ_${name}__`, "readonly"]) + ) + }, + rules: { "no-undef": "error" } + }, eslintConfigPrettier, ...(strictLint ? diff.configs[process.env.CI ? "flat/ci" : "flat/diff"] : []) ); diff --git a/packages/runtime/src/templates/smithers/agents/deepseek.tsx b/packages/runtime/src/templates/smithers/agents/deepseek.tsx index 3e85d0296..b61f11424 100644 --- a/packages/runtime/src/templates/smithers/agents/deepseek.tsx +++ b/packages/runtime/src/templates/smithers/agents/deepseek.tsx @@ -1,4 +1,5 @@ import { readFileSync } from "node:fs"; +import path from "node:path"; import { ClaudeCodeAgent as SmithersClaudeCodeAgent } from "smthrs"; import { workflowControlChildEnvironment, workflowControlCredentialValue } from "./environment"; import { resolveProviderHome } from "./provider-home"; From ef3a8faffa7f35a0a6881f9df80ae0f640a05315 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:56:04 +0000 Subject: [PATCH 028/206] fix(artifacts): the task-manifest gate no longer re-derives which inputs are optional assertSmithersTaskManifestMatchesPlannedGraph re-derived which dependency directories the compiler marks optional and required the manifest to match exactly. That is a second copy of compiler policy: when #1120 narrowed the compiler to review-only opt-in, the copy kept the old rule, and every launch of the default, low-cost, exhaustive and invariant-only profiles failed with "optional dependency artifact directories do not match" before any model ran (#1140). #1160 fixed that by adding a third copy of the review rule to the gate, which leaves the next policy change free to break launch, sync, cloud tasks and report completion again, since all of them run this gate. Keep only the safety property the gate can check without knowing the policy: every optional dependency directory must belong to a producer in a failure_policy: continue group, so a halting producer's output can never become optional. Which consumers opt in stays compiler policy, now in one exported helper (reconcilesPartialResults) that the static compiler and dynamic lowering share. The gate also accepts a review task that keeps a continuing input required, which is what dynamic lowering emits for an empty expansion's source. The #1160 unit test asserted the equality rule itself, so it is rewritten to pin the safety property. A new runtime test plans and compiles every packaged audit profile and runs the gate on the result; against the pre-#1160 gate it rejects default, exhaustive, invariant-only and low-cost, the #1140 regression. Refs #1140, #1160 Co-Authored-By: Claude Opus 5.5 --- .../artifacts/src/smithers-task-manifest.ts | 29 +++++----- .../test/smithers-task-manifest.test.ts | 36 +++++++----- packages/runtime/src/dynamic-runtime.ts | 12 +++- packages/runtime/src/smithers.ts | 10 ++-- packages/runtime/test/runtime.test.ts | 56 +++++++++++++++++++ 5 files changed, 105 insertions(+), 38 deletions(-) diff --git a/packages/artifacts/src/smithers-task-manifest.ts b/packages/artifacts/src/smithers-task-manifest.ts index 8dccba01c..03b8e2b18 100644 --- a/packages/artifacts/src/smithers-task-manifest.ts +++ b/packages/artifacts/src/smithers-task-manifest.ts @@ -1410,7 +1410,11 @@ export function assertSmithersTaskManifestMatchesPlannedGraph( } } - const optionalAttemptIds = new Set( + // Safety only: a producer that halts on failure can never have its output treated as optional. + // Which consumers opt in to a continuing producer's output is compiler policy and is not + // re-derived here: a second copy of that policy rejected every plan the compiler produced once + // the two drifted apart (#1140). + const continuingAttemptIds = new Set( manifest.tasks .filter((task) => { const node = graphNodes.get(task.concreteNodeId)!; @@ -1419,21 +1423,14 @@ export function assertSmithersTaskManifestMatchesPlannedGraph( .map((task) => task.attemptId) ); for (const task of manifest.tasks) { - // Mirrors the runtime compiler: continuation lets independent tasks settle, - // but only the review group reconciles partial results, so only review tasks - // treat inputs from continuing groups as optional. - const reconcilesPartialResults = graphNodes.get(task.concreteNodeId)?.group === "review"; - const expectedOptionalDirectories = reconcilesPartialResults - ? task.dependencyArtifactDirs.filter((directory) => { - const attemptId = directory.split(/[\\/]/u).at(-1); - return attemptId !== undefined && optionalAttemptIds.has(attemptId); - }) - : []; - assertSameStringSet( - task.optionalDependencyArtifactDirs ?? [], - expectedOptionalDirectories, - `Smithers task ${JSON.stringify(task.attemptId)} optional dependency artifact directories` - ); + for (const directory of task.optionalDependencyArtifactDirs ?? []) { + const producer = portablePathBasename(directory) ?? ""; + if (!continuingAttemptIds.has(producer)) { + throw new Error( + `Smithers task ${JSON.stringify(task.attemptId)} marks dependency ${JSON.stringify(producer)} optional, but that producer does not continue on failure` + ); + } + } } } diff --git a/packages/artifacts/test/smithers-task-manifest.test.ts b/packages/artifacts/test/smithers-task-manifest.test.ts index 4590a2679..1bcf1e6cc 100644 --- a/packages/artifacts/test/smithers-task-manifest.test.ts +++ b/packages/artifacts/test/smithers-task-manifest.test.ts @@ -428,7 +428,7 @@ test("rejects missing graph coverage, extra tasks, and graph dependency drift", ); }); -test("only review tasks treat inputs from continuing groups as optional", () => { +test("optional task inputs must come from producers that continue on failure", () => { const continuingGraph = graph(); continuingGraph.groups = { specialists: { defaults: { failure_policy: "continue" } }, @@ -482,21 +482,27 @@ test("only review tasks treat inputs from continuing groups as optional", () => withDependencies(dependent("reviewer", reviewerOptional), "review") ]); - // A continuing specialist's own inputs stay required (#1132's stateful topology). - assert.doesNotThrow(() => - assertSmithersTaskManifestMatchesPlannedGraph(current([], ["/runs/run-1/artifacts/producer"]), continuingGraph) - ); - assert.throws( - () => - assertSmithersTaskManifestMatchesPlannedGraph( - current(["/runs/run-1/artifacts/producer"], ["/runs/run-1/artifacts/producer"]), - continuingGraph - ), - /"consumer" optional dependency artifact directories/u - ); + const producer = "/runs/run-1/artifacts/producer"; + + // Which consumers opt in is compiler policy, so the gate accepts every shape it has emitted: only + // the review task opts in (#1120), every consumer opts in (before #1120), or a review task keeps the + // input required (dynamic lowering of an empty expansion keeps its source required). + const emittedShapes: Array<[consumerOptional: string[], reviewerOptional: string[]]> = [ + [[], [producer]], + [[producer], [producer]], + [[], []] + ]; + for (const [consumerOptional, reviewerOptional] of emittedShapes) { + assert.doesNotThrow(() => + assertSmithersTaskManifestMatchesPlannedGraph(current(consumerOptional, reviewerOptional), continuingGraph) + ); + } + // A producer that halts on failure can never become optional. + const haltingGraph = structuredClone(continuingGraph); + haltingGraph.groups.specialists = {}; assert.throws( - () => assertSmithersTaskManifestMatchesPlannedGraph(current([], []), continuingGraph), - /"reviewer" optional dependency artifact directories/u + () => assertSmithersTaskManifestMatchesPlannedGraph(current([], [producer]), haltingGraph), + /task "reviewer" marks dependency "producer" optional, but that producer does not continue on failure/u ); }); diff --git a/packages/runtime/src/dynamic-runtime.ts b/packages/runtime/src/dynamic-runtime.ts index c71cbc8ae..ca14c1bb6 100644 --- a/packages/runtime/src/dynamic-runtime.ts +++ b/packages/runtime/src/dynamic-runtime.ts @@ -330,6 +330,16 @@ function instantiateDynamicGraphNode(input: { }; } +/** + * Continuation lets independent tasks settle; it does not make a strategy's required inputs + * optional. Only the review group reconciles partial results, so only its tasks treat a continuing + * producer's output as optional (#1120). The compiler and dynamic lowering share this one rule; the + * task-manifest gate checks only that optional inputs come from continuing producers. + */ +export function reconcilesPartialResults(task: Pick): boolean { + return task.metadata.node.group === "review"; +} + function lowerTaskDynamicDependencies( task: CompiledSmithersTask, groups: readonly CompiledSmithersDynamicGroup[], @@ -370,7 +380,7 @@ function lowerTaskDynamicDependencies( dependencies.push(...generated.map((candidate) => candidate.attemptId)); dependencySmithersNodeIds.push(...generated.map((candidate) => candidate.verifierSmithersNodeId)); dependencyArtifactDirs.push(...generated.map((candidate) => candidate.artifactDir)); - if (group.continueOnFail && task.metadata.node.group === "review") { + if (group.continueOnFail && reconcilesPartialResults(task)) { optionalDependencyArtifactDirs.push(...generated.map((candidate) => candidate.artifactDir)); } concreteNodeIds.push(...manifest.items.map((item) => item.node_id)); diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..d45ca7f4e 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -69,6 +69,7 @@ import { routeOwnsCredentialLikeEnvironmentVariable } from "./data-governance.js"; import { archiveDynamicExpansionsForRetry, planDynamicExpansionRetryArchive } from "./dynamic-expansion-retry.js"; +import { reconcilesPartialResults } from "./dynamic-runtime.js"; import { assertControllerSourceDigest, inspectControllerSource, @@ -4077,12 +4078,9 @@ export function compileSmithersWorkflow(input: SmithersCompileInput): CompiledSm const nonBlockingAttemptIdSet = new Set(nonBlockingAttemptIds); const tasks = compiledTasks.map((task) => ({ ...task, - // Continuation lets independent tasks settle. It does not make a strategy's - // required inputs optional; only the review group reconciles partial results. - optionalDependencyArtifactDirs: - task.metadata.node.group === "review" - ? task.dependencyArtifactDirs.filter((directory) => nonBlockingAttemptIdSet.has(path.basename(directory))) - : [] + optionalDependencyArtifactDirs: reconcilesPartialResults(task) + ? task.dependencyArtifactDirs.filter((directory) => nonBlockingAttemptIdSet.has(path.basename(directory))) + : [] })); const smithersDir = path.join(input.runLayout.root, "smithers"); fs.mkdirSync(smithersDir, { recursive: true }); diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..4579780e1 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -32,10 +32,12 @@ import { GOAL_PLAN_JSON_SCHEMA_ID, THREAT_MODEL_JSON_SCHEMA_ID, appendEvent, + assertSmithersTaskManifestMatchesPlannedGraph, createEventRecord, goalPlanJsonSchema, layoutForRunRoot, manifestDigest, + parseSmithersTaskManifestBytes, readPlannedGraphDocument, readRunState, promptArtifactAuthorityPathSelectorId, @@ -52,6 +54,7 @@ import { } from "@ultrafuzz/artifacts"; import { MODAL_NODE_LIFECYCLE_RESERVE_SECONDS, + loadAuditProfileCatalog, parseProjectConfigToml, parseResolvedConfigJsonBytes, resolveConfig, @@ -10509,6 +10512,59 @@ test("a clean scaffold plans the threat-model, goal-plan, and dynamic fanout nod } }); +test("every packaged audit profile compiles into a task plan the launch manifest gate accepts", async () => { + // #1140: after #1120 changed which inputs the compiler marks optional, this gate rejected the + // default, exhaustive, and invariant-only launches before any model ran. + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const xdgCacheHome = path.join(project, "xdg-cache"); + writeShippedDocumentReferenceCaches(xdgCacheHome, loadReferenceCatalog(project)); + writeShippedVulnerabilityDatabaseCache(xdgCacheHome); + const { compileSmithersWorkflow } = await import("../src/smithers.js"); + const previousXdgCacheHome = process.env.XDG_CACHE_HOME; + process.env.XDG_CACHE_HOME = xdgCacheHome; + const rejected: string[] = []; + try { + for (const auditProfile of Object.keys(loadAuditProfileCatalog().profiles)) { + const runId = `packaged-${auditProfile}`; + const plan = await planRun({ projectRoot: project, runId, env: {}, runtimeOverrides: { auditProfile } }); + assert.ok(plan.ok && plan.value, `${auditProfile}: ${JSON.stringify(plan.diagnostics)}`); + const { graph, expanded_graph, layout, rendered_prompts, resolved_config } = plan.value; + const compiled = compileSmithersWorkflow({ + projectRoot: project, + config: resolved_config, + graph: expanded_graph, + runLayout: layout, + workflowName: `ultrafuzz-${runId}`, + renderedPrompts: rendered_prompts + }); + // Launch binds each planned node to its compiled tasks before it runs the gate. + for (const node of graph.nodes) { + const taskNodeIds = compiled.tasks + .filter((task) => task.concreteNodeId === node.id) + .map((task) => task.smithersNodeId); + const [primary] = taskNodeIds; + if (primary !== undefined) node.workflow = { node_id: primary, task_node_ids: taskNodeIds }; + } + try { + assertSmithersTaskManifestMatchesPlannedGraph( + parseSmithersTaskManifestBytes(fs.readFileSync(compiled.tasksPath)), + graph + ); + } catch (error) { + rejected.push(`${auditProfile}: ${error instanceof Error ? error.message : String(error)}`); + } + } + } finally { + if (previousXdgCacheHome === undefined) { + delete process.env.XDG_CACHE_HOME; + } else { + process.env.XDG_CACHE_HOME = previousXdgCacheHome; + } + } + assert.deepEqual(rejected, []); +}); + test("project and runtime topology paths override a profile topology atomically", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From b139d3a276490f4d47be4c75069e0a39f0f98fb7 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:56:16 +0000 Subject: [PATCH 029/206] fix(runtime): record failed launches and report their original error A launch that failed after its run directory existed was either never recorded or recorded where no reader looked: - planRun creates the run directory and then returns later failures (reference materialization, prompt rendering, source-ref publication) without recording them, and startRun's control-lock failure records nothing either. The run stays pending with an empty event journal, and status reports "launch-incomplete ... Launcher liveness is unknown ... wait for it to finish" indefinitely. - startRun's catch does mark the run failed and appends a workflow-submit-failed event with the original error, but a run without a control seal is reported as WORKFLOW_CONTROL_SEAL_MISSING ("may predate sealed runs"), so status, events and why hide the error. - resume of either run takes the lifecycle lock first, whose smithers/ directory a pre-compile failure never created, and fails with an ENOENT lstat error. recordLaunchFailure now writes the existing record (state failed plus the workflow-submit-failed event) for planRun failures after the directory exists, the control-lock failure, and the existing catch, where it replaces a payload wrapper that could throw inside the catch. The event contract has one code, so other codes are kept in the message. For a run with no control seal, the linked-evidence reader and resume report the recorded error as RUN_LAUNCH_FAILED. No new status, seal or schema. The vulnerability-database reference check needs only the graph, and the reference caches can be verified with the existing verifyReferencesCached (previously unused), so both now run before createRunLayout and those errors commit no run directory. The topology-summary invariant moves up with them, since it is a throw that planRun's caller does not catch. Closes #1140 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/cli.md | 9 +++ packages/runtime/src/plan-run.ts | 60 ++++++++++++-------- packages/runtime/src/start-run.ts | 81 ++++++++++++++++++++++----- packages/runtime/test/runtime.test.ts | 79 ++++++++++++++++++++++++++ 5 files changed, 193 insertions(+), 37 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..9042a30ca 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[artifacts] [runtime] [docs]** The launch task-manifest gate no longer re-derives which inputs the compiler marks optional: it checks only that each optional input comes from a producer that continues on failure, and the compiler and dynamic lowering share the one review-only opt-in rule. Before #1160 the gate's older copy of that rule rejected the `default`, `low-cost`, `exhaustive`, and `invariant-only` launches before any model ran; a new test plans and compiles every packaged audit profile and runs the gate on the result. A launch that fails while planning or compiling, after its run directory exists, is now recorded as `failed` with its original error, which `status`, `resume`, and the other linked-workflow commands report as `RUN_LAUNCH_FAILED` instead of a missing control seal or an open-ended incomplete launch. A goal-planning topology without its vulnerability-database reference, and a missing or invalid reference cache, now fail before any run directory is created (#1140, #1160). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..bb4d0f3fb 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -316,6 +316,15 @@ the linked workflow status when it is available; its JSON output retains both `workflow_status` and `ultrafuzz_status` so lifecycle divergence remains explicit. +A launch that fails while planning or compiling, after its run directory +exists, records the run as `failed` with its original error. `status`, +`resume`, and the other commands that read the linked workflow then report +that error as `RUN_LAUNCH_FAILED`. Such a run never reached Smithers and cannot +be resumed; fix the cause and start a new run. A goal-planning topology without +its vulnerability-database reference, and a missing or invalid reference cache +(run `ultrafuzz references sync`), are rejected before the run directory is +created. + `resume` delegates continuation to Smithers with the same Ultrafuzz and Smithers run IDs and automatically accepts changed workflow source. Control seals, link journals, controller generations, graph fingerprints, current diff --git a/packages/runtime/src/plan-run.ts b/packages/runtime/src/plan-run.ts index 23a091dd0..f3e1086ae 100644 --- a/packages/runtime/src/plan-run.ts +++ b/packages/runtime/src/plan-run.ts @@ -43,7 +43,7 @@ import { type PromptConcreteNode, type PromptGraphNode } from "@ultrafuzz/prompts"; -import { loadReferenceCatalog, materializeReferenceArtifacts } from "@ultrafuzz/references"; +import { loadReferenceCatalog, materializeReferenceArtifacts, verifyReferencesCached } from "@ultrafuzz/references"; import { expandTopology, fingerprintGraph, @@ -108,6 +108,8 @@ interface PlanRunHooks { resolvedConfig: PlanRunValue["resolved_config"]; expandedGraph: ExpandedGraph; }): Promise; + /** The run directory now exists, so a later planning failure leaves it for the caller to record. */ + afterLayoutCreated?(layout: RunLayout): void; } export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { @@ -130,6 +132,9 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { if (!validation.ok || !validation.value) { return runtimeFailure(validation.diagnostics); } + if (validation.value.topology === undefined) { + throw new Error("validated run plan is missing its topology summary"); + } const resolved = await loadResolvedProject(input); if (!resolved.config) { @@ -228,6 +233,23 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { if (hasRuntimeErrors(graphDiagnostics)) { return runtimeFailure(graphDiagnostics); } + // Graph-only checks belong before the run directory exists, so an invalid topology commits nothing. + const requiresVulnerabilityDatabase = graph.nodes.some( + (node) => node.logical_id === "threat-model" || node.logical_id === "goal-plan" + ); + const vulnerabilityDatabaseReferenceNode = graph.nodes.find( + (node) => node.logical_id === VULNERABILITY_DATABASE_REFERENCE_NODE_ID && node.kind === "reference" + ); + if (requiresVulnerabilityDatabase && vulnerabilityDatabaseReferenceNode === undefined) { + return runtimeFailure([ + { + code: "VULNERABILITY_DATABASE_REFERENCE_REQUIRED", + message: `topology nodes threat-model/goal-plan require the ${VULNERABILITY_DATABASE_REFERENCE_NODE_ID} pinned reference node`, + severity: "error", + source: "vulnerability-database" + } + ]); + } let controllerSource: ReturnType; try { controllerSource = inspectControllerSource(projectRoot); @@ -337,6 +359,18 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { requestedConcurrency: input.maxConcurrency ?? resolved.config.run.maxParallelAgents, nodes: stateNodes }); + const referenceIds = graph.nodes.flatMap((node) => + node.kind === "reference" && node.reference !== undefined ? [node.reference] : [] + ); + if (referenceIds.length > 0) { + // Materialization reads these caches after the run directory exists; check them first so a + // missing or stale cache does not leave a run behind that can never launch. + try { + verifyReferencesCached(loadReferenceCatalog(projectRoot), referenceIds); + } catch (error) { + return runtimeFailure([diagnosticFromError(error, "references", "REFERENCE_MATERIALIZE_FAILED")]); + } + } let layout; try { @@ -382,6 +416,7 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { forge_guard: forgeGuardMetadata(resolved.config, false) } }); + hooks.afterLayoutCreated?.(layout); writeFileDurable(path.join(layout.root, DATA_GOVERNANCE_PROVENANCE_PATH), governanceBytes); } catch (error) { return runtimeFailure([diagnosticFromError(error, "runtime", "RUN_LAYOUT_INVALID")]); @@ -395,27 +430,11 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { } let vulnerabilityDatabase: MaterializedVulnerabilityDatabaseCatalog | undefined; - const requiresVulnerabilityDatabase = graph.nodes.some( - (node) => node.logical_id === "threat-model" || node.logical_id === "goal-plan" - ); - if (requiresVulnerabilityDatabase) { - const referenceNode = graph.nodes.find( - (node) => node.logical_id === VULNERABILITY_DATABASE_REFERENCE_NODE_ID && node.kind === "reference" - ); - if (referenceNode === undefined) { - return runtimeFailure([ - { - code: "VULNERABILITY_DATABASE_REFERENCE_REQUIRED", - message: `topology nodes threat-model/goal-plan require the ${VULNERABILITY_DATABASE_REFERENCE_NODE_ID} pinned reference node`, - severity: "error", - source: "vulnerability-database" - } - ]); - } + if (requiresVulnerabilityDatabase && vulnerabilityDatabaseReferenceNode !== undefined) { try { vulnerabilityDatabase = materializeVulnerabilityDatabasePlannerCatalog( layout.root, - getNodeArtifactDir(layout, referenceNode.id) + getNodeArtifactDir(layout, vulnerabilityDatabaseReferenceNode.id) ); } catch (error) { return runtimeFailure([ @@ -446,9 +465,6 @@ export async function planRun(input: PlanRunInput, hooks: PlanRunHooks = {}) { } catch (error) { return runtimeFailure([diagnosticFromError(error, "prompts", "PROMPT_TEMPLATE_SNAPSHOT_FAILED")]); } - if (validation.value.topology === undefined) { - throw new Error("validated run plan is missing its topology summary"); - } if (sourceRevision !== undefined) { try { assertLaunchCheckoutRevision(projectRoot, sourceRevision); diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index ba32e842b..f772bce71 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -17,6 +17,7 @@ import { readRegularFileSnapshot, readRunMetadataDocument, readRunState, + replayEvents, safeResolveInside, sensitiveEnvironmentValues, updateRunStatus, @@ -189,12 +190,19 @@ function controllerRefreshInspectionEnvironment( } export async function startRun(input: StartRunInput) { + let createdLayout: RunLayout | undefined; const planned = await planRun(input, { enforceDataGovernance: true, beforeMaterialize: async ({ resolvedConfig, expandedGraph }) => - requiredCommandPreflightDiagnostics(input, resolvedConfig, expandedGraph) + requiredCommandPreflightDiagnostics(input, resolvedConfig, expandedGraph), + afterLayoutCreated: (layout) => { + createdLayout = layout; + } }); if (!planned.ok || !planned.value) { + if (createdLayout !== undefined) { + recordLaunchFailure(createdLayout, planned.diagnostics, sensitiveEnvironmentValues(input.env ?? process.env)); + } return runtimeFailure(planned.diagnostics); } @@ -209,7 +217,9 @@ export async function startRun(input: StartRunInput) { try { releaseControlLock = await acquireWorkflowControlLock(plan.layout); } catch (error) { - return runtimeFailure([smithersDiagnostic(error, "WORKFLOW_CONTROL_PREPARATION_FAILED")]); + const diagnostic = smithersDiagnostic(error, "WORKFLOW_CONTROL_PREPARATION_FAILED"); + recordLaunchFailure(plan.layout, [diagnostic], forbiddenSecretValues); + return runtimeFailure([diagnostic]); } try { const compiled = compileSmithersWorkflow({ @@ -331,13 +341,7 @@ export async function startRun(input: StartRunInput) { ); } catch (error) { const diagnostic = smithersDiagnostic(error, "WORKFLOW_SUBMISSION_FAILED"); - updateRunStatus(plan.layout, "failed", undefined, { forbiddenSecretValues }); - appendEvent(plan.layout, { - eventType: "workflow-submit-failed", - status: "failed", - payload: workflowSubmissionFailureEventPayload(diagnostic), - forbiddenSecretValues - }); + recordLaunchFailure(plan.layout, [diagnostic], forbiddenSecretValues); return runtimeFailure([diagnostic]); } finally { await releaseControlLock(); @@ -485,6 +489,11 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { const layout = layoutForRunRoot(path.join(runsRoot, runId), runId); assertPathInside(runsRoot, layout.root, "run root"); if (fs.existsSync(runsRoot)) assertNoSymlinkComponents(runsRoot, layout.root, "run root"); + // Checked before the lifecycle lock, whose directory a launch that failed before compiling never created. + const launchFailure = pathIsMissing(workflowControlPaths(projectRoot, layout).integrityPath) + ? recordedLaunchFailure(layout) + : undefined; + if (launchFailure !== undefined) return runtimeFailure([launchFailure]); releaseLifecycleLock = await acquireWorkflowLifecycleLock(layout); assertRegularFileInside(layout.root, layout.runMetadataPath, "run metadata"); @@ -833,13 +842,53 @@ export async function pauseRun(input: PauseRunInput) { } } -function workflowSubmissionFailureEventPayload( - diagnostic: RuntimeDiagnostic -): Extract["payload"] { - if (!isWorkflowSubmissionFailureEventPayload(diagnostic)) { - throw new Error("workflow submission diagnostic does not match the current event contract"); +/** + * A launch that fails after its run directory exists must not look pending, running, or resumable: + * mark the run failed and keep the original error, which readers report in place of a missing seal. + * The event contract allows one code, so any other code is kept in the message. + */ +function recordLaunchFailure( + layout: RunLayout, + diagnostics: readonly RuntimeDiagnostic[], + forbiddenSecretValues: readonly string[] +): void { + const errors = diagnostics.filter((diagnostic) => diagnostic.severity === "error"); + const [only] = errors; + const payload = + errors.length === 1 && only !== undefined && isWorkflowSubmissionFailureEventPayload(only) + ? only + : { + code: "WORKFLOW_SUBMISSION_FAILED" as const, + message: errors.map((diagnostic) => `${diagnostic.code}: ${diagnostic.message}`).join("; "), + severity: "error" as const, + source: "workflow" as const, + details: {} + }; + try { + updateRunStatus(layout, "failed", undefined, { forbiddenSecretValues }); + appendEvent(layout, { eventType: "workflow-submit-failed", status: "failed", payload, forbiddenSecretValues }); + } catch { + // Best effort: the caller returns the original diagnostics either way. + } +} + +/** A failed launch's recorded error; callers ask only about runs whose workflow controls were never sealed. */ +function recordedLaunchFailure(layout: RunLayout): RuntimeDiagnostic | undefined { + try { + const failure = replayEvents(layout, Number.MAX_SAFE_INTEGER) + .records.filter((record) => record.event_type === "workflow-submit-failed") + .at(-1); + if (failure?.event_type !== "workflow-submit-failed") return undefined; + return { + code: "RUN_LAUNCH_FAILED", + message: `run ${layout.runId} failed to launch: ${failure.payload.message}. A run whose launch failed cannot be resumed; fix the cause and start a new run`, + severity: "error", + source: "workflow", + path: layout.eventsPath + }; + } catch { + return undefined; } - return diagnostic; } function isWorkflowSubmissionFailureEventPayload( @@ -1496,6 +1545,8 @@ function missingLinkedWorkflowEvidenceDiagnostic( path: controlSealPath }; } + const launchFailure = recordedLaunchFailure(layout); + if (launchFailure !== undefined) return launchFailure; return { code: "WORKFLOW_CONTROL_SEAL_MISSING", message: `run ${layout.runId} lacks the required workflow control seal; it may predate sealed runs or be incomplete and cannot be safely upgraded in place. Preserve its stored artifacts and start a new run with a new run ID`, diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 4579780e1..b6f41ac57 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -25895,6 +25895,85 @@ test("resume continues under a superseded generated manifest without changing ei if (before.ok && migrated.ok) assert.equal(migrated.smithersRunId, before.smithersRunId); }); +test("a launch that fails after creating its run directory is recorded failed and reports its own error", async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + writeSmallTopology(project); + const git = (args: string[]): string => + execFileSync("git", args, { cwd: project, encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] }).trim(); + git(["init", "--quiet", "--initial-branch=main"]); + git(["config", "user.name", "Ultrafuzz Test"]); + git(["config", "user.email", "test@invalid"]); + git(["commit", "--quiet", "--allow-empty", "-m", "earlier"]); + const earlier = git(["rev-parse", "HEAD"]); + git(["add", "--all"]); + git(["commit", "--quiet", "-m", "launch"]); + // A source ref left by an earlier run with the same ID fails planning once the run directory exists. + git(["update-ref", "refs/ultrafuzz/runs/stale-source-ref/source", earlier]); + const env = fakeSmithersEnv(project); + const launches = [ + { runId: "stale-source-ref", error: /RUN_SOURCE_REVISION_PERSIST_FAILED: run source ref .+ different commit/u }, + // Planning accepts this run ID and compilation rejects it, before workflow controls are sealed. + { runId: "Uppercase-Run", error: /"Uppercase-Run" does not produce a current workflow runner run ID/u } + ]; + for (const { runId, error } of launches) { + const launched = await startRun({ projectRoot: project, runId, env }); + assert.equal(launched.ok, false, runId); + const layout = layoutForRunRoot(path.join(project, ".ultrafuzz", "runs", runId), runId); + assert.equal(readRunState(layout).status, "failed", runId); + const health = await getRunHealth({ projectRoot: project, runId, env }); + const resumed = await resumeRun({ projectRoot: project, runId, env }); + for (const result of [health, resumed]) { + assert.equal(result.ok, false, runId); + assert.equal(result.diagnostics[0]?.code, "RUN_LAUNCH_FAILED", JSON.stringify(result.diagnostics)); + assert.match(result.diagnostics[0]?.message ?? "", error); + } + } + assert.equal(fs.existsSync(path.join(project, "smithers-commands.log")), false, "no workflow command ran"); +}); + +test("planning errors that need no run directory leave none behind", async () => { + // A graph-only topology error: goal planning without its pinned vulnerability database. + const graphProject = tempProject(); + assert.equal(initProject({ projectRoot: graphProject, force: true }).ok, true); + writeSmallTopology(graphProject); + const topologyPath = path.join(graphProject, ".ultrafuzz", "topology.yml"); + fs.writeFileSync( + topologyPath, + fs.readFileSync(topologyPath, "utf8").replaceAll("project-discovery\n", "threat-model\n"), + "utf8" + ); + const graphPlan = await planRun({ projectRoot: graphProject, runId: "no-vulnerability-database", env: {} }); + assert.equal(graphPlan.ok, false); + assert.equal( + graphPlan.diagnostics[0]?.code, + "VULNERABILITY_DATABASE_REFERENCE_REQUIRED", + JSON.stringify(graphPlan.diagnostics) + ); + assert.equal(fs.existsSync(path.join(graphProject, ".ultrafuzz", "runs", "no-vulnerability-database")), false); + + // An operator setup error: a pinned reference whose cache was never synced. + const cacheProject = tempProject(); + assert.equal(initProject({ projectRoot: cacheProject, force: true }).ok, true); + writeReferenceTopology(cacheProject); + const previousXdgCacheHome = process.env.XDG_CACHE_HOME; + process.env.XDG_CACHE_HOME = path.join(cacheProject, "empty-cache"); + let cachePlan: Awaited>; + try { + cachePlan = await planRun({ projectRoot: cacheProject, runId: "unsynced-reference", env: {} }); + } finally { + if (previousXdgCacheHome === undefined) { + delete process.env.XDG_CACHE_HOME; + } else { + process.env.XDG_CACHE_HOME = previousXdgCacheHome; + } + } + assert.equal(cachePlan.ok, false); + assert.equal(cachePlan.diagnostics[0]?.code, "MISSING_CACHE", JSON.stringify(cachePlan.diagnostics)); + assert.match(cachePlan.diagnostics[0]?.message ?? "", /ultrafuzz references sync/u); + assert.equal(fs.existsSync(path.join(cacheProject, ".ultrafuzz", "runs", "unsynced-reference")), false); +}); + test("incomplete launch is observable without granting execution authority or claiming liveness", async () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); From f8d72a54a4b4f218a370ef6debee758551713b47 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:41:53 +0000 Subject: [PATCH 030/206] fix(runtime): build host aggregation sources from the sealed attempt authority The host re-check of aggregate-test-files built its source bundles from its own walk of every planned transitive ancestor, keyed by planned node ID, with its own marker, manifest and prerequisite-chain authentication. That walk disagreed with the in-workflow verifier in two ways: - It read every generated-tests producer, including a failure_policy continue strategy or goal that failed before its verifier wrote a marker. The verifier had admitted the aggregation without that optional producer, but the host threw "artifact verification marker does not exist", failing aggregate-test-files and, through it, the final report. - It looked a materialized dynamic producer up under its node ID ("dynamic:threat:") instead of its storage attempt ID, and threw "contains unsafe segment". Take the producers from finalizedDeclaredContractProducers instead: it walks the attempt's sealed ancestor closure, skips optional producers the verifier-persisted admission omitted, maps each attempt to its planned node, and reads each producer through the canonical loadFinalizedNodeOutputSnapshot authority. The module keeps only the bundle projection, attributed to the producer's loop attempt index as the verifier does. The bespoke authentication chain and its tests go away; the canonical reader has its own coverage. The artifact-gates fixture sealer also stops giving deferred dynamic templates workflow bindings, which real graphs never carry and the task manifest validator rejects. Co-Authored-By: Claude Opus 5.5 --- .../src/aggregation-semantic-context.ts | 511 ++---------- packages/runtime/src/artifact-gates.ts | 9 +- .../test/aggregation-semantic-context.test.ts | 725 ------------------ packages/runtime/test/artifact-gates.test.ts | 277 ++++++- 4 files changed, 343 insertions(+), 1179 deletions(-) delete mode 100644 packages/runtime/test/aggregation-semantic-context.test.ts diff --git a/packages/runtime/src/aggregation-semantic-context.ts b/packages/runtime/src/aggregation-semantic-context.ts index abeca8c00..c62c72a4a 100644 --- a/packages/runtime/src/aggregation-semantic-context.ts +++ b/packages/runtime/src/aggregation-semantic-context.ts @@ -1,137 +1,85 @@ -import crypto from "node:crypto"; -import path from "node:path"; -import { isDeepStrictEqual } from "node:util"; - import { - ARTIFACT_MANIFEST_FILE, - ARTIFACT_VERIFICATION_SCHEMA_VERSION, - MAX_GENERATED_TEST_COMPANION_BYTES, - assertArtifactVerificationMarkerSemantics, - assertGeneratedTestManifestSemantics, - assertNoSymlinkComponents, - assertPathInside, - parseStrictJsonBytes, - readPlannedGraphDocument, - readSinglyLinkedRegularFileSnapshotInside, safeResolveInside, - validateArtifactContractBytes, - validateArtifactManifest, - validateArtifactVerificationMarker, - type ArtifactManifest, - type ArtifactVerificationMarker, type GeneratedTestEntry, type GeneratedTestManifest, - type PlannedGraphNodeDocument, type RunLayout, type SemanticAggregationContext, type SemanticAggregationSourceBundleContext, type SemanticAggregationSourceEntryContext } from "@ultrafuzz/artifacts"; -const MAX_AGGREGATION_AUTHORITY_BYTES = 64 * 1024 * 1024; -const MAX_AGGREGATION_PREREQUISITE_MANIFESTS = 4_096; -const MAX_AGGREGATION_PREREQUISITE_EDGES = 16_384; -const MAX_AGGREGATION_PREREQUISITE_BYTES = 64 * 1024 * 1024; -const VERIFICATION_MARKER_DIRECTORY = ".ultrafuzz-verification"; - -interface AuthenticatedPublicationSnapshot { - path: string; - absolutePath: string; - bytes: Buffer; - sha256: string; -} - -interface AuthenticatedProducerAuthority { - marker: ArtifactVerificationMarker; - artifactManifest: ArtifactManifest; - artifactManifestBytes: Buffer; - publications: ReadonlyMap; -} +import type { PlannedGraphNode } from "./types.js"; +import type { VerifiedNodeOutputSnapshot, VerifiedOutputArtifactSnapshot } from "./verified-output.js"; -interface PlannedAttemptAuthority { - node: PlannedGraphNodeDocument; +/** A finalized `ultrafuzz/generated-tests@3` producer the aggregating attempt admitted. */ +export interface AggregationSourceProducer { attemptId: string; - attemptIndex: number; + node: Pick; + authority: Pick; + outputs: readonly VerifiedOutputArtifactSnapshot[]; } /** - * Build the aggregation gate's immutable source authority from the planned - * graph's exact transitive ancestors and their current verifier publications. + * Build the aggregation gate's source authority from the producers the caller + * resolved through the aggregating attempt's sealed ancestor closure and + * verifier-persisted dependency admission. Those are the producers the + * in-workflow verifier admitted, so an optional producer that failed is absent + * rather than fatal, and each source is keyed by its sealed attempt ID. */ export function authenticatedAggregationSemanticContext(input: { layout: RunLayout; - node: PlannedGraphNodeDocument; attemptId: string; + producers: readonly AggregationSourceProducer[]; }): SemanticAggregationContext { - const graph = readPlannedGraphDocument(input.layout.graphPath); - const current = graph.nodes.find((candidate) => candidate.id === input.node.id); - if (current === undefined || !isDeepStrictEqual(current, input.node)) { - throw new Error(`aggregation node ${JSON.stringify(input.node.id)} does not match the current planned graph`); - } - const sourceBundles: SemanticAggregationSourceBundleContext[] = []; - for (const producer of transitiveAncestorNodes(current, graph.nodes)) { - if (producer.kind !== "agentic") continue; - const generatedOutputs = producer.outputs.filter((output) => output.contract === "ultrafuzz/generated-tests@3"); - if (generatedOutputs.length === 0) continue; - for (const attempt of plannedAttempts(producer)) { - const artifactRoot = safeResolveInside( - input.layout.artifactsDir, - attempt.attemptId, - "generated-test aggregation source artifact directory" - ); - const authority = authenticatedProducerAuthority( - input.layout, - graph.nodes, - producer, - attempt.attemptId, - attempt.attemptIndex, - artifactRoot - ); - for (const output of generatedOutputs) { - const markerArtifact = authority.marker.artifacts.find((artifact) => artifact.path === output.path); - if (markerArtifact === undefined) { - throw new Error(`verified producer ${attempt.attemptId} omits generated-test output ${output.path}`); - } - const manifestSnapshot = authority.publications.get(output.path); - if (manifestSnapshot === undefined || markerArtifact.sha256 !== manifestSnapshot.sha256) { - throw new Error(`verified generated-test manifest publication changed ${attempt.attemptId}/${output.path}`); - } - const validation = validateArtifactContractBytes( - "ultrafuzz/generated-tests@3", - manifestSnapshot.bytes, - manifestSnapshot.absolutePath - ); - if (!validation.ok || validation.value === undefined) { - throw new Error(`verified generated-test manifest is no longer valid ${attempt.attemptId}/${output.path}`); - } - const manifest = validation.value as GeneratedTestManifest; - assertGeneratedTestManifestSemantics(manifest); - if (manifest.run_id !== input.layout.runId || manifest.node_id !== producer.logical_id) { - throw new Error(`verified generated-test manifest identity changed ${attempt.attemptId}/${output.path}`); - } - const entries = [ - ...manifest.generated_tests.map((entry) => - authenticatedEntry("generated-test", entry, authority.publications) - ), - ...manifest.support_files.map((entry) => authenticatedEntry("support-file", entry, authority.publications)) - ]; - sourceBundles.push( - Object.freeze({ - strategy: producer.logical_id, - nodeId: manifest.node_id, - sourceAttemptId: attempt.attemptId, - attemptIndex: attempt.attemptIndex, - sourceManifestPath: manifestSnapshot.absolutePath, - sourceManifestRelativePath: output.path, - sourceManifestSha256: manifestSnapshot.sha256, - sourceRunId: manifest.run_id, - framework: manifest.framework, - entries: Object.freeze(entries) - }) + const sourceBundles = input.producers.flatMap((producer) => { + const publications = new Map(producer.authority.publications.map((publication) => [publication.path, publication])); + const entry = ( + kind: SemanticAggregationSourceEntryContext["kind"], + candidate: GeneratedTestEntry + ): SemanticAggregationSourceEntryContext => { + const publication = publications.get(candidate.path); + if ( + publication === undefined || + publication.sha256 !== candidate.sha256 || + publication.bytes.byteLength !== candidate.size_bytes + ) { + throw new Error( + `verified generated-test companion publication changed ${producer.attemptId}/${candidate.path}` ); } - } - } + return Object.freeze({ + kind, + sourceArtifactPath: publication.absolute_path, + sourceRelativePath: candidate.path, + sizeBytes: candidate.size_bytes, + sha256: candidate.sha256, + bytes: Buffer.from(publication.bytes), + ...(candidate.language === undefined ? {} : { language: candidate.language }), + ...(candidate.description === undefined ? {} : { description: candidate.description }), + ...(candidate.provenance === undefined ? {} : { provenance: Object.freeze({ ...candidate.provenance }) }) + }); + }; + return producer.outputs.map((output): SemanticAggregationSourceBundleContext => { + const manifest = output.value as GeneratedTestManifest; + return Object.freeze({ + strategy: producer.node.logical_id, + nodeId: manifest.node_id, + sourceAttemptId: producer.attemptId, + // The verifier attributes a bundle to its producer's loop attempt index, + // which the sealed task manifest binds to the planned node's loop. + attemptIndex: producer.node.loop.attempt_index, + sourceManifestPath: output.absolute_path, + sourceManifestRelativePath: output.path, + sourceManifestSha256: output.sha256, + sourceRunId: manifest.run_id, + framework: manifest.framework, + entries: Object.freeze([ + ...manifest.generated_tests.map((candidate) => entry("generated-test", candidate)), + ...manifest.support_files.map((candidate) => entry("support-file", candidate)) + ]) + }); + }); + }); sourceBundles.sort( (left, right) => left.sourceAttemptId.localeCompare(right.sourceAttemptId) || @@ -142,342 +90,3 @@ export function authenticatedAggregationSemanticContext(input: { sourceBundles: Object.freeze(sourceBundles) }); } - -function transitiveAncestorNodes( - current: PlannedGraphNodeDocument, - nodes: readonly PlannedGraphNodeDocument[] -): PlannedGraphNodeDocument[] { - const byId = new Map(nodes.map((node) => [node.id, node])); - const ancestors = new Map(); - const pending = [...current.depends_on]; - while (pending.length > 0) { - const id = pending.pop()!; - if (ancestors.has(id)) continue; - const node = byId.get(id); - if (node === undefined) throw new Error(`aggregation dependency is absent from the planned graph: ${id}`); - ancestors.set(id, node); - pending.push(...node.depends_on); - } - return [...ancestors.values()].sort((left, right) => left.id.localeCompare(right.id)); -} - -function plannedAttempts(node: PlannedGraphNodeDocument): Array<{ attemptId: string; attemptIndex: number }> { - if (node.model_fanout.length <= 1) { - return [{ attemptId: node.id, attemptIndex: node.model_fanout[0]?.attempt_index ?? node.loop.attempt_index }]; - } - return node.model_fanout.map((model) => ({ - attemptId: `${node.id}__model_${model.model_index}__attempt_${model.attempt_index}`, - attemptIndex: model.attempt_index - })); -} - -function plannedAttemptAuthorities(nodes: readonly PlannedGraphNodeDocument[]): Map { - const authorities = new Map(); - for (const node of nodes) { - for (const attempt of plannedAttempts(node)) { - if (authorities.has(attempt.attemptId)) { - throw new Error(`planned graph repeats artifact attempt authority ${attempt.attemptId}`); - } - authorities.set(attempt.attemptId, { node, ...attempt }); - } - } - return authorities; -} - -function plannedPrerequisiteAttempts( - node: PlannedGraphNodeDocument, - nodesById: ReadonlyMap -): PlannedAttemptAuthority[] { - return node.depends_on - .flatMap((dependencyId) => { - const dependency = nodesById.get(dependencyId); - if (dependency === undefined) { - throw new Error(`producer dependency is absent from the planned graph: ${dependencyId}`); - } - return plannedAttempts(dependency).map((attempt) => ({ node: dependency, ...attempt })); - }) - .sort((left, right) => left.attemptId.localeCompare(right.attemptId)); -} - -function authenticatedProducerAuthority( - layout: RunLayout, - nodes: readonly PlannedGraphNodeDocument[], - producer: PlannedGraphNodeDocument, - attemptId: string, - attemptIndex: number, - artifactRoot: string -): AuthenticatedProducerAuthority { - const markerRoot = path.resolve(layout.root, VERIFICATION_MARKER_DIRECTORY); - assertPathInside(layout.root, markerRoot, "artifact verification marker root"); - assertNoSymlinkComponents(layout.root, markerRoot, "artifact verification marker root"); - const markerPath = safeResolveInside(markerRoot, `${attemptId}.json`, "artifact verification marker"); - const markerBytes = readSinglyLinkedRegularFileSnapshotInside( - markerRoot, - markerPath, - MAX_AGGREGATION_AUTHORITY_BYTES, - "artifact verification marker" - ); - const parsed = parseStrictJsonBytes(markerBytes); - const shape = validateArtifactVerificationMarker(parsed); - if (!shape.ok) throw new Error(`artifact verification marker is invalid for ${attemptId}`); - const marker = parsed as ArtifactVerificationMarker; - assertArtifactVerificationMarkerSemantics(marker); - if ( - marker.schema_version !== ARTIFACT_VERIFICATION_SCHEMA_VERSION || - marker.attempt_id !== attemptId || - marker.node_id !== producer.logical_id || - marker.artifacts.length !== producer.outputs.length - ) { - throw new Error(`artifact verification marker identity changed for ${attemptId}`); - } - const markerArtifacts = new Map(marker.artifacts.map((artifact) => [artifact.path, artifact])); - if (markerArtifacts.size !== marker.artifacts.length) - throw new Error(`artifact verification marker repeats ${attemptId}`); - - const artifactManifestPath = safeResolveInside(artifactRoot, ARTIFACT_MANIFEST_FILE, "producer artifact manifest"); - const artifactManifestBytes = readSinglyLinkedRegularFileSnapshotInside( - artifactRoot, - artifactManifestPath, - MAX_AGGREGATION_AUTHORITY_BYTES, - "producer artifact manifest" - ); - const artifactManifestValue = parseStrictJsonBytes(artifactManifestBytes); - const artifactManifestShape = validateArtifactManifest(artifactManifestValue); - if (!artifactManifestShape.ok) throw new Error(`producer artifact manifest is invalid for ${attemptId}`); - const artifactManifest = artifactManifestValue as ArtifactManifest; - const provenanceMetadata = artifactManifest.provenance.metadata as { concrete_node_id?: string } | undefined; - if ( - artifactManifest.run_id !== layout.runId || - artifactManifest.node_id !== attemptId || - artifactManifest.producer_node_id !== attemptId || - artifactManifest.provenance.producer_node_id !== attemptId || - artifactManifest.provenance.run_id !== layout.runId || - artifactManifest.provenance.logical_node_id !== producer.logical_id || - artifactManifest.provenance.attempt_index !== attemptIndex || - provenanceMetadata?.concrete_node_id !== producer.id || - !isDeepStrictEqual(artifactManifest.output_contracts, producer.outputs) - ) { - throw new Error(`producer artifact manifest identity changed for ${attemptId}`); - } - authenticatePrerequisiteManifestChain(layout, nodes, artifactManifest, attemptId); - - const manifestFiles = new Map(artifactManifest.files.map((entry) => [entry.path, entry])); - const markerPublications = new Map(marker.publications.map((entry) => [entry.path, entry])); - if ( - manifestFiles.size !== artifactManifest.files.length || - markerPublications.size !== marker.publications.length || - manifestFiles.size !== markerPublications.size || - [...manifestFiles.keys()].some((relativePath) => !markerPublications.has(relativePath)) - ) { - throw new Error(`producer artifact manifest and verifier publication sets differ for ${attemptId}`); - } - const publications = new Map(); - for (const [relativePath, publication] of markerPublications) { - const manifestFile = manifestFiles.get(relativePath)!; - const absolutePath = safeResolveInside(artifactRoot, relativePath, "verified producer publication"); - const bytes = readSinglyLinkedRegularFileSnapshotInside( - artifactRoot, - absolutePath, - MAX_AGGREGATION_AUTHORITY_BYTES, - `verified producer publication ${relativePath}` - ); - const sha256 = digest(bytes); - if ( - publication.sha256 !== sha256 || - manifestFile.sha256 !== sha256 || - manifestFile.size_bytes !== bytes.byteLength - ) { - throw new Error(`verified producer publication changed ${attemptId}/${relativePath}`); - } - publications.set(relativePath, Object.freeze({ path: relativePath, absolutePath, bytes, sha256 })); - } - for (const output of producer.outputs) { - const artifact = markerArtifacts.get(output.path); - const expected = { - path: output.path, - contract: output.contract, - contract_digest: output.contract_digest, - ...(output.schema_file === undefined - ? {} - : { - schema_file: output.schema_file, - schema_id: output.schema_id, - schema_sha256: output.schema_sha256, - schema_bundle_sha256: output.schema_bundle_sha256, - validator_build: output.validator_build - }), - primary: output.primary - }; - if ( - artifact === undefined || - !isDeepStrictEqual( - { - path: artifact.path, - contract: artifact.contract, - contract_digest: artifact.contract_digest, - ...(artifact.schema_file === undefined ? {} : { schema_file: artifact.schema_file }), - ...(artifact.schema_id === undefined ? {} : { schema_id: artifact.schema_id }), - ...(artifact.schema_sha256 === undefined ? {} : { schema_sha256: artifact.schema_sha256 }), - ...(artifact.schema_bundle_sha256 === undefined - ? {} - : { schema_bundle_sha256: artifact.schema_bundle_sha256 }), - ...(artifact.validator_build === undefined ? {} : { validator_build: artifact.validator_build }), - primary: artifact.primary - }, - expected - ) - ) { - throw new Error(`artifact verification marker output binding changed ${attemptId}/${output.path}`); - } - const publication = publications.get(output.path); - if (publication === undefined || artifact.sha256 !== publication.sha256) { - throw new Error(`verified producer output publication changed ${attemptId}/${output.path}`); - } - } - return Object.freeze({ - marker, - artifactManifest, - artifactManifestBytes: Buffer.from(artifactManifestBytes), - publications - }); -} - -function authenticatePrerequisiteManifestChain( - layout: RunLayout, - nodes: readonly PlannedGraphNodeDocument[], - rootManifest: ArtifactManifest, - rootAttemptId: string -): void { - const nodesById = new Map(nodes.map((node) => [node.id, node])); - const attemptsById = plannedAttemptAuthorities(nodes); - const rootAuthority = attemptsById.get(rootAttemptId); - if (rootAuthority === undefined) { - throw new Error(`producer attempt is absent from the planned graph: ${rootAttemptId}`); - } - - const pending: Array<{ authority: PlannedAttemptAuthority; expectedSha256: string }> = []; - const scheduledDigests = new Map(); - const schedule = (authority: PlannedAttemptAuthority, expectedSha256: string): void => { - const previous = scheduledDigests.get(authority.attemptId); - if (previous !== undefined) { - if (previous !== expectedSha256) { - throw new Error(`prerequisite artifact manifest has conflicting sealed digests for ${authority.attemptId}`); - } - return; - } - if (scheduledDigests.size >= MAX_AGGREGATION_PREREQUISITE_MANIFESTS) { - throw new Error( - `prerequisite artifact manifest chain exceeds ${MAX_AGGREGATION_PREREQUISITE_MANIFESTS} manifests` - ); - } - scheduledDigests.set(authority.attemptId, expectedSha256); - pending.push({ authority, expectedSha256 }); - }; - - const rootExpectedPrerequisites = plannedPrerequisiteAttempts(rootAuthority.node, nodesById); - const rootExpectedById = new Map(rootExpectedPrerequisites.map((entry) => [entry.attemptId, entry])); - const rootActualIds = rootManifest.prerequisite_manifests.map((entry) => entry.node_id).sort(); - if (!isDeepStrictEqual(rootActualIds, [...rootExpectedById.keys()].sort())) { - throw new Error(`producer artifact manifest prerequisite set changed for ${rootAttemptId}`); - } - for (const prerequisite of rootManifest.prerequisite_manifests) { - schedule(rootExpectedById.get(prerequisite.node_id)!, prerequisite.sha256); - } - - let authenticatedManifestCount = 0; - let authenticatedEdgeCount = rootManifest.prerequisite_manifests.length; - let authenticatedBytes = 0; - if (authenticatedEdgeCount > MAX_AGGREGATION_PREREQUISITE_EDGES) { - throw new Error(`prerequisite artifact manifest chain exceeds ${MAX_AGGREGATION_PREREQUISITE_EDGES} edges`); - } - - while (pending.length > 0) { - const { authority, expectedSha256 } = pending.pop()!; - const remainingBytes = MAX_AGGREGATION_PREREQUISITE_BYTES - authenticatedBytes; - if (remainingBytes <= 0) { - throw new Error(`prerequisite artifact manifest chain exceeds ${MAX_AGGREGATION_PREREQUISITE_BYTES} bytes`); - } - const artifactRoot = safeResolveInside(layout.artifactsDir, authority.attemptId, "prerequisite artifact directory"); - const manifestPath = safeResolveInside(artifactRoot, ARTIFACT_MANIFEST_FILE, "prerequisite artifact manifest"); - const bytes = readSinglyLinkedRegularFileSnapshotInside( - artifactRoot, - manifestPath, - Math.min(MAX_AGGREGATION_AUTHORITY_BYTES, remainingBytes), - `prerequisite artifact manifest ${authority.attemptId}` - ); - authenticatedBytes += bytes.byteLength; - authenticatedManifestCount += 1; - const sha256 = digest(bytes); - if (sha256 !== expectedSha256) { - throw new Error(`prerequisite artifact manifest bytes changed for ${authority.attemptId}`); - } - const value = parseStrictJsonBytes(bytes); - const shape = validateArtifactManifest(value); - if (!shape.ok) throw new Error(`prerequisite artifact manifest is invalid for ${authority.attemptId}`); - const manifest = value as ArtifactManifest; - const provenanceMetadata = manifest.provenance.metadata as { concrete_node_id?: string } | undefined; - if ( - manifest.run_id !== layout.runId || - manifest.node_id !== authority.attemptId || - manifest.producer_node_id !== authority.attemptId || - manifest.provenance.run_id !== layout.runId || - manifest.provenance.producer_node_id !== authority.attemptId || - manifest.provenance.logical_node_id !== authority.node.logical_id || - !isDeepStrictEqual(manifest.output_contracts, authority.node.outputs) || - (authority.node.kind === "agentic" && - (manifest.provenance.attempt_index !== authority.attemptIndex || - provenanceMetadata?.concrete_node_id !== authority.node.id)) - ) { - throw new Error(`prerequisite artifact manifest identity changed for ${authority.attemptId}`); - } - - const expectedPrerequisites = plannedPrerequisiteAttempts(authority.node, nodesById); - const expectedById = new Map(expectedPrerequisites.map((entry) => [entry.attemptId, entry])); - const actualIds = manifest.prerequisite_manifests.map((entry) => entry.node_id).sort(); - if (!isDeepStrictEqual(actualIds, [...expectedById.keys()].sort())) { - throw new Error(`prerequisite artifact manifest prerequisite set changed for ${authority.attemptId}`); - } - authenticatedEdgeCount += manifest.prerequisite_manifests.length; - if (authenticatedEdgeCount > MAX_AGGREGATION_PREREQUISITE_EDGES) { - throw new Error(`prerequisite artifact manifest chain exceeds ${MAX_AGGREGATION_PREREQUISITE_EDGES} edges`); - } - for (const prerequisite of manifest.prerequisite_manifests) { - schedule(expectedById.get(prerequisite.node_id)!, prerequisite.sha256); - } - } - - if (authenticatedManifestCount !== scheduledDigests.size) { - throw new Error(`prerequisite artifact manifest chain authentication was incomplete for ${rootAttemptId}`); - } -} - -function authenticatedEntry( - kind: "generated-test" | "support-file", - entry: GeneratedTestEntry, - publications: ReadonlyMap -): SemanticAggregationSourceEntryContext { - const publication = publications.get(entry.path); - if ( - publication === undefined || - publication.bytes.length > MAX_GENERATED_TEST_COMPANION_BYTES || - publication.bytes.length !== entry.size_bytes || - publication.sha256 !== entry.sha256 - ) { - throw new Error(`verified generated-test companion publication changed ${entry.path}`); - } - return Object.freeze({ - kind, - sourceArtifactPath: publication.absolutePath, - sourceRelativePath: entry.path, - sizeBytes: entry.size_bytes, - sha256: entry.sha256, - bytes: Buffer.from(publication.bytes), - ...(entry.language === undefined ? {} : { language: entry.language }), - ...(entry.description === undefined ? {} : { description: entry.description }), - ...(entry.provenance === undefined ? {} : { provenance: Object.freeze({ ...entry.provenance }) }) - }); -} - -function digest(bytes: Uint8Array): string { - return crypto.createHash("sha256").update(bytes).digest("hex"); -} diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 8bebe8a23..32dbe08c3 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -2237,8 +2237,13 @@ function semanticGateContextForArtifact(input: { input.schemaFilename === "aggregation-manifest.schema.json" ? authenticatedAggregationSemanticContext({ layout: input.layout, - node: input.node, - attemptId: input.attemptId + attemptId: input.attemptId, + producers: finalizedDeclaredContractProducers( + input.layout, + "ultrafuzz/generated-tests@3", + input.node, + input.attemptAuthority + ) }) : undefined; const propertyCampaignEvidence = diff --git a/packages/runtime/test/aggregation-semantic-context.test.ts b/packages/runtime/test/aggregation-semantic-context.test.ts deleted file mode 100644 index ca39e1976..000000000 --- a/packages/runtime/test/aggregation-semantic-context.test.ts +++ /dev/null @@ -1,725 +0,0 @@ -import assert from "node:assert/strict"; -import { temporaryRoot } from "./temporary-root.js"; -import crypto from "node:crypto"; -import fs from "node:fs"; -import path from "node:path"; -import test from "node:test"; - -import { - ARTIFACT_VERIFICATION_SCHEMA_VERSION, - PLANNED_GRAPH_SCHEMA_VERSION, - artifactContractDefinition, - artifactContractSchemaBinding, - createRunLayout, - executeSchemaSemanticGates, - writeArtifact, - writeArtifactManifest, - writeJsonDurable, - type ArtifactContractId, - type ArtifactManifest, - type ArtifactVerificationMarker, - type PlannedGraphDocument, - type PlannedGraphNodeDocument, - type PlannedGraphOutput, - type RunLayout -} from "@ultrafuzz/artifacts"; - -import { authenticatedAggregationSemanticContext } from "../src/aggregation-semantic-context.js"; - -interface AggregationAuthorityFixture { - layout: RunLayout; - aggregationNode: PlannedGraphNodeDocument; - expectedPrerequisiteAttemptIds: string[]; - prerequisiteManifestPaths: string[]; - originManifestPath: string; - unrelatedManifestPath: string; - markerPath: string; - artifactManifestPath: string; - generatedManifestPath: string; - generatedTestPath: string; - supportFilePath: string; -} - -interface RunFileSnapshot { - bytes: Buffer; - dev: bigint; - ino: bigint; - nlink: bigint; - size: bigint; - mtimeNs: bigint; - ctimeNs: bigint; -} - -interface RunTreeSnapshot { - directories: string[]; - files: Map; -} - -test("authenticated aggregation context accepts the exact sealed producer authority", (t) => { - const fixture = createAggregationAuthorityFixture("aggregation-authority-valid"); - t.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - - const before = snapshotRunTree(fixture.layout.root); - const context = authenticatedContext(fixture); - - assert.equal(context.sourceBundles.length, 1); - assert.deepEqual( - context.sourceBundles.map((bundle) => ({ - strategy: bundle.strategy, - sourceAttemptId: bundle.sourceAttemptId, - sourceManifestRelativePath: bundle.sourceManifestRelativePath, - sourceRunId: bundle.sourceRunId, - framework: bundle.framework, - entries: bundle.entries.map((entry) => ({ kind: entry.kind, path: entry.sourceRelativePath })) - })), - [ - { - strategy: "generated-producer", - sourceAttemptId: "generated-producer", - sourceManifestRelativePath: "generated-tests.json", - sourceRunId: "aggregation-authority-valid", - framework: "foundry", - entries: [ - { kind: "generated-test", path: "generated-tests/Property.t.sol" }, - { kind: "support-file", path: "generated-tests/PropertyHelper.sol" } - ] - } - ] - ); - assert.deepEqual(snapshotRunTree(fixture.layout.root), before); -}); - -test("authenticated aggregation source authority rejects hard-linked sealed inputs without mutation", async (t) => { - const cases: Array<{ name: string; select: (fixture: AggregationAuthorityFixture) => string }> = [ - { name: "verification marker", select: (fixture) => fixture.markerPath }, - { name: "artifact manifest", select: (fixture) => fixture.artifactManifestPath }, - { name: "generated-test manifest", select: (fixture) => fixture.generatedManifestPath }, - { name: "declared companion", select: (fixture) => fixture.supportFilePath } - ]; - - for (const [index, attack] of cases.entries()) { - await t.test(attack.name, (subtest) => { - const fixture = createAggregationAuthorityFixture(`aggregation-hardlink-${index}`); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - const target = attack.select(fixture); - const attackerDirectory = path.join(fixture.layout.root, "attacker-links"); - fs.mkdirSync(attackerDirectory); - fs.linkSync(target, path.join(attackerDirectory, `${index}.alias`)); - assert.equal(fs.lstatSync(target, { bigint: true }).nlink, 2n); - - assertAuthorityRejectedWithoutMutation(fixture, /must be a singly linked regular file/u); - assert.equal(fs.lstatSync(target, { bigint: true }).nlink, 2n); - }); - } -}); - -test("authenticated aggregation source authority requires exact marker and artifact-manifest publication sets", async (t) => { - const attacks: Array<{ - name: string; - mutate: (fixture: AggregationAuthorityFixture) => void; - }> = [ - { - name: "marker omits a declared companion", - mutate(fixture) { - mutateMarker(fixture, (marker) => { - marker.publications = marker.publications.filter( - (publication) => publication.path !== "generated-tests/PropertyHelper.sol" - ); - }); - } - }, - { - name: "marker fabricates an extra publication", - mutate(fixture) { - mutateMarker(fixture, (marker) => { - marker.publications.push({ path: "generated-tests/Extra.sol", sha256: "0".repeat(64) }); - }); - } - }, - { - name: "artifact manifest omits a declared companion", - mutate(fixture) { - mutateArtifactManifest(fixture, (manifest) => { - manifest.files = manifest.files.filter((entry) => entry.path !== "generated-tests/PropertyHelper.sol"); - }); - } - }, - { - name: "artifact manifest fabricates an extra publication", - mutate(fixture) { - mutateArtifactManifest(fixture, (manifest) => { - const template = manifest.files.find((entry) => entry.path === "generated-tests/PropertyHelper.sol"); - assert.ok(template); - manifest.files.push({ ...structuredClone(template), path: "generated-tests/Extra.sol" }); - }); - } - } - ]; - - for (const [index, attack] of attacks.entries()) { - await t.test(attack.name, (subtest) => { - const fixture = createAggregationAuthorityFixture(`aggregation-publications-${index}`); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - attack.mutate(fixture); - assertAuthorityRejectedWithoutMutation(fixture, /publication sets differ/u); - }); - } -}); - -test("authenticated aggregation source authority rejects prerequisite fanout attempt-set drift", async (t) => { - const attacks: Array<{ - name: string; - mutate: (fixture: AggregationAuthorityFixture, manifest: ArtifactManifest) => void; - }> = [ - { - name: "one planned prerequisite attempt is omitted", - mutate(fixture, manifest) { - manifest.prerequisite_manifests = manifest.prerequisite_manifests.filter( - (entry) => entry.node_id !== fixture.expectedPrerequisiteAttemptIds[1] - ); - } - }, - { - name: "a planned prerequisite attempt is replaced with a fabricated attempt", - mutate(_fixture, manifest) { - manifest.prerequisite_manifests[1] = { - ...manifest.prerequisite_manifests[1]!, - node_id: "seed__model_9__attempt_9" - }; - } - } - ]; - - for (const [index, attack] of attacks.entries()) { - await t.test(attack.name, (subtest) => { - const fixture = createAggregationAuthorityFixture(`aggregation-prerequisites-${index}`); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - assert.deepEqual( - readJsonFile(fixture.artifactManifestPath).prerequisite_manifests.map( - (entry) => entry.node_id - ), - fixture.expectedPrerequisiteAttemptIds - ); - mutateArtifactManifest(fixture, (manifest) => attack.mutate(fixture, manifest)); - assertAuthorityRejectedWithoutMutation(fixture, /prerequisite set changed/u); - }); - } -}); - -test("authenticated aggregation source authority snapshots the sealed prerequisite manifest chain", async (t) => { - await t.test("hard-linked prerequisite manifest", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-prerequisite-hardlink"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - const manifestPath = fixture.prerequisiteManifestPaths[0]!; - const aliasPath = path.join(fixture.layout.root, "prerequisite-manifest.alias.json"); - fs.linkSync(manifestPath, aliasPath); - - assertAuthorityRejectedWithoutMutation(fixture, /must be a singly linked regular file/u); - assert.equal(fs.lstatSync(manifestPath).nlink, 2); - assert.deepEqual(fs.readFileSync(aliasPath), fs.readFileSync(manifestPath)); - }); - - await t.test("stale prerequisite manifest digest", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-prerequisite-byte-drift"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - const manifestPath = fixture.prerequisiteManifestPaths[0]!; - const manifest = readJsonFile(manifestPath); - manifest.created_at = "2026-08-10T00:00:00.000Z"; - writeJsonForAttack(manifestPath, manifest); - - assertAuthorityRejectedWithoutMutation(fixture, /prerequisite artifact manifest bytes changed/u); - }); -}); - -test("authenticated aggregation source authority reconciles every transitive manifest with the planned graph", async (t) => { - await t.test("stale grandparent bytes", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-stale-grandparent"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - mutateManifestAtPath(fixture.originManifestPath, (manifest) => { - manifest.created_at = "2026-08-10T00:00:01.000Z"; - }); - - assertAuthorityRejectedWithoutMutation(fixture, /prerequisite artifact manifest bytes changed for origin/u); - }); - - await t.test("omitted planned grandparent after attacker reseals descendants", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-omitted-grandparent"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - mutateManifestAtPath(fixture.prerequisiteManifestPaths[0]!, (manifest) => { - manifest.prerequisite_manifests = []; - }); - resealProducerPrerequisiteDigests(fixture); - - assertAuthorityRejectedWithoutMutation( - fixture, - /prerequisite artifact manifest prerequisite set changed for seed__model_0__attempt_0/u - ); - }); - - await t.test("fabricated graph-known nondependency after attacker reseals descendants", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-fabricated-grandparent"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - mutateManifestAtPath(fixture.prerequisiteManifestPaths[0]!, (manifest) => { - manifest.prerequisite_manifests.push({ - node_id: "unrelated-generated-producer", - sha256: digest(fs.readFileSync(fixture.unrelatedManifestPath)) - }); - }); - resealProducerPrerequisiteDigests(fixture); - - assertAuthorityRejectedWithoutMutation( - fixture, - /prerequisite artifact manifest prerequisite set changed for seed__model_0__attempt_0/u - ); - }); - - await t.test("transitive manifest identity drift after attacker reseals the complete chain", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-grandparent-identity"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - mutateManifestAtPath(fixture.originManifestPath, (manifest) => { - manifest.provenance.logical_node_id = "unrelated-generated-producer"; - }); - const originSha256 = digest(fs.readFileSync(fixture.originManifestPath)); - for (const manifestPath of fixture.prerequisiteManifestPaths) { - mutateManifestAtPath(manifestPath, (manifest) => { - manifest.prerequisite_manifests[0]!.sha256 = originSha256; - }); - } - resealProducerPrerequisiteDigests(fixture); - - assertAuthorityRejectedWithoutMutation(fixture, /prerequisite artifact manifest identity changed for origin/u); - }); - - await t.test("diamond ancestors with conflicting sealed digests", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-conflicting-diamond"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - mutateManifestAtPath(fixture.prerequisiteManifestPaths[0]!, (manifest) => { - manifest.prerequisite_manifests[0]!.sha256 = "f".repeat(64); - }); - resealProducerPrerequisiteDigests(fixture); - - assertAuthorityRejectedWithoutMutation( - fixture, - /prerequisite artifact manifest has conflicting sealed digests for origin/u - ); - }); -}); - -test("authenticated aggregation context excludes graph-known generated-test nondependencies", (t) => { - const fixture = createAggregationAuthorityFixture("aggregation-nondependency-producer"); - t.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - - const context = authenticatedContext(fixture); - - assert.deepEqual( - context.sourceBundles.map((bundle) => bundle.sourceAttemptId), - ["generated-producer"] - ); - assert.equal( - fs.existsSync(path.join(fixture.layout.root, ".ultrafuzz-verification", "unrelated-generated-producer.json")), - false - ); -}); - -test("authenticated aggregation context distinguishes valid empty authority from a missing required producer", async (t) => { - await t.test("one authenticated empty bundle remains an explicit typed source bundle", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-empty-bundle"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - makeGeneratedProducerBundleEmpty(fixture); - const before = snapshotRunTree(fixture.layout.root); - - const context = authenticatedContext(fixture); - - assert.equal(context.sourceBundles.length, 1); - assert.equal(context.sourceBundles[0]!.framework, "foundry"); - assert.deepEqual(context.sourceBundles[0]!.entries, []); - assert.deepEqual(snapshotRunTree(fixture.layout.root), before); - }); - - await t.test("an exact graph with no generated-test producers admits the canonical empty aggregation", (subtest) => { - const fixture = createNoGeneratedProducerFixture("aggregation-empty-producer-set"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - const before = snapshotRunTree(fixture.layout.root); - - const context = authenticatedAggregationSemanticContext({ - layout: fixture.layout, - node: fixture.aggregationNode, - attemptId: fixture.aggregationNode.id - }); - const document = { - schema_version: "ultrafuzz.aggregation-manifest.v1", - source_generated_tests: 0, - copied_generated_tests: 0, - source_support_files: 0, - copied_support_files: 0, - source_bundles: [], - files: [], - support_files: [], - skipped_files: [] - }; - const results = executeSchemaSemanticGates("aggregation-manifest.schema.json", { - document, - context: { aggregation: context } - }); - - assert.deepEqual(context.sourceBundles, []); - assert.deepEqual( - results.filter((result) => result.status !== "passed"), - [] - ); - assert.deepEqual(snapshotRunTree(fixture.layout.root), before); - }); - - await t.test("a declared generated-test producer without verifier authority fails closed", (subtest) => { - const fixture = createAggregationAuthorityFixture("aggregation-missing-producer-marker"); - subtest.after(() => fs.rmSync(path.dirname(fixture.layout.root), { recursive: true, force: true })); - fs.rmSync(fixture.markerPath); - - assertAuthorityRejectedWithoutMutation(fixture, /artifact verification marker/u); - }); -}); - -function createNoGeneratedProducerFixture(runId: string): { - layout: RunLayout; - aggregationNode: PlannedGraphNodeDocument; -} { - const outputRoot = temporaryRoot("ultrafuzz-empty-aggregation-authority-"); - const seedNode = plannedNode("seed", [], boundOutput("seed.md", "ultrafuzz/nonempty-markdown@1")); - const aggregationNode = plannedNode( - "aggregate-test-files", - [seedNode.id], - boundOutput("aggregation-manifest.json", "ultrafuzz/aggregation-manifest@1") - ); - const graph: PlannedGraphDocument = { - schema_version: PLANNED_GRAPH_SCHEMA_VERSION, - graph_version: "4", - topology_version: 2, - groups: {}, - nodes: [seedNode, aggregationNode] - }; - const layout = createRunLayout({ outputRoot, runId, graph }); - fs.mkdirSync(path.join(layout.workspacesDir, aggregationNode.id), { recursive: true }); - return { layout, aggregationNode }; -} - -function makeGeneratedProducerBundleEmpty(fixture: AggregationAuthorityFixture): void { - const generatedManifest = readJsonFile<{ - generated_tests: unknown[]; - support_files: unknown[]; - }>(fixture.generatedManifestPath); - generatedManifest.generated_tests = []; - generatedManifest.support_files = []; - writeJsonForAttack(fixture.generatedManifestPath, generatedManifest); - fs.rmSync(fixture.generatedTestPath); - fs.rmSync(fixture.supportFilePath); - - const manifestSha256 = digest(fs.readFileSync(fixture.generatedManifestPath)); - const manifestSize = fs.statSync(fixture.generatedManifestPath).size; - mutateArtifactManifest(fixture, (manifest) => { - manifest.files = manifest.files.filter( - (entry) => entry.path !== "generated-tests/Property.t.sol" && entry.path !== "generated-tests/PropertyHelper.sol" - ); - const generatedManifestEntry = manifest.files.find((entry) => entry.path === "generated-tests.json"); - assert.ok(generatedManifestEntry); - generatedManifestEntry.sha256 = manifestSha256; - generatedManifestEntry.size_bytes = manifestSize; - }); - mutateMarker(fixture, (marker) => { - marker.publications = marker.publications.filter( - (entry) => entry.path !== "generated-tests/Property.t.sol" && entry.path !== "generated-tests/PropertyHelper.sol" - ); - const generatedArtifact = marker.artifacts.find((entry) => entry.path === "generated-tests.json"); - const generatedPublication = marker.publications.find((entry) => entry.path === "generated-tests.json"); - assert.ok(generatedArtifact); - assert.ok(generatedPublication); - generatedArtifact.sha256 = manifestSha256; - generatedPublication.sha256 = manifestSha256; - }); -} - -function resealProducerPrerequisiteDigests(fixture: AggregationAuthorityFixture): void { - const currentDigests = new Map( - fixture.prerequisiteManifestPaths.map((manifestPath) => [ - path.basename(path.dirname(manifestPath)), - digest(fs.readFileSync(manifestPath)) - ]) - ); - mutateArtifactManifest(fixture, (manifest) => { - for (const prerequisite of manifest.prerequisite_manifests) { - const sha256 = currentDigests.get(prerequisite.node_id); - if (sha256 !== undefined) prerequisite.sha256 = sha256; - } - }); -} - -function createAggregationAuthorityFixture(runId: string): AggregationAuthorityFixture { - const outputRoot = temporaryRoot("ultrafuzz-aggregation-authority-"); - const originOutput = boundOutput("origin.md", "ultrafuzz/nonempty-markdown@1"); - const seedOutput = boundOutput("seed.md", "ultrafuzz/nonempty-markdown@1"); - const generatedOutput = boundOutput("generated-tests.json", "ultrafuzz/generated-tests@3"); - const aggregationOutput = boundOutput("aggregation-manifest.json", "ultrafuzz/aggregation-manifest@1"); - const originNode = plannedNode("origin", [], originOutput); - const seedNode: PlannedGraphNodeDocument = { - ...plannedNode("seed", [originNode.id], seedOutput), - model_fanout: [ - { - model_profile_id: "model-zero", - agent_ref: "CodexAgent", - model_index: 0, - loop_index: 0, - attempt_index: 0 - }, - { - model_profile_id: "model-one", - agent_ref: "CodexAgent", - model_index: 1, - loop_index: 0, - attempt_index: 0 - } - ] - }; - const producerNode = plannedNode("generated-producer", [seedNode.id], generatedOutput); - const unrelatedNode = plannedNode("unrelated-generated-producer", [], generatedOutput); - const aggregationNode = plannedNode("aggregate-test-files", [producerNode.id], aggregationOutput); - const graph: PlannedGraphDocument = { - schema_version: PLANNED_GRAPH_SCHEMA_VERSION, - graph_version: "4", - topology_version: 2, - groups: {}, - nodes: [originNode, seedNode, producerNode, unrelatedNode, aggregationNode] - }; - const layout = createRunLayout({ outputRoot, runId, graph }); - writeArtifact(layout, originNode.id, originOutput.path, "origin\n"); - writeArtifactManifest({ - layout, - nodeId: originNode.id, - outputs: [originOutput], - provenance: { - logical_node_id: originNode.logical_id, - attempt_index: 0, - metadata: { concrete_node_id: originNode.id } - } - }); - const originManifestPath = path.join(layout.artifactsDir, originNode.id, "artifact-manifest.json"); - - writeArtifact( - layout, - unrelatedNode.id, - generatedOutput.path, - `${JSON.stringify({ - schema_version: "ultrafuzz.generated-tests.v3", - run_id: runId, - node_id: unrelatedNode.logical_id, - framework: "medusa", - generated_tests: [], - support_files: [] - })}\n` - ); - writeArtifactManifest({ - layout, - nodeId: unrelatedNode.id, - outputs: [generatedOutput], - provenance: { - logical_node_id: unrelatedNode.logical_id, - attempt_index: 0, - metadata: { concrete_node_id: unrelatedNode.id } - } - }); - const unrelatedManifestPath = path.join(layout.artifactsDir, unrelatedNode.id, "artifact-manifest.json"); - - const expectedPrerequisiteAttemptIds = ["seed__model_0__attempt_0", "seed__model_1__attempt_0"]; - const prerequisiteManifestPaths: string[] = []; - for (const [modelIndex, attemptId] of expectedPrerequisiteAttemptIds.entries()) { - writeArtifact(layout, attemptId, seedOutput.path, `seed ${modelIndex}\n`); - writeArtifactManifest({ - layout, - nodeId: attemptId, - outputs: [seedOutput], - prerequisiteNodeIds: [originNode.id], - provenance: { - logical_node_id: seedNode.logical_id, - attempt_index: 0, - model_index: modelIndex, - metadata: { concrete_node_id: seedNode.id } - } - }); - prerequisiteManifestPaths.push(path.join(layout.artifactsDir, attemptId, "artifact-manifest.json")); - } - - const generatedTestContents = Buffer.from("contract Property {}\n", "utf8"); - const supportFileContents = Buffer.from("library PropertyHelper {}\n", "utf8"); - const generatedTestRelativePath = "generated-tests/Property.t.sol"; - const supportFileRelativePath = "generated-tests/PropertyHelper.sol"; - const generatedManifest = { - schema_version: "ultrafuzz.generated-tests.v3", - run_id: runId, - node_id: producerNode.logical_id, - framework: "foundry", - generated_tests: [generatedTestEntry(generatedTestRelativePath, generatedTestContents)], - support_files: [generatedTestEntry(supportFileRelativePath, supportFileContents)] - }; - const generatedTestPath = writeArtifact(layout, producerNode.id, generatedTestRelativePath, generatedTestContents); - const supportFilePath = writeArtifact(layout, producerNode.id, supportFileRelativePath, supportFileContents); - const generatedManifestPath = writeArtifact( - layout, - producerNode.id, - generatedOutput.path, - `${JSON.stringify(generatedManifest)}\n` - ); - const artifactManifest = writeArtifactManifest({ - layout, - nodeId: producerNode.id, - outputs: [generatedOutput], - prerequisiteNodeIds: expectedPrerequisiteAttemptIds, - provenance: { - logical_node_id: producerNode.logical_id, - attempt_index: 0, - metadata: { concrete_node_id: producerNode.id } - } - }); - const generatedManifestPublication = artifactManifest.files.find((entry) => entry.path === generatedOutput.path); - assert.ok(generatedManifestPublication); - const marker: ArtifactVerificationMarker = { - schema_version: ARTIFACT_VERIFICATION_SCHEMA_VERSION, - attempt_id: producerNode.id, - node_id: producerNode.logical_id, - artifacts: [{ ...generatedOutput, sha256: generatedManifestPublication.sha256 }], - publications: artifactManifest.files.map((entry) => ({ path: entry.path, sha256: entry.sha256 })) - }; - const markerPath = path.join(layout.root, ".ultrafuzz-verification", `${producerNode.id}.json`); - writeJsonDurable(markerPath, marker); - - return { - layout, - aggregationNode, - expectedPrerequisiteAttemptIds, - prerequisiteManifestPaths, - originManifestPath, - unrelatedManifestPath, - markerPath, - artifactManifestPath: path.join(layout.artifactsDir, producerNode.id, "artifact-manifest.json"), - generatedManifestPath, - generatedTestPath, - supportFilePath - }; -} - -function plannedNode(id: string, dependsOn: string[], output: PlannedGraphOutput): PlannedGraphNodeDocument { - return { - id, - logical_id: id, - display_name: id, - kind: "agentic", - depends_on: dependsOn, - artifact_dir: `artifacts/${id}`, - outputs: [output], - prompt_id: id, - prompt_path: `prompts/${id}.md`, - loop: { index: 0, count: 1, mode: "parallel", attempt_index: 0 }, - model_fanout: [] - }; -} - -function boundOutput(artifactPath: string, contract: ArtifactContractId): PlannedGraphOutput { - return { - path: artifactPath, - contract, - contract_digest: artifactContractDefinition(contract).digest, - ...(artifactContractSchemaBinding(contract) ?? {}), - primary: true - }; -} - -function generatedTestEntry( - entryPath: string, - contents: Buffer -): { - path: string; - size_bytes: number; - sha256: string; -} { - return { - path: entryPath, - size_bytes: contents.byteLength, - sha256: digest(contents) - }; -} - -function authenticatedContext(fixture: AggregationAuthorityFixture) { - return authenticatedAggregationSemanticContext({ - layout: fixture.layout, - node: fixture.aggregationNode, - attemptId: fixture.aggregationNode.id - }); -} - -function assertAuthorityRejectedWithoutMutation(fixture: AggregationAuthorityFixture, expected: RegExp): void { - const before = snapshotRunTree(fixture.layout.root); - assert.throws(() => authenticatedContext(fixture), expected); - assert.deepEqual(snapshotRunTree(fixture.layout.root), before); -} - -function mutateMarker( - fixture: AggregationAuthorityFixture, - mutate: (marker: ArtifactVerificationMarker) => void -): void { - const marker = readJsonFile(fixture.markerPath); - mutate(marker); - writeJsonForAttack(fixture.markerPath, marker); -} - -function mutateArtifactManifest( - fixture: AggregationAuthorityFixture, - mutate: (manifest: ArtifactManifest) => void -): void { - mutateManifestAtPath(fixture.artifactManifestPath, mutate); -} - -function mutateManifestAtPath(filePath: string, mutate: (manifest: ArtifactManifest) => void): void { - const manifest = readJsonFile(filePath); - mutate(manifest); - writeJsonForAttack(filePath, manifest); -} - -function readJsonFile(filePath: string): T { - return JSON.parse(fs.readFileSync(filePath, "utf8")) as T; -} - -function writeJsonForAttack(filePath: string, value: unknown): void { - fs.writeFileSync(filePath, `${JSON.stringify(value)}\n`, "utf8"); -} - -function snapshotRunTree(root: string): RunTreeSnapshot { - const directories: string[] = []; - const files = new Map(); - const visit = (directory: string): void => { - for (const entry of fs - .readdirSync(directory, { withFileTypes: true }) - .sort((left, right) => left.name.localeCompare(right.name))) { - const absolutePath = path.join(directory, entry.name); - const relativePath = path.relative(root, absolutePath).split(path.sep).join("/"); - if (entry.isDirectory()) { - directories.push(relativePath); - visit(absolutePath); - continue; - } - const stat = fs.lstatSync(absolutePath, { bigint: true }); - assert.equal(stat.isFile(), true, `unexpected non-file in authority fixture: ${relativePath}`); - files.set(relativePath, { - bytes: fs.readFileSync(absolutePath), - dev: stat.dev, - ino: stat.ino, - nlink: stat.nlink, - size: stat.size, - mtimeNs: stat.mtimeNs, - ctimeNs: stat.ctimeNs - }); - } - }; - visit(root); - return { directories, files }; -} - -function digest(bytes: Uint8Array): string { - return crypto.createHash("sha256").update(bytes).digest("hex"); -} diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 703a32d38..142c562da 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -617,7 +617,7 @@ function writeSealedFixtureTaskAuthority( suppliedTasks?: readonly SmithersTaskManifestTask[] ): void { const sealedNodes = nodes.map((node) => { - if (node.kind !== "agentic" || node.workflow !== undefined) return node; + if (node.kind !== "agentic" || node.workflow !== undefined || node.dynamic !== undefined) return node; const attempts = fixtureAttemptIds(node); return { ...node, @@ -3625,6 +3625,281 @@ for (const markerAuthority of ["dangling leaf", "symlinked root"] as const) { }); } +const AGGREGATION_FIXTURE_GROUPS = { + strategies: { defaults: { failure_policy: "continue" } }, + goals: { defaults: { failure_policy: "continue" } }, + review: {} +}; +const AGGREGATION_DYNAMIC_STORAGE_ID = "dynamic-goal-template-0123456789abcdef0123456789abcdef"; + +interface AggregationFixtureSource { + attemptId: string; + logicalNodeId: string; + manifestPath: string; + manifestSha256: string; + testRelativePath: string; + testPath: string; + testBytes: Buffer; +} + +function aggregationFixtureNode( + id: string, + options: { group?: string; dependsOn?: string[]; outputs?: PlannedGraphNode["outputs"] } = {} +): PlannedGraphNode { + return { + ...plannedNode([]), + id, + logical_id: id, + display_name: id, + ...(options.group === undefined ? {} : { group: options.group }), + depends_on: options.dependsOn ?? [], + artifact_dir: `artifacts/${id}`, + prompt_id: id, + prompt_path: `strategies/${id}.md`, + outputs: options.outputs ?? [boundOutput("generated-tests.json", "ultrafuzz/generated-tests@3", true)] + }; +} + +/** A goal template and one materialized goal, with the identity split `materializeDynamicRuntime` writes. */ +function aggregationDynamicGoalNodes(sourceNodeId: string): { + template: PlannedGraphNode; + generated: PlannedGraphNode; +} { + const template: PlannedGraphNode = { + ...aggregationFixtureNode("goal-template", { group: "goals", dependsOn: [sourceNodeId] }), + dynamic: { + from: { node: sourceNodeId, path: "goal-plan.md" }, + key: "id", + node_id: "dynamic:threat:{{ item.id }}", + status: "pending" + } + }; + const generated: PlannedGraphNode = { + ...template, + id: "dynamic:threat:t1", + display_name: "goal-template: t1", + artifact_dir: `artifacts/${AGGREGATION_DYNAMIC_STORAGE_ID}`, + artifact_dirs: [`artifacts/${AGGREGATION_DYNAMIC_STORAGE_ID}`], + model_fanout: [ + { + attempt_id: AGGREGATION_DYNAMIC_STORAGE_ID, + model_profile_id: "default", + agent_ref: "CodexAgent", + model_name: "gpt-test", + reasoning_effort: "high", + model_index: 0, + loop_index: 0, + attempt_index: 0 + } + ], + workflow: { + node_id: `node:${AGGREGATION_DYNAMIC_STORAGE_ID}`, + task_node_ids: [`node:${AGGREGATION_DYNAMIC_STORAGE_ID}`] + }, + dynamic: undefined, + dynamic_generated: { + group_node_id: template.id, + source_node_id: sourceNodeId, + source_attempt_id: sourceNodeId, + expansion_key: "t1", + item_sha256: "a".repeat(64), + storage_id: AGGREGATION_DYNAMIC_STORAGE_ID, + manifest_path: `dynamic-expansions/${template.id}.json` + } + }; + return { template, generated }; +} + +/** Publish and controller-finalize a generated-tests producer attempt holding one test file. */ +function finalizeGeneratedTestsProducer( + layout: ReturnType, + node: PlannedGraphNode, + attemptId = node.id +): AggregationFixtureSource { + const testRelativePath = `generated-tests/${node.logical_id}.t.sol`; + const testBytes = Buffer.from(`contract GeneratedBy${attemptId.length} {}\n`, "utf8"); + registerArtifactNode(layout, attemptId, node.outputs); + const testPath = writeArtifactFile(layout, attemptId, testRelativePath, testBytes); + const manifestPath = writeArtifactFile( + layout, + attemptId, + "generated-tests.json", + JSON.stringify({ + schema_version: "ultrafuzz.generated-tests.v3", + run_id: layout.runId, + node_id: node.logical_id, + framework: "foundry", + generated_tests: [ + { + path: testRelativePath, + size_bytes: testBytes.byteLength, + sha256: createHash("sha256").update(testBytes).digest("hex") + } + ], + support_files: [] + }) + ); + finalizeArtifactNode(layout, attemptId, node.outputs, { concreteNodeId: node.id, logicalNodeId: node.logical_id }, [ + testRelativePath + ]); + return { + attemptId, + logicalNodeId: node.logical_id, + manifestPath, + manifestSha256: createHash("sha256").update(fs.readFileSync(manifestPath)).digest("hex"), + testRelativePath, + testPath, + testBytes + }; +} + +/** The aggregation manifest an agent writes after copying every listed source bundle into its workspace. */ +function writeAggregationManifestCopying( + layout: ReturnType, + aggregationNode: PlannedGraphNode, + sources: readonly AggregationFixtureSource[] +): void { + const workspace = path.join(layout.workspacesDir, aggregationNode.id); + fs.mkdirSync(workspace, { recursive: true }); + const bundleIdentity = (source: AggregationFixtureSource) => ({ + strategy: source.logicalNodeId, + node_id: source.logicalNodeId, + source_attempt_id: source.attemptId, + attempt_index: 0, + source_manifest_path: source.manifestPath, + source_manifest_relative_path: "generated-tests.json", + source_manifest_sha256: source.manifestSha256 + }); + const files = sources.map((source) => { + const destinationRelativePath = `test/foundry/${source.logicalNodeId}/attempt-0/${path.basename(source.testPath)}`; + const destinationPath = path.join(workspace, ...destinationRelativePath.split("/")); + fs.mkdirSync(path.dirname(destinationPath), { recursive: true }); + fs.writeFileSync(destinationPath, source.testBytes); + return { + ...bundleIdentity(source), + source_artifact_path: source.testPath, + source_relative_path: source.testRelativePath, + destination_path: destinationPath, + destination_relative_path: destinationRelativePath, + size_bytes: source.testBytes.byteLength, + sha256: createHash("sha256").update(source.testBytes).digest("hex") + }; + }); + writeArtifactFile( + layout, + aggregationNode.id, + "aggregation.json", + JSON.stringify({ + schema_version: "ultrafuzz.aggregation-manifest.v1", + source_generated_tests: sources.length, + copied_generated_tests: sources.length, + source_support_files: 0, + copied_support_files: 0, + source_bundles: sources.map((source) => ({ + ...bundleIdentity(source), + source_run_id: layout.runId, + framework: "foundry", + generated_test_count: 1, + support_file_count: 0, + disposition: "copied" + })), + files, + support_files: [], + skipped_files: [] + }) + ); +} + +/** Run the host gate the way finalization does: on the persisted graph node with its sealed task set. */ +function verifyAggregationAttempt( + layout: ReturnType, + aggregationTask: SmithersTaskManifestTask, + tasks: SmithersTaskManifestTask[], + admittedDependencyAttemptIds: string[] +): ReturnType { + const graph = JSON.parse(fs.readFileSync(layout.graphPath, "utf8")) as PlannedGraph; + const node = graph.nodes.find((candidate) => candidate.id === aggregationTask.concreteNodeId); + assert.ok(node); + return verifyRuntimeRequiredArtifactsForAttempt( + layout, + node, + aggregationTask.attemptId, + { task: aggregationTask, tasks, admittedDependencyAttemptIds }, + authenticatedSnapshotsForNode(layout, node, aggregationTask.attemptId) + ); +} + +test("host aggregation intake skips a failed optional generated-tests producer the verifier did not admit", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-aggregation-failed-optional" }); + const verified = aggregationFixtureNode("strategy-a", { group: "strategies" }); + const failed = aggregationFixtureNode("strategy-b", { group: "strategies" }); + const aggregationNode = aggregationFixtureNode("aggregate-test-files", { + group: "review", + dependsOn: [verified.id, failed.id], + outputs: [boundOutput("aggregation.json", "ultrafuzz/aggregation-manifest@1", true)] + }); + const nodes = [verified, failed, aggregationNode]; + writePlannedGraph(layout, nodes, AGGREGATION_FIXTURE_GROUPS); + const verifiedTask = sealedTaskForNode(layout, verified); + const failedTask = sealedTaskForNode(layout, failed); + const aggregationTask: SmithersTaskManifestTask = { + ...sealedTaskForNode(layout, aggregationNode, [verifiedTask, failedTask]), + optionalDependencyArtifactDirs: [verifiedTask.artifactDir, failedTask.artifactDir] + }; + const tasks = [verifiedTask, failedTask, aggregationTask]; + const source = finalizeGeneratedTestsProducer(layout, verified); + // The continue-policy strategy failed before its verifier wrote a marker. + registerArtifactNode(layout, failedTask.attemptId, failed.outputs); + updateNodeState(layout, failedTask.attemptId, { status: "failed" }); + writeSealedFixtureTaskAuthority(layout, nodes, tasks); + + writeAggregationManifestCopying(layout, aggregationNode, [source]); + const result = verifyAggregationAttempt(layout, aggregationTask, tasks, [verifiedTask.attemptId]); + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); + + // The admitted producer is still part of the authority: omitting it fails. + writeAggregationManifestCopying(layout, aggregationNode, []); + const omitted = verifyAggregationAttempt(layout, aggregationTask, tasks, [verifiedTask.attemptId]); + assert.equal(omitted.ok, false); + assert.ok( + omitted.diagnostics.some((diagnostic) => diagnostic.message.includes("omits authenticated source bundle")), + JSON.stringify(omitted.diagnostics) + ); +}); + +test("host aggregation intake keys a directly consumed dynamic producer by its storage attempt ID", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-aggregation-direct-dynamic" }); + const source = aggregationFixtureNode("goal-plan", { + outputs: [boundOutput("goal-plan.md", "ultrafuzz/nonempty-markdown@1", true)] + }); + const { template, generated } = aggregationDynamicGoalNodes(source.id); + const aggregationNode = aggregationFixtureNode("aggregate-test-files", { + group: "review", + dependsOn: [generated.id], + outputs: [boundOutput("aggregation.json", "ultrafuzz/aggregation-manifest@1", true)] + }); + const nodes = [source, template, generated, aggregationNode]; + writePlannedGraph(layout, nodes, AGGREGATION_FIXTURE_GROUPS); + const sourceTask = sealedTaskForNode(layout, source); + const generatedTask = sealedTaskForNode(layout, generated, [sourceTask], AGGREGATION_DYNAMIC_STORAGE_ID); + const aggregationTask: SmithersTaskManifestTask = { + ...sealedTaskForNode(layout, aggregationNode, [generatedTask]), + optionalDependencyArtifactDirs: [generatedTask.artifactDir] + }; + const tasks = [sourceTask, generatedTask, aggregationTask]; + writeDeclaredArtifactNode(layout, sourceTask.attemptId, source.outputs, { "goal-plan.md": "# Goal plan\n" }); + const dynamicSource = finalizeGeneratedTestsProducer(layout, generated, AGGREGATION_DYNAMIC_STORAGE_ID); + writeSealedFixtureTaskAuthority(layout, nodes, tasks); + + writeAggregationManifestCopying(layout, aggregationNode, [dynamicSource]); + const result = verifyAggregationAttempt(layout, aggregationTask, tasks, [ + sourceTask.attemptId, + AGGREGATION_DYNAMIC_STORAGE_ID + ]); + + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); +}); + test("review lifecycle and strategy gates authenticate every dedupe, triage, and severity transition", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-review-lifecycle" }); const artifactPaths = { From 21924548e15ff6ed8565c824015897e0247a6c67 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:42:29 +0000 Subject: [PATCH 031/206] fix(runtime): accept the sealed closure dynamic lowering gives deeper descendants assertExactSealedAttemptAuthority requires a consumer's sealed dependencyArtifactDirs to equal the attempts of every ancestor in the current planned graph. After a dynamic group expands, the lowered graph reaches the generated attempts from every descendant of the group's direct dependents, but lowerTaskDynamicDependencies adds those attempts only to the direct dependents' sealed tasks. Every deeper descendant therefore failed each contextual host gate that consults the sealed authority, even though the in-workflow verifier had admitted it through the same sealed closure: sealed Smithers attempt "" artifact ancestor closure does not match the exact planned set; missing: ".../artifacts/dynamic-..." In the default topology dedupe-findings is the direct dependent of the goal groups, so triage, severity-classification, aggregate-test-files and final-report are such descendants. I reproduced the closure shape with startRun plus materializeDynamicRuntime on a planner -> dynamic fan-out -> join -> triage-shaped grandchild topology: the grandchild's sealed closure is [planner, join], the lowered graph's is [planner, join, dynamic:threat:], and verifyRequiredArtifactsForAttempt failed with the message above. Expect a generated ancestor's attempts only when the sealed closure names them. Static ancestors must still all be present and nothing outside the planned closure is accepted, so the existing fan-in omitted-ancestor and unrelated-ancestor checks keep failing as before. Changing the lowering itself would rewrite in-flight runs' sealed task plans and widen what descendants must admit, so it is left alone here. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/artifact-gates.ts | 18 +++-- packages/runtime/test/artifact-gates.test.ts | 75 ++++++++++++++++++++ 2 files changed, 87 insertions(+), 6 deletions(-) diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 32dbe08c3..248e32a21 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -3279,13 +3279,19 @@ function assertExactSealedAttemptAuthority( } const ancestorNodeIds = plannedAncestorIds(graph, plannedConsumer); - const expectedAncestorAttempts = graph.nodes - .filter((node) => ancestorNodeIds.has(node.id)) - .flatMap(plannedAttemptIdsForAuthority); - const expectedAncestorDirectories = expectedAncestorAttempts.map((attemptId) => - getNodeArtifactDir(layout, attemptId) - ); const actualAncestorDirectories = authority.task.dependencyArtifactDirs.map((directory) => path.resolve(directory)); + const sealedAncestorDirectories = new Set(actualAncestorDirectories); + // Dynamic lowering extends only a group's direct dependents, so a deeper + // descendant's sealed closure omits the generated attempts the lowered graph + // reaches through them. Gate contexts read ancestors through the sealed + // closure, exactly as the in-workflow verifier admitted them. + const expectedAncestorDirectories = graph.nodes + .filter((node) => ancestorNodeIds.has(node.id)) + .flatMap((node) => + plannedAttemptIdsForAuthority(node) + .map((attemptId) => getNodeArtifactDir(layout, attemptId)) + .filter((directory) => node.dynamic_generated === undefined || sealedAncestorDirectories.has(directory)) + ); assertExactStringSet( actualAncestorDirectories, expectedAncestorDirectories, diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 142c562da..17e084a3f 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -3900,6 +3900,81 @@ test("host aggregation intake keys a directly consumed dynamic producer by its s assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); }); +test("host aggregation intake follows the sealed closure below a dynamic group's direct dependent", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-aggregation-transitive-dynamic" }); + const source = aggregationFixtureNode("goal-plan", { + outputs: [boundOutput("goal-plan.md", "ultrafuzz/nonempty-markdown@1", true)] + }); + const { template, generated } = aggregationDynamicGoalNodes(source.id); + const join = aggregationFixtureNode("goal-join", { + group: "review", + dependsOn: [generated.id], + outputs: [boundOutput("goal-join.md", "ultrafuzz/nonempty-markdown@1", true)] + }); + const strategy = aggregationFixtureNode("strategy-a", { group: "strategies" }); + const unrelated = aggregationFixtureNode("strategy-z", { group: "strategies" }); + const aggregationNode = aggregationFixtureNode("aggregate-test-files", { + group: "review", + dependsOn: [join.id, strategy.id], + outputs: [boundOutput("aggregation.json", "ultrafuzz/aggregation-manifest@1", true)] + }); + const nodes = [source, template, generated, join, strategy, unrelated, aggregationNode]; + writePlannedGraph(layout, nodes, AGGREGATION_FIXTURE_GROUPS); + const sourceTask = sealedTaskForNode(layout, source); + const generatedTask = sealedTaskForNode(layout, generated, [sourceTask], AGGREGATION_DYNAMIC_STORAGE_ID); + const joinTask: SmithersTaskManifestTask = { + ...sealedTaskForNode(layout, join, [generatedTask]), + optionalDependencyArtifactDirs: [generatedTask.artifactDir] + }; + const strategyTask = sealedTaskForNode(layout, strategy); + const unrelatedTask = sealedTaskForNode(layout, unrelated); + // Dynamic lowering extends only a group's direct dependents, so the sealed + // closure of this grandchild keeps its compile-time ancestors and never + // names the generated attempt that the lowered planned graph reaches. + const aggregationTaskWithClosure = (closure: readonly SmithersTaskManifestTask[]): SmithersTaskManifestTask => ({ + ...smithersTaskForNode({ + layout, + node: aggregationNode, + attemptId: aggregationNode.id, + dependencies: [joinTask.attemptId, strategyTask.attemptId], + dependencyArtifactDirs: closure.map((task) => task.artifactDir) + }), + optionalDependencyArtifactDirs: closure + .filter((task) => task.metadata.node.group === "strategies") + .map((task) => task.artifactDir) + }); + const aggregationTask = aggregationTaskWithClosure([sourceTask, joinTask, strategyTask]); + const tasks = [sourceTask, generatedTask, joinTask, strategyTask, unrelatedTask, aggregationTask]; + writeDeclaredArtifactNode(layout, sourceTask.attemptId, source.outputs, { "goal-plan.md": "# Goal plan\n" }); + finalizeGeneratedTestsProducer(layout, generated, AGGREGATION_DYNAMIC_STORAGE_ID); + const strategySource = finalizeGeneratedTestsProducer(layout, strategy); + writeSealedFixtureTaskAuthority(layout, nodes, tasks); + + writeAggregationManifestCopying(layout, aggregationNode, [strategySource]); + const admitted = [sourceTask.attemptId, joinTask.attemptId, strategyTask.attemptId]; + const result = verifyAggregationAttempt(layout, aggregationTask, tasks, admitted); + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); + + // Only generated attempts may be absent; the closure still may not reach + // outside the planned ancestors. + const widenedTask = aggregationTaskWithClosure([sourceTask, joinTask, strategyTask, unrelatedTask]); + const widened = verifyAggregationAttempt( + layout, + widenedTask, + tasks.map((task) => (task === aggregationTask ? widenedTask : task)), + admitted + ); + assert.equal(widened.ok, false); + assert.ok( + widened.diagnostics.some((diagnostic) => + diagnostic.message.includes( + `artifact ancestor closure does not match the exact planned set; unexpected: ${JSON.stringify(unrelatedTask.artifactDir)}` + ) + ), + JSON.stringify(widened.diagnostics) + ); +}); + test("review lifecycle and strategy gates authenticate every dedupe, triage, and severity transition", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-review-lifecycle" }); const artifactPaths = { From 24b025e0a63a346588933f8a0983639267577b8b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:43:17 +0000 Subject: [PATCH 032/206] fix(runtime): delete the host-only severity matrix gate severity-matrix.ts ran only in the host re-check, keyed on a node whose logical ID is severity-classification or that declares report@3, and required machine-readable severity, impact and likelihood on every record. severity-classified-findings.schema.json requires those fields only for true positives, and the severity prompt keeps false-positive and non-production records in the file while forbidding a guess at their impact or likelihood. A compliant file with one false-positive record therefore passed JSON Schema and every registry gate inside the attempt, then failed host finalization with three SEVERITY_MATRIX_FIELD_MISSING errors, which blocks aggregate-test-files and final-report behind it. Everything else the module checked is already enforced where the verifier runs it too: both schemas are additionalProperties: false with High/Medium/Low enums (legacy aliases, lowercase levels, wrappers and non-object records), the severity-classification-matrix registry gate checks the matrix whenever all three fields are present, and in strict severity-handoff mode report rows must copy those fields exactly from the matrix-checked classified findings. One check has no replacement: report rows authored in bounded classification mode (topologies without a severity-classification node, such as smoke) are no longer matrix-checked anywhere. That mismatch now lets a report through instead of costing the run its final report; if it should be enforced, it belongs in the bounded report gate where the verifier can retry it. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/artifact-gates.ts | 37 ----- packages/runtime/src/index.ts | 1 - packages/runtime/src/severity-matrix.ts | 148 ------------------ packages/runtime/test/artifact-gates.test.ts | 51 ++++++ packages/runtime/test/severity-matrix.test.ts | 102 ------------ 5 files changed, 51 insertions(+), 288 deletions(-) delete mode 100644 packages/runtime/src/severity-matrix.ts delete mode 100644 packages/runtime/test/severity-matrix.test.ts diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 248e32a21..709644224 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -99,7 +99,6 @@ import { parseRuntimeDocumentBytes } from "./runtime-document-codec.js"; import type { PlannedGraph, PlannedGraphNode, RuntimeDiagnostic } from "./types.js"; import { topologyRuntimeBudgetForTimeout } from "./topology-runtime-budget.js"; import { diagnosticFromError } from "./utils.js"; -import { validateSeverityMatrixArtifact, type SeverityArtifactKind } from "./severity-matrix.js"; import { deriveWorkspacePatchGitFacts } from "./workspace-handoff.js"; import { loadFinalizedNodeOutputSnapshot, type VerifiedOutputArtifactSnapshot } from "./verified-output.js"; import { renderCoverageEvidenceMarkdownSection } from "./final-report-markdown.js"; @@ -422,7 +421,6 @@ export function verifyRequiredArtifactsForAttempt( diagnostics.push(diagnosticFromError(error, "artifact-gates", "REQUIRED_ARTIFACT_INVALID")); } } - diagnostics.push(...verifySeverityMatrixArtifacts(artifactDir, node, authenticated)); try { diagnostics.push(...verifyInvariantEvidenceArtifacts(layout, artifactDir, node, attemptAuthority, authenticated)); } catch (error) { @@ -3831,41 +3829,6 @@ function sealedArtifactSchemaPath(layout: RunLayout, schemaFile: string): string return schemaPath; } -function verifySeverityMatrixArtifacts( - artifactDir: string, - node: PlannedGraphNode, - authenticated?: AuthenticatedArtifactGateSnapshots -): RuntimeDiagnostic[] { - const artifact = severityArtifactForNode(node); - if (artifact === undefined) { - return []; - } - const artifactPath = safeResolveInside(artifactDir, artifact.path, "severity artifact output"); - try { - const document = parseCurrentArtifactJson(artifactDir, artifactPath, authenticated); - if (document === undefined) return []; - return validateSeverityMatrixArtifact({ - artifact: document, - artifactPath, - kind: artifact.kind - }); - } catch (error) { - return [diagnosticFromError(error, "severity-matrix", "SEVERITY_ARTIFACT_READ_FAILED")]; - } -} - -function severityArtifactForNode(node: PlannedGraphNode): { kind: SeverityArtifactKind; path: string } | undefined { - const logicalId = node.logical_id ?? node.id; - if (logicalId === "severity-classification") { - return { kind: "severity-classification", path: "severity-classified-findings.json" }; - } - const reportOutputs = node.outputs.filter((output) => output.contract === "ultrafuzz/report@3"); - if (reportOutputs.length === 1) { - return { kind: "final-report", path: reportOutputs[0]!.path }; - } - return undefined; -} - const RECON_MAX_TEST_LIMIT = "18446744073709551615"; const RECON_STATEFUL_SEQUENCE_LENGTH = 100; const CAMPAIGN_HOST_FORCE_KILL_GRACE_SECONDS = 300; diff --git a/packages/runtime/src/index.ts b/packages/runtime/src/index.ts index 856193a23..cb8766908 100644 --- a/packages/runtime/src/index.ts +++ b/packages/runtime/src/index.ts @@ -29,7 +29,6 @@ export * from "./runtime-semantic-gates.js"; export * from "./schema-registry.js"; export * from "./semantic-artifact-context.js"; export * from "./semantic-gates.js"; -export * from "./severity-matrix.js"; export * from "./smithers.js"; export * from "./smithers-package.js"; export * from "./smithers-attempt-authority.js"; diff --git a/packages/runtime/src/severity-matrix.ts b/packages/runtime/src/severity-matrix.ts deleted file mode 100644 index 237d0a08d..000000000 --- a/packages/runtime/src/severity-matrix.ts +++ /dev/null @@ -1,148 +0,0 @@ -import type { RuntimeDiagnostic } from "./types.js"; - -export type SeverityLevel = "High" | "Medium" | "Low"; -export type SeverityArtifactKind = "severity-classification" | "final-report"; - -const LEVELS: readonly SeverityLevel[] = ["High", "Medium", "Low"]; -const LEGACY_FIELDS = ["final_severity", "impact_level", "likelihood_level", "severity_classification"] as const; - -export function expectedSeverityFromMatrix(impact: unknown, likelihood: unknown): SeverityLevel | undefined { - if (!isSeverityLevel(impact) || !isSeverityLevel(likelihood)) { - return undefined; - } - if (impact === "Low") { - return "Low"; - } - if (impact === "Medium") { - return likelihood === "Low" ? "Low" : "Medium"; - } - return likelihood === "Low" ? "Medium" : "High"; -} - -export function validateSeverityMatrixArtifact(input: { - artifact: unknown; - artifactPath: string; - kind: SeverityArtifactKind; -}): RuntimeDiagnostic[] { - const records = recordsForArtifact(input.artifact, input.kind); - if (records === undefined) { - return [ - { - code: "SEVERITY_ARTIFACT_SHAPE_INVALID", - message: - input.kind === "final-report" - ? "report.json must contain an issues array" - : "severity-classified-findings.json must be an array", - severity: "error", - source: "severity-matrix", - path: input.artifactPath - } - ]; - } - - return records.flatMap(({ value, path }) => { - if (!isRecord(value)) { - return [ - { - code: "SEVERITY_RECORD_SHAPE_INVALID", - message: `${path} must be an object`, - severity: "error" as const, - source: "severity-matrix", - path: `${input.artifactPath}#/${path}` - } - ]; - } - return validateRecord(value, path, input.artifactPath); - }); -} - -function recordsForArtifact( - artifact: unknown, - kind: SeverityArtifactKind -): Array<{ value: unknown; path: string }> | undefined { - if (kind === "final-report") { - if (!isRecord(artifact) || !Array.isArray(artifact.issues)) { - return undefined; - } - return artifact.issues.map((value, index) => ({ value, path: `issues.${index}` })); - } - if (!Array.isArray(artifact)) { - return undefined; - } - return artifact.map((value, index) => ({ value, path: `${index}` })); -} - -function validateRecord( - record: Record, - recordPath: string, - artifactPath: string -): RuntimeDiagnostic[] { - const diagnostics: RuntimeDiagnostic[] = []; - - for (const field of LEGACY_FIELDS) { - if (field in record) { - diagnostics.push({ - code: "SEVERITY_FIELD_ALIAS_UNSUPPORTED", - message: `${recordPath}.${field} is not supported; use the canonical severity, impact, and likelihood fields`, - severity: "error", - source: "severity-matrix", - path: `${artifactPath}#/${recordPath}.${field}` - }); - } - } - - const severity = validateLevel(record, "severity", recordPath, artifactPath, diagnostics); - const impact = validateLevel(record, "impact", recordPath, artifactPath, diagnostics); - const likelihood = validateLevel(record, "likelihood", recordPath, artifactPath, diagnostics); - const expected = expectedSeverityFromMatrix(impact, likelihood); - if (severity !== undefined && expected !== undefined && severity !== expected) { - diagnostics.push({ - code: "SEVERITY_MATRIX_MISMATCH", - message: `${recordPath} has severity ${severity}, impact ${impact}, likelihood ${likelihood}; expected ${expected}`, - severity: "error", - source: "severity-matrix", - path: `${artifactPath}#/${recordPath}` - }); - } - - return diagnostics; -} - -function validateLevel( - record: Record, - field: "severity" | "impact" | "likelihood", - recordPath: string, - artifactPath: string, - diagnostics: RuntimeDiagnostic[] -): SeverityLevel | undefined { - const value = record[field]; - if (value === undefined || value === null) { - diagnostics.push({ - code: "SEVERITY_MATRIX_FIELD_MISSING", - message: `${recordPath} has no machine-readable ${field}`, - severity: "error", - source: "severity-matrix", - path: `${artifactPath}#/${recordPath}` - }); - return undefined; - } - if (!isSeverityLevel(value)) { - diagnostics.push({ - code: "SEVERITY_LEVEL_INVALID", - message: `${recordPath}.${field} must be one of ${LEVELS.join(", ")}, got ${JSON.stringify(value)}`, - severity: "error", - source: "severity-matrix", - path: `${artifactPath}#/${recordPath}.${field}` - }); - return undefined; - } - return value; -} - -function isSeverityLevel(value: unknown): value is SeverityLevel { - return value === "High" || value === "Medium" || value === "Low"; -} - -function isRecord(value: unknown): value is Record { - return value !== null && typeof value === "object" && !Array.isArray(value); -} diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 17e084a3f..07cbec291 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -2508,6 +2508,57 @@ test("severity classification gates preserve triaged fields and enforce the fina ); }); +test("severity classification keeps a false-positive record without severity, impact, or likelihood", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-severity-false-positive" }); + const triageNode: PlannedGraphNode = { + ...plannedNode([]), + id: "triage", + logical_id: "triage", + artifact_dir: "artifacts/triage", + outputs: [boundOutput("triaged-findings.json", "ultrafuzz/triaged-findings@1", true)] + }; + const severityNode: PlannedGraphNode = { + ...plannedNode([]), + id: "severity-classification", + logical_id: "severity-classification", + depends_on: [triageNode.id], + artifact_dir: "artifacts/severity-classification", + outputs: [boundOutput("severity-classified-findings.json", "ultrafuzz/severity-classified-findings@1", true)] + }; + const nodes = [triageNode, severityNode]; + writePlannedGraph(layout, nodes); + const triageTask = sealedTaskForNode(layout, triageNode); + const severityTask = sealedTaskForNode(layout, severityNode, [triageTask]); + const tasks = [triageTask, severityTask]; + // The schema requires the severity fields only for true positives, and the + // prompt forbids guessing them for records it keeps but does not promote. + const falsePositive = currentFinding("finding-unreachable", { + status: "false-positive", + triage_classification: "false-positive", + notes: [ + "triage_reason=the state is unreachable through the public path", + "demotion_reason=no public entrypoint reaches the failing state" + ] + }); + writeDeclaredArtifactNode(layout, triageTask.attemptId, triageNode.outputs, { + "triaged-findings.json": JSON.stringify([falsePositive]) + }); + writeDeclaredArtifactNode(layout, severityTask.attemptId, severityNode.outputs, { + "severity-classified-findings.json": JSON.stringify([falsePositive]) + }); + writeSealedFixtureTaskAuthority(layout, nodes, tasks); + + const result = verifyRuntimeRequiredArtifactsForAttempt( + layout, + severityNode, + severityTask.attemptId, + { task: severityTask, tasks }, + authenticatedSnapshotsForNode(layout, severityNode, severityTask.attemptId) + ); + + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); +}); + test("triage gates preserve every deduped finding and upstream note", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-triage-preservation" }); const dedupeNode: PlannedGraphNode = { diff --git a/packages/runtime/test/severity-matrix.test.ts b/packages/runtime/test/severity-matrix.test.ts deleted file mode 100644 index e48164450..000000000 --- a/packages/runtime/test/severity-matrix.test.ts +++ /dev/null @@ -1,102 +0,0 @@ -import assert from "node:assert/strict"; -import test from "node:test"; - -import { expectedSeverityFromMatrix, validateSeverityMatrixArtifact } from "../src/severity-matrix.js"; - -test("severity matrix accepts only exact canonical levels", () => { - assert.equal(expectedSeverityFromMatrix("High", "Low"), "Medium"); - assert.equal(expectedSeverityFromMatrix("Medium", "Low"), "Low"); - assert.equal(expectedSeverityFromMatrix("High", "Medium"), "High"); - assert.equal(expectedSeverityFromMatrix("Low", "High"), "Low"); - assert.equal(expectedSeverityFromMatrix("high", "Low"), undefined); - assert.equal(expectedSeverityFromMatrix("High ", "Low"), undefined); -}); - -test("final report validation accepts canonical matrix fields and ignores a preliminary severity guess", () => { - const diagnostics = validateSeverityMatrixArtifact({ - kind: "final-report", - artifactPath: "/tmp/report.json", - artifact: { - issues: [ - { - severity: "Medium", - severity_guess: "High", - impact: "High", - likelihood: "Low" - } - ] - } - }); - - assert.deepEqual(diagnostics, []); -}); - -test("final report validation blocks matrix-inconsistent issues", () => { - const diagnostics = validateSeverityMatrixArtifact({ - kind: "final-report", - artifactPath: "/tmp/report.json", - artifact: { - issues: [ - { severity: "High", impact: "High", likelihood: "Low" }, - { severity: "Medium", impact: "Medium", likelihood: "Low" } - ] - } - }); - - assert.equal(diagnostics.filter((diagnostic) => diagnostic.code === "SEVERITY_MATRIX_MISMATCH").length, 2); - assert.match(diagnostics[0]?.message ?? "", /expected Medium/); - assert.match(diagnostics[1]?.message ?? "", /expected Low/); -}); - -test("severity validation rejects lowercase levels and legacy field aliases", () => { - const diagnostics = validateSeverityMatrixArtifact({ - kind: "severity-classification", - artifactPath: "/tmp/severity-classified-findings.json", - artifact: [ - { - severity: "medium", - impact: "High", - likelihood: "Low", - final_severity: "Medium", - impact_level: "High", - likelihood_level: "Low", - severity_classification: { impact: "High", likelihood: "Low" } - } - ] - }); - - assert.ok(diagnostics.some((diagnostic) => diagnostic.code === "SEVERITY_LEVEL_INVALID")); - assert.equal(diagnostics.filter((diagnostic) => diagnostic.code === "SEVERITY_FIELD_ALIAS_UNSUPPORTED").length, 4); -}); - -test("severity classification rejects wrappers and does not parse matrix values from notes", () => { - const wrapped = validateSeverityMatrixArtifact({ - kind: "severity-classification", - artifactPath: "/tmp/severity-classified-findings.json", - artifact: { findings: [{ severity: "Medium", impact: "High", likelihood: "Low" }] } - }); - assert.deepEqual( - wrapped.map((diagnostic) => diagnostic.code), - ["SEVERITY_ARTIFACT_SHAPE_INVALID"] - ); - - const notesOnly = validateSeverityMatrixArtifact({ - kind: "severity-classification", - artifactPath: "/tmp/severity-classified-findings.json", - artifact: [{ severity: "Medium", notes: ["impact=High", "likelihood=Low"] }] - }); - assert.equal(notesOnly.filter((diagnostic) => diagnostic.code === "SEVERITY_MATRIX_FIELD_MISSING").length, 2); -}); - -test("severity validation reports non-object records instead of silently dropping them", () => { - const diagnostics = validateSeverityMatrixArtifact({ - kind: "severity-classification", - artifactPath: "/tmp/severity-classified-findings.json", - artifact: [null] - }); - - assert.deepEqual( - diagnostics.map((diagnostic) => diagnostic.code), - ["SEVERITY_RECORD_SHAPE_INVALID"] - ); -}); From 84b4a7994a3d74f5dd7a3cb4ddd1a9a6567b5b32 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:45:07 +0000 Subject: [PATCH 033/206] fix(runtime): report unscoped coverage prose scores as warnings The host re-check scans report.md, report.json strings and the coverage producer's Markdown for natural-language coverage scores and rejected any percentage or fraction that does not name an exact declaration-completeness scope on the same rendered line. Ordinary campaign prose matches it: "Recon reached 85% line coverage on Vault.sol." in a report.json note, or "Handlers reachable: 7/9" and "Retry 1/2 reproduced the same revert." in coverage-report.md. The in-workflow verifier has no equivalent check and agents cannot self-check it, so a compliant final report or coverage node passed its attempt and then failed host finalization. As a stop-gap, emit UNSCOPED_COVERAGE_SCORE, _PERCENTAGE and _FRACTION as warnings. Whether to delete the scanner is a separate decision. The typed coverage_evidence deep-equality check, the canonical scoped section projection (including a duplicated or contradicting scoped score anywhere in the document), the stylesheet and select-control checks and the 2,048-candidate scan limit keep their error severity. Tests that asserted these prose findings were fatal now assert that the scanner still reports them, as warnings, and that the node passes when nothing else is wrong. Co-Authored-By: Claude Opus 5.5 --- docs/reference/artifacts-reports.md | 9 +- packages/runtime/src/artifact-gates.ts | 9 +- packages/runtime/test/artifact-gates.test.ts | 154 +++++++++++-------- 3 files changed, 105 insertions(+), 67 deletions(-) diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index 2683bb70e..b1a62f68c 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -737,9 +737,12 @@ Blockers: ``` Repeat blocker and evidence rows in artifact order. Runtime publication compares -this section with the typed handoff and rejects missing, duplicated, reordered, -or bare coverage scores. Raw `covg-eval` output is for iteration only and -defines neither published declaration-completeness view. +this section with the typed handoff and rejects missing, duplicated, or +reordered scores. A coverage score elsewhere in the Markdown or in `report.json` +text that names no exact declaration-completeness scope is reported as a +warning rather than failing publication, although a document with more than +2,048 score candidates still fails the scan limit. Raw `covg-eval` output is +for iteration only and defines neither published declaration-completeness view. Current-run `report.md` contains concise links to `THREAT_MODEL.md`, `threat-model.json`, and `goal-plan.json`, plus source-node provenance for each diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 709644224..165bb268d 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -7454,7 +7454,7 @@ function unscopedCoverageScoreDiagnostics( { code: "UNSCOPED_COVERAGE_SCORE", message: "Published coverage scores must name an exact declaration-completeness scope on the same line", - severity: "error" as const, + severity: "warning" as const, source: "coverage-evidence", path: `${artifactPath}:${occurrence.line}` }, @@ -7512,13 +7512,18 @@ function reportCoverageScoreDiagnostic( path: `${artifactPath}:${occurrence.line}` }; } + // Prose scores are advisory. This natural-language scan runs only on the + // host, after the in-workflow verifier accepted the attempt, and it matches + // ordinary sentences such as "Recon reached 85% line coverage" or "Handlers + // reachable: 7/9". The typed coverage evidence and its canonical Markdown + // section remain errors when they disagree. const percentage = occurrence.kind === "percentage"; return { code: percentage ? "UNSCOPED_COVERAGE_PERCENTAGE" : "UNSCOPED_COVERAGE_FRACTION", message: percentage ? "Coverage percentages must name an exact declaration-completeness scope on the same rendered line" : "Coverage fractions must name an exact declaration-completeness scope on the same rendered line", - severity: "error", + severity: "warning", source: "coverage-evidence", path: `${artifactPath}:${occurrence.line}` }; diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 07cbec291..67cbf706c 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -12969,10 +12969,16 @@ test("coverage gate binds selected and unselected ranges to the trusted producti ]) { publish(evidence, `${scopedMarkdown}\n## Notes\n\n${unscopedProducerScore}\n`); const hiddenProducerScope = verifyRequiredArtifactsForAttempt(layout, node, node.id); - assert.equal(hiddenProducerScope.ok, false, unscopedProducerScore); + // Prose scores are advisory; the typed evidence and canonical section bind the result. + assert.equal( + hiddenProducerScope.ok, + true, + `${unscopedProducerScore}: ${JSON.stringify(hiddenProducerScope.diagnostics)}` + ); assert.ok( - hiddenProducerScope.diagnostics.some((diagnostic) => - /^UNSCOPED_COVERAGE_(?:FRACTION|PERCENTAGE)$/u.test(diagnostic.code) + hiddenProducerScope.diagnostics.some( + (diagnostic) => + /^UNSCOPED_COVERAGE_(?:FRACTION|PERCENTAGE)$/u.test(diagnostic.code) && diagnostic.severity === "warning" ), JSON.stringify(hiddenProducerScope.diagnostics) ); @@ -13304,13 +13310,22 @@ test("coverage gate binds selected and unselected ranges to the trusted producti JSON.stringify(contradictoryScopedFraction.diagnostics) ); - publish(evidence, `${scopedMarkdown}\n## Notes\n\nRetry 1/2 reproduced the same revert.\n`); - const ordinaryFraction = verifyRequiredArtifactsForAttempt(layout, node, node.id); - assert.equal(ordinaryFraction.ok, false); - assert.ok( - ordinaryFraction.diagnostics.some((diagnostic) => diagnostic.code === "UNSCOPED_COVERAGE_SCORE"), - JSON.stringify(ordinaryFraction.diagnostics) - ); + for (const ordinaryProse of [ + "Retry 1/2 reproduced the same revert.", + "Handlers reachable: 7/9", + "Target: 90% of Recon-selected declarations.", + "After iteration 2 we covered 41 of 57 functions." + ]) { + publish(evidence, `${scopedMarkdown}\n## Notes\n\n${ordinaryProse}\n`); + const ordinary = verifyRequiredArtifactsForAttempt(layout, node, node.id); + assert.equal(ordinary.ok, true, `${ordinaryProse}: ${JSON.stringify(ordinary.diagnostics)}`); + assert.ok( + ordinary.diagnostics.some( + (diagnostic) => diagnostic.code === "UNSCOPED_COVERAGE_SCORE" && diagnostic.severity === "warning" + ), + `${ordinaryProse}: ${JSON.stringify(ordinary.diagnostics)}` + ); + } const staleGoal = structuredClone(goal); staleGoal.current_measurement.covered_ranges = 0; @@ -14116,6 +14131,24 @@ test("coverage gate authenticates Vyper declaration boundaries in the Recon sele ); }); +/** Natural-language coverage scores are reported only as advisory warnings, never as gate errors. */ +function assertAdvisoryCoverageScore( + result: ReturnType, + code: "UNSCOPED_COVERAGE_PERCENTAGE" | "UNSCOPED_COVERAGE_FRACTION", + label: string +): void { + assert.ok( + result.diagnostics.some((diagnostic) => diagnostic.code === code && diagnostic.severity === "warning"), + `${label}: ${JSON.stringify(result.diagnostics)}` + ); + assert.ok( + result.diagnostics.every( + (diagnostic) => !diagnostic.code.startsWith("UNSCOPED_COVERAGE_") || diagnostic.severity === "warning" + ), + `${label}: ${JSON.stringify(result.diagnostics)}` + ); +} + test("final report preserves typed coverage evidence and its canonical Markdown projection", () => { const lcovArtifactPath = "reports/2026/08/coverage-input.lcov"; const layout = createRunLayout({ @@ -14495,16 +14528,12 @@ test("final report preserves typed coverage evidence and its canonical Markdown ]) { writeArtifact(layout, reportNode.id, "report.md", `${scopedMarkdown}\n## Notes\n\n${mixedScore}\n`); const mixed = verifyRequiredArtifactsForAttempt(layout, reportNode, reportNode.id); - assert.equal(mixed.ok, false, mixedScore); const normalizedMixedScore = mixedScore.normalize("NFKC").replace(/[\u2044\u2215\u29f8]/gu, "/"); - assert.ok( - mixed.diagnostics.some( - (diagnostic) => - diagnostic.code === - (/%|٪|&(?:percnt|#0*37|#x0*25);|\bpct\b\.?|\bper[ -]?cent(?:age)?\b/iu.test(normalizedMixedScore) - ? "UNSCOPED_COVERAGE_PERCENTAGE" - : "UNSCOPED_COVERAGE_FRACTION") - ), + assertAdvisoryCoverageScore( + mixed, + /%|٪|&(?:percnt|#0*37|#x0*25);|\bpct\b\.?|\bper[ -]?cent(?:age)?\b/iu.test(normalizedMixedScore) + ? "UNSCOPED_COVERAGE_PERCENTAGE" + : "UNSCOPED_COVERAGE_FRACTION", mixedScore ); } @@ -14617,11 +14646,8 @@ test("final report preserves typed coverage evidence and its canonical Markdown ]) { writeArtifact(layout, reportNode.id, "report.md", `${scopedMarkdown}\n## Notes\n\n${crossRenderedLineScope}\n`); const crossLine = verifyRequiredArtifactsForAttempt(layout, reportNode, reportNode.id); - assert.equal(crossLine.ok, false, crossRenderedLineScope); - assert.ok( - crossLine.diagnostics.some((diagnostic) => diagnostic.code === "UNSCOPED_COVERAGE_PERCENTAGE"), - `${crossRenderedLineScope}: ${JSON.stringify(crossLine.diagnostics)}` - ); + assert.equal(crossLine.ok, true, `${crossRenderedLineScope}: ${JSON.stringify(crossLine.diagnostics)}`); + assertAdvisoryCoverageScore(crossLine, "UNSCOPED_COVERAGE_PERCENTAGE", crossRenderedLineScope); } for (const implicitlyVisibleScore of [ @@ -14653,11 +14679,7 @@ test("final report preserves typed coverage evidence and its canonical Markdown ]) { writeArtifact(layout, reportNode.id, "report.md", `${scopedMarkdown}\n## Notes\n\n${implicitlyVisibleScore}\n`); const visibleAfterImplicitClose = verifyRequiredArtifactsForAttempt(layout, reportNode, reportNode.id); - assert.equal(visibleAfterImplicitClose.ok, false, implicitlyVisibleScore); - assert.ok( - visibleAfterImplicitClose.diagnostics.some((diagnostic) => diagnostic.code === "UNSCOPED_COVERAGE_PERCENTAGE"), - `${implicitlyVisibleScore}: ${JSON.stringify(visibleAfterImplicitClose.diagnostics)}` - ); + assertAdvisoryCoverageScore(visibleAfterImplicitClose, "UNSCOPED_COVERAGE_PERCENTAGE", implicitlyVisibleScore); } for (const paragraphClosingTag of [ @@ -14676,11 +14698,12 @@ test("final report preserves typed coverage evidence and its canonical Markdown const implicitlyVisibleScore = ` && …`, `… | tee ` or `…; echo …`, or a property failure in a campaign declared at a nested path, is no longer rejected after the verifier passed. The `--seq-len 100` / `sequence_length` rules moved into the shared gate and now fail inside the attempt, `verifyRequiredArtifactsForAttempt` requires the sealed attempt authority every production caller already passed, and the implementation-selection gate reads `config.resolved.toml` with the config parser instead of line regexes (#998). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 5358806a39143860a0270405d34e0578e257852c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:51:33 +0000 Subject: [PATCH 046/206] refactor(prompts): delete the unused prompt rename feature and duplicate template resolver renamePromptId, renameConcreteNodeId and rewriteArtifactPathForPromptId (rename.ts), their render.ts helpers (renamePromptArtifactReferences and three private rewriters) and the frontmatter round-trip support they needed (serializePromptDocument, diffPromptIdentity, the allowUnknownFields option, unknownFrontmatter/rawFrontmatter and the invalid-rename error code) had no caller outside their own tests. getPrompt had no caller at all. TemplateOccurrence.rawName existed only so rename could preserve whitespace. Output-contract templates had their own three-candidate resolver whose middle candidate, dist/prompts, is the layout the build stopped producing in #796. They now resolve through builtInPromptRoot(), the packaged-first lookup the agent preamble templates already use, so both kinds of prompt asset follow one rule. The packaged-layout render test, which copies the real dist below a directory with no repository above it, still covers the sealed-snapshot case; the installed-layout test that rebuilt the old dist/prompts layout is removed and its content assertions are folded into the packaged-layout test. Co-Authored-By: Claude Opus 5.5 --- packages/prompts/src/catalog.ts | 8 - packages/prompts/src/frontmatter.ts | 74 +------ packages/prompts/src/index.ts | 1 - packages/prompts/src/rename.ts | 247 ---------------------- packages/prompts/src/render.ts | 108 +--------- packages/prompts/test/frontmatter.test.ts | 24 +-- packages/prompts/test/rename.test.ts | 95 --------- packages/prompts/test/render.test.ts | 49 +---- 8 files changed, 10 insertions(+), 596 deletions(-) delete mode 100644 packages/prompts/src/rename.ts delete mode 100644 packages/prompts/test/rename.test.ts diff --git a/packages/prompts/src/catalog.ts b/packages/prompts/src/catalog.ts index 576a92ca9..b339be718 100644 --- a/packages/prompts/src/catalog.ts +++ b/packages/prompts/src/catalog.ts @@ -88,14 +88,6 @@ export function loadPromptCatalog(options: LoadPromptCatalogOptions = {}): Promp }; } -export function getPrompt(catalog: PromptCatalog, id: string): PromptCatalogEntry { - const entry = catalog.entries.get(id); - if (!entry) { - throw new PromptError("missing-template-variable", `prompt id \`${id}\` was not found`); - } - return entry; -} - export function builtInPromptRelativePaths(): string[] { return discoverBuiltInPromptRelativePaths(builtInPromptRoot()); } diff --git a/packages/prompts/src/frontmatter.ts b/packages/prompts/src/frontmatter.ts index acc00fbb1..4a2639834 100644 --- a/packages/prompts/src/frontmatter.ts +++ b/packages/prompts/src/frontmatter.ts @@ -1,4 +1,4 @@ -import { parse as parseYaml, stringify as stringifyYaml } from "yaml"; +import { parse as parseYaml } from "yaml"; export const PROMPT_FRONTMATTER_FIELDS = ["id", "display_name"] as const; @@ -12,12 +12,6 @@ export interface PromptFrontmatter { export interface ParsedPromptDocument { frontmatter: PromptFrontmatter; body: string; - rawFrontmatter?: string; - unknownFrontmatter?: Record; -} - -export interface ParsePromptFrontmatterOptions { - allowUnknownFields?: boolean; } export type PromptErrorCode = @@ -27,7 +21,6 @@ export type PromptErrorCode = | "invalid-frontmatter" | "invalid-prompt-path" | "invalid-render-input" - | "invalid-rename" | "missing-template-variable" | "not-ancestor" | "symlink-prompt-path" @@ -47,10 +40,7 @@ export class PromptError extends Error { } } -export function parsePromptFrontmatter( - markdown: string, - options: ParsePromptFrontmatterOptions = {} -): ParsedPromptDocument { +export function parsePromptFrontmatter(markdown: string): ParsedPromptDocument { const frontmatterBlock = readFrontmatterBlock(markdown); if (!frontmatterBlock) { return { @@ -76,16 +66,12 @@ export function parsePromptFrontmatter( throw new PromptError("invalid-frontmatter", "prompt frontmatter must be a YAML mapping"); } - const unknownFrontmatter: Record = {}; for (const key of Object.keys(parsed)) { if (key === "category") { throw new PromptError("invalid-frontmatter", "`category` is no longer supported in prompt frontmatter"); } if (!PROMPT_FRONTMATTER_FIELDS.includes(key as PromptFrontmatterField)) { - if (!options.allowUnknownFields) { - throw new PromptError("invalid-frontmatter", `unsupported prompt frontmatter field: ${key}`); - } - unknownFrontmatter[key] = parsed[key]; + throw new PromptError("invalid-frontmatter", `unsupported prompt frontmatter field: ${key}`); } } @@ -105,59 +91,7 @@ export function parsePromptFrontmatter( } return { frontmatter, - body: frontmatterBlock.body, - rawFrontmatter: frontmatterBlock.raw, - ...(Object.keys(unknownFrontmatter).length > 0 ? { unknownFrontmatter } : {}) - }; -} - -export function serializePromptDocument( - frontmatter: PromptFrontmatter, - body: string, - unknownFrontmatter: Record = {} -): string { - const mapping: Record = {}; - for (const key of PROMPT_FRONTMATTER_FIELDS) { - const value = frontmatter[key]; - if (value !== undefined) { - mapping[key] = value; - } - } - for (const [key, value] of Object.entries(unknownFrontmatter)) { - if (!(key in mapping)) { - mapping[key] = value; - } - } - - if (Object.keys(mapping).length === 0) { - return body; - } - - const yaml = stringifyYaml(mapping, { sortMapEntries: true }).trimEnd(); - const normalizedBody = body.startsWith("\n") ? body.slice(1) : body; - return `---\n${yaml}\n---\n\n${normalizedBody}`; -} - -export interface PromptIdentityDiff { - idChanged: boolean; - displayNameChanged: boolean; - executionIdentityChanged: boolean; - displayNameOnly: boolean; -} - -export function diffPromptIdentity(beforeMarkdown: string, afterMarkdown: string): PromptIdentityDiff { - const before = parsePromptFrontmatter(beforeMarkdown); - const after = parsePromptFrontmatter(afterMarkdown); - - const idChanged = before.frontmatter.id !== after.frontmatter.id; - const displayNameChanged = before.frontmatter.display_name !== after.frontmatter.display_name; - const executionIdentityChanged = idChanged || before.body !== after.body; - - return { - idChanged, - displayNameChanged, - executionIdentityChanged, - displayNameOnly: displayNameChanged && !executionIdentityChanged + body: frontmatterBlock.body }; } diff --git a/packages/prompts/src/index.ts b/packages/prompts/src/index.ts index 50b5497c0..0502495bf 100644 --- a/packages/prompts/src/index.ts +++ b/packages/prompts/src/index.ts @@ -1,5 +1,4 @@ export * from "./catalog.js"; export * from "./frontmatter.js"; export * from "./render.js"; -export * from "./rename.js"; export * from "./scaffold.js"; diff --git a/packages/prompts/src/rename.ts b/packages/prompts/src/rename.ts deleted file mode 100644 index 6f1a20d98..000000000 --- a/packages/prompts/src/rename.ts +++ /dev/null @@ -1,247 +0,0 @@ -import path from "node:path"; -import { parse as parseYaml, stringify as stringifyYaml } from "yaml"; -import { isSafePromptId, parsePromptFrontmatter, PromptError, serializePromptDocument } from "./frontmatter.js"; -import { renamePromptArtifactReferences } from "./render.js"; - -export interface PromptFileSnapshot { - path: string; - contents: string; -} - -export interface PromptTopologyNode { - id: string; - prompt?: string; - dependsOn?: string[]; - depends_on?: string[]; - outputs?: Array<{ path: string; contract: string; primary?: boolean }>; - [key: string]: unknown; -} - -export interface PromptTopologyDocument { - version?: number; - nodes: PromptTopologyNode[]; - [key: string]: unknown; -} - -export interface PromptConcreteNodeSnapshot { - id: string; - logicalId?: string; - logical_id?: string; - dependsOn?: string[]; - depends_on?: string[]; - artifactDir?: string; - artifact_dir?: string; - [key: string]: unknown; -} - -export interface RenamePromptIdInput { - oldId: string; - newId: string; - topology: PromptTopologyDocument | string; - promptFiles: PromptFileSnapshot[]; - concreteNodes?: PromptConcreteNodeSnapshot[]; -} - -export interface RenamePromptIdResult { - topology: PromptTopologyDocument; - topologyYaml?: string; - promptFiles: PromptFileSnapshot[]; - concreteNodes: PromptConcreteNodeSnapshot[]; - concreteIdMap: Record; - renamedPromptPaths: Record; -} - -export function renamePromptId(input: RenamePromptIdInput): RenamePromptIdResult { - validateRename(input.oldId, input.newId); - - const parsedTopology = - typeof input.topology === "string" ? (parseYaml(input.topology) as PromptTopologyDocument) : clone(input.topology); - if (!parsedTopology || !Array.isArray(parsedTopology.nodes)) { - throw new PromptError("invalid-rename", "topology must contain a nodes array"); - } - if (parsedTopology.nodes.some((node) => node.id === input.newId && node.id !== input.oldId)) { - throw new PromptError("invalid-rename", `topology already contains node id \`${input.newId}\``); - } - validatePromptFileRenameConflicts(input.promptFiles, input.oldId, input.newId); - - const renamedPromptPaths: Record = {}; - const promptFiles = input.promptFiles.map((file) => - renamePromptFile(file, input.oldId, input.newId, renamedPromptPaths) - ); - const topology = renameTopology(parsedTopology, input.oldId, input.newId, renamedPromptPaths); - const { concreteNodes, concreteIdMap } = renameConcreteNodes(input.concreteNodes ?? [], input.oldId, input.newId); - - return { - topology, - ...(typeof input.topology === "string" ? { topologyYaml: stringifyYaml(topology, { sortMapEntries: false }) } : {}), - promptFiles, - concreteNodes, - concreteIdMap, - renamedPromptPaths - }; -} - -export function renameConcreteNodeId(concreteId: string, oldId: string, newId: string): string { - if (concreteId === oldId) { - return newId; - } - if (concreteId.startsWith(`${oldId}-`)) { - return `${newId}${concreteId.slice(oldId.length)}`; - } - return concreteId; -} - -export function rewriteArtifactPathForPromptId(value: string, oldId: string, newId: string): string { - return value - .split(/[\\/]/g) - .map((segment) => rewriteArtifactSegment(segment, oldId, newId)) - .join(value.includes("\\") ? "\\" : "/"); -} - -function renamePromptFile( - file: PromptFileSnapshot, - oldId: string, - newId: string, - renamedPromptPaths: Record -): PromptFileSnapshot { - const parsed = parsePromptFrontmatter(file.contents); - const fallbackId = path.basename(file.path).replace(/\.(md|mdx)$/i, ""); - const ownsRenamedId = parsed.frontmatter.id === oldId || fallbackId === oldId; - const nextFrontmatter = { ...parsed.frontmatter }; - if (ownsRenamedId) { - nextFrontmatter.id = newId; - } - - const nextBody = renamePromptArtifactReferences(parsed.body, oldId, newId); - const contents = serializePromptDocument(nextFrontmatter, nextBody, parsed.unknownFrontmatter); - const nextPath = ownsRenamedId ? renamePromptPath(file.path, oldId, newId) : file.path; - if (nextPath !== file.path) { - renamedPromptPaths[file.path] = nextPath; - } - return { - path: nextPath, - contents - }; -} - -function renamePromptPath(promptPath: string, oldId: string, newId: string): string { - const extension = path.extname(promptPath); - const dirname = path.dirname(promptPath); - const basename = path.basename(promptPath, extension); - if (basename !== oldId) { - return promptPath; - } - return path.posix.join(dirname.split(path.sep).join("/"), `${newId}${extension}`); -} - -function renameTopology( - topology: PromptTopologyDocument, - oldId: string, - newId: string, - renamedPromptPaths: Record -): PromptTopologyDocument { - return { - ...topology, - nodes: topology.nodes.map((node) => { - const renamed: PromptTopologyNode = { - ...node, - id: node.id === oldId ? newId : node.id - }; - if (node.prompt) { - renamed.prompt = renamedPromptPaths[node.prompt] ?? renamePromptPath(node.prompt, oldId, newId); - } - if (node.dependsOn) { - renamed.dependsOn = node.dependsOn.map((dependency) => (dependency === oldId ? newId : dependency)); - } - if (node.depends_on) { - renamed.depends_on = node.depends_on.map((dependency) => (dependency === oldId ? newId : dependency)); - } - if (node.outputs) { - renamed.outputs = node.outputs.map((output) => ({ - ...output, - path: rewriteArtifactPathForPromptId(output.path, oldId, newId) - })); - } - return renamed; - }) - }; -} - -function renameConcreteNodes( - concreteNodes: PromptConcreteNodeSnapshot[], - oldId: string, - newId: string -): { concreteNodes: PromptConcreteNodeSnapshot[]; concreteIdMap: Record } { - const concreteIdMap: Record = {}; - const renamed = concreteNodes.map((node) => { - const nextId = renameConcreteNodeId(node.id, oldId, newId); - if (nextId !== node.id) { - concreteIdMap[node.id] = nextId; - } - const next: PromptConcreteNodeSnapshot = { - ...node, - id: nextId - }; - if (node.logicalId === oldId) { - next.logicalId = newId; - } - if (node.logical_id === oldId) { - next.logical_id = newId; - } - if (node.dependsOn) { - next.dependsOn = node.dependsOn.map((dependency) => renameConcreteNodeId(dependency, oldId, newId)); - } - if (node.depends_on) { - next.depends_on = node.depends_on.map((dependency) => renameConcreteNodeId(dependency, oldId, newId)); - } - if (node.artifactDir) { - next.artifactDir = rewriteArtifactPathForPromptId(node.artifactDir, oldId, newId); - } - if (node.artifact_dir) { - next.artifact_dir = rewriteArtifactPathForPromptId(node.artifact_dir, oldId, newId); - } - return next; - }); - - return { - concreteNodes: renamed, - concreteIdMap - }; -} - -function validateRename(oldId: string, newId: string): void { - if (!isSafePromptId(oldId) || !isSafePromptId(newId)) { - throw new PromptError("invalid-rename", "prompt rename IDs must be safe IDs"); - } - if (oldId === newId) { - throw new PromptError("invalid-rename", "prompt rename requires different old and new IDs"); - } -} - -function validatePromptFileRenameConflicts(promptFiles: PromptFileSnapshot[], oldId: string, newId: string): void { - for (const file of promptFiles) { - const parsed = parsePromptFrontmatter(file.contents); - const fallbackId = path.basename(file.path).replace(/\.(md|mdx)$/i, ""); - if (parsed.frontmatter.id === newId || (fallbackId === newId && parsed.frontmatter.id !== oldId)) { - throw new PromptError("invalid-rename", `prompt file already contains id \`${newId}\`: ${file.path}`); - } - } -} - -function rewriteArtifactSegment(segment: string, oldId: string, newId: string): string { - if (segment === oldId) { - return newId; - } - if (segment.startsWith(`${oldId}-`)) { - return `${newId}${segment.slice(oldId.length)}`; - } - const extensionIndex = segment.lastIndexOf("."); - if (extensionIndex > 0 && segment.slice(0, extensionIndex) === oldId) { - return `${newId}${segment.slice(extensionIndex)}`; - } - return segment; -} - -function clone(value: T): T { - return structuredClone(value); -} diff --git a/packages/prompts/src/render.ts b/packages/prompts/src/render.ts index ae91a889a..0a24888b9 100644 --- a/packages/prompts/src/render.ts +++ b/packages/prompts/src/render.ts @@ -1,6 +1,5 @@ -import { mkdirSync, readFileSync, statSync, writeFileSync } from "node:fs"; +import { mkdirSync, readFileSync, writeFileSync } from "node:fs"; import path from "node:path"; -import { fileURLToPath } from "node:url"; import { findingNoteKeyPromptVocabulary, @@ -199,7 +198,6 @@ type ArtifactProducer = export interface TemplateOccurrence { name: string; - rawName: string; start: number; end: number; } @@ -625,36 +623,11 @@ function loadOutputContractTemplate(relativePath: string): string { if (cached !== undefined) { return cached; } - const template = readFileSync(path.join(outputContractTemplateRoot(), relativePath), "utf8"); + const template = readFileSync(path.join(builtInPromptRoot(), "_templates", "output-contract", relativePath), "utf8"); outputContractTemplateCache.set(relativePath, template); return template; } -function outputContractTemplateRoot(): string { - const here = path.dirname(fileURLToPath(import.meta.url)); - const candidates = [ - // This build now copies the prompt tree to dist/assets/prompts, so the - // packaged templates live here. Without this candidate the only path that - // still resolved was the repository fallback below, which exists in a source - // checkout but not in a sealed execution snapshot -- where the package sits - // at modules/@ultrafuzz/prompts/dist and "../../../" reaches modules/. That - // rendered fine at submission and threw the first time a workflow - // re-rendered mid-run, taking the whole run down at WORKFLOW_RENDER_FAILED. - path.join(here, "assets", "prompts", "_templates", "output-contract"), - path.join(here, "prompts", "_templates", "output-contract"), - path.resolve(here, "../../../.ultrafuzz/prompts/_templates/output-contract") - ]; - const found = candidates.find((candidate) => { - try { - return statSync(candidate).isDirectory(); - } catch { - return false; - } - }); - if (found === undefined) throw new Error(`unable to locate the packaged output contract templates from ${here}`); - return found; -} - const agentPreambleTemplateCache = new Map(); /** Load one trusted global-agent prompt fragment from the packaged MDX assets. */ @@ -700,38 +673,6 @@ export function writeRenderedPrompt(result: PromptRenderResult): string { return result.renderedPromptPath; } -export function renamePromptArtifactReferences(template: string, oldLogicalId: string, newLogicalId: string): string { - validateArtifactReferenceId(oldLogicalId); - validateArtifactReferenceId(newLogicalId); - validatePromptVariables(template); - - let rewritten = ""; - let consumed = 0; - for (const occurrence of findTemplateOccurrences(template)) { - rewritten += template.slice(consumed, occurrence.start); - const replacement = renameTemplateVariableName(occurrence.name, oldLogicalId, newLogicalId); - const leadingWhitespace = occurrence.rawName.length - occurrence.rawName.trimStart().length; - const trailingWhitespace = occurrence.rawName.length - occurrence.rawName.trimEnd().length; - rewritten += "{{"; - rewritten += occurrence.rawName.slice(0, leadingWhitespace); - rewritten += replacement; - rewritten += occurrence.rawName.slice(occurrence.rawName.length - trailingWhitespace); - rewritten += "}}"; - - consumed = occurrence.end; - const producer = parseArtifactProducer(occurrence.name); - if (producer?.kind === "current" || (producer?.kind === "logical" && producer.logicalId === oldLogicalId)) { - const suffix = rewriteArtifactPathSuffix(template.slice(consumed), oldLogicalId, newLogicalId); - if (suffix.consumed > 0) { - rewritten += suffix.value; - consumed += suffix.consumed; - } - } - } - rewritten += template.slice(consumed); - return rewritten; -} - export function validateArtifactRelativePath(relativePath: string): void { const parts = relativePath.split("/"); if ( @@ -771,14 +712,12 @@ function findTemplateOccurrences(template: string): TemplateOccurrence[] { if (end === -1) { throw new PromptError("unclosed-template-variable", "template variable is missing a closing delimiter"); } - const rawName = template.slice(start + 2, end); - const name = rawName.trim(); + const name = template.slice(start + 2, end).trim(); if (name === "") { throw new PromptError("empty-template-variable", "template variable name cannot be empty"); } occurrences.push({ name, - rawName, start, end: end + 2 }); @@ -1390,44 +1329,3 @@ function ensureInsidePath(root: string, candidate: string, label: string): void } throw new PromptError("invalid-render-input", `${label} must stay inside ${root}: ${candidate}`); } - -function renameTemplateVariableName(name: string, oldLogicalId: string, newLogicalId: string): string { - if (name === `artifact_path:${oldLogicalId}`) { - return `artifact_path:${newLogicalId}`; - } - if (name === `artifact_handoff:${oldLogicalId}`) { - return `artifact_handoff:${newLogicalId}`; - } - return name; -} - -function rewriteArtifactPathSuffix( - suffix: string, - oldLogicalId: string, - newLogicalId: string -): { value: string; consumed: number } { - if (!suffix.startsWith("/")) { - return { value: "", consumed: 0 }; - } - const end = findSuffixEnd(suffix); - const value = suffix - .slice(0, end) - .split("/") - .map((segment) => rewriteArtifactPathSegment(segment, oldLogicalId, newLogicalId)) - .join("/"); - return { value, consumed: end }; -} - -function rewriteArtifactPathSegment(segment: string, oldLogicalId: string, newLogicalId: string): string { - if (segment === oldLogicalId) { - return newLogicalId; - } - if (segment.startsWith(`${oldLogicalId}-`)) { - return `${newLogicalId}${segment.slice(oldLogicalId.length)}`; - } - const extensionIndex = segment.lastIndexOf("."); - if (extensionIndex > 0 && segment.slice(0, extensionIndex) === oldLogicalId) { - return `${newLogicalId}${segment.slice(extensionIndex)}`; - } - return segment; -} diff --git a/packages/prompts/test/frontmatter.test.ts b/packages/prompts/test/frontmatter.test.ts index 6983ac2a0..7a9aa1ab8 100644 --- a/packages/prompts/test/frontmatter.test.ts +++ b/packages/prompts/test/frontmatter.test.ts @@ -1,5 +1,5 @@ import { describe, expect, it } from "vitest"; -import { diffPromptIdentity, parsePromptFrontmatter, PromptError } from "../src/index.js"; +import { parsePromptFrontmatter, PromptError } from "../src/index.js"; describe("prompt frontmatter", () => { it("parses Markdown-compatible prompt frontmatter", () => { @@ -12,31 +12,11 @@ describe("prompt frontmatter", () => { expect(parsed.body).toBe("# Body"); }); - it("rejects unsupported additional fields by default", () => { + it("rejects unsupported additional fields", () => { expect(() => parsePromptFrontmatter("---\nid: x\nowner: user\n---\nBody")).toThrow(PromptError); }); - it("can preserve documented unknown frontmatter when explicitly allowed", () => { - const parsed = parsePromptFrontmatter("---\nid: x\nowner: user\n---\nBody", { - allowUnknownFields: true - }); - - expect(parsed.unknownFrontmatter).toEqual({ owner: "user" }); - }); - it("rejects removed category frontmatter", () => { expect(() => parsePromptFrontmatter("---\nid: x\ncategory: custom\n---\nBody")).toThrow(/category/); }); - - it("treats display_name changes as labels only", () => { - const before = "---\nid: boundary-tests\ndisplay_name: Boundary Tests\n---\nBody"; - const after = "---\nid: boundary-tests\ndisplay_name: Boundary Tests v2\n---\nBody"; - - expect(diffPromptIdentity(before, after)).toEqual({ - idChanged: false, - displayNameChanged: true, - executionIdentityChanged: false, - displayNameOnly: true - }); - }); }); diff --git a/packages/prompts/test/rename.test.ts b/packages/prompts/test/rename.test.ts deleted file mode 100644 index 1030a79c6..000000000 --- a/packages/prompts/test/rename.test.ts +++ /dev/null @@ -1,95 +0,0 @@ -import { describe, expect, it } from "vitest"; -import { diffPromptIdentity, renamePromptId } from "../src/index.js"; - -describe("prompt ID rename sync", () => { - it("synchronizes topology ids, dependencies, prompt files, artifact refs, and concrete ids", () => { - const result = renamePromptId({ - oldId: "boundary-tests", - newId: "edge-tests", - topology: { - version: 1, - nodes: [ - { - id: "boundary-tests", - prompt: "strategies/boundary-tests.mdx", - depends_on: ["base-test-setup"], - outputs: [ - { - path: "boundary-tests/generated-tests.json", - contract: "ultrafuzz/generated-tests@3", - primary: true - } - ] - }, - { - id: "dedupe-findings", - prompt: "review/dedupe-findings.mdx", - depends_on: ["boundary-tests"] - } - ] - }, - promptFiles: [ - { - path: "strategies/boundary-tests.mdx", - contents: - "---\nid: boundary-tests\ndisplay_name: Boundary Tests\n---\nRead {{artifact_handoff:base-test-setup}} and {{artifact_path:boundary-tests}}/boundary-tests.json" - }, - { - path: "review/dedupe-findings.mdx", - contents: - "---\nid: dedupe-findings\n---\nUse {{artifact_handoff:boundary-tests}} and {{artifact_path:boundary-tests}}/generated-tests.json" - } - ], - concreteNodes: [ - { - id: "boundary-tests-0", - logicalId: "boundary-tests", - artifactDir: "/runs/run-1/artifacts/boundary-tests-0" - }, - { - id: "dedupe-findings", - logicalId: "dedupe-findings", - dependsOn: ["boundary-tests-0"] - } - ] - }); - - expect(result.topology.nodes[0]?.id).toBe("edge-tests"); - expect(result.topology.nodes[0]?.prompt).toBe("strategies/edge-tests.mdx"); - expect(result.topology.nodes[0]?.outputs?.map((output) => output.path)).toEqual([ - "edge-tests/generated-tests.json" - ]); - expect(result.topology.nodes[1]?.depends_on).toEqual(["edge-tests"]); - expect(result.promptFiles[0]?.path).toBe("strategies/edge-tests.mdx"); - expect(result.promptFiles[0]?.contents).toContain("id: edge-tests"); - expect(result.promptFiles[0]?.contents).toContain("{{artifact_path:edge-tests}}/edge-tests.json"); - expect(result.promptFiles[1]?.contents).toContain("{{artifact_handoff:edge-tests}}"); - expect(result.concreteIdMap).toEqual({ "boundary-tests-0": "edge-tests-0" }); - expect(result.concreteNodes[0]?.artifactDir).toBe("/runs/run-1/artifacts/edge-tests-0"); - expect(result.concreteNodes[1]?.dependsOn).toEqual(["edge-tests-0"]); - }); - - it("rejects frontmatter rename conflicts", () => { - expect(() => - renamePromptId({ - oldId: "boundary-tests", - newId: "edge-tests", - topology: { - version: 1, - nodes: [{ id: "boundary-tests", prompt: "strategies/boundary-tests.md" }] - }, - promptFiles: [ - { path: "strategies/boundary-tests.md", contents: "---\nid: boundary-tests\n---\nBody" }, - { path: "strategies/edge-tests.md", contents: "---\nid: edge-tests\n---\nBody" } - ] - }) - ).toThrow(/already contains id/); - }); - - it("does not change execution identity for display_name-only edits", () => { - const before = "---\nid: boundary-tests\ndisplay_name: Boundary Tests\n---\nBody"; - const after = "---\nid: boundary-tests\ndisplay_name: Edge Boundary Tests\n---\nBody"; - - expect(diffPromptIdentity(before, after).displayNameOnly).toBe(true); - }); -}); diff --git a/packages/prompts/test/render.test.ts b/packages/prompts/test/render.test.ts index 8921c2ae9..6f59c11c7 100644 --- a/packages/prompts/test/render.test.ts +++ b/packages/prompts/test/render.test.ts @@ -1,15 +1,4 @@ -import { - copyFileSync, - cpSync, - existsSync, - mkdirSync, - mkdtempSync, - readFileSync, - realpathSync, - rmSync, - symlinkSync -} from "node:fs"; -import { execFileSync } from "node:child_process"; +import { cpSync, existsSync, mkdirSync, mkdtempSync, readFileSync, realpathSync, rmSync, symlinkSync } from "node:fs"; import os from "node:os"; import path from "node:path"; import { fileURLToPath, pathToFileURL } from "node:url"; @@ -208,42 +197,6 @@ describe("prompt rendering", () => { const rendered = packaged.renderPrompt(input).renderedMarkdown; expect(rendered).toContain("For every output declared with `Contract: ultrafuzz/findings@2`"); - }); - - it("renders output-contract guidance and prompt partials from an installed package layout", async () => { - const tmp = mkdtempSync(path.join(realpathSync(os.tmpdir()), "ufz-installed-prompts-")); - tmpDirs.push(tmp); - const repoRoot = fileURLToPath(new URL("../../..", import.meta.url)); - const appRoot = path.join(tmp, "app"); - const packageRoot = path.join(appRoot, "node_modules", "@ultrafuzz", "prompts"); - const distRoot = path.join(packageRoot, "dist"); - mkdirSync(packageRoot, { recursive: true }); - execFileSync( - "pnpm", - [ - "exec", - "tsc", - "-p", - path.join(repoRoot, "packages", "prompts", "tsconfig.json"), - "--outDir", - distRoot, - "--tsBuildInfoFile", - path.join(distRoot, ".tsbuildinfo") - ], - { cwd: repoRoot, stdio: "pipe" } - ); - cpSync(path.join(repoRoot, ".ultrafuzz", "prompts"), path.join(distRoot, "prompts"), { recursive: true }); - copyFileSync(path.join(repoRoot, "packages", "prompts", "package.json"), path.join(packageRoot, "package.json")); - linkPromptDependencies(repoRoot, path.join(appRoot, "node_modules")); - - const installed = (await import( - `${pathToFileURL(path.join(distRoot, "render.js")).href}?installed-layout=${Date.now()}` - )) as { renderPrompt: typeof renderPrompt }; - const input = baseRenderInput(tmp); - input.prompt = `${input.prompt}\n{{coverage_evidence_markdown_projection}}`; - const rendered = installed.renderPrompt(input).renderedMarkdown; - - expect(rendered).toContain("For every output declared with `Contract: ultrafuzz/findings@2`"); expect(rendered).toContain("Validation command: `ultrafuzz json validate --schema"); expect(rendered).toContain("Contract validation command: `ultrafuzz artifact validate"); expect(rendered).toContain("- : `/`"); From dd12e8bbefa7b9973fed744c7e86b7bc214861ae Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:51:45 +0000 Subject: [PATCH 047/206] refactor(config): delete the prompt-metadata layer and helpers with no production caller - The prompt-metadata layer: resolveConfig applied createDefaultPromptMetadataLayer(), which always returned {}, and an input.promptMetadata that only one test supplied. PromptMetadataLayer, ResolveConfigInput.promptMetadata and the "prompt-metadata" diagnostic source go with it. - restoreRedactedConfig and assertNoRedactionPlaceholders: nothing in runtime, cli, modal or evals calls them. Runs launch from the unredacted resolved config; the redacted copy and its manifest are written for readers of the run directory, and no command restores values from them. docs/config.md no longer describes a restore step or launch guard that does not exist. The "redaction" diagnostic source goes with them. - applyDefaultProfileOverrides: a one-line wrapper only tests called. Its test now drives applyModelProfileOverrides, which validate.ts and plan-run.ts use. - assertResolvedConfigZod and resolvedConfigSchemaEntry: no callers. - resolvedConfigValidatorsAgree: moved into the schema test, its only user. redactResolvedConfig and serializeRedactedResolvedConfigToml, which plan-run.ts and init.ts use, are unchanged. Co-Authored-By: Claude Opus 5.5 --- docs/config.md | 6 +- packages/config/src/config-schema-registry.ts | 17 +-- packages/config/src/defaults.ts | 5 - packages/config/src/model-profiles.ts | 4 - packages/config/src/redaction.ts | 118 +----------------- packages/config/src/resolve.ts | 41 ------ packages/config/src/resolved-config-schema.ts | 13 -- packages/config/src/types.ts | 15 +-- packages/config/test/config.test.ts | 71 ++--------- .../test/resolved-config-schema.test.ts | 5 +- 10 files changed, 20 insertions(+), 275 deletions(-) diff --git a/docs/config.md b/docs/config.md index 4ccc55e3e..355ea291c 100644 --- a/docs/config.md +++ b/docs/config.md @@ -458,10 +458,8 @@ trusted local execution model remains unchanged. ## Redaction Run artifacts store redacted resolved config and a redaction manifest. -Sensitive model values are redacted before persistence. Launch guards for -literal redaction placeholders may fail before workflow launch when enabled, -and manifest entries mark values that must be restored from current config -before launch. +Sensitive model values are redacted before persistence. The manifest records +which values were redacted; no command restores values from it. ## Eval suites diff --git a/packages/config/src/config-schema-registry.ts b/packages/config/src/config-schema-registry.ts index 1683a58bc..3dfd63ef1 100644 --- a/packages/config/src/config-schema-registry.ts +++ b/packages/config/src/config-schema-registry.ts @@ -15,11 +15,7 @@ import { type SchemaRegistryEntry } from "@ultrafuzz/artifacts"; -import { - RESOLVED_CONFIG_JSON_SCHEMA_ID, - RESOLVED_CONFIG_SCHEMA_FILENAME, - resolvedConfigZodSchema -} from "./resolved-config-schema.js"; +import { RESOLVED_CONFIG_JSON_SCHEMA_ID, RESOLVED_CONFIG_SCHEMA_FILENAME } from "./resolved-config-schema.js"; import type { ResolvedConfig } from "./types.js"; import { isRecord } from "@ultrafuzz/artifacts"; @@ -142,13 +138,6 @@ export function configSchemaBundleDigest(): string { return schemaRegistryBundleDigest(configSchemaRegistry()); } -export function resolvedConfigSchemaEntry(): SchemaRegistryEntry { - const entry = configSchemaRegistry().find((candidate) => candidate.id === RESOLVED_CONFIG_JSON_SCHEMA_ID); - if (entry === undefined) - throw new Error(`registered config schema is unavailable: ${RESOLVED_CONFIG_JSON_SCHEMA_ID}`); - return entry; -} - export function validateResolvedConfigJson(value: unknown): JsonSchemaValidationResult { const validator = configValidator().getSchema(RESOLVED_CONFIG_JSON_SCHEMA_ID); if (validator === undefined) @@ -196,10 +185,6 @@ export function serializeResolvedConfigJsonBytes(config: ResolvedConfig): Buffer return bytes; } -export function resolvedConfigValidatorsAgree(value: unknown): boolean { - return validateResolvedConfigJson(value).ok === resolvedConfigZodSchema.safeParse(value).success; -} - function configValidator(): ReturnType { if (cachedValidator !== undefined) return cachedValidator; const registry = configSchemaRegistry(); diff --git a/packages/config/src/defaults.ts b/packages/config/src/defaults.ts index 62d6e3285..cab776c99 100644 --- a/packages/config/src/defaults.ts +++ b/packages/config/src/defaults.ts @@ -15,7 +15,6 @@ import { type ModelProfile, type PermissionConfig, type ProjectConfigInput, - type PromptMetadataLayer, type RetryConfig, type ResolvedConfig, type RunConfig @@ -37,10 +36,6 @@ export function createDefaultResolvedConfig(): ResolvedConfig { return cloneResolvedConfig(DEFAULT_CONFIG); } -export function createDefaultPromptMetadataLayer(): PromptMetadataLayer { - return {}; -} - export function synthesizeDefaultModelProfile(agent = DEFAULT_AGENT): ModelProfile { return { id: DEFAULT_MODEL_PROFILE_ID, diff --git a/packages/config/src/model-profiles.ts b/packages/config/src/model-profiles.ts index 26d9f7ca4..3bffb3a01 100644 --- a/packages/config/src/model-profiles.ts +++ b/packages/config/src/model-profiles.ts @@ -66,10 +66,6 @@ export function validProfileId(id: string): boolean { return profileIdSchema.safeParse(id).success; } -export function applyDefaultProfileOverrides(config: ResolvedConfig, overrides: DefaultProfileOverrides): void { - applyModelProfileOverrides(config, config.models.default, overrides); -} - export function applyModelProfileOverrides( config: ResolvedConfig, profileId: string, diff --git a/packages/config/src/redaction.ts b/packages/config/src/redaction.ts index b382f670c..462cd4505 100644 --- a/packages/config/src/redaction.ts +++ b/packages/config/src/redaction.ts @@ -1,20 +1,7 @@ import { cloneResolvedConfig } from "./defaults.js"; import { serializeResolvedConfigToml, type SerializeResolvedConfigTomlOptions } from "./resolve.js"; -import { - SENSITIVE_REDACTION_PLACEHOLDER, - hasRedactionPlaceholder, - isSensitiveSecretValue, - redactSecretsInText -} from "@ultrafuzz/security"; -import { - diagnostic, - fail, - hasErrors, - ok, - type ConfigDiagnostic, - type ConfigResult, - type ResolvedConfig -} from "./types.js"; +import { SENSITIVE_REDACTION_PLACEHOLDER, isSensitiveSecretValue, redactSecretsInText } from "@ultrafuzz/security"; +import type { ConfigDiagnostic, ResolvedConfig } from "./types.js"; export const REDACTION_PLACEHOLDER = SENSITIVE_REDACTION_PLACEHOLDER; export const CONFIG_REDACTIONS_SCHEMA_VERSION = "ultrafuzz.config-redactions.v2" as const; @@ -64,54 +51,6 @@ export function serializeRedactedResolvedConfigToml( return serializeResolvedConfigToml("config" in redacted ? redacted.config : redacted, options); } -export function restoreRedactedConfig( - redactedConfig: ResolvedConfig, - currentConfig: ResolvedConfig, - manifest: RedactionManifest -): ConfigResult { - const restored = cloneResolvedConfig(redactedConfig); - const diagnostics: ConfigDiagnostic[] = []; - - for (const entry of manifest.entries) { - const currentValue = getPath(currentConfig, entry.path); - if (typeof currentValue !== "string" || currentValue.length === 0 || hasRedactionPlaceholder(currentValue)) { - diagnostics.push( - diagnostic( - "CONFIG_REDACTION_RESTORE_MISSING", - `redacted value at ${entry.key} must be restored from current config before workflow launch`, - entry.path, - "redaction" - ) - ); - continue; - } - setPath(restored, entry.path, currentValue); - } - - diagnostics.push(...assertNoRedactionPlaceholders(restored).diagnostics); - if (hasErrors(diagnostics)) { - return fail(diagnostics); - } - return ok(restored); -} - -export function assertNoRedactionPlaceholders(config: ResolvedConfig): ConfigResult { - const diagnostics: ConfigDiagnostic[] = []; - visitStrings(config, [], (path, value) => { - if (hasRedactionPlaceholder(value)) { - diagnostics.push( - diagnostic( - "CONFIG_REDACTION_PLACEHOLDER_PRESENT", - `redaction placeholder at ${path.join(".")} cannot be passed to workflow launch`, - path, - "redaction" - ) - ); - } - }); - return diagnostics.length > 0 ? fail(diagnostics) : ok(undefined); -} - export function redactDiagnostics(diagnostics: ConfigDiagnostic[]): ConfigDiagnostic[] { return diagnostics.map((entry) => ({ ...entry, @@ -139,59 +78,6 @@ function redactSensitiveScalar( } } -function visitStrings(value: unknown, path: string[], visitor: (path: string[], value: string) => void): void { - if (typeof value === "string") { - visitor(path, value); - return; - } - if (Array.isArray(value)) { - value.forEach((entry, index) => { - visitStrings(entry, path.concat(String(index)), visitor); - }); - return; - } - if (value && typeof value === "object") { - for (const [key, child] of Object.entries(value)) { - visitStrings(child, path.concat(key), visitor); - } - } -} - -function getPath(value: unknown, path: string[]): unknown { - let current = value; - for (const segment of path) { - if (Array.isArray(current)) { - current = current[Number(segment)]; - } else if (current && typeof current === "object") { - current = (current as Record)[segment]; - } else { - return undefined; - } - } - return current; -} - -function setPath(value: unknown, path: string[], nextValue: string): void { - let current = value as Record; - for (const segment of path.slice(0, -1)) { - const child = current[segment]; - if (Array.isArray(child)) { - current = child as unknown as Record; - } else { - current = child as Record; - } - } - const last = path[path.length - 1]; - if (last === undefined) { - return; - } - if (Array.isArray(current)) { - current[Number(last)] = nextValue; - } else { - current[last] = nextValue; - } -} - function formatPath(path: string[]): string { return path.join("."); } diff --git a/packages/config/src/resolve.ts b/packages/config/src/resolve.ts index c431a1ced..797c33fd5 100644 --- a/packages/config/src/resolve.ts +++ b/packages/config/src/resolve.ts @@ -5,7 +5,6 @@ import { DEFAULT_MODEL_PROFILE_ID, MAX_TIMEOUT_SECONDS, cloneResolvedConfig, - createDefaultPromptMetadataLayer, createDefaultResolvedConfig } from "./defaults.js"; import { @@ -33,7 +32,6 @@ import { type ConfigResult, type PermissionConfig, type ProjectConfigInput, - type PromptMetadataLayer, type ResolveConfigInput, type ResolvedConfig, type RuntimeConfigOverrides @@ -71,10 +69,6 @@ export function resolveConfig(input: ResolveConfigInput = {}): ConfigResult): void { if (layer.mode !== undefined) { config.execution.mode = layer.mode; diff --git a/packages/config/src/resolved-config-schema.ts b/packages/config/src/resolved-config-schema.ts index e7d1e095b..76cd63583 100644 --- a/packages/config/src/resolved-config-schema.ts +++ b/packages/config/src/resolved-config-schema.ts @@ -365,16 +365,3 @@ export const resolvedConfigZodSchema: z.ZodType = z }); } }); - -export function assertResolvedConfigZod( - value: unknown, - label = "resolved configuration" -): asserts value is ResolvedConfig { - const result = resolvedConfigZodSchema.safeParse(value); - if (result.success) return; - const summary = result.error.issues - .slice(0, 10) - .map((issue) => `${issue.path.join(".") || "/"}: ${issue.message}`) - .join("; "); - throw new Error(`${label} does not match ${RESOLVED_CONFIG_JSON_SCHEMA_ID}: ${summary}`); -} diff --git a/packages/config/src/types.ts b/packages/config/src/types.ts index 772f9c6f4..511339ab0 100644 --- a/packages/config/src/types.ts +++ b/packages/config/src/types.ts @@ -3,14 +3,7 @@ import type { AuditProfileSettings, DynamicStrategiesEnumerator } from "./audit- export type DiagnosticSeverity = "error" | "warning"; export type ConfigDiagnosticSource = - | "defaults" - | "audit-profile" - | "prompt-metadata" - | "project-toml" - | "environment" - | "runtime" - | "validation" - | "redaction"; + "defaults" | "audit-profile" | "project-toml" | "environment" | "runtime" | "validation"; export interface ConfigDiagnostic { code: string; @@ -240,11 +233,6 @@ export interface AuditProfileResolution { export type AuditProfileSettingOrigin = "default" | "audit-profile" | "project-config" | "environment" | "runtime-override"; -export interface PromptMetadataLayer { - models?: Record & { id?: string }>; - run?: Partial; -} - export interface ProjectConfigInput { schemaVersion?: typeof PROJECT_CONFIG_SCHEMA_VERSION; auditProfile?: string; @@ -304,7 +292,6 @@ export interface RuntimeConfigOverrides extends ProjectConfigInput { export interface ResolveConfigInput { projectConfig?: ProjectConfigInput; - promptMetadata?: PromptMetadataLayer; env?: Record; runtimeOverrides?: RuntimeConfigOverrides; } diff --git a/packages/config/test/config.test.ts b/packages/config/test/config.test.ts index 2c1a976d5..5fd260040 100644 --- a/packages/config/test/config.test.ts +++ b/packages/config/test/config.test.ts @@ -11,8 +11,7 @@ import { MODAL_NODE_MAX_INNER_TIMEOUT_SECONDS, MODAL_SANDBOX_MAX_LIFETIME_SECONDS, REDACTION_PLACEHOLDER, - assertNoRedactionPlaceholders, - applyDefaultProfileOverrides, + applyModelProfileOverrides, invariantPropertyPrioritySelection, loadProjectConfig, parseProjectConfigToml, @@ -20,7 +19,6 @@ import { redactResolvedConfig, resolveExecutionResources, resolveConfig, - restoreRedactedConfig, serializeRedactedResolvedConfigToml, type ConfigDiagnostic, type ProjectConfigInput, @@ -547,7 +545,7 @@ credential_env = ["MODAL_TOKEN_ID", "MODAL_TOKEN_SECRET"] } }); - it("applies defaults, prompt metadata, project TOML, env, then runtime overrides", () => { + it("applies defaults, project TOML, env, then runtime overrides", () => { const project = parseProjectConfigToml(` schema_version = "ultrafuzz.config.v2" dynamic_strategies_enumerator = 5 @@ -578,15 +576,6 @@ config_dir = "teams/codex" if (!project.ok) return; const resolved = resolveConfig({ - promptMetadata: { - run: { defaultTimeoutSeconds: 900 }, - models: { - "prompt-model": { - agent: "CodexAgent", - model: "gpt-5.5" - } - } - }, projectConfig: project.value, env: { ULTRAFUZZ_MAX_PARALLEL_AGENTS: "7", @@ -820,7 +809,7 @@ output_dir = ".ultrafuzz/custom-runs" }); describe("redaction", () => { - it("redacts sensitive model values while preserving restore requirements", () => { + it("redacts sensitive model values from the persisted config and records them in the manifest", () => { const resolved = resolveConfig({ env: {}, projectConfig: { @@ -845,50 +834,7 @@ describe("redaction", () => { expect(toml).not.toContain("sk-test-secret"); expect(redacted.manifest.entries.map((entry) => entry.key)).toContain("models.profiles.secret-model.model"); expect(redacted.manifest.entries[0]?.requiredForWorkflowLaunch).toBe(true); - - const restored = restoreRedactedConfig(redacted.config, resolved.value, redacted.manifest); - expect(restored.ok).toBe(true); - if (!restored.ok) return; - expect(restored.value.models.profiles["secret-model"]?.model).toBe("sk-test-secret"); - expect(assertNoRedactionPlaceholders(restored.value).ok).toBe(true); - }); - - it("rejects literal redaction placeholders before workflow launch", () => { - const resolved = resolveConfig({ env: {} }); - expect(resolved.ok).toBe(true); - if (!resolved.ok) return; - resolved.value.models.profiles.default!.model = "[redacted]"; - const checked = assertNoRedactionPlaceholders(resolved.value); - expect(checked.ok).toBe(false); - expect(checked.diagnostics[0]?.code).toBe("CONFIG_REDACTION_PLACEHOLDER_PRESENT"); - }); - - it("fails restoration when a required redacted value is unavailable", () => { - const resolved = resolveConfig({ - env: {}, - projectConfig: { - models: { - profiles: { - "secret-model": { - agent: "CodexAgent", - model: "sk-test-secret" - } - }, - default: "secret-model" - } - } - }); - expect(resolved.ok).toBe(true); - if (!resolved.ok) throw new Error(JSON.stringify(resolved.diagnostics, null, 2)); - - const redacted = redactResolvedConfig(resolved.value); - const current = structuredClone(resolved.value); - delete current.models.profiles["secret-model"]!.model; - - const restored = restoreRedactedConfig(redacted.config, current, redacted.manifest); - expect(restored.ok).toBe(false); - expect(restored.diagnostics.map((entry) => entry.code)).toContain("CONFIG_REDACTION_RESTORE_MISSING"); - expect(restored.diagnostics[0]?.message).toContain("before workflow launch"); + expect(resolved.value.models.profiles["secret-model"]?.model).toBe("sk-test-secret"); }); }); @@ -1054,7 +1000,7 @@ describe("model profile and triage validation", () => { const agentOnly = resolveConfig({ env: {} }); expect(agentOnly.ok).toBe(true); if (!agentOnly.ok) return; - applyDefaultProfileOverrides(agentOnly.value, { agent: "ClaudeAgent" }); + applyModelProfileOverrides(agentOnly.value, agentOnly.value.models.default, { agent: "ClaudeAgent" }); expect(agentOnly.value.models.profiles.default).toMatchObject({ agent: "ClaudeAgent" }); @@ -1064,7 +1010,10 @@ describe("model profile and triage validation", () => { const pinned = resolveConfig({ env: {} }); expect(pinned.ok).toBe(true); if (!pinned.ok) return; - applyDefaultProfileOverrides(pinned.value, { agent: "ClaudeAgent", model: "claude-sonnet-5" }); + applyModelProfileOverrides(pinned.value, pinned.value.models.default, { + agent: "ClaudeAgent", + model: "claude-sonnet-5" + }); expect(pinned.value.models.profiles.default).toMatchObject({ agent: "ClaudeAgent", model: "claude-sonnet-5" @@ -1074,7 +1023,7 @@ describe("model profile and triage validation", () => { const benchmark = resolveConfig({ env: {} }); expect(benchmark.ok).toBe(true); if (!benchmark.ok) return; - applyDefaultProfileOverrides(benchmark.value, { + applyModelProfileOverrides(benchmark.value, benchmark.value.models.default, { agent: "CodexAgent", model: "gpt-5.6-luna", reasoning: "high" diff --git a/packages/config/test/resolved-config-schema.test.ts b/packages/config/test/resolved-config-schema.test.ts index f83d529cd..51a0544bd 100644 --- a/packages/config/test/resolved-config-schema.test.ts +++ b/packages/config/test/resolved-config-schema.test.ts @@ -17,7 +17,6 @@ import { parseProjectConfigToml, parseResolvedConfigJsonBytes, resolvedConfigJsonSchema, - resolvedConfigValidatorsAgree, resolvedConfigZodSchema, resolveConfig, serializeResolvedConfigJsonBytes, @@ -320,6 +319,10 @@ describe("resolved config JSON contract", () => { ); }); +function resolvedConfigValidatorsAgree(value: unknown): boolean { + return validateResolvedConfigJson(value).ok === resolvedConfigZodSchema.safeParse(value).success; +} + function fixturePath(filename: string): string { return path.join(path.dirname(fileURLToPath(import.meta.url)), "fixtures", filename); } From 1785e1c675066caff2416e9cd08e62c2bb19232b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:51:53 +0000 Subject: [PATCH 048/206] ci: delete benchmark CI scripts orphaned by #1131 #1131 deleted eval-benchmarks.yml, the only workflow that invoked scripts/ci/modal-benchmark-control-window.mjs (create/deadline) and scripts/ci/validate-threat-model-benchmark-gate.mjs. Since then the control-window script was kept alive only by a knip entry and a block of ci-config.test.ts that exercised it, and the threat-model gate only by its own test, which still ran in every PR's test:ci-scripts step. Delete both scripts, the gate's test, the knip entry, and the control-window block of the EVMBench full-mode manifest test; the manifest assertions around it stay. Co-Authored-By: Claude Opus 5.5 --- knip.jsonc | 6 +- packages/modal/test/ci-config.test.ts | 92 ---- scripts/ci/modal-benchmark-control-window.mjs | 175 ------- .../validate-threat-model-benchmark-gate.mjs | 445 ------------------ ...lidate-threat-model-benchmark-gate.test.ts | 333 ------------- 5 files changed, 1 insertion(+), 1050 deletions(-) delete mode 100644 scripts/ci/modal-benchmark-control-window.mjs delete mode 100644 scripts/ci/validate-threat-model-benchmark-gate.mjs delete mode 100644 scripts/ci/validate-threat-model-benchmark-gate.test.ts diff --git a/knip.jsonc b/knip.jsonc index 09983944f..ed1a3da7a 100644 --- a/knip.jsonc +++ b/knip.jsonc @@ -2,11 +2,7 @@ "$schema": "https://unpkg.com/knip@6/schema.json", "workspaces": { ".": { - "entry": [ - "scripts/validate-audit-profile-package.mjs", - "scripts/ci/prepare-modal-benchmarks.mjs", - "scripts/ci/modal-benchmark-control-window.mjs" - ] + "entry": ["scripts/validate-audit-profile-package.mjs", "scripts/ci/prepare-modal-benchmarks.mjs"] }, "packages/dashboard": { "entry": ["frontend/public/theme-bootstrap.js"] diff --git a/packages/modal/test/ci-config.test.ts b/packages/modal/test/ci-config.test.ts index 6df0d84eb..0836379ab 100644 --- a/packages/modal/test/ci-config.test.ts +++ b/packages/modal/test/ci-config.test.ts @@ -203,98 +203,6 @@ describe("public Modal benchmark configuration", () => { expect(publicBenchmarkMaxRuntimeSeconds("full")).toBe(15_000); expect(manifest.control_timeout_seconds).toBe(37_500); expect(manifest.control_timeout_seconds * 1000).toBeLessThan(MODAL_PUBLIC_FULL_SANDBOX_TIMEOUT_MS); - const controlWindowPath = path.join(output, "control-window.json"); - execFileSync( - process.execPath, - [ - path.join(workspace, "scripts/ci/modal-benchmark-control-window.mjs"), - "create", - path.join(output, "manifest.json"), - controlWindowPath, - "d".repeat(40), - "https://github.com/monad-developers/ultrafuzz", - "23456-1", - "full" - ], - { cwd: workspace } - ); - const controlWindow = JSON.parse(fs.readFileSync(controlWindowPath, "utf8")) as { - schema_version: string; - candidate_commit: string; - repository: string; - generation: string; - mode: string; - manifest_sha256: string; - control_timeout_seconds: number; - started_at_epoch_seconds: number; - deadline_at_epoch_seconds: number; - }; - expect(controlWindow).toEqual( - expect.objectContaining({ - schema_version: "ultrafuzz.modal.ci-control-window.v1", - candidate_commit: "d".repeat(40), - repository: "https://github.com/monad-developers/ultrafuzz", - generation: "23456-1", - mode: "full", - control_timeout_seconds: 37_500, - manifest_sha256: expect.stringMatching(/^[0-9a-f]{64}$/u) - }) - ); - expect(controlWindow.deadline_at_epoch_seconds - controlWindow.started_at_epoch_seconds).toBe(37_500); - const restoredDeadline = Number( - execFileSync( - process.execPath, - [ - path.join(workspace, "scripts/ci/modal-benchmark-control-window.mjs"), - "deadline", - path.join(output, "manifest.json"), - controlWindowPath, - "d".repeat(40), - "https://github.com/monad-developers/ultrafuzz", - "23456-1", - "full" - ], - { cwd: workspace, encoding: "utf8" } - ).trim() - ); - expect(restoredDeadline).toBe(controlWindow.deadline_at_epoch_seconds); - const wrongRepository = spawnSync( - process.execPath, - [ - path.join(workspace, "scripts/ci/modal-benchmark-control-window.mjs"), - "deadline", - path.join(output, "manifest.json"), - controlWindowPath, - "d".repeat(40), - "https://github.com/monad-developers/other", - "23456-1", - "full" - ], - { cwd: workspace, encoding: "utf8" } - ); - expect(wrongRepository.status).not.toBe(0); - expect(wrongRepository.stderr).toMatch(/manifest identity does not match/u); - const tamperedWindowPath = path.join(output, "control-window-tampered.json"); - fs.writeFileSync( - tamperedWindowPath, - `${JSON.stringify({ ...controlWindow, deadline_at_epoch_seconds: restoredDeadline + 1 })}\n` - ); - const tamperedDeadline = spawnSync( - process.execPath, - [ - path.join(workspace, "scripts/ci/modal-benchmark-control-window.mjs"), - "deadline", - path.join(output, "manifest.json"), - tamperedWindowPath, - "d".repeat(40), - "https://github.com/monad-developers/ultrafuzz", - "23456-1", - "full" - ], - { cwd: workspace, encoding: "utf8" } - ); - expect(tamperedDeadline.status).not.toBe(0); - expect(tamperedDeadline.stderr).toMatch(/deadline is invalid/u); expect(manifest.concurrency.max_parallel_eval_rows_per_sandbox).toBe(maxParallel); expect(manifest.concurrency.max_parallel_workflow_nodes_per_row).toBe( publicBenchmarkMaxParallelWorkflowNodes("full") diff --git a/scripts/ci/modal-benchmark-control-window.mjs b/scripts/ci/modal-benchmark-control-window.mjs deleted file mode 100644 index 4d98ee36b..000000000 --- a/scripts/ci/modal-benchmark-control-window.mjs +++ /dev/null @@ -1,175 +0,0 @@ -import { createHash } from "node:crypto"; -import fs from "node:fs"; -import path from "node:path"; -import { pathToFileURL } from "node:url"; - -import { parseStrictJsonBytes, readRegularFileSnapshot } from "../../packages/artifacts/dist/index.js"; -import { MODAL_BENCHMARK_CONTROL_MANIFEST_SCHEMA_ID } from "../../packages/modal/dist/modal-contracts.js"; -import { readModalDocument } from "../../packages/modal/dist/modal-documents.js"; - -const SCHEMA_VERSION = "ultrafuzz.modal.ci-control-window.v1"; -const MAX_CONTROL_BYTES = 1024 * 1024; -const FULL_COMMIT = /^[0-9a-f]{40}$/u; -const GENERATION = /^[1-9][0-9]*-[1-9][0-9]*$/u; -const REPOSITORY = /^https:\/\/github\.com\/[A-Za-z0-9_.-]+\/[A-Za-z0-9_.-]+$/u; -const SHA256 = /^[0-9a-f]{64}$/u; -const WINDOW_KEYS = [ - "schema_version", - "candidate_commit", - "repository", - "generation", - "mode", - "manifest_sha256", - "control_timeout_seconds", - "started_at_epoch_seconds", - "deadline_at_epoch_seconds" -]; - -export function createModalBenchmarkControlWindow(input) { - const identity = expectedIdentity(input); - const manifest = readManifest(input.manifestPath, identity); - const startedAt = input.nowEpochSeconds ?? Math.floor(Date.now() / 1000); - if (!Number.isSafeInteger(startedAt) || startedAt <= 0) { - throw new Error("Modal benchmark control start must be a positive epoch second"); - } - const deadlineAt = startedAt + manifest.control_timeout_seconds; - if (!Number.isSafeInteger(deadlineAt)) throw new Error("Modal benchmark control deadline exceeds the safe range"); - const window = { - schema_version: SCHEMA_VERSION, - candidate_commit: identity.candidateCommit, - repository: identity.repository, - generation: identity.generation, - mode: identity.mode, - manifest_sha256: manifestDigest(input.manifestPath), - control_timeout_seconds: manifest.control_timeout_seconds, - started_at_epoch_seconds: startedAt, - deadline_at_epoch_seconds: deadlineAt - }; - const outputPath = path.resolve(input.outputPath); - fs.writeFileSync(outputPath, `${JSON.stringify(window, null, 2)}\n`, { flag: "wx", mode: 0o600 }); - return window; -} - -export function readModalBenchmarkControlWindow(input) { - const identity = expectedIdentity(input); - const manifest = readManifest(input.manifestPath, identity); - const windowPath = path.resolve(input.windowPath); - const value = parseStrictJsonBytes(boundedRegularFile(windowPath, "Modal benchmark control window"), { - maxBytes: MAX_CONTROL_BYTES, - maxDepth: 8, - maxItems: 32, - maxProperties: 32 - }); - if (value === null || typeof value !== "object" || Array.isArray(value)) { - throw new Error("Modal benchmark control window must be an object"); - } - const keys = Object.keys(value).sort(); - if (JSON.stringify(keys) !== JSON.stringify([...WINDOW_KEYS].sort())) { - throw new Error("Modal benchmark control window has unexpected fields"); - } - if ( - value.schema_version !== SCHEMA_VERSION || - value.candidate_commit !== identity.candidateCommit || - value.repository !== identity.repository || - value.generation !== identity.generation || - value.mode !== identity.mode - ) { - throw new Error("Modal benchmark control window identity does not match the producer attempt"); - } - if (!SHA256.test(value.manifest_sha256) || value.manifest_sha256 !== manifestDigest(input.manifestPath)) { - throw new Error("Modal benchmark control window is not bound to the restored manifest"); - } - if ( - !Number.isSafeInteger(value.control_timeout_seconds) || - value.control_timeout_seconds !== manifest.control_timeout_seconds || - !Number.isSafeInteger(value.started_at_epoch_seconds) || - value.started_at_epoch_seconds <= 0 || - !Number.isSafeInteger(value.deadline_at_epoch_seconds) || - value.deadline_at_epoch_seconds !== value.started_at_epoch_seconds + value.control_timeout_seconds - ) { - throw new Error("Modal benchmark control window deadline is invalid"); - } - return value; -} - -function readManifest(manifestPath, identity) { - const absolutePath = path.resolve(manifestPath); - boundedRegularFile(absolutePath, "Modal benchmark control manifest"); - const manifest = readModalDocument(absolutePath, MODAL_BENCHMARK_CONTROL_MANIFEST_SCHEMA_ID).value; - if ( - manifest.candidate_commit !== identity.candidateCommit || - manifest.repository !== identity.repository || - manifest.generation !== identity.generation || - manifest.mode !== identity.mode - ) { - throw new Error("Modal benchmark control manifest identity does not match the producer attempt"); - } - if (!Number.isSafeInteger(manifest.control_timeout_seconds) || manifest.control_timeout_seconds < 300) { - throw new Error("Modal benchmark control timeout is invalid"); - } - return manifest; -} - -function expectedIdentity(input) { - if (!FULL_COMMIT.test(input.candidateCommit ?? "")) throw new Error("candidate commit must be a full lowercase SHA"); - if (!REPOSITORY.test(input.repository ?? "")) throw new Error("repository must be a canonical public GitHub URL"); - if (!GENERATION.test(input.generation ?? "")) throw new Error("generation must be a GitHub run-attempt pair"); - if (input.mode !== "smoke" && input.mode !== "full") throw new Error("benchmark mode must be smoke or full"); - return { - candidateCommit: input.candidateCommit, - repository: input.repository, - generation: input.generation, - mode: input.mode - }; -} - -function manifestDigest(manifestPath) { - return createHash("sha256") - .update(boundedRegularFile(path.resolve(manifestPath), "Modal benchmark control manifest")) - .digest("hex"); -} - -function boundedRegularFile(filePath, label) { - let stat; - try { - stat = fs.lstatSync(filePath); - } catch (error) { - throw new Error(`${label} is unavailable`, { cause: error }); - } - if (stat.isSymbolicLink() || !stat.isFile() || stat.size < 1 || stat.size > MAX_CONTROL_BYTES) { - throw new Error(`${label} must be a bounded regular file`); - } - return readRegularFileSnapshot(filePath, MAX_CONTROL_BYTES); -} - -function main(args) { - const [command, manifestPath, windowPath, candidateCommit, repository, generation, mode] = args; - if (args.length !== 7 || (command !== "create" && command !== "deadline")) { - throw new Error( - "usage: modal-benchmark-control-window.mjs " - ); - } - if (command === "create") { - console.log( - JSON.stringify( - createModalBenchmarkControlWindow({ - manifestPath, - outputPath: windowPath, - candidateCommit, - repository, - generation, - mode - }) - ) - ); - return; - } - console.log( - readModalBenchmarkControlWindow({ manifestPath, windowPath, candidateCommit, repository, generation, mode }) - .deadline_at_epoch_seconds - ); -} - -if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) { - main(process.argv.slice(2)); -} diff --git a/scripts/ci/validate-threat-model-benchmark-gate.mjs b/scripts/ci/validate-threat-model-benchmark-gate.mjs deleted file mode 100644 index cf8096cca..000000000 --- a/scripts/ci/validate-threat-model-benchmark-gate.mjs +++ /dev/null @@ -1,445 +0,0 @@ -import crypto from "node:crypto"; -import fs from "node:fs"; -import path from "node:path"; -import { pathToFileURL } from "node:url"; - -const MAX_MANIFEST_BYTES = 1024 * 1024; -const MAX_BUNDLE_BYTES = 256 * 1024 * 1024; -const SAFE_ID = /^[A-Za-z0-9][A-Za-z0-9._-]{0,127}$/u; -const FULL_COMMIT = /^[0-9a-f]{40}$/u; -const SHA256 = /^[0-9a-f]{64}$/u; -const CANONICAL_TARGET_IDS = ["stableswap-ng-vyper", "venus-isolated-pools-hardhat", "very-liquid-vaults-foundry"]; -const COMPLETENESS_FIELDS = ["nodes", "lineage", "concurrency_evidence", "plan_evidence"]; -const NODE_STATUSES = [ - "pending", - "ready", - "runnable", - "running", - "succeeded", - "failed", - "skipped", - "timed-out", - "reused-from-prior-run", - "invalidated" -]; -const SUCCESSFUL_NODE_STATUSES = new Set(["succeeded", "reused-from-prior-run"]); -const REQUIRED_ROW_ARTIFACTS = [ - ["threat-model", "THREAT_MODEL.md", "markdown"], - ["threat-model", "threat-model.json", "json"], - ["goal-plan", "goal-plan.json", "json"], - ["goal-plan", "vulnerability-db-manifest.json", "json"] -]; - -/** - * Validate the structural claims made by the opt-in threat-model gate. - * - * Quality, cost, retry, token and wall-clock values deliberately remain - * telemetry. The checks here are limited to whether the production fanout ran, - * matched the planner-owned expectation from #373, retained its reviewable - * evidence and preserved independently attributable goal lanes. - */ -export function validateThreatModelBenchmarkGate(manifestValue, pairBundles) { - const manifest = record(manifestValue, "threat-model benchmark manifest"); - if (manifest.mode !== "threat-model" || manifest.benchmark !== "ultrafuzz-bench") { - throw new Error("threat-model gate requires the threat-model Ultrafuzz-bench manifest"); - } - if (typeof manifest.candidate_commit !== "string" || !FULL_COMMIT.test(manifest.candidate_commit)) { - throw new Error("threat-model gate manifest has an invalid candidate commit"); - } - const targets = manifestTargets(manifest.targets); - const targetIds = [...targets.keys()].sort(); - if (JSON.stringify(targetIds) !== JSON.stringify(CANONICAL_TARGET_IDS)) { - throw new Error("threat-model gate must use exactly the three canonical Ultrafuzz-bench targets"); - } - if (manifest.matrix_rows_per_pair !== targets.size) { - throw new Error("threat-model gate manifest must schedule one row per canonical target"); - } - const concurrency = record(manifest.concurrency, "threat-model gate concurrency"); - const requestedConcurrency = positiveInteger( - concurrency.max_parallel_workflow_nodes_per_row, - "threat-model gate requested workflow concurrency" - ); - if (requestedConcurrency < 2) { - throw new Error("threat-model gate requested workflow concurrency must exceed one"); - } - if (!Array.isArray(manifest.pairs) || manifest.pairs.length !== 1) { - throw new Error("threat-model gate must contain exactly one checked-in runner pair"); - } - if (!Array.isArray(pairBundles) || pairBundles.length !== 1) { - throw new Error("threat-model gate must provide exactly one collected runner bundle"); - } - const pair = manifestPair(manifest.pairs[0]); - const supplied = record(pairBundles[0], "threat-model gate pair bundle"); - if (supplied.pair !== pair.pair) throw new Error("threat-model gate bundle does not match its manifest pair"); - const bundle = record(supplied.bundle, `threat-model gate bundle ${pair.pair}`); - validateBundleIdentity(bundle, manifest, pair, targets); - const files = verifiedBundleFiles(bundle.files); - const summary = jsonBundleFile(files, "eval/summary.json"); - const rows = summaryRows(summary, targetIds); - for (const row of rows) { - validateRowArtifacts(files, row.row_id); - validateRowExpansion(row, requestedConcurrency); - } - return { - pair_count: 1, - target_count: targets.size, - row_count: rows.length, - expected_dynamic_child_count: rows.reduce((total, row) => total + row.expansion.plan.expected_child_count, 0) - }; -} - -export function validateThreatModelBenchmarkGateFiles(manifestPath, resultsRoot) { - const manifest = readJsonRegular(manifestPath, MAX_MANIFEST_BYTES, "threat-model benchmark manifest"); - const pairs = Array.isArray(manifest?.pairs) ? manifest.pairs : []; - const root = path.resolve(resultsRoot); - const pairBundles = pairs.map((value, index) => { - const pair = manifestPair(value, index); - const bundlePath = regularFileInside( - root, - [pair.pair, pair.model_slug, "public-results.json"], - MAX_BUNDLE_BYTES, - `threat-model gate bundle ${pair.pair}` - ); - return { pair: pair.pair, bundle: readJsonRegular(bundlePath, MAX_BUNDLE_BYTES, `bundle ${pair.pair}`) }; - }); - return validateThreatModelBenchmarkGate(manifest, pairBundles); -} - -function manifestTargets(value) { - if (!Array.isArray(value) || value.length !== CANONICAL_TARGET_IDS.length) { - throw new Error("threat-model gate manifest must contain three targets"); - } - const targets = new Map(); - for (const [index, entry] of value.entries()) { - const target = record(entry, `threat-model gate target ${index}`); - const id = safeId(target.id, `threat-model gate target ${index} ID`); - if ( - typeof target.repository !== "string" || - !target.repository.startsWith("https://github.com/") || - typeof target.revision !== "string" || - !FULL_COMMIT.test(target.revision) || - typeof target.framework !== "string" || - !SAFE_ID.test(target.framework) - ) { - throw new Error(`threat-model gate target ${id} has invalid immutable identity`); - } - if (targets.has(id)) throw new Error(`threat-model gate repeats target ${id}`); - targets.set(id, { - id, - repository: target.repository, - revision: target.revision, - framework: target.framework - }); - } - return targets; -} - -function manifestPair(value, index = 0) { - const pair = record(value, `threat-model gate pair ${index}`); - const pairId = safeId(pair.pair, `threat-model gate pair ${index} ID`); - const modelSlug = safeId(pair.model_slug, `threat-model gate pair ${index} model slug`); - if ( - pair.mode !== "threat-model" || - pair.lane !== "threat-model" || - pair.benchmark !== "ultrafuzz-bench" || - pair.provider !== "openai" || - !modelSlug.startsWith("benchmark-threat-model-") - ) { - throw new Error("threat-model gate pair does not use its checked-in OpenAI runner profile"); - } - return { pair: pairId, model_slug: modelSlug }; -} - -function validateBundleIdentity(bundle, manifest, pair, targets) { - if ( - bundle.status !== "succeeded" || - bundle.benchmark !== "ultrafuzz-bench" || - bundle.lane !== "threat-model" || - bundle.model_slug !== pair.model_slug || - bundle.model !== "gpt-5.6-luna" || - bundle.reasoning !== "high" || - bundle.candidate_commit !== manifest.candidate_commit - ) { - throw new Error("threat-model gate bundle does not represent a successful pinned Luna/high run"); - } - if ( - bundle.executed_case_count !== targets.size || - bundle.graded_case_count !== targets.size || - !Array.isArray(bundle.targets) || - bundle.targets.length !== targets.size - ) { - throw new Error("threat-model gate bundle does not contain three executed and graded targets"); - } - const seen = new Set(); - for (const entry of bundle.targets) { - const target = record(entry, "threat-model gate bundle target"); - const expected = targets.get(target.id); - if ( - expected === undefined || - seen.has(target.id) || - target.repository !== expected.repository || - target.revision !== expected.revision || - target.framework !== expected.framework || - target.status !== "succeeded" - ) { - throw new Error("threat-model gate bundle target identity or status does not match the manifest"); - } - seen.add(target.id); - } -} - -function verifiedBundleFiles(value) { - if (!Array.isArray(value)) throw new Error("threat-model gate bundle files must be an array"); - const files = new Map(); - for (const [index, entry] of value.entries()) { - const file = record(entry, `threat-model gate bundle file ${index}`); - if (typeof file.path !== "string" || file.path === "" || files.has(file.path)) { - throw new Error("threat-model gate bundle contains an invalid or duplicate file path"); - } - if ( - typeof file.contents_base64 !== "string" || - !Number.isSafeInteger(file.size_bytes) || - file.size_bytes < 0 || - typeof file.sha256 !== "string" || - !SHA256.test(file.sha256) - ) { - throw new Error(`threat-model gate bundle file ${file.path} has invalid integrity metadata`); - } - const contents = Buffer.from(file.contents_base64, "base64"); - if ( - contents.toString("base64") !== file.contents_base64 || - contents.byteLength !== file.size_bytes || - crypto.createHash("sha256").update(contents).digest("hex") !== file.sha256 - ) { - throw new Error(`threat-model gate bundle file ${file.path} failed its integrity check`); - } - files.set(file.path, contents); - } - return files; -} - -function summaryRows(summaryValue, expectedTargetIds) { - const summary = record(summaryValue, "threat-model gate eval summary"); - if (!Array.isArray(summary.rows) || summary.rows.length !== expectedTargetIds.length) { - throw new Error("threat-model gate eval summary must contain one row per canonical target"); - } - const expected = new Set(expectedTargetIds); - const seenTargets = new Set(); - const seenRows = new Set(); - return summary.rows.map((value, index) => { - const row = record(value, `threat-model gate summary row ${index}`); - const rowId = safeId(row.row_id, `threat-model gate summary row ${index} ID`); - const targetId = safeId(row.target_id, `threat-model gate summary row ${index} target`); - if (!expected.has(targetId) || seenTargets.has(targetId) || seenRows.has(rowId)) { - throw new Error("threat-model gate eval summary has duplicate or unexpected row identity"); - } - seenTargets.add(targetId); - seenRows.add(rowId); - return { ...row, row_id: rowId, target_id: targetId }; - }); -} - -function validateRowArtifacts(files, rowId) { - for (const [producer, name, kind] of REQUIRED_ROW_ARTIFACTS) { - const artifactPath = `reports/${rowId}/artifacts/${producer}/${name}`; - const contents = files.get(artifactPath); - if (contents === undefined || contents.byteLength === 0) { - throw new Error(`threat-model gate row ${rowId} is missing retained artifact ${artifactPath}`); - } - if (kind === "json") { - let value; - try { - value = JSON.parse(contents.toString("utf8")); - } catch (error) { - throw new Error(`threat-model gate retained artifact ${artifactPath} is not valid JSON`, { cause: error }); - } - record(value, `threat-model gate retained artifact ${artifactPath}`); - } else if (contents.toString("utf8").trim() === "") { - throw new Error(`threat-model gate retained artifact ${artifactPath} is empty`); - } - } -} - -function validateRowExpansion(row, requestedConcurrency) { - const label = `threat-model gate row ${row.row_id}`; - const expansion = record(row.expansion, `${label} expansion`); - for (const field of COMPLETENESS_FIELDS) assertComplete(expansion[field], `${label} ${field}`); - if (expansion.truncated !== false) throw new Error(`${label} expansion evidence is truncated`); - const plan = record(expansion.plan, `${label} expansion plan`); - const expected = positiveInteger(plan.expected_child_count, `${label} expected child count`); - const maxDynamicNodes = positiveInteger(plan.max_dynamic_nodes, `${label} max dynamic nodes`); - const laneCount = positiveInteger(plan.lane_count, `${label} goal lane count`); - if (expected > maxDynamicNodes) throw new Error(`${label} expected child count exceeds its recorded limit`); - const comparison = record(expansion.expected_vs_actual, `${label} expected-versus-actual expansion`); - if ( - comparison.matches !== true || - comparison.delta !== 0 || - comparison.expected_child_count !== expected || - comparison.actual_dynamic_node_count !== expected || - expansion.dynamic_node_count !== expected - ) { - throw new Error(`${label} dynamic child count does not match the planner-owned expectation`); - } - if (!Array.isArray(expansion.dynamic_nodes) || expansion.dynamic_nodes.length !== expected) { - throw new Error(`${label} does not expose every expected dynamic child`); - } - const dynamicIds = new Set(); - for (const value of expansion.dynamic_nodes) { - const node = record(value, `${label} dynamic child`); - if ( - typeof node.node_id !== "string" || - node.node_id === "" || - dynamicIds.has(node.node_id) || - typeof node.source_node_id !== "string" || - node.source_node_id === "" || - !SUCCESSFUL_NODE_STATUSES.has(node.status) || - node.timed_out !== false - ) { - throw new Error(`${label} has an incomplete, failed or unattributed dynamic child`); - } - dynamicIds.add(node.node_id); - } - if (!Array.isArray(expansion.goal_lanes) || expansion.goal_lanes.length !== laneCount) { - throw new Error(`${label} does not expose every planner-owned goal lane`); - } - const laneIds = new Set(); - const observedLaneNodeIds = new Set(); - for (const value of expansion.goal_lanes) { - const lane = record(value, `${label} goal lane`); - if (typeof lane.lane_id !== "string" || lane.lane_id === "" || laneIds.has(lane.lane_id)) { - throw new Error(`${label} has an invalid or duplicate goal lane`); - } - laneIds.add(lane.lane_id); - const plannedIds = nonemptyUniqueStrings(lane.planned_node_ids, `${label} lane ${lane.lane_id} planned nodes`); - const observedPlannedIds = nonemptyUniqueStrings( - lane.observed_planned_node_ids, - `${label} lane ${lane.lane_id} observed planned nodes` - ); - const observedIds = nonemptyUniqueStrings(lane.observed_node_ids, `${label} lane ${lane.lane_id} observed nodes`); - if (!sameStringSet(plannedIds, observedPlannedIds)) { - throw new Error(`${label} lane ${lane.lane_id} did not resolve its exact planner-owned node set`); - } - if (observedIds.some((nodeId) => observedLaneNodeIds.has(nodeId))) { - throw new Error(`${label} assigns one observed node to more than one goal lane`); - } - for (const nodeId of observedIds) observedLaneNodeIds.add(nodeId); - if ( - lane.observed_node_count !== plannedIds.length || - observedIds.length !== plannedIds.length || - lane.failed !== false || - !Array.isArray(lane.failed_node_ids) || - lane.failed_node_ids.length !== 0 || - !Array.isArray(lane.timed_out_node_ids) || - lane.timed_out_node_ids.length !== 0 - ) { - throw new Error(`${label} lane ${lane.lane_id} was not independently and successfully observed`); - } - assertSuccessfulStatusCounts(lane.status_counts, observedIds.length, `${label} lane ${lane.lane_id}`); - } - if ([...dynamicIds].some((nodeId) => !observedLaneNodeIds.has(nodeId))) { - throw new Error(`${label} has a dynamic child outside its planner-owned goal lanes`); - } - const observedConcurrency = record(expansion.concurrency, `${label} observed concurrency`); - const effective = positiveInteger(observedConcurrency.effective, `${label} effective concurrency`); - nonnegativeInteger(observedConcurrency.ready_queue_depth, `${label} ready queue depth`); - nonnegativeInteger(observedConcurrency.active_work, `${label} active work`); - if (observedConcurrency.requested !== requestedConcurrency || effective > requestedConcurrency || effective < 2) { - throw new Error(`${label} did not demonstrate independently scheduled dynamic work`); - } -} - -function assertComplete(value, label) { - const evidence = record(value, label); - if (evidence.status !== "complete" || evidence.reason !== null) { - throw new Error(`${label} is not complete`); - } -} - -function assertSuccessfulStatusCounts(value, expectedCount, label) { - const counts = record(value, `${label} status counts`); - let total = 0; - for (const status of NODE_STATUSES) { - const count = nonnegativeInteger(counts[status], `${label} ${status} count`); - if (!SUCCESSFUL_NODE_STATUSES.has(status) && count !== 0) { - throw new Error(`${label} has non-successful node status ${status}`); - } - total += count; - } - if (total !== expectedCount) throw new Error(`${label} status counts do not match its observed nodes`); -} - -function nonemptyUniqueStrings(value, label) { - if (!Array.isArray(value) || value.length === 0 || value.some((entry) => typeof entry !== "string" || entry === "")) { - throw new Error(`${label} must be a non-empty string array`); - } - if (new Set(value).size !== value.length) throw new Error(`${label} contains duplicates`); - return value; -} - -function sameStringSet(left, right) { - const rightSet = new Set(right); - return left.length === right.length && left.every((value) => rightSet.has(value)); -} - -function jsonBundleFile(files, filePath) { - const contents = files.get(filePath); - if (contents === undefined) throw new Error(`threat-model gate bundle is missing ${filePath}`); - try { - return JSON.parse(contents.toString("utf8")); - } catch (error) { - throw new Error(`threat-model gate bundle ${filePath} is not valid JSON`, { cause: error }); - } -} - -function readJsonRegular(filePath, maxBytes, label) { - const resolved = path.resolve(filePath); - const stat = fs.lstatSync(resolved); - if (!stat.isFile() || stat.isSymbolicLink() || stat.size > maxBytes) { - throw new Error(`${label} must be a bounded regular file`); - } - return JSON.parse(fs.readFileSync(resolved, "utf8")); -} - -function regularFileInside(root, parts, maxBytes, label) { - const resolved = path.resolve(root, ...parts); - if (resolved !== root && !resolved.startsWith(`${root}${path.sep}`)) - throw new Error(`${label} escapes its results root`); - const stat = fs.lstatSync(resolved); - if (!stat.isFile() || stat.isSymbolicLink() || stat.size > maxBytes) { - throw new Error(`${label} must be a bounded regular file`); - } - return resolved; -} - -function record(value, label) { - if (value === null || typeof value !== "object" || Array.isArray(value)) - throw new Error(`${label} must be an object`); - return value; -} - -function safeId(value, label) { - if (typeof value !== "string" || !SAFE_ID.test(value)) throw new Error(`${label} is invalid`); - return value; -} - -function positiveInteger(value, label) { - if (!Number.isSafeInteger(value) || value <= 0) throw new Error(`${label} must be a positive integer`); - return value; -} - -function nonnegativeInteger(value, label) { - if (!Number.isSafeInteger(value) || value < 0) throw new Error(`${label} must be a non-negative integer`); - return value; -} - -function main(args) { - if (args.length !== 2) { - throw new Error( - "usage: validate-threat-model-benchmark-gate.mjs " - ); - } - const result = validateThreatModelBenchmarkGateFiles(args[0], args[1]); - process.stdout.write(`${JSON.stringify(result)}\n`); -} - -if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) main(process.argv.slice(2)); diff --git a/scripts/ci/validate-threat-model-benchmark-gate.test.ts b/scripts/ci/validate-threat-model-benchmark-gate.test.ts deleted file mode 100644 index 8f102d641..000000000 --- a/scripts/ci/validate-threat-model-benchmark-gate.test.ts +++ /dev/null @@ -1,333 +0,0 @@ -import { afterEach, describe, expect, it } from "bun:test"; -import crypto from "node:crypto"; -import fs from "node:fs"; -import os from "node:os"; -import path from "node:path"; - -import { - validateThreatModelBenchmarkGate, - validateThreatModelBenchmarkGateFiles -} from "./validate-threat-model-benchmark-gate.mjs"; - -const roots: string[] = []; -const candidate = "a".repeat(40); -const modelSlug = "benchmark-threat-model-gpt-5-6-luna-high"; -const pairId = `ultrafuzz-bench-${modelSlug}`; -const targets = [ - { - id: "very-liquid-vaults-foundry", - repository: "https://github.com/rheo-xyz/very-liquid-vaults", - revision: "b".repeat(40), - framework: "foundry" - }, - { - id: "venus-isolated-pools-hardhat", - repository: "https://github.com/code-423n4/2023-05-venus", - revision: "c".repeat(40), - framework: "hardhat" - }, - { - id: "stableswap-ng-vyper", - repository: "https://github.com/curvefi/stableswap-ng", - revision: "d".repeat(40), - framework: "vyper" - } -]; - -afterEach(() => { - for (const root of roots.splice(0)) fs.rmSync(root, { recursive: true, force: true }); -}); - -describe("threat-model benchmark structural gate", () => { - it("accepts complete planner-independent fanout evidence for the canonical cohort", () => { - const fixture = gateFixture(); - expect(validateThreatModelBenchmarkGate(fixture.manifest, fixture.pairBundles)).toEqual({ - pair_count: 1, - target_count: 3, - row_count: 3, - expected_dynamic_child_count: 6 - }); - }); - - it("reads the exact collected bundle path named by the launch manifest", () => { - const fixture = gateFixture(); - const root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ultrafuzz-threat-model-gate-")); - roots.push(root); - const manifestPath = path.join(root, "control", "manifest.json"); - const resultsRoot = path.join(root, "results"); - const bundlePath = path.join(resultsRoot, pairId, modelSlug, "public-results.json"); - fs.mkdirSync(path.dirname(manifestPath), { recursive: true }); - fs.mkdirSync(path.dirname(bundlePath), { recursive: true }); - fs.writeFileSync(manifestPath, `${JSON.stringify(fixture.manifest)}\n`); - fs.writeFileSync(bundlePath, `${JSON.stringify(fixture.pairBundles[0]!.bundle)}\n`); - - expect(validateThreatModelBenchmarkGateFiles(manifestPath, resultsRoot)).toMatchObject({ - target_count: 3, - row_count: 3 - }); - }); - - it("fails a mismatched or unavailable planner-owned expectation", () => { - const mismatch = gateFixture(); - firstRow(mismatch).expansion.expected_vs_actual = { - expected_child_count: 2, - actual_dynamic_node_count: 1, - delta: -1, - matches: false - }; - expect(() => validateFixture(mismatch)).toThrow(/does not match the planner-owned expectation/u); - - const unavailable = gateFixture(); - firstRow(unavailable).expansion.plan_evidence = { status: "unavailable", reason: "goal-plan-unavailable" }; - expect(() => validateFixture(unavailable)).toThrow(/plan_evidence is not complete/u); - }); - - it("fails missing lineage and incomplete, failed, or timed-out goal lanes", () => { - const missingLineage = gateFixture(); - firstRow(missingLineage).expansion.dynamic_nodes[0].source_node_id = null; - expect(() => validateFixture(missingLineage)).toThrow(/failed or unattributed dynamic child/u); - - const missingLane = gateFixture(); - const lane = firstRow(missingLane).expansion.goal_lanes[0]; - lane.observed_node_count = 0; - lane.observed_node_ids = []; - expect(() => validateFixture(missingLane)).toThrow(/observed nodes must be a non-empty string array/u); - - const failedLane = gateFixture(); - const failed = firstRow(failedLane).expansion.goal_lanes[0]; - failed.failed = true; - failed.failed_node_ids = [failed.observed_node_ids[0]]; - failed.status_counts.succeeded = 0; - failed.status_counts.failed = 1; - expect(() => validateFixture(failedLane)).toThrow(/was not independently and successfully observed/u); - - const timedOut = gateFixture(); - firstRow(timedOut).expansion.dynamic_nodes[0].status = "timed-out"; - firstRow(timedOut).expansion.dynamic_nodes[0].timed_out = true; - expect(() => validateFixture(timedOut)).toThrow(/failed or unattributed dynamic child/u); - - const swappedAcrossLanes = gateFixture(); - const firstObserved = firstRow(swappedAcrossLanes).expansion.goal_lanes[0].observed_planned_node_ids; - const secondObserved = firstRow(swappedAcrossLanes).expansion.goal_lanes[1].observed_planned_node_ids; - firstRow(swappedAcrossLanes).expansion.goal_lanes[0].observed_planned_node_ids = secondObserved; - firstRow(swappedAcrossLanes).expansion.goal_lanes[1].observed_planned_node_ids = firstObserved; - expect(() => validateFixture(swappedAcrossLanes)).toThrow(/exact planner-owned node set/u); - - const reusedAcrossLanes = gateFixture(); - firstRow(reusedAcrossLanes).expansion.goal_lanes[1].observed_node_ids = [ - firstRow(reusedAcrossLanes).expansion.goal_lanes[0].observed_node_ids[0] - ]; - expect(() => validateFixture(reusedAcrossLanes)).toThrow(/more than one goal lane/u); - - const childOutsideLanes = gateFixture(); - firstRow(childOutsideLanes).expansion.goal_lanes[0].observed_node_ids = ["unrelated-succeeded-node"]; - expect(() => validateFixture(childOutsideLanes)).toThrow(/outside its planner-owned goal lanes/u); - }); - - it("fails missing review artifacts, noncanonical targets, and serialized fanout", () => { - const missingArtifact = gateFixture(); - missingArtifact.pairBundles[0].bundle.files = missingArtifact.pairBundles[0].bundle.files.filter( - (file) => !file.path.endsWith("/artifacts/goal-plan/goal-plan.json") - ); - expect(() => validateFixture(missingArtifact)).toThrow(/missing retained artifact/u); - - const wrongTarget = gateFixture(); - wrongTarget.manifest.targets[0].id = "replacement-target"; - expect(() => validateFixture(wrongTarget)).toThrow(/exactly the three canonical/u); - - const serialized = gateFixture(); - firstRow(serialized).expansion.concurrency.effective = 1; - expect(() => validateFixture(serialized)).toThrow(/did not demonstrate independently scheduled/u); - }); - - it("rejects configured or observed serial concurrency even for a one-child fanout", () => { - const configuredSerial = gateFixture(); - configuredSerial.manifest.concurrency.max_parallel_workflow_nodes_per_row = 1; - expect(() => validateFixture(configuredSerial)).toThrow(/requested workflow concurrency must exceed one/u); - - const observedSerial = gateFixture(); - const row = firstRow(observedSerial); - row.expansion.dynamic_node_count = 1; - row.expansion.dynamic_nodes = row.expansion.dynamic_nodes.slice(0, 1); - row.expansion.plan.expected_child_count = 1; - row.expansion.plan.lane_count = 2; - row.expansion.expected_vs_actual = { - expected_child_count: 1, - actual_dynamic_node_count: 1, - delta: 0, - matches: true - }; - row.expansion.goal_lanes = [row.expansion.goal_lanes[0], row.expansion.goal_lanes[2]]; - row.expansion.concurrency.effective = 1; - expect(() => validateFixture(observedSerial)).toThrow(/did not demonstrate independently scheduled/u); - }); - - it("keeps quality, cost, retry and wall-clock values as telemetry", () => { - const fixture = gateFixture(); - const row = firstRow(fixture); - row.precision = 0; - row.recall = 0; - row.f1_score = 0; - row.expansion.retried_node_count = 99; - for (const lane of row.expansion.goal_lanes) { - lane.total_tokens = null; - lane.cost_usd = null; - lane.wall_time_seconds = null; - lane.cost_evidence = { status: "unavailable", reason: "usage-ledger-unavailable" }; - } - row.expansion.lane_cost_evidence = { status: "unavailable", reason: "usage-ledger-unavailable" }; - - expect(() => validateFixture(fixture)).not.toThrow(); - }); -}); - -function gateFixture() { - const manifest = { - candidate_commit: candidate, - mode: "threat-model", - benchmark: "ultrafuzz-bench", - targets: structuredClone(targets), - matrix_rows_per_pair: 3, - concurrency: { max_parallel_workflow_nodes_per_row: 8 }, - pairs: [ - { - pair: pairId, - benchmark: "ultrafuzz-bench", - mode: "threat-model", - lane: "threat-model", - model_slug: modelSlug, - provider: "openai" - } - ] - }; - const rows = targets.map((target, index) => scoreRow(target.id, index)); - const summary = { rows }; - const files = [bundleFile("eval/summary.json", JSON.stringify(summary))]; - for (const row of rows) { - files.push( - bundleFile(`reports/${row.row_id}/artifacts/threat-model/THREAT_MODEL.md`, "# Threat model\n"), - bundleFile(`reports/${row.row_id}/artifacts/threat-model/threat-model.json`, '{"threats":[]}\n'), - bundleFile(`reports/${row.row_id}/artifacts/goal-plan/goal-plan.json`, '{"goals":[]}\n'), - bundleFile(`reports/${row.row_id}/artifacts/goal-plan/vulnerability-db-manifest.json`, '{"sha256":"abc"}\n') - ); - } - const bundle = { - status: "succeeded", - benchmark: "ultrafuzz-bench", - lane: "threat-model", - model_slug: modelSlug, - model: "gpt-5.6-luna", - reasoning: "high", - candidate_commit: candidate, - executed_case_count: 3, - graded_case_count: 3, - targets: targets.map((target) => ({ ...target, status: "succeeded" })), - files - }; - return { manifest, pairBundles: [{ pair: pairId, bundle }], summary }; -} - -function scoreRow(targetId: string, index: number) { - const dynamicIds = [`dynamic-threat-${index}`, `dynamic-class-${index}`]; - const lanes = [ - goalLane(`threat-${index}`, dynamicIds[0]), - goalLane(`class-${index}`, dynamicIds[1]), - goalLane("goal-roaming", `goal-roaming-${index}`) - ]; - return { - row_id: `row-${index}`, - target_id: targetId, - precision: 1, - recall: 1, - f1_score: 1, - expansion: { - dynamic_node_count: 2, - dynamic_nodes: dynamicIds.map((nodeId) => ({ - node_id: nodeId, - source_node_id: "goal-plan", - status: "succeeded", - timed_out: false - })), - plan: { - expected_child_count: 2, - threat_count: 1, - applicable_class_count: 1, - max_dynamic_nodes: 2048, - lane_count: 3 - }, - expected_vs_actual: { - expected_child_count: 2, - actual_dynamic_node_count: 2, - delta: 0, - matches: true - }, - goal_lanes: lanes, - truncated: false, - nodes: complete(), - lineage: complete(), - concurrency_evidence: complete(), - plan_evidence: complete(), - lane_cost_evidence: complete(), - concurrency: { requested: 8, effective: 6, ready_queue_depth: 0, active_work: 0 }, - retried_node_count: 0 - } - }; -} - -function goalLane(laneId: string, nodeId: string) { - return { - lane_id: laneId, - planned_node_ids: [nodeId], - observed_planned_node_ids: [nodeId], - observed_node_ids: [nodeId], - observed_node_count: 1, - status_counts: statusCounts(), - failed: false, - failed_node_ids: [], - timed_out_node_ids: [], - total_tokens: 100, - cost_usd: 0.01, - wall_time_seconds: 1, - cost_evidence: complete() - }; -} - -function statusCounts() { - return { - pending: 0, - ready: 0, - runnable: 0, - running: 0, - succeeded: 1, - failed: 0, - skipped: 0, - "timed-out": 0, - "reused-from-prior-run": 0, - invalidated: 0 - }; -} - -function complete() { - return { status: "complete", reason: null }; -} - -function bundleFile(filePath: string, contents: string) { - const bytes = Buffer.from(contents, "utf8"); - return { - path: filePath, - size_bytes: bytes.byteLength, - sha256: crypto.createHash("sha256").update(bytes).digest("hex"), - contents_base64: bytes.toString("base64") - }; -} - -function firstRow(fixture: ReturnType) { - return fixture.summary.rows[0]!; -} - -function validateFixture(fixture: ReturnType): unknown { - const files = fixture.pairBundles[0]!.bundle.files; - const index = files.findIndex((file) => file.path === "eval/summary.json"); - files[index] = bundleFile("eval/summary.json", JSON.stringify(fixture.summary)); - return validateThreatModelBenchmarkGate(fixture.manifest, fixture.pairBundles); -} From 5ccd0a8748686eada78a80c5a4e9f1af7cc41b01 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:52:05 +0000 Subject: [PATCH 049/206] refactor(modal): delete unreferenced exports and move test-only helpers into tests Deleted, no reference anywhere: MAX_MODAL_DOCUMENT_BYTES, MODAL_SMOKE_STOP_PATH (smoke-worker.ts builds its own stop path) and DEFAULT_NODE_TIMEOUT_SECONDS. Deleted with the tests that only exercised them: - createTrackedSourceArchive; production archives the exact candidate with createExactCandidateSourceArchive. - persistentWorkspaceRoot. - locateModalResumeWorkspace, documented as "retained for tests"; worker.ts uses findModalResumeWorkspace, and the removed "locates one exact linked durable run" test duplicated the existing tolerant-lookup test. The two finalize tests unwrap findModalResumeWorkspace through a local helper. Moved into tests, their only users: - DEFAULT_BENCHMARK_MODELS, a stale model list no production code reads, is now test/model-spec-fixtures.ts, typed as a const tuple so indexed fixtures need no non-null assertion. - modalBenchmarkConfigValidatorsAgree, WORKER_RESULT_ALLOWED_KEYS (the test now states the expected persisted key set), and publicEvalFailureDiagnosticLogPayload, which composed the same two calls runCommand already makes inline. - PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS and its full-lane twin, aliases of the evals lane limits that only a test imported. Un-exported, used only inside their own module: five public-worker constants and WORKER_STDERR_TAIL_BYTES. The source-constant reader in prepare-eval-history-publication.mjs matches top-level const declarations with or without export, and its test against the real public-worker.ts still passes. Co-Authored-By: Claude Opus 5.5 --- packages/modal/src/config.ts | 9 --- packages/modal/src/defaults.ts | 68 ------------------- packages/modal/src/layout.ts | 4 -- packages/modal/src/modal-documents.ts | 3 - packages/modal/src/public-worker.ts | 22 ++----- packages/modal/src/resume.ts | 10 --- packages/modal/src/runner.ts | 12 ---- packages/modal/src/smoke.ts | 1 - packages/modal/src/worker-diagnostics.ts | 2 +- packages/modal/src/worker-result.ts | 16 ----- packages/modal/test/config.test.ts | 28 ++++---- packages/modal/test/layout.test.ts | 2 - packages/modal/test/model-spec-fixtures.ts | 69 ++++++++++++++++++++ packages/modal/test/public-worker.test.ts | 16 +++-- packages/modal/test/resume.test.ts | 29 +++----- packages/modal/test/runner.test.ts | 20 ------ packages/modal/test/worker-lineage.test.ts | 5 +- packages/modal/test/worker-result.test.ts | 21 ++++-- packages/modal/test/workspace-config.test.ts | 8 +-- 19 files changed, 134 insertions(+), 211 deletions(-) create mode 100644 packages/modal/test/model-spec-fixtures.ts diff --git a/packages/modal/src/config.ts b/packages/modal/src/config.ts index cd055d46a..dacb87aef 100644 --- a/packages/modal/src/config.ts +++ b/packages/modal/src/config.ts @@ -3,8 +3,6 @@ import { resolveConfig } from "@ultrafuzz/config"; import { MODAL_BENCHMARK_CONFIG_SCHEMA_ID, type StrictModalBenchmarkConfigDocument } from "./modal-contracts.js"; import { assertModalDocumentValue, readModalDocument } from "./modal-documents.js"; -import { validateModalJsonSchema } from "./modal-schema-registry.js"; -import { modalBenchmarkConfigZodSchema } from "./benchmark-config-zod.js"; import { MODAL_PUBLIC_FULL_SANDBOX_TIMEOUT_MS, MODAL_PUBLIC_SANDBOX_TIMEOUT_MS, @@ -62,13 +60,6 @@ export function fingerprintModalModel(model: ModalModelSpec): string { .digest("hex"); } -export function modalBenchmarkConfigValidatorsAgree(value: unknown): boolean { - return ( - validateModalJsonSchema(MODAL_BENCHMARK_CONFIG_SCHEMA_ID, value).ok === - modalBenchmarkConfigZodSchema.safeParse(value).success - ); -} - /** Refuse an execution envelope that cannot contain even one complete campaign. * This is separate from document validation: current-schema configs remain inspectable. */ diff --git a/packages/modal/src/defaults.ts b/packages/modal/src/defaults.ts index fa9e77462..183820094 100644 --- a/packages/modal/src/defaults.ts +++ b/packages/modal/src/defaults.ts @@ -27,7 +27,6 @@ export const EVAL_WATCH_TIMEOUT_SECONDS = (MODAL_SANDBOX_TIMEOUT_MS - EVAL_POST_ export const MODAL_PRE_MODEL_RETRY_LIMIT = 3; export const MODAL_PRE_MODEL_RETRY_BASE_DELAY_MS = 1_000; export const MODAL_PRE_MODEL_RETRY_MAX_DELAY_MS = 4_000; -export const DEFAULT_NODE_TIMEOUT_SECONDS = 2 * 60 * 60; export const DEFAULT_MODAL_APP = "ultrafuzz-evals"; export const DEFAULT_MODAL_IMAGE = "ultrafuzz-security-runner:latest"; export const MODAL_BENCHMARK_SANDBOX_RESOURCES = { @@ -50,70 +49,3 @@ export interface ModalModelSpec { reasoning: string; auth_mode: ModelAuthMode; } - -export const DEFAULT_BENCHMARK_MODELS: readonly ModalModelSpec[] = [ - { - slug: "gpt-5-5", - model: "gpt-5.5", - provider: "openai", - agent: "CodexAgent", - reasoning: "xhigh", - auth_mode: "subscription" - }, - { - slug: "gpt-5-6-sol", - model: "gpt-5.6-sol", - provider: "openai", - agent: "CodexAgent", - reasoning: "xhigh", - auth_mode: "subscription" - }, - { - slug: "gpt-5-6-terra", - model: "gpt-5.6-terra", - provider: "openai", - agent: "CodexAgent", - reasoning: "xhigh", - auth_mode: "subscription" - }, - { - slug: "gpt-5-6-luna", - model: "gpt-5.6-luna", - provider: "openai", - agent: "CodexAgent", - reasoning: "xhigh", - auth_mode: "subscription" - }, - { - slug: "claude-fable-5", - model: "claude-fable-5", - provider: "anthropic", - agent: "ClaudeAgent", - reasoning: "max", - auth_mode: "subscription" - }, - { - slug: "claude-opus-4-8", - model: "claude-opus-4-8", - provider: "anthropic", - agent: "ClaudeAgent", - reasoning: "max", - auth_mode: "subscription" - }, - { - slug: "kimi-k3", - model: "kimi-k3", - provider: "kimi", - agent: "KimiAgent", - reasoning: "max", - auth_mode: "subscription" - }, - { - slug: "deepseek-v4-pro", - model: "deepseek-v4-pro", - provider: "deepseek", - agent: "DeepSeekAgent", - reasoning: "max", - auth_mode: "api-key" - } -] as const; diff --git a/packages/modal/src/layout.ts b/packages/modal/src/layout.ts index 87088dd8c..e337fca05 100644 --- a/packages/modal/src/layout.ts +++ b/packages/modal/src/layout.ts @@ -14,10 +14,6 @@ export function persistentDataRoot(runId: string, slug: string): string { return path.posix.join("/data", runId, slug); } -export function persistentWorkspaceRoot(runId: string, slug: string): string { - return path.posix.join(persistentDataRoot(runId, slug), "workspace"); -} - export function modalVolumeName(runId: string, slug: string): string { const identity = `${runId}\0${slug}`; const suffix = createHash("sha256").update(identity).digest("hex").slice(0, 12); diff --git a/packages/modal/src/modal-documents.ts b/packages/modal/src/modal-documents.ts index 6d7444434..544d2f7ae 100644 --- a/packages/modal/src/modal-documents.ts +++ b/packages/modal/src/modal-documents.ts @@ -6,7 +6,6 @@ import path from "node:path"; import { assertNoSymlinkComponents, assertPathInside, - DEFAULT_MAX_JSON_INSTANCE_BYTES, parseStrictJsonBytes, readRegularFileSnapshot, validateRegisteredJsonBytesSync, @@ -24,8 +23,6 @@ import { } from "./modal-schema-registry.js"; import { assertModalDocumentSemantics } from "./modal-semantic-gates.js"; -export const MAX_MODAL_DOCUMENT_BYTES = DEFAULT_MAX_JSON_INSTANCE_BYTES; - export interface ModalDocumentSnapshot { readonly schema_id: SchemaId; readonly schema_sha256: string; diff --git a/packages/modal/src/public-worker.ts b/packages/modal/src/public-worker.ts index e45472acc..131cbdfcf 100644 --- a/packages/modal/src/public-worker.ts +++ b/packages/modal/src/public-worker.ts @@ -9,8 +9,6 @@ import { adaptBenchmarkManifestToEvalSuite, benchmarkLaneConcurrency, benchmarkLaneSelectedTargetIds, - BENCHMARK_FULL_MAX_PARALLEL_RUNS, - BENCHMARK_SMOKE_MAX_PARALLEL_RUNS, boundedEvalId, evalSuiteInputDocument, evalRunRoot, @@ -84,13 +82,11 @@ import { OperationalDispositionError, runNamingUnhandledFailure } from "./termin const ULTRAFUZZ_ROOT = "/opt/ultrafuzz"; const BAKED_CANDIDATE_ARCHIVE = "/opt/ultrafuzz-source.tgz"; -export const PUBLIC_SMITHERS_SEED_ROOT = "/opt/ultrafuzz-smithers-seed"; +const PUBLIC_SMITHERS_SEED_ROOT = "/opt/ultrafuzz-smithers-seed"; const CLI = path.join(ULTRAFUZZ_ROOT, "packages/cli/dist/index.js"); const PUBLIC_BUNDLE_FILE = "public-results.json"; const PUBLIC_WORKSPACE_ROOT = "/tmp/ultrafuzz-public-workspace"; -export const PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS = BENCHMARK_SMOKE_MAX_PARALLEL_RUNS; -export const PUBLIC_FULL_BENCHMARK_MAX_PARALLEL_EVAL_ROWS = BENCHMARK_FULL_MAX_PARALLEL_RUNS; -export const PUBLIC_BENCHMARK_PREPARATION_PARALLELISM = 8; +const PUBLIC_BENCHMARK_PREPARATION_PARALLELISM = 8; export const PUBLIC_BENCHMARK_EVAL_CLEANUP_SECONDS = 5 * 60; export const PUBLIC_BENCHMARK_SCORE_PER_WAVE_TIMEOUT_SECONDS = 45 * 60; export const PUBLIC_BENCHMARK_REPORT_TIMEOUT_SECONDS = 5 * 60; @@ -99,12 +95,12 @@ export const PUBLIC_BENCHMARK_PREPARATION_TIMEOUT_SECONDS = 20 * 60; // whose runner is OpenRouter runs one eval row at a time and one workflow node inside it. // Every trusted policy re-derivation reads this constant, so the launch, cleanup, and // publication guardrails agree with the suite the Modal worker actually runs. -export const PUBLIC_BENCHMARK_OPENROUTER_MAX_PARALLEL = 1; +const PUBLIC_BENCHMARK_OPENROUTER_MAX_PARALLEL = 1; // The smoke graph has four sequential agent stages. Each stage may use both of // its 1,800-second attempts, so retain ten minutes beyond the four-hour // topology bound for workflow transitions and final synchronization. -export const PUBLIC_BENCHMARK_SMOKE_MAX_RUNTIME_SECONDS = 4 * 60 * 60 + 10 * 60; -export const PUBLIC_BENCHMARK_THREAT_MODEL_MAX_RUNTIME_SECONDS = 15_000; +const PUBLIC_BENCHMARK_SMOKE_MAX_RUNTIME_SECONDS = 4 * 60 * 60 + 10 * 60; +const PUBLIC_BENCHMARK_THREAT_MODEL_MAX_RUNTIME_SECONDS = 15_000; // A manual full row retains the packaged specialist timeouts, including the // 7,200-second invariant campaign. This is a bounded row execution budget, not // a guarantee that every topology node can consume its worst-case timeout. @@ -1664,14 +1660,6 @@ export function publicEvalFailureEnvelopeDiagnostics(stdout: string): WorkerDiag return failedEvalDiagnostics(parsed.diagnostics); } -export function publicEvalFailureDiagnosticLogPayload( - stdout: string, - forbiddenSecretValues: readonly string[] -): string | undefined { - const diagnostics = publicEvalFailureEnvelopeDiagnostics(stdout); - return diagnostics === undefined ? undefined : workerDiagnosticLogPayload(diagnostics, forbiddenSecretValues); -} - export function publicEvalFailureDiagnosticLogPayloadFromRecords( records: readonly EvalRunRecord[], forbiddenSecretValues: readonly string[] diff --git a/packages/modal/src/resume.ts b/packages/modal/src/resume.ts index a332871a1..b29527add 100644 --- a/packages/modal/src/resume.ts +++ b/packages/modal/src/resume.ts @@ -396,16 +396,6 @@ async function readdirIfMissing(directoryPath: string): Promise { }); } -/** - * The strict form. Retained for tests and for any caller that has already established the run must be - * resumable; `worker.ts` deliberately uses the tolerant lookup above instead. - */ -export async function locateModalResumeWorkspace(workRoot: string): Promise { - const found = await findModalResumeWorkspace(workRoot); - if (found.kind === "not-started") throw new Error(found.reason); - return found.workspace; -} - export async function finalizeModalEvalRunRecord( workspace: ModalResumeWorkspace, state: ModalResumeRunState, diff --git a/packages/modal/src/runner.ts b/packages/modal/src/runner.ts index dcb9e797e..254e39877 100644 --- a/packages/modal/src/runner.ts +++ b/packages/modal/src/runner.ts @@ -2959,18 +2959,6 @@ export function modalSecurityToolchainCommands(): string[] { ]; } -export function createTrackedSourceArchive( - repoRoot: string, - archive = path.join(realpathSync(os.tmpdir()), `ultrafuzz-modal-source-${String(process.pid)}.tgz`) -): string { - const trackedFiles = execFileSync("git", ["ls-files", "-z"], { cwd: repoRoot }); - if (trackedFiles.length === 0) throw new Error(`no Git-tracked source files found under ${repoRoot}`); - execFileSync("tar", ["--null", "-czf", archive, "-C", repoRoot, "--files-from=-"], { - input: trackedFiles - }); - return archive; -} - /** * Create a self-contained, shallow Git checkout for the exact candidate HEAD. * The worker extracts this immutable archive instead of cloning the candidate diff --git a/packages/modal/src/smoke.ts b/packages/modal/src/smoke.ts index 4384b1963..900570297 100644 --- a/packages/modal/src/smoke.ts +++ b/packages/modal/src/smoke.ts @@ -4,7 +4,6 @@ import { remoteAuthDir, remoteAuthPath } from "./layout.js"; export const MODAL_SMOKE_RESULT_SCHEMA_VERSION = "ultrafuzz.modal.smoke-result.v1" as const; export const MODAL_SMOKE_ENTRY_PATH = "/opt/ultrafuzz/packages/modal/dist/smoke-worker.js"; export const MODAL_SMOKE_DATA_ROOT = "/data/ultrafuzz-modal-smoke"; -export const MODAL_SMOKE_STOP_PATH = `${MODAL_SMOKE_DATA_ROOT}/fresh-stop`; export type ModalSmokePhase = "fresh" | "resume"; export type ModalSmokeFailureStage = diff --git a/packages/modal/src/worker-diagnostics.ts b/packages/modal/src/worker-diagnostics.ts index 856b8180c..d61d30e28 100644 --- a/packages/modal/src/worker-diagnostics.ts +++ b/packages/modal/src/worker-diagnostics.ts @@ -4,7 +4,7 @@ import { redactSecretsInText, SENSITIVE_REDACTION_PLACEHOLDER } from "@ultrafuzz * Bytes of child stderr retained while the rest is discarded, so a non-zero exit can name its own reason * without persisting provider responses or benchmark contents. */ -export const WORKER_STDERR_TAIL_BYTES = 8_192; +const WORKER_STDERR_TAIL_BYTES = 8_192; /** * Byte bound on one diagnostic message. It is the collector's bound, not this module's: diff --git a/packages/modal/src/worker-result.ts b/packages/modal/src/worker-result.ts index 3e04dc1b7..0687874d2 100644 --- a/packages/modal/src/worker-result.ts +++ b/packages/modal/src/worker-result.ts @@ -24,22 +24,6 @@ import { export const WORKER_RESULT_SCHEMA_VERSION = "ultrafuzz.modal.worker-result.v2" as const; -export const WORKER_RESULT_ALLOWED_KEYS = [ - "schema_version", - "result_type", - "generation", - "launch_generation", - "attempt", - "model_work_started", - "counts", - "checkpoint", - "exit_category", - "runtime_ms", - "usage", - "pricing", - "diagnostic_code" -] as const; - export const WORKER_DIAGNOSTIC_CODES = [ "worker-live", "worker-finished", diff --git a/packages/modal/test/config.test.ts b/packages/modal/test/config.test.ts index 76cad1563..98b9f2481 100644 --- a/packages/modal/test/config.test.ts +++ b/packages/modal/test/config.test.ts @@ -13,11 +13,9 @@ import { MODAL_GIT_URL_PATTERN_SOURCE, MODAL_HTTPS_URL_PATTERN_SOURCE, modalBenchmarkConfigZodSchema, - modalBenchmarkConfigValidatorsAgree, parseModalBenchmarkConfig } from "../src/config.js"; import { - DEFAULT_BENCHMARK_MODELS, MODAL_BENCHMARK_SCHEMA_VERSION, MODAL_PUBLIC_FULL_SANDBOX_TIMEOUT_MS, MODAL_PUBLIC_SANDBOX_TIMEOUT_MS, @@ -27,10 +25,18 @@ import { MODAL_BENCHMARK_CONFIG_SCHEMA_ID } from "../src/modal-contracts.js"; import { ModalDocumentValidationError } from "../src/modal-documents.js"; import { modalBenchmarkConfigJsonSchema, validateModalJsonSchema } from "../src/modal-schema-registry.js"; import { ModalSemanticValidationError } from "../src/modal-semantic-gates.js"; +import { MODEL_SPEC_FIXTURES } from "./model-spec-fixtures.js"; + +function modalBenchmarkConfigValidatorsAgree(value: unknown): boolean { + return ( + validateModalJsonSchema(MODAL_BENCHMARK_CONFIG_SCHEMA_ID, value).ok === + modalBenchmarkConfigZodSchema.safeParse(value).success + ); +} function minimalConfig(): Record { return { - ...commonConfig("example-run", [...DEFAULT_BENCHMARK_MODELS]), + ...commonConfig("example-run", [...MODEL_SPEC_FIXTURES]), target: { repo: "https://example.invalid/target.git", ref: "0123456789abcdef" }, ground_truth: { repo: "https://example.invalid/ground-truth.git", @@ -60,7 +66,7 @@ function commonConfig(runId: string, models: unknown[]): Record } function minimalPublicConfig() { - const model = DEFAULT_BENCHMARK_MODELS[0]!; + const model = MODEL_SPEC_FIXTURES[0]; return { ...commonConfig("public-run", [model]), public_benchmark: { @@ -232,7 +238,7 @@ describe("Modal benchmark config", () => { it("bounds new public sandboxes without cutting off the accepted full-lane envelope", () => { const privateConfig = parseModalBenchmarkConfig(minimalConfig()); - const model = DEFAULT_BENCHMARK_MODELS[0]!; + const model = MODEL_SPEC_FIXTURES[0]; const publicConfig = parseModalBenchmarkConfig({ ...commonConfig("public-run", [model]), public_benchmark: { @@ -307,7 +313,7 @@ describe("Modal benchmark config", () => { label: "duplicate projected model slugs", value: { ...minimalConfig(), - models: [DEFAULT_BENCHMARK_MODELS[0], DEFAULT_BENCHMARK_MODELS[0]] + models: [MODEL_SPEC_FIXTURES[0], MODEL_SPEC_FIXTURES[0]] } }, { @@ -349,7 +355,7 @@ describe("Modal benchmark config", () => { expect(() => parseModalBenchmarkConfig({ ...minimalConfig(), - models: [{ ...DEFAULT_BENCHMARK_MODELS[0], agent: "ClaudeAgent" }] + models: [{ ...MODEL_SPEC_FIXTURES[0], agent: "ClaudeAgent" }] }) ).toThrow(); }); @@ -634,8 +640,8 @@ describe("Modal benchmark config", () => { const first = fingerprintModalConfigFile(file); fs.writeFileSync(file, `${JSON.stringify(minimalConfig(), null, 2)}\n`); expect(fingerprintModalConfigFile(file)).not.toBe(first); - expect(fingerprintModalModel(DEFAULT_BENCHMARK_MODELS[0]!)).not.toBe( - fingerprintModalModel({ ...DEFAULT_BENCHMARK_MODELS[0]!, reasoning: "different" }) + expect(fingerprintModalModel(MODEL_SPEC_FIXTURES[0])).not.toBe( + fingerprintModalModel({ ...MODEL_SPEC_FIXTURES[0], reasoning: "different" }) ); }); @@ -660,7 +666,7 @@ describe("Modal benchmark config", () => { file, `${JSON.stringify({ ...config, - models: [DEFAULT_BENCHMARK_MODELS[0], DEFAULT_BENCHMARK_MODELS[0]] + models: [MODEL_SPEC_FIXTURES[0], MODEL_SPEC_FIXTURES[0]] })}\n` ); expect(() => loadModalBenchmarkConfig(file)).toThrow(/trusted semantic gates/u); @@ -672,7 +678,7 @@ describe("Modal benchmark config", () => { valid, { ...valid, unexpected: true }, { ...valid, loops: "3" }, - { ...valid, models: [DEFAULT_BENCHMARK_MODELS[0], DEFAULT_BENCHMARK_MODELS[0]] }, + { ...valid, models: [MODEL_SPEC_FIXTURES[0], MODEL_SPEC_FIXTURES[0]] }, { ...valid, benchmark_execution: { excluded_node_ids: ["boundary-tests", "boundary-tests"] } diff --git a/packages/modal/test/layout.test.ts b/packages/modal/test/layout.test.ts index 7e28e57fe..c1ba0fe10 100644 --- a/packages/modal/test/layout.test.ts +++ b/packages/modal/test/layout.test.ts @@ -14,7 +14,6 @@ import { REMOTE_LAUNCH_READY_PATH, REMOTE_LINEAGE_PATH, modalVolumeName, - persistentWorkspaceRoot, remoteAuthPath, resolvePersistentRemoteRoot } from "../src/layout.js"; @@ -22,7 +21,6 @@ import { modalEvalRunCommand } from "../src/resume.js"; describe("Modal storage layout", () => { it("persists workspaces while keeping config and auth ephemeral", () => { - expect(persistentWorkspaceRoot("run-1", "model-1")).toBe("/data/run-1/model-1/workspace"); expect(REMOTE_CONFIG_PATH).toBe("/run/ultrafuzz-config/benchmark.json"); expect(REMOTE_LINEAGE_PATH).toBe("/run/ultrafuzz-config/lineage.json"); expect(REMOTE_LAUNCH_READY_PATH).toBe("/run/ultrafuzz-config/launch-ready"); diff --git a/packages/modal/test/model-spec-fixtures.ts b/packages/modal/test/model-spec-fixtures.ts new file mode 100644 index 000000000..70d725d78 --- /dev/null +++ b/packages/modal/test/model-spec-fixtures.ts @@ -0,0 +1,69 @@ +import type { ModalModelSpec } from "../src/defaults.js"; + +/** Model specs used as ordinary configuration fixtures; production has no default model list. */ +export const MODEL_SPEC_FIXTURES = [ + { + slug: "gpt-5-5", + model: "gpt-5.5", + provider: "openai", + agent: "CodexAgent", + reasoning: "xhigh", + auth_mode: "subscription" + }, + { + slug: "gpt-5-6-sol", + model: "gpt-5.6-sol", + provider: "openai", + agent: "CodexAgent", + reasoning: "xhigh", + auth_mode: "subscription" + }, + { + slug: "gpt-5-6-terra", + model: "gpt-5.6-terra", + provider: "openai", + agent: "CodexAgent", + reasoning: "xhigh", + auth_mode: "subscription" + }, + { + slug: "gpt-5-6-luna", + model: "gpt-5.6-luna", + provider: "openai", + agent: "CodexAgent", + reasoning: "xhigh", + auth_mode: "subscription" + }, + { + slug: "claude-fable-5", + model: "claude-fable-5", + provider: "anthropic", + agent: "ClaudeAgent", + reasoning: "max", + auth_mode: "subscription" + }, + { + slug: "claude-opus-4-8", + model: "claude-opus-4-8", + provider: "anthropic", + agent: "ClaudeAgent", + reasoning: "max", + auth_mode: "subscription" + }, + { + slug: "kimi-k3", + model: "kimi-k3", + provider: "kimi", + agent: "KimiAgent", + reasoning: "max", + auth_mode: "subscription" + }, + { + slug: "deepseek-v4-pro", + model: "deepseek-v4-pro", + provider: "deepseek", + agent: "DeepSeekAgent", + reasoning: "max", + auth_mode: "api-key" + } +] as const satisfies readonly ModalModelSpec[]; diff --git a/packages/modal/test/public-worker.test.ts b/packages/modal/test/public-worker.test.ts index 2a37dd3b5..11ec2aec5 100644 --- a/packages/modal/test/public-worker.test.ts +++ b/packages/modal/test/public-worker.test.ts @@ -43,11 +43,9 @@ import { MAX_PUBLIC_OPTIONAL_ROW_ARTIFACT_FILES, PUBLIC_OPTIONAL_ROW_ARTIFACTS, PUBLIC_BENCHMARK_EVAL_CLEANUP_SECONDS, - PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS, PUBLIC_BENCHMARK_PREPARATION_TIMEOUT_SECONDS, PUBLIC_BENCHMARK_REPORT_TIMEOUT_SECONDS, PUBLIC_BENCHMARK_SCORE_PER_WAVE_TIMEOUT_SECONDS, - PUBLIC_FULL_BENCHMARK_MAX_PARALLEL_EVAL_ROWS, PublicEvalDiagnosticsBuildError, PublicWorkerCommandInterruptedError, publicBenchmarkMaxParallelEvalRows, @@ -60,8 +58,8 @@ import { publicEvalRunId, preparePublicEvalSuite, publicCommandExitFailureCause, - publicEvalFailureDiagnosticLogPayload, publicEvalFailureDiagnosticLogPayloadFromRecords, + publicEvalFailureEnvelopeDiagnostics, publicEvalModelWorkEvidence, publicEvalCommandLeftFinalJournal, publicEvalRunErrorCanBePublished, @@ -1222,6 +1220,14 @@ it("continues after the eval command reports one publishable failed datapoint", ).toBe(false); }); +function publicEvalFailureDiagnosticLogPayload( + stdout: string, + forbiddenSecretValues: readonly string[] +): string | undefined { + const diagnostics = publicEvalFailureEnvelopeDiagnostics(stdout); + return diagnostics === undefined ? undefined : workerDiagnosticLogPayload(diagnostics, forbiddenSecretValues); +} + it("publishes only bounded redacted workflow-submission messages from eval JSON", () => { const secret = "sk-fixture-secret-value"; const payload = publicEvalFailureDiagnosticLogPayload( @@ -1857,7 +1863,7 @@ it("bounds public provider fan-out by mode", () => { ...smokeBaseSuite, run: { ...smokeBaseSuite.run, - max_parallel_runs: PUBLIC_BENCHMARK_MAX_PARALLEL_EVAL_ROWS, + max_parallel_runs: 3, max_parallel_targets: 4 } }); @@ -1877,7 +1883,7 @@ it("bounds public provider fan-out by mode", () => { expect(fullSuite.targets).toHaveLength(40); expect(fullSuite.run).toEqual({ ...fullBaseSuite.run, - max_parallel_runs: PUBLIC_FULL_BENCHMARK_MAX_PARALLEL_EVAL_ROWS, + max_parallel_runs: 20, max_parallel_targets: 8 }); expect(publicBenchmarkMaxParallelEvalRows("smoke")).toBe(3); diff --git a/packages/modal/test/resume.test.ts b/packages/modal/test/resume.test.ts index ba908623c..61de64f8b 100644 --- a/packages/modal/test/resume.test.ts +++ b/packages/modal/test/resume.test.ts @@ -15,14 +15,14 @@ import { import { findModalResumeWorkspace, - locateModalResumeWorkspace, modalDurableResumeCommand, modalDurableRunAdvanced, modalDurableRunNeedsResume, modalEvalRunCommand, NonResumableTerminalRunError, finalizeModalEvalRunRecord, - readModalDurableRunState + readModalDurableRunState, + type ModalResumeWorkspace } from "../src/resume.js"; import { currentRunState } from "./current-artifact-fixtures.js"; @@ -213,6 +213,12 @@ function fixture() { return { workRoot, target, control, evalRunId, evalDir }; } +async function resumableWorkspace(workRoot: string): Promise { + const found = await findModalResumeWorkspace(workRoot); + if (found.kind === "not-started") throw new Error(found.reason); + return found.workspace; +} + describe("Modal durable evaluation resume", () => { it("uses durable resume without any node reset path", () => { expect(modalDurableResumeCommand("/opt/tool/cli.js", "durable-run-one", "/workspace/target")).toEqual([ @@ -334,21 +340,6 @@ describe("Modal durable evaluation resume", () => { expect(modalDurableRunAdvanced(before, { ...before, run_id: "durable-run-two", status: "running" })).toBe(false); }); - it("locates one exact linked durable run and rejects ambiguity", async () => { - const value = fixture(); - await expect(locateModalResumeWorkspace(value.workRoot)).resolves.toEqual({ - target: value.target, - control: value.control, - evalRunId: value.evalRunId, - productRunId: "durable-run-one" - }); - fs.appendFileSync( - path.join(value.evalDir, "runs.jsonl"), - `${JSON.stringify(currentEvalRunRecord(value.target, value.evalRunId, { rowId: "row-two", runId: "durable-run-two" }))}\n` - ); - await expect(locateModalResumeWorkspace(value.workRoot)).rejects.toThrow("exactly one linked durable run"); - }); - it("validates the exact current eval manifest and refuses historical, mismatched, or symlinked manifests", async () => { const value = fixture(); const manifestPath = path.join(value.evalDir, "eval.json"); @@ -662,7 +653,7 @@ describe("Modal durable evaluation resume", () => { it("finalizes succeeded and genuine task outcomes without resetting completed nodes", async () => { const value = fixture(); - const workspace = await locateModalResumeWorkspace(value.workRoot); + const workspace = await resumableWorkspace(value.workRoot); await finalizeModalEvalRunRecord( workspace, { run_id: "durable-run-one", status: "succeeded", started_at: T0, finished_at: T1 }, @@ -698,7 +689,7 @@ describe("Modal durable evaluation resume", () => { it("fails closed for operational terminal states and unrelated runs", async () => { const value = fixture(); - const workspace = await locateModalResumeWorkspace(value.workRoot); + const workspace = await resumableWorkspace(value.workRoot); const rejected = finalizeModalEvalRunRecord( workspace, { run_id: "durable-run-one", status: "failed" }, diff --git a/packages/modal/test/runner.test.ts b/packages/modal/test/runner.test.ts index d668dc906..3059a1874 100644 --- a/packages/modal/test/runner.test.ts +++ b/packages/modal/test/runner.test.ts @@ -57,7 +57,6 @@ import { createModalBenchmarkSandbox, createModalLaunchSandbox, createExactCandidateSourceArchive, - createTrackedSourceArchive, finishReservedModalLaunch, hasExactPublicDiagnosticCollectionConfig, isModalRecoveryResultComplete, @@ -515,25 +514,6 @@ describe("Modal image source staging", () => { expect(standaloneDockerfile).toContain("RUN command -v zstd"); }); - it("archives tracked files only", () => { - const root = mkdtempSync(path.join(fs.realpathSync(tmpdir()), "ultrafuzz-modal-archive-")); - execFileSync("git", ["init", "--quiet"], { cwd: root }); - fs.writeFileSync(path.join(root, ".gitignore"), ".private/\n", "utf8"); - fs.writeFileSync(path.join(root, "tracked.txt"), "tracked\n", "utf8"); - fs.writeFileSync(path.join(root, "untracked.txt"), "untracked\n", "utf8"); - fs.mkdirSync(path.join(root, ".private")); - fs.writeFileSync(path.join(root, ".private", "benchmark.json"), "private\n", "utf8"); - execFileSync("git", ["add", ".gitignore", "tracked.txt"], { cwd: root }); - const archive = path.join(root, "source.tgz"); - - createTrackedSourceArchive(root, archive); - const entries = execFileSync("tar", ["-tzf", archive], { encoding: "utf8" }).trim().split("\n"); - - expect(entries).toEqual(expect.arrayContaining([".gitignore", "tracked.txt"])); - expect(entries).not.toContain("untracked.txt"); - expect(entries).not.toContain(".private/benchmark.json"); - }); - it("bakes a clean shallow Git checkout at the exact candidate commit", () => { const root = mkdtempSync(path.join(fs.realpathSync(tmpdir()), "ultrafuzz-modal-candidate-")); execFileSync("git", ["init", "--quiet"], { cwd: root }); diff --git a/packages/modal/test/worker-lineage.test.ts b/packages/modal/test/worker-lineage.test.ts index f6bc3aae0..54d171cd7 100644 --- a/packages/modal/test/worker-lineage.test.ts +++ b/packages/modal/test/worker-lineage.test.ts @@ -5,7 +5,7 @@ import path from "node:path"; import { afterEach, describe, expect, it } from "vitest"; import { fingerprintModalModel } from "../src/config.js"; -import { DEFAULT_BENCHMARK_MODELS, MODAL_WORKER_LINEAGE_SCHEMA_VERSION } from "../src/defaults.js"; +import { MODAL_WORKER_LINEAGE_SCHEMA_VERSION } from "../src/defaults.js"; import type { ModalWorkerLineage } from "../src/launch-state.js"; import { CheckpointIncompatibleError, @@ -15,6 +15,7 @@ import { readModalWorkerLineage } from "../src/worker-lineage.js"; import { emptyWorkerCheckpoint, runWithTerminalPersistence, WorkerResultWriter } from "../src/worker-result.js"; +import { MODEL_SPEC_FIXTURES } from "./model-spec-fixtures.js"; const roots: string[] = []; @@ -24,7 +25,7 @@ afterEach(() => { describe("persistent Modal worker lineage", () => { it("derives the worker model only from the exact configured lineage fingerprint", () => { - const model = DEFAULT_BENCHMARK_MODELS[0]!; + const model = MODEL_SPEC_FIXTURES[0]; const current = { ...lineage(), model_fingerprint: fingerprintModalModel(model) }; expect(modelForModalWorkerLineage({ models: [model] }, current)).toBe(model); diff --git a/packages/modal/test/worker-result.test.ts b/packages/modal/test/worker-result.test.ts index c27fe4586..d94ebacde 100644 --- a/packages/modal/test/worker-result.test.ts +++ b/packages/modal/test/worker-result.test.ts @@ -12,7 +12,6 @@ import { emptyWorkerCheckpoint, readWorkerCheckpoint, runWithTerminalPersistence, - WORKER_RESULT_ALLOWED_KEYS, WORKER_RESULT_SCHEMA_VERSION, WorkerResultWriter, type WorkerResultContract @@ -232,12 +231,20 @@ describe("strict worker result contracts", () => { }); expect(persistedStatus).toEqual(persistedResult); expect(persistedResult).toEqual(terminal); - expect(Object.keys(persistedResult).every((key) => new Set(WORKER_RESULT_ALLOWED_KEYS).has(key))).toBe( - true - ); - expect(Object.keys(persistedResult).sort()).toEqual( - WORKER_RESULT_ALLOWED_KEYS.filter((key) => key !== "pricing").sort() - ); + expect(Object.keys(persistedResult).sort()).toEqual([ + "attempt", + "checkpoint", + "counts", + "diagnostic_code", + "exit_category", + "generation", + "launch_generation", + "model_work_started", + "result_type", + "runtime_ms", + "schema_version", + "usage" + ]); expect(Object.keys(persistedResult.counts).sort()).toEqual(["failed", "remaining", "succeeded"]); expect(Object.keys(persistedResult.checkpoint).sort()).toEqual(["age_ms", "digest"]); expect(fs.statSync(path.join(root, "result.json")).mode & 0o777).toBe(0o600); diff --git a/packages/modal/test/workspace-config.test.ts b/packages/modal/test/workspace-config.test.ts index 519cdea53..f4080f204 100644 --- a/packages/modal/test/workspace-config.test.ts +++ b/packages/modal/test/workspace-config.test.ts @@ -5,13 +5,13 @@ import { parseProjectConfigToml, resolveConfig } from "@ultrafuzz/config"; import { describe, expect, it } from "vitest"; import { parse } from "yaml"; -import { DEFAULT_BENCHMARK_MODELS } from "../src/defaults.js"; import { PUBLIC_FULL_BENCHMARK_MAX_RUNTIME_SECONDS } from "../src/public-worker.js"; import { modalTargetToml } from "../src/workspace-config.js"; +import { MODEL_SPEC_FIXTURES } from "./model-spec-fixtures.js"; describe("Modal target model profiles", () => { it("overrides both explicit default and benchmark profiles with the selected model", () => { - const model = DEFAULT_BENCHMARK_MODELS[4]!; + const model = MODEL_SPEC_FIXTURES[4]; const config = modalTargetToml(model, 7_200); expect(config).toMatch(/^schema_version = "ultrafuzz\.config\.v2"$/mu); @@ -27,7 +27,7 @@ describe("Modal target model profiles", () => { }); it("selects the packaged smoke audit profile for smoke target preparation", () => { - const config = modalTargetToml(DEFAULT_BENCHMARK_MODELS[0]!, 900, "smoke"); + const config = modalTargetToml(MODEL_SPEC_FIXTURES[0], 900, "smoke"); expect(config).toMatch(/^schema_version = "ultrafuzz\.config\.v2"$/mu); expect(config).toContain('audit_profile = "smoke"'); @@ -39,7 +39,7 @@ describe("Modal target model profiles", () => { }); it("selects the packaged exhaustive audit profile for full-lane target preparation", () => { - const config = modalTargetToml(DEFAULT_BENCHMARK_MODELS[0]!, 1_800, "exhaustive"); + const config = modalTargetToml(MODEL_SPEC_FIXTURES[0], 1_800, "exhaustive"); expect(config).toContain('audit_profile = "exhaustive"'); expect(config).toContain("dynamic_strategies_enumerator = 3"); From 9da181677ded2111015a467cedde18aa9dac8d95 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:52:45 +0000 Subject: [PATCH 050/206] refactor(dashboard): drop the always-false command capabilities commandCapabilities() returned restartWholeRun, doctor, triage, merge, restartFromNode, rerunSelectedNode and arbitraryShell as the constant false. The frontend only declared them in its CommandCapabilities type: capabilityForCommand never maps a command to any of them, so nothing read them. Remove them from the server, the dashboard HTTP schema's commandCapabilities definition, and the frontend type. The schema edit is required because that definition is closed (required plus additionalProperties: false). The dashboard schema feeds neither the validator build identity, which hashes only the artifacts validator modules and the ajv versions, nor the artifact schema bundle digest that planned outputs and the trusted CLI identity bind to. The flow test's two assertions that doctor and merge are false are removed. Every dashboard HTTP response in the tests is still validated against the schema, so a server that emitted one of these keys again would fail the flow test with "must NOT have additional properties" (checked by re-adding doctor: false to the compiled server). Co-Authored-By: Claude Opus 5.5 --- packages/dashboard/frontend/src/main.tsx | 7 ------- .../schema/dashboard-http.schema.json | 18 ++---------------- packages/dashboard/src/index.ts | 9 +-------- packages/dashboard/test/dashboard.test.ts | 2 -- 4 files changed, 3 insertions(+), 33 deletions(-) diff --git a/packages/dashboard/frontend/src/main.tsx b/packages/dashboard/frontend/src/main.tsx index b3894a5f9..d85821538 100644 --- a/packages/dashboard/frontend/src/main.tsx +++ b/packages/dashboard/frontend/src/main.tsx @@ -223,18 +223,11 @@ type CommandCapabilities = { referencesStatus: boolean; referencesSync: boolean; referencesUpdate: boolean; - restartWholeRun: boolean; status: boolean; - doctor: boolean; config: boolean; report: boolean; - triage: boolean; - merge: boolean; materialize: boolean; clean: boolean; - restartFromNode: boolean; - rerunSelectedNode: boolean; - arbitraryShell: boolean; }; type DashboardFlowNode = Node; diff --git a/packages/dashboard/schema/dashboard-http.schema.json b/packages/dashboard/schema/dashboard-http.schema.json index a40ddfebe..a1e244889 100644 --- a/packages/dashboard/schema/dashboard-http.schema.json +++ b/packages/dashboard/schema/dashboard-http.schema.json @@ -430,14 +430,7 @@ "materialize", "clean", "config", - "restartWholeRun", - "status", - "doctor", - "triage", - "merge", - "restartFromNode", - "rerunSelectedNode", - "arbitraryShell" + "status" ], "properties": { "validate": { "type": "boolean" }, @@ -455,14 +448,7 @@ "materialize": { "type": "boolean" }, "clean": { "type": "boolean" }, "config": { "type": "boolean" }, - "restartWholeRun": { "type": "boolean" }, - "status": { "type": "boolean" }, - "doctor": { "type": "boolean" }, - "triage": { "type": "boolean" }, - "merge": { "type": "boolean" }, - "restartFromNode": { "type": "boolean" }, - "rerunSelectedNode": { "type": "boolean" }, - "arbitraryShell": { "type": "boolean" } + "status": { "type": "boolean" } } }, "findingItem": { diff --git a/packages/dashboard/src/index.ts b/packages/dashboard/src/index.ts index bd68cbbb2..2adc81997 100644 --- a/packages/dashboard/src/index.ts +++ b/packages/dashboard/src/index.ts @@ -1595,14 +1595,7 @@ class DashboardApp { materialize: hasRun, clean: hasRun, config: true, - restartWholeRun: false, - status: hasRun, - doctor: false, - triage: false, - merge: false, - restartFromNode: false, - rerunSelectedNode: false, - arbitraryShell: false + status: hasRun }; } diff --git a/packages/dashboard/test/dashboard.test.ts b/packages/dashboard/test/dashboard.test.ts index 1769793cd..0b4d2f1f8 100644 --- a/packages/dashboard/test/dashboard.test.ts +++ b/packages/dashboard/test/dashboard.test.ts @@ -81,8 +81,6 @@ test("serves logical topology flow with expanded attempt details", async () => { assert.ok(flow.nodes.every((node) => node.id === node.data.logicalNodeId)); assert.equal(flow.capabilities.runNewCampaign, true); assert.equal(flow.capabilities.referencesStatus, true); - assert.equal(flow.capabilities.doctor, false); - assert.equal(flow.capabilities.merge, false); } finally { await handle.close(); } From 23359f74bb50ccc4405927edae812acd21ade421 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:52:45 +0000 Subject: [PATCH 051/206] chore(runtime): drop the unused run-tests selectors packages/runtime/package.json calls scripts/run-tests.mjs with no selector or with "supporting"; nothing passes "materialize", "clean" or "smithers". A single test file can still be selected by name through the existing fallback. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/scripts/run-tests.mjs | 8 +------- 1 file changed, 1 insertion(+), 7 deletions(-) diff --git a/packages/runtime/scripts/run-tests.mjs b/packages/runtime/scripts/run-tests.mjs index 4cf057a0b..8eb8d5b61 100644 --- a/packages/runtime/scripts/run-tests.mjs +++ b/packages/runtime/scripts/run-tests.mjs @@ -12,13 +12,7 @@ const selectorFiles = new Map([ .filter((entry) => entry.endsWith(".test.js") && entry !== "runtime.test.js") .sort() .map((entry) => path.join("dist-test/test", entry)) - ], - [ - "materialize", - ["dist-test/test/materialize.test.js", "dist-test/test/clean.test.js", "dist-test/test/runtime.test.js"] - ], - ["clean", ["dist-test/test/clean.test.js", "dist-test/test/runtime.test.js"]], - ["smithers", ["dist-test/test/runtime.test.js"]] + ] ]); const testFiles = From e6c395c1ad65258a64960f8481db1694f1941699 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:52:46 +0000 Subject: [PATCH 052/206] docs(changelog): record the dead-code deletions Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..c347b82f3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[prompts] [config] [modal] [dashboard] [ci]** Deletes code that no production path calls: the prompt rename feature (`renamePromptId`, `renamePromptArtifactReferences`, `diffPromptIdentity`, `serializePromptDocument`, `getPrompt`), the config prompt-metadata layer, `restoreRedactedConfig`, `assertNoRedactionPlaceholders`, `applyDefaultProfileOverrides` and `assertResolvedConfigZod`, unreferenced or test-only Modal exports, and `scripts/ci/validate-threat-model-benchmark-gate.mjs` and `scripts/ci/modal-benchmark-control-window.mjs`, whose only callers were the workflows #1131 deleted. Two observable changes remain: output-contract templates now resolve through the same packaged-first lookup as the other prompt assets, so the renderer no longer probes a `dist/prompts` layout that no build produces, and the dashboard's `/api/flow` `capabilities` object drops seven always-false commands the frontend never read (#462). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 4c64f5835e79322023d13e7f54a9c32713c435d7 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:08:39 +0000 Subject: [PATCH 053/206] refactor(artifacts): delete the 48 semantic gates nothing dispatches Of the 173 registered artifact semantic gates, 48 were only reachable through executeSchemaSemanticGates for runtime documents that no caller passes (planned-graph, smithers-task-manifest, artifact-verification, artifact-manifest, agent-source-proof, run-metadata, run-plan, source-run, config-redactions, invariant-suite-manifest, the single-finding subschema, the run-state fingerprint and the analysis-bundle file digest). Production dispatches gates only for artifact-contract schemas (workflow.tsx, artifact-gates.ts, `artifact validate`), report.schema.json (modal), the validator preflight, the invariant source proof, the attempt/usage ledgers, and a fixed list of by-name calls; none of those names any of the 48. The runtime validates those documents with their JSON Schemas and the TypeScript assertions it does call (assertPlannedGraph and assertSealedPlannedGraph, assertSmithersTaskManifestMatchesPlannedGraph, assertRunMetadataDocument, assertArtifactVerificationMarkerSemantics, ...). The planned-graph gates duplicated assertPlannedGraphSemantics, whose own test covers each of the 12 rules, and at least one copy had drifted: planned-graph-artifact-dir-identity required artifacts/ for every node, while assertPlannedGraphSemantics accepts the storage/attempt directory of dynamically generated nodes. Delete the registrations, their exclusive handlers and helpers, their names in ARTIFACT_SCHEMA_METADATA, the context fields only they read (plannedGraph, runtimeState, git.refs) and the runtime line that built the unused plannedGraph context. Tests that only exercised the deleted gates go with them; the single-finding provenance case now runs through the live findings@2 gate. Schema bytes, the schema-bundle digest, every contract digest and VALIDATOR_BUILD_IDENTITY are byte-identical to main (compared from both builds), so the schema bindings sealed into existing runs still match. Refs #462 Co-Authored-By: Claude Opus 5.5 --- .../artifacts/src/artifact-schema-metadata.ts | 83 +- packages/artifacts/src/semantic-gates.ts | 1070 ----------------- .../artifacts/test/semantic-gates.test.ts | 674 +---------- packages/runtime/src/artifact-gates.ts | 1 - 4 files changed, 25 insertions(+), 1803 deletions(-) diff --git a/packages/artifacts/src/artifact-schema-metadata.ts b/packages/artifacts/src/artifact-schema-metadata.ts index 971b55b32..3732040f9 100644 --- a/packages/artifacts/src/artifact-schema-metadata.ts +++ b/packages/artifacts/src/artifact-schema-metadata.ts @@ -64,11 +64,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "adminConfigBoundaryMatrixSchema", ["admin-config-surface-id-uniqueness", "admin-config-surface-joins"] ), - "agent-source-proof.schema.json": runtime("agentSourceProofJsonSchema", undefined, [ - "agent-source-proof-ref-uniqueness", - "agent-source-proof-dependency-lineage", - "agent-source-proof-commit-binding" - ]), + "agent-source-proof.schema.json": runtime("agentSourceProofJsonSchema"), "aggregation-manifest.schema.json": artifact( "ultrafuzz/aggregation-manifest@1", "aggregationManifestJsonSchema", @@ -83,8 +79,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ ] ), "analysis-bundle.schema.json": runtime("analysisBundleManifestJsonSchema", "analysisBundleManifestSchema", [ - "analysis-bundle-path-order", - "analysis-bundle-file-digest" + "analysis-bundle-path-order" ]), "analysis-bundle-accounting-summary.schema.json": runtime( "analysisAccountingSummaryJsonSchema", @@ -116,19 +111,8 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "analysisTerminalStatusSchema", ["analysis-bundle-terminal-status-reconciliation"] ), - "artifact-manifest.schema.json": runtime("artifactManifestJsonSchema", undefined, [ - "artifact-manifest-file-path-uniqueness", - "artifact-manifest-output-path-uniqueness", - "artifact-manifest-prerequisite-node-uniqueness", - "artifact-manifest-file-digest" - ]), - "artifact-verification.schema.json": runtime("artifactVerificationJsonSchema", undefined, [ - "artifact-verification-artifact-path-uniqueness", - "artifact-verification-publication-path-uniqueness", - "artifact-verification-exactly-one-primary", - "artifact-verification-publication-digest-correspondence", - "artifact-verification-plan-contract-identity" - ]), + "artifact-manifest.schema.json": runtime("artifactManifestJsonSchema"), + "artifact-verification.schema.json": runtime("artifactVerificationJsonSchema"), "audited-differential-lanes.schema.json": artifact( "ultrafuzz/audited-differential-lanes@1", "auditedDifferentialLanesJsonSchema", @@ -147,10 +131,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "campaignSummarySchema", ["campaign-summary-backend-uniqueness", "campaign-summary-count-coupling"] ), - "config-redactions.schema.json": runtime("configRedactionsJsonSchema", undefined, [ - "config-redactions-path-key-equality", - "config-redactions-path-uniqueness" - ]), + "config-redactions.schema.json": runtime("configRedactionsJsonSchema"), "coverage-goal.schema.json": artifact("ultrafuzz/coverage-goal@2", "coverageGoalJsonSchema", "coverageGoalSchema", [ "coverage-goal-reconciliation" ]), @@ -241,11 +222,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "finding.schema.json": { role: "subschema", contractIds: [], - semanticGates: [ - "finding-campaign-provenance-coherence", - "finding-evidence-span-consistency", - "finding-projected-reference-uniqueness" - ], + semanticGates: [], typescriptExport: "findingJsonSchema", zodParser: "findingSchema" }, @@ -294,11 +271,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "invariant-source-proof-path-uniqueness", "invariant-source-proof-git-binding" ]), - "invariant-suite-manifest.schema.json": runtime("invariantSuiteManifestJsonSchema", undefined, [ - "invariant-suite-file-path-uniqueness", - "invariant-suite-tombstone-uniqueness", - "invariant-suite-file-tombstone-disjointness" - ]), + "invariant-suite-manifest.schema.json": runtime("invariantSuiteManifestJsonSchema"), "json-validator-preflight-success.schema.json": runtime("jsonValidatorPreflightSuccessJsonSchema", undefined, [ "json-validator-preflight-current-identity" ]), @@ -337,20 +310,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "lensPropertiesSchema", ["property-lens-id-uniqueness"] ), - "planned-graph.schema.json": runtime("plannedGraphJsonSchema", undefined, [ - "planned-graph-node-id-uniqueness", - "planned-graph-dependency-join", - "planned-graph-acyclicity", - "planned-graph-output-path-uniqueness", - "planned-graph-exactly-one-primary", - "planned-graph-model-fanout-uniqueness", - "planned-graph-workflow-task-uniqueness", - "planned-graph-workflow-node-join", - "planned-graph-artifact-dir-identity", - "planned-graph-loop-coupling", - "planned-graph-contract-identity", - "planned-graph-model-loop-coupling" - ]), + "planned-graph.schema.json": runtime("plannedGraphJsonSchema"), "reference-expectations.schema.json": artifact( "ultrafuzz/reference-expectations@2", "referenceExpectationsJsonSchema", @@ -383,17 +343,10 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "report-severity-classification-preservation", "report-property-provenance-join" ]), - "run-plan.schema.json": runtime("runPlanJsonSchema", undefined, ["run-plan-attempt-id-uniqueness"]), - "run-metadata.schema.json": runtime("runMetadataJsonSchema", undefined, [ - "run-metadata-workflow-id-equality", - "run-metadata-current-segment-equality", - "run-metadata-accounting-workflow-identity" - ]), - "run-state.schema.json": runtime("runStateJsonSchema", "runStateSchema", [ - "run-state-fingerprint", - "run-state-node-key-equality" - ]), - "source-run.schema.json": runtime("sourceRunJsonSchema", undefined, ["source-run-not-self"]), + "run-plan.schema.json": runtime("runPlanJsonSchema"), + "run-metadata.schema.json": runtime("runMetadataJsonSchema"), + "run-state.schema.json": runtime("runStateJsonSchema", "runStateSchema", ["run-state-node-key-equality"]), + "source-run.schema.json": runtime("sourceRunJsonSchema"), "terminal-disposition.schema.json": runtime("terminalDispositionJsonSchema", "terminalDispositionSchema"), "selected-strategies.schema.json": artifact( "ultrafuzz/selected-strategies@1", @@ -419,17 +372,7 @@ export const ARTIFACT_SCHEMA_METADATA = Object.freeze({ "severity-classification-upstream-preservation" ] ), - "smithers-task-manifest.schema.json": runtime("smithersTaskManifestJsonSchema", undefined, [ - "smithers-task-attempt-id-uniqueness", - "smithers-task-workflow-id-uniqueness", - "smithers-task-document-identity", - "smithers-task-pinned-submodule-expectation", - "smithers-task-dependency-join", - "smithers-task-dependency-acyclicity", - "smithers-task-planned-graph-coverage", - "smithers-task-planned-graph-identity", - "smithers-task-planned-graph-dependency-join" - ]), + "smithers-task-manifest.schema.json": runtime("smithersTaskManifestJsonSchema"), "strategy-detections.schema.json": artifact( "ultrafuzz/strategy-detections@1", "strategyDetectionsJsonSchema", diff --git a/packages/artifacts/src/semantic-gates.ts b/packages/artifacts/src/semantic-gates.ts index c0bcc42a1..cce642516 100644 --- a/packages/artifacts/src/semantic-gates.ts +++ b/packages/artifacts/src/semantic-gates.ts @@ -3,7 +3,6 @@ import fs from "node:fs"; import path from "node:path"; import { isDeepStrictEqual } from "node:util"; -import { artifactContractDefinition, artifactContractSchemaBinding } from "./artifact-contracts.js"; import { ARTIFACT_SCHEMA_METADATA, type ArtifactSchemaFilename } from "./artifact-schema-metadata.js"; import { artifactMetadataCompletenessIssues, metadataOmission } from "./artifact-validation.js"; import { findingNoteAssignmentIssue } from "./findings-schema.js"; @@ -38,7 +37,6 @@ export interface SemanticFilesystemContext { export interface SemanticGitContext { commit: string; tree: string; - refs?: Readonly>; baseCommit?: string; baseTree?: string; resultTree?: string; @@ -164,22 +162,12 @@ export interface SemanticArtifactSetContext { differentialArtifacts?: SemanticDifferentialArtifactsContext; } -export interface SemanticPlannedGraphContext { - node?: unknown; - document?: unknown; -} - export interface SemanticAttemptLedgerContext { entries: readonly unknown[]; /** Trusted entries from a source run that may be referenced by reuse evidence. */ sourceEntries?: readonly unknown[]; } -export interface SemanticRuntimeStateContext { - graphFingerprint: string; - configFingerprint: string; -} - export interface SemanticArtifactIdentityContext { runId: string; nodeId: string; @@ -287,10 +275,8 @@ export interface SemanticGateContext { filesystem?: SemanticFilesystemContext; git?: SemanticGitContext; artifactSet?: SemanticArtifactSetContext; - plannedGraph?: SemanticPlannedGraphContext; artifactIdentity?: SemanticArtifactIdentityContext; attemptLedger?: SemanticAttemptLedgerContext; - runtimeState?: SemanticRuntimeStateContext; usageLedger?: SemanticUsageLedgerContext; eventLog?: SemanticEventLogContext; validatorPreflight?: SemanticValidatorPreflightContext; @@ -3046,34 +3032,6 @@ function appendBigintEqualityIssue( } } -function artifactVerificationDigestIssues(document: unknown): SemanticGateIssue[] { - const publications = new Map(); - for (const row of arrayAt(document, ["publications"])) { - const rowPath = stringField(row, "path"); - const digest = stringField(row, "sha256"); - if (rowPath !== undefined && digest !== undefined) publications.set(rowPath, digest); - } - const issues: SemanticGateIssue[] = []; - for (const [index, artifact] of arrayAt(document, ["artifacts"]).entries()) { - const artifactPath = stringField(artifact, "path"); - const digest = stringField(artifact, "sha256"); - if (artifactPath !== undefined && publications.get(artifactPath) !== digest) { - issues.push( - issue( - `$.artifacts[${index}].sha256`, - `Publication digest does not correspond to artifact ${JSON.stringify(artifactPath)}` - ) - ); - } - } - return issues; -} - -function exactlyOnePrimaryIssues(document: unknown, key: string): SemanticGateIssue[] { - const count = arrayAt(document, [key]).filter((row) => booleanField(row, "primary") === true).length; - return count === 1 ? [] : [issue(`$.${key}`, `Exactly one ${key} entry must be primary; found ${count}`)]; -} - function attemptOrderIssues(document: unknown): SemanticGateIssue[] { const lifecycle = at(document, ["lifecycle"]); const started = stringField(lifecycle, "started_at"); @@ -4889,74 +4847,6 @@ function externalizedStateJoinIssues(document: unknown): SemanticGateIssue[] { return issues; } -function findingProjectedReferenceIssues(document: unknown): SemanticGateIssue[] { - if (!isRecord(document)) return []; - const groups: Array<{ - items: readonly unknown[]; - path: string; - project: (row: Readonly>) => string | undefined; - label: string; - }> = [ - { - items: arrayAt(document, ["family_variants"]), - path: "$.family_variants", - project: (row) => stringField(row, "id"), - label: "family variant ID" - }, - { - items: arrayAt(document, ["family_variants"]), - path: "$.family_variants", - project: (row) => stringField(row, "dedupe_key"), - label: "family variant dedupe key" - }, - { - items: arrayAt(document, ["related_findings"]), - path: "$.related_findings", - project: (row) => stringField(row, "id"), - label: "related finding ID" - }, - { - items: arrayAt(document, ["lifecycle", "source_artifacts"]), - path: "$.lifecycle.source_artifacts", - project: (row) => { - const values = [row.path, row.node_id, row.finding_id]; - return values.some((value) => value === undefined) ? undefined : JSON.stringify(values); - }, - label: "lifecycle source reference" - }, - { - items: arrayAt(document, ["lifecycle", "strategy_hits"]), - path: "$.lifecycle.strategy_hits", - project: (row) => - JSON.stringify([ - row.strategy, - row.attempt_index ?? null, - row.model_id ?? null, - row.model_index ?? null, - row.loop_index ?? null - ]), - label: "strategy hit identity" - } - ]; - const issues = projectedUniquenessIssues(groups); - const contributions = arrayAt(document, ["contributing_backend_failures"]); - const seen = new Set(); - for (const [index, contribution] of contributions.entries()) { - const key = - typeof contribution === "string" - ? JSON.stringify([null, contribution]) - : isRecord(contribution) - ? JSON.stringify([contribution.fuzzer_backend, contribution.failure_id]) - : undefined; - if (key === undefined) continue; - if (seen.has(key)) { - issues.push(issue(`$.contributing_backend_failures[${index}]`, `Duplicate contributing backend failure ${key}`)); - } - seen.add(key); - } - return issues; -} - function findingEvidenceSpanIssues(document: unknown, findingPath = "$"): SemanticGateIssue[] { if (!isRecord(document)) return []; const issues = evidenceArraySpanIssues(arrayAt(document, ["evidence"]), `${findingPath}.evidence`); @@ -5094,208 +4984,6 @@ function invariantLedgerJoinIssues(document: unknown): SemanticGateIssue[] { return issues; } -function plannedNodes(document: unknown): readonly unknown[] { - return arrayAt(document, ["nodes"]); -} - -function plannedNodeIdIssues(document: unknown): SemanticGateIssue[] { - return uniqueFieldGate([["nodes"]], "id", "planned graph node ID")(document, {}); -} - -function plannedDependencyJoinIssues(document: unknown): SemanticGateIssue[] { - const ids = new Set(plannedNodes(document).flatMap((node) => stringField(node, "id") ?? "")); - ids.delete(""); - const issues: SemanticGateIssue[] = []; - for (const [nodeIndex, node] of plannedNodes(document).entries()) { - const nodeId = stringField(node, "id"); - for (const [dependencyIndex, dependency] of stringArray(at(node, ["depends_on"])).entries()) { - if (!ids.has(dependency)) { - issues.push( - issue( - `$.nodes[${nodeIndex}].depends_on[${dependencyIndex}]`, - `Unknown planned dependency ${JSON.stringify(dependency)}` - ) - ); - } else if (dependency === nodeId) { - issues.push( - issue(`$.nodes[${nodeIndex}].depends_on[${dependencyIndex}]`, "A planned node cannot depend on itself") - ); - } - } - } - return issues; -} - -function plannedAcyclicityIssues(document: unknown): SemanticGateIssue[] { - const nodes = new Map(); - for (const node of plannedNodes(document)) { - const id = stringField(node, "id"); - if (id !== undefined) nodes.set(id, node); - } - const visiting = new Set(); - const visited = new Set(); - let cycle: string | undefined; - const visit = (nodeId: string): void => { - if (cycle !== undefined || visited.has(nodeId)) return; - if (visiting.has(nodeId)) { - cycle = nodeId; - return; - } - visiting.add(nodeId); - for (const dependency of stringArray(at(nodes.get(nodeId), ["depends_on"]))) { - if (nodes.has(dependency)) visit(dependency); - } - visiting.delete(nodeId); - visited.add(nodeId); - }; - for (const nodeId of nodes.keys()) visit(nodeId); - return cycle === undefined - ? [] - : [issue("$.nodes", `Planned graph contains a dependency cycle at ${JSON.stringify(cycle)}`)]; -} - -function plannedOutputPathIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => - projectedUniquenessIssues([ - { - items: arrayAt(node, ["outputs"]), - path: `$.nodes[${nodeIndex}].outputs`, - project: (row) => stringField(row, "path"), - label: "planned output path" - } - ]) - ); -} - -function plannedPrimaryIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => { - const count = arrayAt(node, ["outputs"]).filter((output) => booleanField(output, "primary") === true).length; - return count === 1 - ? [] - : [ - issue( - `$.nodes[${nodeIndex}].outputs`, - `Planned node must identify exactly one primary output; found ${count}` - ) - ]; - }); -} - -function plannedModelFanoutIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => - projectedUniquenessIssues([ - { - items: arrayAt(node, ["model_fanout"]), - path: `$.nodes[${nodeIndex}].model_fanout`, - project: (row) => JSON.stringify([row.model_profile_id, row.model_index, row.loop_index, row.attempt_index]), - label: "model-fanout identity" - } - ]) - ); -} - -function plannedWorkflowTaskIssues(document: unknown): SemanticGateIssue[] { - const seen = new Set(); - const issues: SemanticGateIssue[] = []; - for (const [nodeIndex, node] of plannedNodes(document).entries()) { - for (const [taskIndex, taskId] of stringArray(at(node, ["workflow", "task_node_ids"])).entries()) { - if (seen.has(taskId)) { - issues.push( - issue( - `$.nodes[${nodeIndex}].workflow.task_node_ids[${taskIndex}]`, - `Duplicate workflow task ID ${JSON.stringify(taskId)}` - ) - ); - } - seen.add(taskId); - } - } - return issues; -} - -function plannedWorkflowJoinIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => { - const workflow = at(node, ["workflow"]); - if (!isRecord(workflow)) return []; - const nodeId = stringField(workflow, "node_id"); - return nodeId !== undefined && !stringArray(workflow.task_node_ids).includes(nodeId) - ? [issue(`$.nodes[${nodeIndex}].workflow.node_id`, "Workflow node_id must be present in task_node_ids")] - : []; - }); -} - -function plannedArtifactDirIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => { - const id = stringField(node, "id"); - const artifactDir = stringField(node, "artifact_dir"); - return id !== undefined && artifactDir !== `artifacts/${id}` - ? [issue(`$.nodes[${nodeIndex}].artifact_dir`, "artifact_dir must be derived from the planned node ID")] - : []; - }); -} - -function plannedLoopIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => { - const loop = at(node, ["loop"]); - const index = numberField(loop, "index"); - const count = numberField(loop, "count"); - const attempt = numberField(loop, "attempt_index"); - return index !== undefined && count !== undefined && (index >= count || attempt !== index) - ? [issue(`$.nodes[${nodeIndex}].loop`, "Planned loop coordinates are inconsistent")] - : []; - }); -} - -function plannedContractIdentityIssues(document: unknown): SemanticGateIssue[] { - const issues: SemanticGateIssue[] = []; - for (const [nodeIndex, node] of plannedNodes(document).entries()) { - for (const [outputIndex, output] of arrayAt(node, ["outputs"]).entries()) { - const contract = stringField(output, "contract"); - if (contract === undefined) continue; - let definition: ReturnType; - try { - definition = artifactContractDefinition(contract as Parameters[0]); - } catch { - continue; - } - if (stringField(output, "contract_digest") !== definition.digest) { - issues.push( - issue( - `$.nodes[${nodeIndex}].outputs[${outputIndex}].contract_digest`, - "Planned output contract digest changed" - ) - ); - } - const binding = artifactContractSchemaBinding(contract as Parameters[0]); - const bindingFields = [ - "schema_file", - "schema_id", - "schema_sha256", - "schema_bundle_sha256", - "validator_build" - ] as const; - if ( - (binding === undefined && bindingFields.some((field) => isRecord(output) && output[field] !== undefined)) || - (binding !== undefined && bindingFields.some((field) => isRecord(output) && output[field] !== binding[field])) - ) { - issues.push(issue(`$.nodes[${nodeIndex}].outputs[${outputIndex}]`, "Planned output schema binding changed")); - } - } - } - return issues; -} - -function plannedModelLoopIssues(document: unknown): SemanticGateIssue[] { - return plannedNodes(document).flatMap((node, nodeIndex) => { - const loopIndex = numberField(at(node, ["loop"]), "index"); - return arrayAt(node, ["model_fanout"]).flatMap((model, modelIndex) => - numberField(model, "loop_index") === loopIndex - ? [] - : [issue(`$.nodes[${nodeIndex}].model_fanout[${modelIndex}].loop_index`, "Model is bound to another loop")] - ); - }); -} - function propertySourceProjectedIssues(document: unknown): SemanticGateIssue[] { const seen = new Set(); const issues: SemanticGateIssue[] = []; @@ -5345,404 +5033,6 @@ function runStateNodeKeyIssues(document: unknown): SemanticGateIssue[] { ); } -function smithersTasks(document: unknown): readonly unknown[] { - return arrayAt(document, ["tasks"]); -} - -function smithersWorkflowIdentityIssues(document: unknown): SemanticGateIssue[] { - return [ - ...uniqueFieldGate([["tasks"]], "smithersNodeId", "Smithers workflow node ID")(document, {}), - ...uniqueFieldGate([["tasks"]], "verifierSmithersNodeId", "Smithers verifier node ID")(document, {}) - ]; -} - -function sameUnknownArray(left: unknown, right: unknown): boolean { - return Array.isArray(left) && Array.isArray(right) && JSON.stringify(left) === JSON.stringify(right); -} - -function smithersDocumentIdentityIssues(document: unknown): SemanticGateIssue[] { - const runId = stringField(document, "run_id"); - const workflowName = stringField(document, "workflow_name"); - const issues: SemanticGateIssue[] = []; - for (const [index, task] of smithersTasks(document).entries()) { - if (!isRecord(task)) continue; - const taskPath = `$.tasks[${index}]`; - const attemptId = stringField(task, "attemptId"); - if (attemptId !== undefined && stringField(task, "smithersNodeId") !== `node:${attemptId}`) { - issues.push(issue(`${taskPath}.smithersNodeId`, "Smithers workflow node ID must be derived from attemptId")); - } - if (attemptId !== undefined && stringField(task, "verifierSmithersNodeId") !== `verify:${attemptId}`) { - issues.push( - issue(`${taskPath}.verifierSmithersNodeId`, "Smithers verifier node ID must be derived from attemptId") - ); - } - const metadata = at(task, ["metadata"]); - if (isRecord(metadata)) { - const metadataRun = at(metadata, ["run"]); - if ( - stringField(metadataRun, "ultrafuzzRunId") !== runId || - stringField(metadataRun, "smithersWorkflowName") !== workflowName - ) { - issues.push(issue(`${taskPath}.metadata.run`, "Smithers task run metadata does not match its document")); - } - const metadataNode = at(metadata, ["node"]); - if ( - stringField(metadataNode, "attemptId") !== attemptId || - stringField(metadataNode, "concreteNodeId") !== stringField(task, "concreteNodeId") || - stringField(metadataNode, "logicalNodeId") !== stringField(task, "logicalNodeId") - ) { - issues.push(issue(`${taskPath}.metadata.node`, "Smithers task node metadata does not match its envelope")); - } - const metadataModel = at(metadata, ["model"]); - if ( - stringField(metadataModel, "agentRef") !== stringField(task, "agentRef") || - (isRecord(metadataModel) ? metadataModel.modelName : undefined) !== task.modelName || - (isRecord(metadataModel) ? metadataModel.reasoningEffort : undefined) !== task.reasoningEffort - ) { - issues.push(issue(`${taskPath}.metadata.model`, "Smithers task model metadata does not match its envelope")); - } - if (!sameUnknownArray(task.dependencies, at(metadata, ["dependencies", "attemptIds"]))) { - issues.push( - issue(`${taskPath}.metadata.dependencies.attemptIds`, "Dependency attempt metadata does not match") - ); - } - if (!sameUnknownArray(task.dependencySmithersNodeIds, at(metadata, ["dependencies", "smithersNodeIds"]))) { - issues.push( - issue(`${taskPath}.metadata.dependencies.smithersNodeIds`, "Dependency workflow metadata does not match") - ); - } - const timeout = at(metadata, ["timeout"]); - const retryPolicy = at(metadata, ["retryPolicy"]); - const timeoutMs = numberField(task, "timeoutMs"); - const retries = numberField(task, "retries"); - if ( - numberField(timeout, "milliseconds") !== timeoutMs || - numberField(timeout, "heartbeatTimeoutMs") !== numberField(task, "heartbeatTimeoutMs") || - numberField(retryPolicy, "smithersRetries") !== retries || - (retries !== undefined && numberField(retryPolicy, "maxAttempts") !== retries + 1) || - (timeoutMs !== undefined && numberField(timeout, "seconds") !== Math.ceil(timeoutMs / 1_000)) - ) { - issues.push(issue(`${taskPath}.metadata.timeout`, "Smithers timeout or retry metadata does not match")); - } - const execution = at(task, ["execution"]); - const metadataExecution = at(metadata, ["execution"]); - if ( - stringField(execution, "mode") !== stringField(metadataExecution, "mode") || - (isRecord(execution) ? execution.provider : undefined) !== - (isRecord(metadataExecution) ? metadataExecution.provider : undefined) || - JSON.stringify(at(execution, ["resources"])) !== JSON.stringify(at(metadataExecution, ["resources"])) || - stringField(at(metadata, ["artifacts"]), "dir") !== stringField(task, "artifactDir") - ) { - issues.push(issue(`${taskPath}.metadata.execution`, "Smithers execution or artifact metadata does not match")); - } - } - } - return issues; -} - -function smithersPinnedSubmoduleIssues(document: unknown): SemanticGateIssue[] { - const expectation = at(document, ["pinned_submodules"]); - if (expectation === null || !isRecord(expectation)) return []; - const issues: SemanticGateIssue[] = []; - const roots = stringArray(at(expectation, ["top_level_roots"])); - const gitlinks = arrayAt(expectation, ["recursive_gitlinks"]); - const gitlinkPaths = gitlinks.flatMap((entry) => stringField(entry, "path") ?? []); - issues.push(...pinnedSubmodulePortablePathIssues(expectation, "$.pinned_submodules")); - const canonical = (values: readonly string[]): boolean => - new Set(values).size === values.length && values.every((value, index) => index === 0 || values[index - 1]! < value); - if (!canonical(roots)) { - issues.push(issue("$.pinned_submodules.top_level_roots", "Pinned submodule roots must be unique and ordered")); - } - if (!canonical(gitlinkPaths)) { - issues.push( - issue("$.pinned_submodules.recursive_gitlinks", "Pinned submodule gitlinks must be unique and path-ordered") - ); - } - for (const [index, root] of roots.entries()) { - if (!gitlinkPaths.includes(root)) { - issues.push(issue(`$.pinned_submodules.top_level_roots[${index}]`, "Pinned submodule root is not a gitlink")); - } - if (roots.some((candidate, candidateIndex) => candidateIndex !== index && root.startsWith(`${candidate}/`))) { - issues.push(issue(`$.pinned_submodules.top_level_roots[${index}]`, "Pinned submodule roots overlap")); - } - } - const entryCount = numberField(expectation, "entry_count"); - const fileCount = numberField(expectation, "file_count"); - if (entryCount !== undefined && fileCount !== undefined && fileCount > entryCount) { - issues.push(issue("$.pinned_submodules.file_count", "Pinned submodule file count exceeds entry count")); - } - return issues; -} - -function smithersDependencyJoinIssues(document: unknown): SemanticGateIssue[] { - const tasks = smithersTasks(document); - const byAttempt = new Map( - tasks.flatMap((task) => { - const id = stringField(task, "attemptId"); - return id === undefined ? [] : [[id, task] as const]; - }) - ); - const byVerifier = new Map( - tasks.flatMap((task) => { - const id = stringField(task, "verifierSmithersNodeId"); - return id === undefined ? [] : [[id, task] as const]; - }) - ); - const issues: SemanticGateIssue[] = []; - for (const [taskIndex, task] of tasks.entries()) { - const attemptId = stringField(task, "attemptId"); - const dependencies = stringArray(at(task, ["dependencies"])); - const joined = new Set(); - for (const [dependencyIndex, verifierId] of stringArray(at(task, ["dependencySmithersNodeIds"])).entries()) { - const dependency = byVerifier.get(verifierId); - const dependencyAttempt = stringField(dependency, "attemptId"); - if (dependency === undefined) { - issues.push( - issue( - `$.tasks[${taskIndex}].dependencySmithersNodeIds[${dependencyIndex}]`, - `Unknown verifier dependency ${JSON.stringify(verifierId)}` - ) - ); - } else if (dependencyAttempt !== undefined && !dependencies.includes(dependencyAttempt)) { - issues.push( - issue( - `$.tasks[${taskIndex}].dependencySmithersNodeIds[${dependencyIndex}]`, - "Verifier dependency is absent from dependency attempts" - ) - ); - } else if (dependencyAttempt !== undefined) { - joined.add(dependencyAttempt); - } - } - for (const [dependencyIndex, dependencyId] of dependencies.entries()) { - if (dependencyId === attemptId) { - issues.push( - issue(`$.tasks[${taskIndex}].dependencies[${dependencyIndex}]`, "A Smithers task cannot depend on itself") - ); - } else if (byAttempt.has(dependencyId) && !joined.has(dependencyId)) { - issues.push( - issue( - `$.tasks[${taskIndex}].dependencies[${dependencyIndex}]`, - "Task dependency is missing its verifier workflow dependency" - ) - ); - } - } - } - return issues; -} - -function smithersDependencyAcyclicityIssues(document: unknown): SemanticGateIssue[] { - const byVerifier = new Map( - smithersTasks(document).flatMap((task) => { - const id = stringField(task, "verifierSmithersNodeId"); - return id === undefined ? [] : [[id, task] as const]; - }) - ); - const visiting = new Set(); - const visited = new Set(); - let cycle: string | undefined; - const visit = (task: unknown): void => { - const id = stringField(task, "attemptId"); - if (id === undefined || cycle !== undefined || visited.has(id)) return; - if (visiting.has(id)) { - cycle = id; - return; - } - visiting.add(id); - for (const verifier of stringArray(at(task, ["dependencySmithersNodeIds"]))) { - const dependency = byVerifier.get(verifier); - if (dependency !== undefined) visit(dependency); - } - visiting.delete(id); - visited.add(id); - }; - for (const task of smithersTasks(document)) visit(task); - return cycle === undefined - ? [] - : [issue("$.tasks", `Smithers task dependencies contain a cycle at ${JSON.stringify(cycle)}`)]; -} - -function smithersPlannedAttemptIds(node: unknown): string[] { - const id = stringField(node, "id"); - if (id === undefined) return []; - const models = arrayAt(node, ["model_fanout"]); - if (models.length <= 1) return [id]; - return models.flatMap((model) => { - const modelIndex = numberField(model, "model_index"); - const attemptIndex = numberField(model, "attempt_index"); - return modelIndex === undefined || attemptIndex === undefined - ? [] - : [`${id}__model_${modelIndex}__attempt_${attemptIndex}`]; - }); -} - -function smithersGraphNodes(context: SemanticGateContext): readonly unknown[] { - return arrayAt(context.plannedGraph!.document, ["nodes"]); -} - -function smithersPlannedCoverageIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const tasks = smithersTasks(document); - const issues: SemanticGateIssue[] = []; - for (const [nodeIndex, node] of smithersGraphNodes(context).entries()) { - const id = stringField(node, "id"); - const matching = tasks.filter((task) => stringField(task, "concreteNodeId") === id); - if (stringField(node, "kind") === "reference") { - if (matching.length > 0) - issues.push(issue(`$.nodes[${nodeIndex}]`, "Reference planned nodes cannot have Smithers tasks")); - continue; - } - const expected = smithersPlannedAttemptIds(node); - const actual = matching.flatMap((task) => stringField(task, "attemptId") ?? "").filter((idValue) => idValue !== ""); - if (!sameStringSet(actual, expected)) { - issues.push(issue(`$.nodes[${nodeIndex}]`, `Smithers tasks do not cover planned node ${JSON.stringify(id)}`)); - } - } - return issues; -} - -function smithersPlannedIdentityIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const nodes = new Map( - smithersGraphNodes(context).flatMap((node) => { - const id = stringField(node, "id"); - return id === undefined ? [] : [[id, node] as const]; - }) - ); - const issues: SemanticGateIssue[] = []; - for (const [taskIndex, task] of smithersTasks(document).entries()) { - const node = nodes.get(stringField(task, "concreteNodeId") ?? ""); - if (node === undefined || stringField(node, "kind") !== "agentic") { - issues.push( - issue(`$.tasks[${taskIndex}].concreteNodeId`, "Smithers task does not join to an agentic planned node") - ); - continue; - } - const metadata = at(task, ["metadata"]); - const plannedLoop = at(node, ["loop"]); - const metadataLoop = at(metadata, ["loop"]); - if ( - stringField(task, "logicalNodeId") !== stringField(node, "logical_id") || - stringField(at(metadata, ["node"]), "logicalNodeId") !== stringField(node, "logical_id") || - stringField(at(metadata, ["node"]), "label") !== stringField(node, "display_name") || - numberField(metadataLoop, "index") !== numberField(plannedLoop, "index") || - numberField(metadataLoop, "count") !== numberField(plannedLoop, "count") || - stringField(metadataLoop, "mode") !== stringField(plannedLoop, "mode") || - numberField(metadataLoop, "attemptIndex") !== numberField(plannedLoop, "attempt_index") - ) { - issues.push(issue(`$.tasks[${taskIndex}].metadata`, "Smithers task identity does not match its planned node")); - } - const outputFields = [ - ["path", "path"], - ["contract", "contract"], - ["contractDigest", "contract_digest"], - ["schemaFile", "schema_file"], - ["schemaId", "schema_id"], - ["schemaSha256", "schema_sha256"], - ["schemaBundleSha256", "schema_bundle_sha256"], - ["validatorBuild", "validator_build"], - ["primary", "primary"] - ] as const; - const actualOutputs = arrayAt(metadata, ["artifacts", "outputs"]); - const plannedOutputs = arrayAt(node, ["outputs"]); - if ( - actualOutputs.length !== plannedOutputs.length || - plannedOutputs.some((output, outputIndex) => - outputFields.some( - ([actualField, plannedField]) => - !isRecord(actualOutputs[outputIndex]) || - !isRecord(output) || - actualOutputs[outputIndex]![actualField] !== output[plannedField] - ) - ) - ) { - issues.push( - issue( - `$.tasks[${taskIndex}].metadata.artifacts.outputs`, - "Smithers output contracts differ from the planned node" - ) - ); - } - } - return issues; -} - -function smithersPlannedDependencyJoinIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const nodes = new Map( - smithersGraphNodes(context).flatMap((node) => { - const id = stringField(node, "id"); - return id === undefined ? [] : [[id, node] as const]; - }) - ); - const issues: SemanticGateIssue[] = []; - for (const [taskIndex, task] of smithersTasks(document).entries()) { - const node = nodes.get(stringField(task, "concreteNodeId") ?? ""); - if (node === undefined) continue; - const dynamicDependencyIds = new Set(stringArray(at(node, ["dynamic_dependencies"]))); - const dependencyIds = stringArray(at(node, ["depends_on"])); - const pendingDynamicDependencyIds: string[] = []; - const dependencyNodes = dependencyIds.flatMap((id) => { - const dependency = nodes.get(id); - if (dependency === undefined) { - issues.push( - issue( - `$.tasks[${taskIndex}].metadata.dependencies.concreteNodeIds`, - `Smithers planned dependency node ${JSON.stringify(id)} is missing` - ) - ); - return []; - } - const dynamicStatus = at(dependency, ["dynamic", "status"]); - if (dynamicStatus === "pending" && dynamicDependencyIds.has(id)) { - pendingDynamicDependencyIds.push(id); - return []; - } - if (dynamicStatus === "expanded" && dynamicDependencyIds.has(id)) { - issues.push( - issue( - `$.tasks[${taskIndex}].metadata.dependencies.concreteNodeIds`, - `Smithers task retains expanded dynamic dependency placeholder ${JSON.stringify(id)}` - ) - ); - return []; - } - return [dependency]; - }); - const expectedAttempts = dependencyNodes.flatMap(smithersPlannedAttemptIds); - const actualAttempts = stringArray(at(task, ["dependencies"])).filter((id) => id !== "meta-start"); - if (!sameStringSet(actualAttempts, expectedAttempts)) { - issues.push( - issue(`$.tasks[${taskIndex}].dependencies`, "Smithers dependency attempts do not match planned dependencies") - ); - } - const expectedNodes = dependencyNodes.flatMap((dependency) => { - const id = stringField(dependency, "id"); - return id === undefined ? [] : [id]; - }); - const actualNodes = stringArray(at(task, ["metadata", "dependencies", "concreteNodeIds"])).filter( - (id) => id !== "__start__" - ); - const compiledExpectedNodes = [...expectedNodes, ...pendingDynamicDependencyIds]; - if (!sameStringSet(actualNodes, expectedNodes) && !sameStringSet(actualNodes, compiledExpectedNodes)) { - issues.push( - issue( - `$.tasks[${taskIndex}].metadata.dependencies.concreteNodeIds`, - "Smithers concrete dependencies do not match the plan" - ) - ); - } - const expectedVerifiers = dependencyNodes - .filter((dependency) => stringField(dependency, "kind") === "agentic") - .flatMap(smithersPlannedAttemptIds) - .map((id) => `verify:${id}`); - if (!sameStringSet(stringArray(at(task, ["dependencySmithersNodeIds"])), expectedVerifiers)) { - issues.push( - issue(`$.tasks[${taskIndex}].dependencySmithersNodeIds`, "Smithers workflow dependencies do not match the plan") - ); - } - } - return issues; -} - function workspacePatchPathIssues(document: unknown): SemanticGateIssue[] { const included = arrayAt(document, ["files"]); const excluded = arrayAt(document, ["excluded_files"]); @@ -5797,73 +5087,6 @@ function resolveArtifactFile(rootDirectory: string, relativePath: string): strin return candidate !== root && candidate.startsWith(`${root}${path.sep}`) ? candidate : undefined; } -function sha256File(filePath: string): string { - return crypto.createHash("sha256").update(fs.readFileSync(filePath)).digest("hex"); -} - -function filesystemManifestIssues( - document: unknown, - context: SemanticGateContext, - rowsPath: readonly string[] -): SemanticGateIssue[] { - const root = context.filesystem!.rootDirectory; - const issues: SemanticGateIssue[] = []; - for (const [index, row] of arrayAt(document, rowsPath).entries()) { - const relativePath = stringField(row, "path"); - const expectedDigest = stringField(row, "sha256"); - if (relativePath === undefined) continue; - const filePath = resolveArtifactFile(root, relativePath); - const snapshot = context.filesystem!.files?.get(relativePath); - if (context.filesystem!.files !== undefined) { - const rowPath = `${displayPath(rowsPath)}[${index}]`; - if (filePath === undefined || snapshot === undefined) { - issues.push( - issue(`${rowPath}.path`, `Referenced file is missing or nonregular: ${JSON.stringify(relativePath)}`) - ); - continue; - } - if ( - expectedDigest !== undefined && - crypto.createHash("sha256").update(snapshot).digest("hex") !== expectedDigest - ) { - issues.push( - issue(`${rowPath}.sha256`, `Referenced file digest does not match ${JSON.stringify(relativePath)}`) - ); - } - const expectedSize = numberField(row, "size_bytes"); - if (expectedSize !== undefined && expectedSize !== snapshot.byteLength) { - issues.push( - issue(`${rowPath}.size_bytes`, `Referenced file size does not match ${JSON.stringify(relativePath)}`) - ); - } - continue; - } - let stats: fs.Stats | undefined; - try { - if (filePath !== undefined) stats = fs.lstatSync(filePath); - } catch { - // Reported below as a missing/nonregular file. - } - const rowPath = `${displayPath(rowsPath)}[${index}]`; - if (filePath === undefined || stats === undefined || !stats.isFile() || stats.isSymbolicLink()) { - issues.push( - issue(`${rowPath}.path`, `Referenced file is missing or nonregular: ${JSON.stringify(relativePath)}`) - ); - continue; - } - if (expectedDigest !== undefined && sha256File(filePath) !== expectedDigest) { - issues.push(issue(`${rowPath}.sha256`, `Referenced file digest does not match ${JSON.stringify(relativePath)}`)); - } - const expectedSize = numberField(row, "size_bytes"); - if (expectedSize !== undefined && expectedSize !== stats.size) { - issues.push( - issue(`${rowPath}.size_bytes`, `Referenced file size does not match ${JSON.stringify(relativePath)}`) - ); - } - } - return issues; -} - function generatedTestFileIntegrityIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { const root = context.filesystem!.rootDirectory; const resourceIssues = generatedTestBundleResourceBoundsIssues(document); @@ -6090,94 +5313,6 @@ function generatedTestIdentityIssues(document: unknown, context: SemanticGateCon return issues; } -function agentSourceProofGitIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const git = context.git!; - const issues: SemanticGateIssue[] = []; - if (stringField(document, "commit") !== git.commit) - issues.push(issue("$.commit", "Source proof commit does not match Git")); - if (stringField(document, "tree") !== git.tree) issues.push(issue("$.tree", "Source proof tree does not match Git")); - for (const [index, ref] of arrayAt(document, ["refs"]).entries()) { - const name = stringField(ref, "name"); - const object = stringField(ref, "object"); - if (name !== undefined && git.refs?.[name] !== object) { - issues.push(issue(`$.refs[${index}].object`, `Source proof ref ${JSON.stringify(name)} does not match Git`)); - } - } - return issues; -} - -function agentSourceProofDependencyIssues(document: unknown): SemanticGateIssue[] { - const dependencies = at(document, ["dependencies"]); - if (dependencies === null || !isRecord(dependencies)) return []; - const issues: SemanticGateIssue[] = []; - issues.push(...pinnedSubmodulePortablePathIssues(dependencies, "$.dependencies")); - if ( - stringField(dependencies, "source_commit") !== stringField(document, "commit") || - stringField(dependencies, "source_tree") !== stringField(document, "tree") - ) { - issues.push(issue("$.dependencies", "Pinned dependency source identity does not match the source proof")); - } - const roots = stringArray(at(dependencies, ["top_level_roots"])); - const gitlinks = arrayAt(dependencies, ["recursive_gitlinks"]); - const gitlinkPaths = gitlinks.flatMap((entry) => { - const entryPath = stringField(entry, "path"); - return entryPath === undefined ? [] : [entryPath]; - }); - const canonical = (values: readonly string[]): boolean => - new Set(values).size === values.length && values.every((value, index) => index === 0 || values[index - 1]! < value); - if (!canonical(roots)) { - issues.push(issue("$.dependencies.top_level_roots", "Pinned dependency roots must be unique and ordered")); - } - if (!canonical(gitlinkPaths)) { - issues.push(issue("$.dependencies.recursive_gitlinks", "Pinned dependency gitlinks must be unique and ordered")); - } - for (const [index, root] of roots.entries()) { - if (!gitlinkPaths.includes(root)) { - issues.push(issue(`$.dependencies.top_level_roots[${index}]`, "Pinned dependency root is not a gitlink")); - } - if (roots.some((candidate, candidateIndex) => candidateIndex !== index && root.startsWith(`${candidate}/`))) { - issues.push(issue(`$.dependencies.top_level_roots[${index}]`, "Pinned dependency roots overlap")); - } - } - const entryCount = numberField(dependencies, "entry_count"); - const fileCount = numberField(dependencies, "file_count"); - if (entryCount !== undefined && fileCount !== undefined && fileCount > entryCount) { - issues.push(issue("$.dependencies.file_count", "Pinned dependency file count exceeds entry count")); - } - return issues; -} - -function pinnedSubmodulePortablePathIssues(expectation: unknown, basePath: string): SemanticGateIssue[] { - const candidates = [ - ...stringArray(at(expectation, ["top_level_roots"])).map((value, index) => ({ - value, - path: `${basePath}.top_level_roots[${index}]` - })), - ...arrayAt(expectation, ["recursive_gitlinks"]).flatMap((entry, index) => { - const value = stringField(entry, "path"); - return value === undefined ? [] : [{ value, path: `${basePath}.recursive_gitlinks[${index}].path` }]; - }) - ]; - return candidates.flatMap(({ value, path: issuePath }) => { - const segments = value.split("/"); - return value.includes("\\") || - path.posix.isAbsolute(value) || - path.posix.normalize(value) !== value || - Buffer.byteLength(value, "utf8") > 4_096 || - segments.length > 128 || - segments.some( - (segment) => - segment.length === 0 || - segment === "." || - segment === ".." || - segment === ".git" || - /^[A-Za-z]:/u.test(segment) - ) - ? [issue(issuePath, "Pinned submodule path is not portable and bounded")] - : []; - }); -} - function invariantSourceProofGitIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { const git = context.git!; const issues: SemanticGateIssue[] = []; @@ -6208,38 +5343,6 @@ function workspacePatchGitIssues(document: unknown, context: SemanticGateContext ); } -function artifactVerificationPlanIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const planned = context.plannedGraph!.node; - const issues: SemanticGateIssue[] = []; - if (stringField(document, "node_id") !== stringField(planned, "id")) { - issues.push(issue("$.node_id", "Verification marker node_id does not match the planned node")); - } - const actual = arrayAt(document, ["artifacts"]); - const expected = arrayAt(planned, ["outputs"]); - if (actual.length !== expected.length) { - issues.push(issue("$.artifacts", "Verification marker artifact count does not match planned outputs")); - return issues; - } - for (const [index, output] of expected.entries()) { - const artifact = actual[index]; - for (const field of [ - "path", - "contract", - "contract_digest", - "schema_file", - "schema_id", - "schema_sha256", - "schema_bundle_sha256", - "validator_build", - "primary" - ] as const) { - if (isRecord(artifact) && isRecord(output) && artifact[field] === output[field]) continue; - issues.push(issue(`$.artifacts[${index}].${field}`, `Verification marker ${field} does not match the plan`)); - } - } - return issues; -} - function campaignSummaryCountIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { const campaigns = context.artifactSet!.campaigns!; const findings = context.artifactSet!.findings!; @@ -7418,18 +6521,6 @@ function attemptSourceEventJoinIssues(document: unknown, context: SemanticGateCo return issues; } -function runStateFingerprintIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { - const expected = context.runtimeState!; - const issues: SemanticGateIssue[] = []; - if (stringField(document, "graph_fingerprint") !== expected.graphFingerprint) { - issues.push(issue("$.graph_fingerprint", "Run-state graph fingerprint does not match runtime state")); - } - if (stringField(document, "config_fingerprint") !== expected.configFingerprint) { - issues.push(issue("$.config_fingerprint", "Run-state config fingerprint does not match runtime state")); - } - return issues; -} - function usageEventOrderIssues(document: unknown, context: SemanticGateContext): SemanticGateIssue[] { const entries = [...context.usageLedger!.entries]; if (!entries.includes(document)) entries.push(document); @@ -7512,13 +6603,6 @@ const gateSpecifications = { "selected-strategies-metadata-completeness": documentGate(artifactMetadataCompletenessIssues), "admin-config-surface-id-uniqueness": documentGate(uniqueFieldGate([["surfaces"]], "surface_id", "admin surface ID")), "admin-config-surface-joins": documentGate(adminConfigJoinIssues), - "agent-source-proof-commit-binding": contextualGate( - "git", - ["git.commit", "git.tree", "git.refs"], - agentSourceProofGitIssues - ), - "agent-source-proof-dependency-lineage": documentGate(agentSourceProofDependencyIssues), - "agent-source-proof-ref-uniqueness": documentGate(uniqueFieldGate([["refs"]], "name", "source proof ref name")), "aggregation-count-coupling": documentGate(aggregationCountIssues), "aggregation-source-bundle-reconciliation": documentGate(aggregationBundleIssues), "aggregation-resource-bounds": documentGate(aggregationResourceIssues), @@ -7529,9 +6613,6 @@ const gateSpecifications = { ), "aggregation-destination-path-uniqueness": documentGate(aggregationDestinationIssues), "aggregation-source-entry-uniqueness": documentGate(aggregationSourceEntryIssues), - "analysis-bundle-file-digest": contextualGate("filesystem", ["filesystem.rootDirectory"], (document, context) => - filesystemManifestIssues(document, context, ["files"]) - ), "analysis-bundle-accounting-reconciliation": documentGate(analysisBundleAccountingIssues), "analysis-bundle-attempt-order": documentGate(analysisBundleAttemptOrderIssues), "analysis-bundle-evaluation-count-reconciliation": documentGate(analysisBundleEvaluationCountIssues), @@ -7544,31 +6625,6 @@ const gateSpecifications = { "analysis-bundle-path-order": documentGate(analysisBundlePathIssues), "analysis-bundle-recovery-reconciliation": documentGate(analysisBundleRecoveryIssues), "analysis-bundle-terminal-status-reconciliation": documentGate(analysisBundleTerminalStatusIssues), - "artifact-manifest-file-digest": contextualGate("filesystem", ["filesystem.rootDirectory"], (document, context) => - filesystemManifestIssues(document, context, ["files"]) - ), - "artifact-manifest-file-path-uniqueness": documentGate(uniqueFieldGate([["files"]], "path", "artifact file path")), - "artifact-manifest-output-path-uniqueness": documentGate( - uniqueFieldGate([["output_contracts"]], "path", "artifact output path") - ), - "artifact-manifest-prerequisite-node-uniqueness": documentGate( - uniqueFieldGate([["prerequisite_manifests"]], "node_id", "prerequisite node ID") - ), - "artifact-verification-artifact-path-uniqueness": documentGate( - uniqueFieldGate([["artifacts"]], "path", "verified artifact path") - ), - "artifact-verification-exactly-one-primary": documentGate((document) => - exactlyOnePrimaryIssues(document, "artifacts") - ), - "artifact-verification-plan-contract-identity": contextualGate( - "cross-artifact", - ["plannedGraph.node"], - artifactVerificationPlanIssues - ), - "artifact-verification-publication-digest-correspondence": documentGate(artifactVerificationDigestIssues), - "artifact-verification-publication-path-uniqueness": documentGate( - uniqueFieldGate([["publications"]], "path", "publication path") - ), "attempt-order": documentGate(attemptOrderIssues), "attempt-failure-message-byte-length": documentGate(attemptFailureMessageByteLengthIssues), "attempt-outcome-digest-coupling": documentGate(attemptOutcomeDigestIssues), @@ -7597,27 +6653,6 @@ const gateSpecifications = { ["artifactSet.campaigns", "artifactSet.findings"], campaignSummaryCountIssues ), - "config-redactions-path-key-equality": documentGate((document) => - arrayAt(document, ["entries"]).flatMap((entry, index) => { - const pathValue = at(entry, ["path"]); - const projected = Array.isArray(pathValue) ? pathValue.join(".") : undefined; - const key = stringField(entry, "key"); - return projected !== undefined && key !== projected - ? [issue(`$.entries[${index}].key`, "Configuration redaction key does not equal its projected path")] - : []; - }) - ), - "config-redactions-path-uniqueness": documentGate((document) => { - const seen = new Set(); - return arrayAt(document, ["entries"]).flatMap((entry, index) => { - const pathValue = at(entry, ["path"]); - if (!Array.isArray(pathValue)) return []; - const projected = JSON.stringify(pathValue); - const duplicate = seen.has(projected); - seen.add(projected); - return duplicate ? [issue(`$.entries[${index}].path`, "Duplicate configuration redaction path")] : []; - }); - }), "dependency-id-uniqueness": documentGate(uniqueFieldGate([["dependencies"]], "dependency_id", "dependency ID")), "dependency-row-joins": documentGate(dependencyJoinIssues), "differential-gap-review-lane-reconciliation": contextualGate( @@ -7737,9 +6772,6 @@ const gateSpecifications = { ["artifactSet.reviewStage"], lifecycleReviewStageIssues ), - "finding-evidence-span-consistency": documentGate((document) => findingEvidenceSpanIssues(document)), - "finding-campaign-provenance-coherence": documentGate((document) => findingCampaignProvenanceIssues(document)), - "finding-projected-reference-uniqueness": documentGate(findingProjectedReferenceIssues), "findings-campaign-provenance-coherence": documentGate(findingArrayCampaignProvenanceIssues), "findings-evidence-span-consistency": documentGate(findingArrayEvidenceSpanIssues), "findings-id-uniqueness": documentGate(uniqueFieldGate([[]], "id", "finding ID")), @@ -7775,33 +6807,6 @@ const gateSpecifications = { "invariant-source-proof-path-uniqueness": documentGate( uniqueFieldGate([["files"]], "path", "invariant source proof path") ), - "invariant-suite-file-path-uniqueness": documentGate( - uniqueFieldGate([["files"]], "path", "invariant suite file path") - ), - "invariant-suite-file-tombstone-disjointness": documentGate((document) => { - const files = new Set(arrayAt(document, ["files"]).flatMap((row) => stringField(row, "path") ?? "")); - return stringArray(at(document, ["tombstones"])).flatMap((tombstone, index) => - files.has(tombstone) - ? [ - issue( - `$.tombstones[${index}]`, - `Invariant suite path is both present and tombstoned ${JSON.stringify(tombstone)}` - ) - ] - : [] - ); - }), - "invariant-suite-tombstone-uniqueness": documentGate((document) => { - const tombstones = stringArray(at(document, ["tombstones"])); - const seen = new Set(); - return tombstones.flatMap((tombstone, index) => { - const duplicate = seen.has(tombstone); - seen.add(tombstone); - return duplicate - ? [issue(`$.tombstones[${index}]`, `Duplicate invariant suite tombstone ${JSON.stringify(tombstone)}`)] - : []; - }); - }), "json-validator-preflight-current-identity": contextualGate( "runtime-state", [ @@ -7813,18 +6818,6 @@ const gateSpecifications = { ], jsonValidatorPreflightIdentityIssues ), - "planned-graph-acyclicity": documentGate(plannedAcyclicityIssues), - "planned-graph-artifact-dir-identity": documentGate(plannedArtifactDirIssues), - "planned-graph-contract-identity": documentGate(plannedContractIdentityIssues), - "planned-graph-dependency-join": documentGate(plannedDependencyJoinIssues), - "planned-graph-exactly-one-primary": documentGate(plannedPrimaryIssues), - "planned-graph-loop-coupling": documentGate(plannedLoopIssues), - "planned-graph-model-fanout-uniqueness": documentGate(plannedModelFanoutIssues), - "planned-graph-model-loop-coupling": documentGate(plannedModelLoopIssues), - "planned-graph-node-id-uniqueness": documentGate(plannedNodeIdIssues), - "planned-graph-output-path-uniqueness": documentGate(plannedOutputPathIssues), - "planned-graph-workflow-node-join": documentGate(plannedWorkflowJoinIssues), - "planned-graph-workflow-task-uniqueness": documentGate(plannedWorkflowTaskIssues), "property-campaign-coverage-metric-uniqueness": documentGate( uniqueFieldGate([["coverage", "metrics"]], "name", "property campaign coverage metric") ), @@ -7920,43 +6913,6 @@ const gateSpecifications = { ["artifactSet.propertyCatalog", "artifactSet.implementedProperties"], reportPropertyJoinIssues ), - "run-metadata-accounting-workflow-identity": documentGate((document) => { - const workflowRunId = stringField(at(document, ["workflow"]), "run_id"); - const accounting = at(document, ["accounting"]); - if (accounting === undefined) return []; - const accountingRunId = stringField(accounting, "workflow_run_id"); - const currentRunId = stringField(at(accounting, ["current"]), "workflow_run_id"); - return workflowRunId !== accountingRunId || workflowRunId !== currentRunId - ? [issue("$.accounting.workflow_run_id", "Accounting identity does not equal the active workflow run")] - : []; - }), - "run-metadata-current-segment-equality": documentGate((document) => { - const accounting = at(document, ["accounting"]); - if (accounting === undefined) return []; - const segments = arrayAt(accounting, ["segments"]); - return isDeepStrictEqual(at(accounting, ["current"]), segments.at(-1)) - ? [] - : [issue("$.accounting.current", "Current accounting does not equal the final segment")]; - }), - "run-metadata-workflow-id-equality": documentGate((document) => { - const workflow = at(document, ["workflow"]); - const ids = stringArray(at(document, ["workflow_ids"])); - if (workflow === undefined) { - return ids.length === 0 ? [] : [issue("$.workflow_ids", "Unlinked metadata carries workflow IDs")]; - } - const runId = stringField(workflow, "run_id"); - return ids.length === 1 && ids[0] === runId - ? [] - : [issue("$.workflow_ids", "Workflow IDs do not equal the active workflow run")]; - }), - "run-plan-attempt-id-uniqueness": documentGate( - uniqueFieldGate([["rendered_prompts"]], "attempt_id", "rendered prompt attempt ID") - ), - "run-state-fingerprint": contextualGate( - "runtime-state", - ["runtimeState.graphFingerprint", "runtimeState.configFingerprint"], - runStateFingerprintIssues - ), "run-state-node-key-equality": documentGate(runStateNodeKeyIssues), "selected-strategy-id-uniqueness": documentGate( uniqueFieldGate([["strategies"]], "strategy_id", "selected strategy ID") @@ -7978,32 +6934,6 @@ const gateSpecifications = { ["artifactSet.triagedFindings"], severityClassificationPreservationIssues ), - "smithers-task-attempt-id-uniqueness": documentGate(uniqueFieldGate([["tasks"]], "attemptId", "Smithers attempt ID")), - "smithers-task-workflow-id-uniqueness": documentGate(smithersWorkflowIdentityIssues), - "smithers-task-document-identity": documentGate(smithersDocumentIdentityIssues), - "smithers-task-pinned-submodule-expectation": documentGate(smithersPinnedSubmoduleIssues), - "smithers-task-dependency-join": documentGate(smithersDependencyJoinIssues), - "smithers-task-dependency-acyclicity": documentGate(smithersDependencyAcyclicityIssues), - "smithers-task-planned-graph-coverage": contextualGate( - "cross-artifact", - ["plannedGraph.document"], - smithersPlannedCoverageIssues - ), - "smithers-task-planned-graph-identity": contextualGate( - "cross-artifact", - ["plannedGraph.document"], - smithersPlannedIdentityIssues - ), - "smithers-task-planned-graph-dependency-join": contextualGate( - "cross-artifact", - ["plannedGraph.document"], - smithersPlannedDependencyJoinIssues - ), - "source-run-not-self": documentGate((document) => - stringField(document, "run_id") === stringField(document, "source_run_id") - ? [issue("$.source_run_id", "Source run must differ from the destination run")] - : [] - ), "semantic-red-registry-lane-reconciliation": contextualGate( "cross-artifact", ["artifactSet.differentialArtifacts.laneResults"], diff --git a/packages/artifacts/test/semantic-gates.test.ts b/packages/artifacts/test/semantic-gates.test.ts index dbdded71b..50641afef 100644 --- a/packages/artifacts/test/semantic-gates.test.ts +++ b/packages/artifacts/test/semantic-gates.test.ts @@ -11,7 +11,6 @@ import { MAX_SEMANTIC_GATE_ISSUES, SEMANTIC_GATE_REGISTRY, SEMANTIC_GATE_SCOPES, - artifactContractDefinition, executeOfflineSchemaSemanticGates, executeSchemaSemanticGates, executeSemanticGate, @@ -45,79 +44,6 @@ function reportWithCompletion(completion: unknown, runId = "run-a"): unknown { return { run_metadata: { run_id: runId }, completion }; } -const validPlannedOutput = { - path: "report.md", - contract: "ultrafuzz/nonempty-markdown@1", - contract_digest: artifactContractDefinition("ultrafuzz/nonempty-markdown@1").digest, - primary: true -}; - -const validPlannedNode = { - id: "node-a", - artifact_dir: "artifacts/node-a", - depends_on: [] as string[], - outputs: [validPlannedOutput], - loop: { index: 0, count: 1, attempt_index: 0 }, - model_fanout: [] as unknown[] -}; - -const validPlannedGraph = { nodes: [validPlannedNode] }; - -const smithersManifestOutput = { - path: validPlannedOutput.path, - contract: validPlannedOutput.contract, - contractDigest: validPlannedOutput.contract_digest, - primary: validPlannedOutput.primary -}; - -const smithersIdentityTask = { - attemptId: "attempt-a", - concreteNodeId: "node-a", - logicalNodeId: "logical-a", - smithersNodeId: "node:attempt-a", - verifierSmithersNodeId: "verify:attempt-a", - agentRef: "agent-a", - agentChain: [{ profileId: "profile-a", agentRef: "agent-a", role: "primary" }], - dependencies: [] as string[], - dependencySmithersNodeIds: [] as string[], - timeoutMs: 1_000, - heartbeatTimeoutMs: 500, - retries: 0, - artifactDir: "artifacts/node-a", - execution: { mode: "local", resources: { cpu: 1, memoryMiB: 512, timeoutSeconds: 1 } }, - metadata: { - run: { ultrafuzzRunId: "run-a", smithersWorkflowName: "workflow-a" }, - node: { - attemptId: "attempt-a", - concreteNodeId: "node-a", - logicalNodeId: "logical-a", - label: "Node A" - }, - model: { - agentRef: "agent-a", - agentChain: [{ profileId: "profile-a", agentRef: "agent-a", role: "primary" }] - }, - dependencies: { attemptIds: [] as string[], smithersNodeIds: [] as string[], concreteNodeIds: [] as string[] }, - timeout: { milliseconds: 1_000, seconds: 1, heartbeatTimeoutMs: 500 }, - retryPolicy: { maxAttempts: 1, sameAgentAttempts: 1, smithersRetries: 0 }, - execution: { mode: "local", resources: { cpu: 1, memoryMiB: 512, timeoutSeconds: 1 } }, - artifacts: { dir: "artifacts/node-a", outputs: [smithersManifestOutput] }, - loop: { index: 0, count: 1, mode: "parallel", attemptIndex: 0 } - } -}; - -const pinnedSubmoduleExpectation = { - schema_version: "ultrafuzz.pinned-submodules-expectation.v1", - source_commit: "a".repeat(40), - source_tree: "b".repeat(40), - manifest_sha256: "c".repeat(64), - top_level_roots: ["vendor/dependency"], - recursive_gitlinks: [{ path: "vendor/dependency", commit: "d".repeat(40), tree: "e".repeat(40) }], - entry_count: 1, - file_count: 0, - total_file_bytes: 0 -}; - const analysisBundleAccountingGatePositive = { run_count: 1, accounted_run_count: 1, @@ -450,18 +376,6 @@ const fixtures = { positive: { surfaces: [{ surface_id: "a" }], coverage_notes: [{ surface_id: "a" }] }, negative: { surfaces: [{ surface_id: "a" }], coverage_notes: [{ surface_id: "missing" }] } }, - "agent-source-proof-ref-uniqueness": { - positive: { refs: [{ name: "refs/heads/a" }] }, - negative: { refs: [{ name: "refs/heads/a" }, { name: "refs/heads/a" }] } - }, - "agent-source-proof-dependency-lineage": { - positive: { commit: "a".repeat(40), tree: "b".repeat(40), dependencies: pinnedSubmoduleExpectation }, - negative: { - commit: "a".repeat(40), - tree: "b".repeat(40), - dependencies: { ...pinnedSubmoduleExpectation, source_commit: "f".repeat(40) } - } - }, "aggregation-count-coupling": { positive: { source_generated_tests: 2, @@ -689,34 +603,6 @@ const fixtures = { finished_at: "2026-01-01T00:00:01Z" } }, - "artifact-manifest-file-path-uniqueness": { - positive: { files: [{ path: "a" }] }, - negative: { files: [{ path: "a" }, { path: "a" }] } - }, - "artifact-manifest-output-path-uniqueness": { - positive: { output_contracts: [{ path: "a" }] }, - negative: { output_contracts: [{ path: "a" }, { path: "a" }] } - }, - "artifact-manifest-prerequisite-node-uniqueness": { - positive: { prerequisite_manifests: [{ node_id: "a" }] }, - negative: { prerequisite_manifests: [{ node_id: "a" }, { node_id: "a" }] } - }, - "artifact-verification-artifact-path-uniqueness": { - positive: { artifacts: [{ path: "a" }] }, - negative: { artifacts: [{ path: "a" }, { path: "a" }] } - }, - "artifact-verification-exactly-one-primary": { - positive: { artifacts: [{ primary: true }] }, - negative: { artifacts: [{ primary: false }] } - }, - "artifact-verification-publication-digest-correspondence": { - positive: { artifacts: [{ path: "a", sha256: "1" }], publications: [{ path: "a", sha256: "1" }] }, - negative: { artifacts: [{ path: "a", sha256: "1" }], publications: [{ path: "a", sha256: "2" }] } - }, - "artifact-verification-publication-path-uniqueness": { - positive: { publications: [{ path: "a" }] }, - negative: { publications: [{ path: "a" }, { path: "a" }] } - }, "attempt-order": { positive: { lifecycle: { started_at: "2026-01-01T00:00:00Z", finished_at: "2026-01-01T00:00:01Z" } }, negative: { lifecycle: { started_at: "2026-01-01T00:00:01Z", finished_at: "2026-01-01T00:00:00Z" } } @@ -741,20 +627,6 @@ const fixtures = { positive: { backend_results: [{ fuzzer_backend: "a" }] }, negative: { backend_results: [{ fuzzer_backend: "a" }, { fuzzer_backend: "a" }] } }, - "config-redactions-path-key-equality": { - positive: { - entries: [{ path: ["models", "profiles", "default", "model"], key: "models.profiles.default.model" }] - }, - negative: { entries: [{ path: ["models", "profiles", "default", "model"], key: "wrong.path" }] } - }, - "config-redactions-path-uniqueness": { - positive: { - entries: [{ path: ["models", "profiles", "a", "model"] }, { path: ["models", "profiles", "b", "model"] }] - }, - negative: { - entries: [{ path: ["models", "profiles", "a", "model"] }, { path: ["models", "profiles", "a", "model"] }] - } - }, "coverage-evidence-reconciliation": { positive: validCoverageEvidence, negative: { @@ -932,47 +804,6 @@ const fixtures = { positive: { records: [{ dedupe_key: "a" }] }, negative: { records: [{ dedupe_key: "a" }, { dedupe_key: "a" }] } }, - "finding-campaign-provenance-coherence": { - positive: { - property_ids: ["property-1"], - fuzzer_backend: "recon", - contributing_backend_failures: [ - { - fuzzer_backend: "recon", - failure_id: "failure-1", - raw_result_ref: "recon-fuzzer-results.json" - } - ], - deduplication: { pre_dedup_count: 1 } - }, - negative: { - property_ids: ["property-1"], - fuzzer_backend: "medusa", - contributing_backend_failures: [ - { - fuzzer_backend: "recon", - failure_id: "failure-1", - raw_result_ref: "recon-fuzzer-results.json" - } - ], - deduplication: { pre_dedup_count: 1 } - } - }, - "finding-evidence-span-consistency": { - positive: { - evidence: [{ line: 4, end_line: 8 }, { line_ranges: [{ line: 10, end_line: 12 }, { line: 14 }] }] - }, - negative: { evidence: [{ line: 8, end_line: 4 }] } - }, - "finding-projected-reference-uniqueness": { - positive: { family_variants: [{ id: "a", dedupe_key: "a" }] }, - negative: { - family_variants: [ - { id: "a", dedupe_key: "a" }, - { id: "a", dedupe_key: "b" } - ] - } - }, "findings-id-uniqueness": { positive: [{ id: "a" }], negative: [{ id: "a" }, { id: "a" }] @@ -1074,97 +905,6 @@ const fixtures = { positive: { files: [{ path: "a" }] }, negative: { files: [{ path: "a" }, { path: "a" }] } }, - "invariant-suite-file-path-uniqueness": { - positive: { files: [{ path: "a" }] }, - negative: { files: [{ path: "a" }, { path: "a" }] } - }, - "invariant-suite-file-tombstone-disjointness": { - positive: { files: [{ path: "a" }], tombstones: ["b"] }, - negative: { files: [{ path: "a" }], tombstones: ["a"] } - }, - "invariant-suite-tombstone-uniqueness": { - positive: { tombstones: ["a"] }, - negative: { tombstones: ["a", "a"] } - }, - "planned-graph-acyclicity": { - positive: validPlannedGraph, - negative: { - nodes: [ - { ...validPlannedNode, id: "a", depends_on: ["b"] }, - { ...validPlannedNode, id: "b", depends_on: ["a"] } - ] - } - }, - "planned-graph-artifact-dir-identity": { - positive: validPlannedGraph, - negative: { nodes: [{ ...validPlannedNode, artifact_dir: "artifacts/other" }] } - }, - "planned-graph-contract-identity": { - positive: validPlannedGraph, - negative: { - nodes: [{ ...validPlannedNode, outputs: [{ ...validPlannedOutput, contract_digest: "0".repeat(64) }] }] - } - }, - "planned-graph-dependency-join": { - positive: validPlannedGraph, - negative: { nodes: [{ ...validPlannedNode, depends_on: ["missing"] }] } - }, - "planned-graph-exactly-one-primary": { - positive: validPlannedGraph, - negative: { nodes: [{ ...validPlannedNode, outputs: [{ ...validPlannedOutput, primary: false }] }] } - }, - "planned-graph-loop-coupling": { - positive: validPlannedGraph, - negative: { nodes: [{ ...validPlannedNode, loop: { index: 1, count: 1, attempt_index: 1 } }] } - }, - "planned-graph-model-fanout-uniqueness": { - positive: validPlannedGraph, - negative: { - nodes: [ - { - ...validPlannedNode, - model_fanout: [ - { model_profile_id: "m", model_index: 0, loop_index: 0, attempt_index: 0 }, - { model_profile_id: "m", model_index: 0, loop_index: 0, attempt_index: 0 } - ] - } - ] - } - }, - "planned-graph-model-loop-coupling": { - positive: validPlannedGraph, - negative: { - nodes: [ - { - ...validPlannedNode, - model_fanout: [{ model_profile_id: "m", model_index: 0, loop_index: 1, attempt_index: 0 }] - } - ] - } - }, - "planned-graph-node-id-uniqueness": { - positive: validPlannedGraph, - negative: { nodes: [validPlannedNode, { ...validPlannedNode }] } - }, - "planned-graph-output-path-uniqueness": { - positive: validPlannedGraph, - negative: { - nodes: [{ ...validPlannedNode, outputs: [validPlannedOutput, { ...validPlannedOutput, primary: false }] }] - } - }, - "planned-graph-workflow-node-join": { - positive: { nodes: [{ ...validPlannedNode, workflow: { node_id: "a", task_node_ids: ["a"] } }] }, - negative: { nodes: [{ ...validPlannedNode, workflow: { node_id: "a", task_node_ids: ["b"] } }] } - }, - "planned-graph-workflow-task-uniqueness": { - positive: { nodes: [{ ...validPlannedNode, workflow: { node_id: "a", task_node_ids: ["a"] } }] }, - negative: { - nodes: [ - { ...validPlannedNode, workflow: { node_id: "a", task_node_ids: ["a"] } }, - { ...validPlannedNode, id: "node-b", workflow: { node_id: "a", task_node_ids: ["a"] } } - ] - } - }, "property-campaign-coverage-metric-uniqueness": { positive: { coverage: { metrics: [{ name: "branches" }] } }, negative: { coverage: { metrics: [{ name: "branches" }, { name: "branches" }] } } @@ -1308,28 +1048,6 @@ const fixtures = { } } }, - "run-metadata-accounting-workflow-identity": { - positive: { - workflow: { run_id: "workflow-a" }, - accounting: { workflow_run_id: "workflow-a", current: { workflow_run_id: "workflow-a" } } - }, - negative: { - workflow: { run_id: "workflow-a" }, - accounting: { workflow_run_id: "workflow-b", current: { workflow_run_id: "workflow-b" } } - } - }, - "run-metadata-current-segment-equality": { - positive: { accounting: { current: { total_tokens: 2 }, segments: [{ total_tokens: 1 }, { total_tokens: 2 }] } }, - negative: { accounting: { current: { total_tokens: 1 }, segments: [{ total_tokens: 1 }, { total_tokens: 2 }] } } - }, - "run-metadata-workflow-id-equality": { - positive: { workflow_ids: ["workflow-a"], workflow: { run_id: "workflow-a" } }, - negative: { workflow_ids: ["workflow-b"], workflow: { run_id: "workflow-a" } } - }, - "run-plan-attempt-id-uniqueness": { - positive: { rendered_prompts: [{ attempt_id: "a" }, { attempt_id: "b" }] }, - negative: { rendered_prompts: [{ attempt_id: "a" }, { attempt_id: "a" }] } - }, "run-state-node-key-equality": { positive: { nodes: { a: { node_id: "a" } } }, negative: { nodes: { a: { node_id: "b" } } } @@ -1360,75 +1078,6 @@ const fixtures = { positive: [{ severity: "Medium", impact: "High", likelihood: "Low" }], negative: [{ severity: "High", impact: "High", likelihood: "Low" }] }, - "smithers-task-attempt-id-uniqueness": { - positive: { tasks: [{ attemptId: "a" }] }, - negative: { tasks: [{ attemptId: "a" }, { attemptId: "a" }] } - }, - "smithers-task-workflow-id-uniqueness": { - positive: { tasks: [{ smithersNodeId: "node:a", verifierSmithersNodeId: "verify:a" }] }, - negative: { - tasks: [ - { smithersNodeId: "node:a", verifierSmithersNodeId: "verify:a" }, - { smithersNodeId: "node:a", verifierSmithersNodeId: "verify:b" } - ] - } - }, - "smithers-task-document-identity": { - positive: { run_id: "run-a", workflow_name: "workflow-a", tasks: [smithersIdentityTask] }, - negative: { - run_id: "run-a", - workflow_name: "workflow-a", - tasks: [{ ...smithersIdentityTask, smithersNodeId: "node:wrong" }] - } - }, - "smithers-task-pinned-submodule-expectation": { - positive: { pinned_submodules: pinnedSubmoduleExpectation, tasks: [{ execution: { mode: "cloud" } }] }, - negative: { - pinned_submodules: { ...pinnedSubmoduleExpectation, file_count: 2 }, - tasks: [{ execution: { mode: "cloud" } }] - } - }, - "smithers-task-dependency-join": { - positive: { - tasks: [ - { attemptId: "a", verifierSmithersNodeId: "verify:a", dependencies: [], dependencySmithersNodeIds: [] }, - { - attemptId: "b", - verifierSmithersNodeId: "verify:b", - dependencies: ["a"], - dependencySmithersNodeIds: ["verify:a"] - } - ] - }, - negative: { - tasks: [ - { - attemptId: "b", - verifierSmithersNodeId: "verify:b", - dependencies: ["missing"], - dependencySmithersNodeIds: ["verify:missing"] - } - ] - } - }, - "smithers-task-dependency-acyclicity": { - positive: { - tasks: [ - { attemptId: "a", verifierSmithersNodeId: "verify:a", dependencySmithersNodeIds: [] }, - { attemptId: "b", verifierSmithersNodeId: "verify:b", dependencySmithersNodeIds: ["verify:a"] } - ] - }, - negative: { - tasks: [ - { attemptId: "a", verifierSmithersNodeId: "verify:a", dependencySmithersNodeIds: ["verify:b"] }, - { attemptId: "b", verifierSmithersNodeId: "verify:b", dependencySmithersNodeIds: ["verify:a"] } - ] - } - }, - "source-run-not-self": { - positive: { run_id: "run-new", source_run_id: "run-source" }, - negative: { run_id: "run-same", source_run_id: "run-same" } - }, "strategy-detection-dedupe-key-uniqueness": { positive: [{ dedupe_key: "a" }], negative: [{ dedupe_key: "a" }, { dedupe_key: "a" }] @@ -1563,12 +1212,6 @@ test("severity and report vocabulary gates reject renamed reachability tokens at }); test("property references alone do not claim fuzzer campaign provenance", () => { - assert.equal( - executeSemanticGate("finding-campaign-provenance-coherence", { - document: { property_ids: ["property-1"] } - }).status, - "passed" - ); assert.equal( executeSemanticGate("findings-campaign-provenance-coherence", { document: [{ property_ids: ["property-1"] }] @@ -1576,24 +1219,26 @@ test("property references alone do not claim fuzzer campaign provenance", () => "passed" ); assert.equal( - executeSemanticGate("finding-campaign-provenance-coherence", { - document: { property_ids: ["property-1"], fuzzer_backend: "recon" } + executeSemanticGate("findings-campaign-provenance-coherence", { + document: [{ property_ids: ["property-1"], fuzzer_backend: "recon" }] }).status, "passed" ); assert.equal( - executeSemanticGate("finding-campaign-provenance-coherence", { - document: { property_ids: ["property-1"], fuzzer_backend: "recon", deduplication: { pre_dedup_count: 1 } } + executeSemanticGate("findings-campaign-provenance-coherence", { + document: [{ property_ids: ["property-1"], fuzzer_backend: "recon", deduplication: { pre_dedup_count: 1 } }] }).status, "passed" ); assert.equal( - executeSemanticGate("finding-campaign-provenance-coherence", { - document: { - property_ids: ["property-1"], - fuzzer_backend: "recon", - contributing_backend_failures: [{ fuzzer_backend: "recon", failure_id: "failure-1" }] - } + executeSemanticGate("findings-campaign-provenance-coherence", { + document: [ + { + property_ids: ["property-1"], + fuzzer_backend: "recon", + contributing_backend_failures: [{ fuzzer_backend: "recon", failure_id: "failure-1" }] + } + ] }).status, "failed" ); @@ -1755,13 +1400,6 @@ test("workspace patch path gate reports exact nonduplicated field diagnostics", test("canonical finding span semantics run for every embedding schema", () => { const cases = [ - { - filename: "finding.schema.json", - gate: "finding-evidence-span-consistency", - positive: { evidence: [{ line: 3, end_line: 5 }] }, - negative: { evidence: [{ line: 5, end_line: 3 }] }, - path: "$.evidence[0].end_line" - }, { filename: "findings.schema.json", gate: "findings-evidence-span-consistency", @@ -2642,12 +2280,10 @@ test("every contextual registration executes real positive and negative checks", const root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ultrafuzz-semantic-gates-")); try { fs.mkdirSync(path.join(root, "generated-tests")); - fs.writeFileSync(path.join(root, "artifact.json"), "artifact\n"); fs.writeFileSync(path.join(root, "generated-tests", "test.sol"), "test\n"); fs.writeFileSync(path.join(root, "generated-tests", "helper.sol"), "helper\n"); fs.writeFileSync(path.join(root, "generated-tests", "binary.dat"), Buffer.from([0xff])); fs.writeFileSync(path.join(root, "copied.sol"), "test\n"); - const digest = crypto.createHash("sha256").update("artifact\n").digest("hex"); const contentDigest = crypto.createHash("sha256").update("snapshot", "utf8").digest("hex"); const aggregationSourceBytes = Buffer.from("test\n", "utf8"); const aggregationSourceDigest = crypto.createHash("sha256").update(aggregationSourceBytes).digest("hex"); @@ -2723,16 +2359,6 @@ test("every contextual registration executes real positive and negative checks", Exclude, { positive: unknown; negative: unknown; context: SemanticGateContext } > = { - "agent-source-proof-commit-binding": { - positive: { commit: "c", tree: "t", refs: [{ name: "r", object: "o" }] }, - negative: { commit: "wrong", tree: "t", refs: [{ name: "r", object: "o" }] }, - context: { git: { commit: "c", tree: "t", refs: { r: "o" } } } - }, - "analysis-bundle-file-digest": { - positive: { files: [{ path: "artifact.json", sha256: digest, size_bytes: 9 }] }, - negative: { files: [{ path: "artifact.json", sha256: "0".repeat(64), size_bytes: 9 }] }, - context: { filesystem: { rootDirectory: root } } - }, "analysis-bundle-inclusion-omission-coverage": { positive: { omissions: [ @@ -2795,16 +2421,6 @@ test("every contextual registration executes real positive and negative checks", } } }, - "artifact-manifest-file-digest": { - positive: { files: [{ path: "artifact.json", sha256: digest, size_bytes: 9 }] }, - negative: { files: [{ path: "missing.json", sha256: digest, size_bytes: 9 }] }, - context: { filesystem: { rootDirectory: root } } - }, - "artifact-verification-plan-contract-identity": { - positive: { node_id: "a", artifacts: [{ ...validPlannedOutput }] }, - negative: { node_id: "b", artifacts: [{ ...validPlannedOutput }] }, - context: { plannedGraph: { node: { id: "a", outputs: [{ ...validPlannedOutput }] } } } - }, "attempt-reuse-source-link": { positive: { workflow_run_id: "workflow-current", @@ -3543,108 +3159,6 @@ test("every contextual registration executes real positive and negative checks", } } }, - "run-state-fingerprint": { - positive: { graph_fingerprint: "g", config_fingerprint: "c" }, - negative: { graph_fingerprint: "wrong", config_fingerprint: "c" }, - context: { runtimeState: { graphFingerprint: "g", configFingerprint: "c" } } - }, - "smithers-task-planned-graph-coverage": { - positive: { - tasks: [{ attemptId: "node-a", concreteNodeId: "node-a" }] - }, - negative: { tasks: [] }, - context: { - plannedGraph: { - document: { - nodes: [ - { - ...validPlannedNode, - logical_id: "logical-a", - display_name: "Node A", - kind: "agentic" - } - ] - } - } - } - }, - "smithers-task-planned-graph-identity": { - positive: { - tasks: [ - { - attemptId: "node-a", - concreteNodeId: "node-a", - logicalNodeId: "logical-a", - metadata: { - node: { logicalNodeId: "logical-a", label: "Node A" }, - loop: { index: 0, count: 1, mode: "parallel", attemptIndex: 0 }, - artifacts: { outputs: [smithersManifestOutput] } - } - } - ] - }, - negative: { - tasks: [ - { - attemptId: "node-a", - concreteNodeId: "node-a", - logicalNodeId: "wrong", - metadata: { - node: { logicalNodeId: "wrong", label: "Node A" }, - loop: { index: 0, count: 1, mode: "parallel", attemptIndex: 0 }, - artifacts: { outputs: [smithersManifestOutput] } - } - } - ] - }, - context: { - plannedGraph: { - document: { - nodes: [ - { - ...validPlannedNode, - logical_id: "logical-a", - display_name: "Node A", - kind: "agentic", - loop: { index: 0, count: 1, mode: "parallel", attempt_index: 0 } - } - ] - } - } - } - }, - "smithers-task-planned-graph-dependency-join": { - positive: { - tasks: [ - { - concreteNodeId: "node-a", - dependencies: ["node-b"], - dependencySmithersNodeIds: ["verify:node-b"], - metadata: { dependencies: { concreteNodeIds: ["node-b"] } } - } - ] - }, - negative: { - tasks: [ - { - concreteNodeId: "node-a", - dependencies: [], - dependencySmithersNodeIds: [], - metadata: { dependencies: { concreteNodeIds: [] } } - } - ] - }, - context: { - plannedGraph: { - document: { - nodes: [ - { ...validPlannedNode, id: "node-a", depends_on: ["node-b"], kind: "agentic" }, - { ...validPlannedNode, id: "node-b", depends_on: [], kind: "agentic" } - ] - } - } - } - }, "usage-ledger-event-order": { positive: { workflow_run_id: "workflow-a", source_event_sequence: 2, control_generation: "a".repeat(64) }, negative: { workflow_run_id: "workflow-a", source_event_sequence: 0, control_generation: "a".repeat(64) }, @@ -3734,170 +3248,6 @@ test("every contextual registration executes real positive and negative checks", } }); -test("Smithers planned dependency semantics omit unresolved dynamic groups only", () => { - const pendingGraph = { - nodes: [ - { - ...validPlannedNode, - id: "consumer", - depends_on: ["fanout"], - dynamic_dependencies: ["fanout"], - kind: "agentic" - }, - { - ...validPlannedNode, - id: "fanout", - depends_on: [], - kind: "agentic", - dynamic: { status: "pending" } - } - ] - }; - const pendingTask = { - tasks: [ - { - concreteNodeId: "consumer", - dependencies: [], - dependencySmithersNodeIds: [], - metadata: { dependencies: { concreteNodeIds: [] } } - } - ] - }; - - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: pendingTask, - context: { plannedGraph: { document: pendingGraph } } - }).status, - "passed" - ); - - const compiledPendingTask = { - tasks: [ - { - ...structuredClone(pendingTask.tasks[0]!), - metadata: { dependencies: { concreteNodeIds: ["fanout"] } } - } - ] - }; - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: compiledPendingTask, - context: { plannedGraph: { document: pendingGraph } } - }).status, - "passed" - ); - - const twoPendingGraph = { - nodes: [ - { - ...structuredClone(pendingGraph.nodes[0]!), - depends_on: ["fanout", "fanout-second"], - dynamic_dependencies: ["fanout", "fanout-second"] - }, - structuredClone(pendingGraph.nodes[1]!), - { - ...structuredClone(pendingGraph.nodes[1]!), - id: "fanout-second" - } - ] - }; - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: compiledPendingTask, - context: { plannedGraph: { document: twoPendingGraph } } - }).status, - "failed" - ); - - const materializedGraph = { - nodes: [ - { ...structuredClone(pendingGraph.nodes[0]!), depends_on: ["generated"] }, - { - ...structuredClone(pendingGraph.nodes[1]!), - dynamic: { status: "expanded" } - }, - { ...validPlannedNode, id: "generated", depends_on: [], kind: "agentic" } - ] - }; - const materializedTask = { - tasks: [ - { - concreteNodeId: "consumer", - dependencies: ["generated"], - dependencySmithersNodeIds: ["verify:generated"], - metadata: { dependencies: { concreteNodeIds: ["generated"] } } - } - ] - }; - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: materializedTask, - context: { plannedGraph: { document: materializedGraph } } - }).status, - "passed" - ); - - const staleExpandedGraph = structuredClone(materializedGraph); - staleExpandedGraph.nodes.find((node) => node.id === "consumer")!.depends_on = ["fanout"]; - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: pendingTask, - context: { plannedGraph: { document: staleExpandedGraph } } - }).status, - "failed" - ); - - const alignedStaleExpandedTask = { - tasks: [ - { - ...structuredClone(pendingTask.tasks[0]!), - dependencies: ["fanout"], - dependencySmithersNodeIds: ["verify:fanout"], - metadata: { dependencies: { concreteNodeIds: ["fanout"] } } - } - ] - }; - const alignedStaleExpanded = executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: alignedStaleExpandedTask, - context: { plannedGraph: { document: staleExpandedGraph } } - }); - assert.equal(alignedStaleExpanded.status, "failed"); - assert.ok( - alignedStaleExpanded.status === "failed" && - alignedStaleExpanded.issues.some((entry) => - /retains expanded dynamic dependency placeholder/u.test(entry.message) - ) - ); - - const ordinaryGraph = { - nodes: [ - { ...structuredClone(pendingGraph.nodes[0]!), depends_on: ["fanout", "ordinary"] }, - structuredClone(pendingGraph.nodes[1]!), - { ...validPlannedNode, id: "ordinary", depends_on: [], kind: "agentic" } - ] - }; - assert.equal( - executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: pendingTask, - context: { plannedGraph: { document: ordinaryGraph } } - }).status, - "failed" - ); - - const missingGraph = { - nodes: [{ ...structuredClone(pendingGraph.nodes[0]!), depends_on: ["missing"] }] - }; - const missing = executeSemanticGate("smithers-task-planned-graph-dependency-join", { - document: pendingTask, - context: { plannedGraph: { document: missingGraph } } - }); - assert.equal(missing.status, "failed"); - assert.ok( - missing.status === "failed" && missing.issues.some((entry) => /planned dependency node/u.test(entry.message)) - ); -}); - test("strict final reports preserve dropped false positives as exactly one non-production row", () => { const finding = { id: "finding-a", diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 8bebe8a23..3231b0d4f 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -2220,7 +2220,6 @@ function semanticGateContextForArtifact(input: { rootDirectory: input.artifactDir, ...(input.authenticated === undefined ? {} : { files: input.authenticated.publications }) }, - plannedGraph: { node: input.node }, artifactIdentity: { runId: input.layout.runId, nodeId: input.node.logical_id ?? input.node.id, From 6c3c69e015129fd3b63ba31fec345b66a3a73985 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:21:29 +0000 Subject: [PATCH 054/206] refactor(artifacts,security): delete code with no production caller Each symbol below had no caller in any package source, script or workflow template (git grep -w over the whole repository), and none ever appeared in the workflow template's history: - finding-provenance.ts (normalizeFindings, buildFindingSourceExpectations, readFindings, ...): only threat-goal-artifacts tests called it. - writeGeneratedTestManifest/readGeneratedTestManifest and their seven private helpers: agents write generated-tests.json and the host only validates it, so only artifacts.test.ts exercised the writer. - appendLineDurable and truncateDurable (safe-paths.ts), and buildRunArtifactIndex with its two types (manifests.ts). - loadOrCreateRunState, validateWorkflowContract, and 29 other workflow-contracts exports that occurred only at their declaration. - Declaration-only exports in findings.ts, findings-schema.ts, finding-note-vocabulary.ts, goal-plan.ts, invariant-ledger.ts, invariant-source-proof.ts, property-provenance.ts, state-schema.ts, runtime-schemas.ts, report-observation.ts, threat-model.ts, usage-ledger.ts and artifact-contract-ids.ts. - security: resolvePathInside and its helpers/types, the UnsafeModeAudit* types, and PolicyResult.audit (no caller ever passed it; it was always []). Tests of the deleted code go with it. The two usage-ledger tests that used appendLineDurable to inject a bad line now use fs.appendFileSync, and the stale-size/hard-link test keeps its appendBytesDurableAt assertions. writeArtifact stays: it has no production caller either, but ~350 test call sites across three suites use it as a fixture helper, and moving it into each suite would add code rather than remove it. Schema bytes, contract digests and VALIDATOR_BUILD_IDENTITY are unchanged. Refs #462 Co-Authored-By: Claude Opus 5.5 --- .../artifacts/src/artifact-contract-ids.ts | 4 - .../artifacts/src/finding-note-vocabulary.ts | 22 - packages/artifacts/src/finding-provenance.ts | 506 ------------- packages/artifacts/src/findings-schema.ts | 18 +- packages/artifacts/src/findings.ts | 10 - packages/artifacts/src/generated-tests.ts | 311 +------- packages/artifacts/src/goal-plan.ts | 4 - packages/artifacts/src/index.ts | 1 - packages/artifacts/src/invariant-ledger.ts | 7 +- .../artifacts/src/invariant-source-proof.ts | 7 +- packages/artifacts/src/manifests.ts | 44 -- packages/artifacts/src/property-provenance.ts | 15 +- packages/artifacts/src/report-observation.ts | 1 - packages/artifacts/src/runtime-schemas.ts | 23 - packages/artifacts/src/safe-paths.ts | 71 -- packages/artifacts/src/state-schema.ts | 21 +- packages/artifacts/src/state.ts | 15 - packages/artifacts/src/threat-model.ts | 1 - packages/artifacts/src/usage-ledger.ts | 1 - packages/artifacts/src/workflow-contracts.ts | 41 -- packages/artifacts/test/artifacts.test.ts | 673 +----------------- .../artifacts/test/events-tail-repair.test.ts | 7 +- .../test/threat-goal-artifacts.test.ts | 432 ----------- packages/security/src/path-policy.ts | 90 --- packages/security/src/types.ts | 23 +- packages/security/test/path-policy.test.ts | 32 +- 26 files changed, 14 insertions(+), 2366 deletions(-) delete mode 100644 packages/artifacts/src/finding-provenance.ts diff --git a/packages/artifacts/src/artifact-contract-ids.ts b/packages/artifacts/src/artifact-contract-ids.ts index 67cb07fca..f01440849 100644 --- a/packages/artifacts/src/artifact-contract-ids.ts +++ b/packages/artifacts/src/artifact-contract-ids.ts @@ -62,7 +62,3 @@ export const JSON_ARTIFACT_CONTRACT_IDS = ARTIFACT_CONTRACT_IDS.filter( export function isArtifactContractId(value: unknown): value is ArtifactContractId { return typeof value === "string" && (ARTIFACT_CONTRACT_IDS as readonly string[]).includes(value); } - -export function isJsonArtifactContractId(value: unknown): value is JsonArtifactContractId { - return typeof value === "string" && (JSON_ARTIFACT_CONTRACT_IDS as readonly string[]).includes(value); -} diff --git a/packages/artifacts/src/finding-note-vocabulary.ts b/packages/artifacts/src/finding-note-vocabulary.ts index e8c2fc687..1073f9057 100644 --- a/packages/artifacts/src/finding-note-vocabulary.ts +++ b/packages/artifacts/src/finding-note-vocabulary.ts @@ -26,14 +26,6 @@ export const FINDING_NOTE_KEYS = [ "impact" ] as const; -/** ASCII-only, case-insensitive forms used by both the runtime and portable - * JSON-Schema grammar. Unicode compatibility folds are deliberately excluded: - * JSON Schema has no portable equivalent, and identifiers in the authority are - * ASCII exact apart from case. */ -export const FINDING_NOTE_KEYS_ASCII_CASE_INSENSITIVE_PATTERN = `(?:${FINDING_NOTE_KEYS.map((key) => - key.replace(/[A-Za-z]/gu, (character) => `[${character.toLowerCase()}${character.toUpperCase()}]`) -).join("|")})`; - export const FINDING_REPORT_ASSIGNMENT_KEY_PATTERN = "[-_0-9A-Za-z]{1,128}"; const findingReportAssignmentKey = new RegExp(`^${FINDING_REPORT_ASSIGNMENT_KEY_PATTERN}$`, "u"); @@ -129,14 +121,6 @@ export const FINDING_REPORT_MULTITERM_METADATA_KEY_PATTERNS = FINDING_REPORT_MET FINDING_REPORT_METADATA_TERM_PATTERNS.map((right) => `[-_0-9A-Za-z]*${left}[-_0-9A-Za-z]*${right}[-_0-9A-Za-z]*`) ); -/** A bounded ASCII identifier containing a report-metadata term. This defines - * a family, rather than a blacklist of guessed aliases, so producer-local - * renames such as `helper_summary` and `reachability_note` fail closed. */ -export const FINDING_REPORT_METADATA_KEY_PATTERNS = FINDING_REPORT_METADATA_TERM_GROUPS.map( - (_, index) => - `(?=${FINDING_REPORT_ASSIGNMENT_KEY_PATTERN}\\s*={1,2})[-_0-9A-Za-z]*${FINDING_REPORT_METADATA_TERM_PATTERNS[index]}[-_0-9A-Za-z]*` -); - export function isFindingReportMetadataKey(key: string): boolean { if (!findingReportAssignmentKey.test(key)) return false; const normalized = key.replace(/[A-Z]/gu, (character) => character.toLowerCase()); @@ -154,12 +138,6 @@ export function isFindingReportMetadataAliasKey(key: string): boolean { return new Set(terms).size >= 2 || (hasAliasBase && hasRenameMarker); } -export function isFindingReportEvidenceAssignmentKey(key: string): boolean { - return FINDING_REPORT_EVIDENCE_ASSIGNMENT_KEYS.includes( - key as (typeof FINDING_REPORT_EVIDENCE_ASSIGNMENT_KEYS)[number] - ); -} - export function findingReachabilityPromptVocabulary(): string { const reachabilityKey = FINDING_NOTE_KEYS[0]; return FINDING_REACHABILITY_VALUES.map((value) => `- \`${reachabilityKey}=${value}\``).join("\n"); diff --git a/packages/artifacts/src/finding-provenance.ts b/packages/artifacts/src/finding-provenance.ts deleted file mode 100644 index 483f4f8be..000000000 --- a/packages/artifacts/src/finding-provenance.ts +++ /dev/null @@ -1,506 +0,0 @@ -import fs from "node:fs"; -import path from "node:path"; - -import { FINDINGS_FILE, FINDINGS_SCHEMA_VERSION } from "./findings.js"; -import { - ArtifactPathError, - readJsonFile, - safeResolveInside, - validateNodeReference, - writeJsonDurable -} from "./safe-paths.js"; - -export interface FindingProvenance { - nodeId?: string; - producerNodeId?: string; - strategy?: string; - attemptIndex?: number; - modelId?: string; - model?: string; - modelIndex?: number; - loopIndex?: number; -} - -export interface NormalizeFindingsInput { - artifactDir: string; - relativePath?: string; - nodeId?: string; - provenance?: FindingProvenance; - preserveSourceNodes?: boolean; - requireSourceNodes?: boolean; - allowedSourceNodes?: readonly string[]; - sourceExpectations?: readonly FindingSourceExpectation[]; - requireSourceExpectation?: boolean; -} - -export interface FindingSourceExpectation { - finding_keys: readonly string[]; - source_nodes: readonly string[]; - dedupe_keys: readonly string[]; - family_ids: readonly string[]; - finding_ids: readonly string[]; - lifecycle_record?: boolean; -} - -export interface UpstreamFindingSource { - node_id: string; - artifact_path: string; - finding: unknown; -} - -export interface FindingsNormalizeReport { - schema_version: typeof FINDINGS_SCHEMA_VERSION; - source_path: string; - normalized_path: string; - count: number; - findings: Array>; -} - -export class FindingsValidationError extends Error { - constructor(message: string) { - super(message); - this.name = "FindingsValidationError"; - } -} - -/** - * Attach controller-owned provenance without repairing producer-visible finding - * fields. The caller still performs the current strict findings-contract check. - */ -export function normalizeFindings(input: NormalizeFindingsInput): FindingsNormalizeReport { - const artifactDir = path.resolve(input.artifactDir); - const relativePath = input.relativePath ?? FINDINGS_FILE; - const sourcePath = resolveFindingsSource(artifactDir, relativePath); - const raw = readJsonFile(sourcePath); - if (!Array.isArray(raw)) { - throw new FindingsValidationError(`${relativePath} must contain a findings array`); - } - const findings = raw.map((value, index) => normalizeFindingProvenance(value, index, input)); - const normalizedPath = safeResolveInside(artifactDir, relativePath, "normalized findings path"); - writeJsonDurable(normalizedPath, findings); - return { - schema_version: FINDINGS_SCHEMA_VERSION, - source_path: sourcePath, - normalized_path: normalizedPath, - count: findings.length, - findings - }; -} - -export function readFindings(artifactDir: string): Array> { - const value = readJsonFile(path.join(artifactDir, FINDINGS_FILE)); - if (!Array.isArray(value) || !value.every(isPlainRecord)) { - throw new FindingsValidationError(`${FINDINGS_FILE} must contain a findings array`); - } - return value; -} - -export function findingIdentityKeys(value: unknown): string[] { - const identity = findingIdentity(value); - return uniqueNonEmptyStrings([...identity.dedupeKeys, ...identity.familyIds, ...identity.findingIds]); -} - -/** - * Derive authoritative discovery-source expectations from verified upstream - * findings. A many-to-one lifecycle record binds its output dedupe key to the - * exact union of its authenticated source artifacts. - */ -export function buildFindingSourceExpectations(input: { - upstream: readonly UpstreamFindingSource[]; - lifecycleLedger?: unknown; - requireLifecycleCoverage?: boolean; -}): FindingSourceExpectation[] { - const upstream = input.upstream.map((entry, index) => normalizeUpstreamFindingSource(entry, index)); - const expectations: FindingSourceExpectation[] = upstream.map((entry) => ({ - finding_keys: entry.keys, - source_nodes: entry.sourceNodes, - dedupe_keys: entry.identity.dedupeKeys, - family_ids: entry.identity.familyIds, - finding_ids: entry.identity.findingIds - })); - if (input.lifecycleLedger === undefined) return expectations; - - if (!isPlainRecord(input.lifecycleLedger) || !Array.isArray(input.lifecycleLedger.records)) { - throw new FindingsValidationError("finding lifecycle ledger must contain a records array"); - } - const covered = new Set(); - const lifecycleDedupeKeys = new Set(); - for (const [recordIndex, rawRecord] of input.lifecycleLedger.records.entries()) { - if (!isPlainRecord(rawRecord)) { - throw new FindingsValidationError(`finding lifecycle record ${recordIndex} must be an object`); - } - const dedupeKey = nonEmptyString(rawRecord.dedupe_key); - if (dedupeKey === undefined) { - throw new FindingsValidationError(`finding lifecycle record ${recordIndex} requires dedupe_key`); - } - if (lifecycleDedupeKeys.has(dedupeKey)) { - throw new FindingsValidationError(`finding lifecycle ledger contains duplicate dedupe_key: ${dedupeKey}`); - } - lifecycleDedupeKeys.add(dedupeKey); - const identity = findingIdentity(rawRecord); - const keys = findingIdentityKeys(rawRecord); - if (keys.length === 0) { - throw new FindingsValidationError(`finding lifecycle record ${recordIndex} has no stable finding key`); - } - if (!Array.isArray(rawRecord.source_artifacts) || rawRecord.source_artifacts.length === 0) { - throw new FindingsValidationError(`finding lifecycle record ${recordIndex} has no source_artifacts`); - } - const sourceNodes: string[] = []; - for (const [sourceIndex, rawSource] of rawRecord.source_artifacts.entries()) { - if (!isPlainRecord(rawSource)) { - throw new FindingsValidationError( - `finding lifecycle record ${recordIndex} source_artifact ${sourceIndex} must be an object` - ); - } - const findingId = requiredLifecycleString(rawSource, "finding_id", recordIndex, sourceIndex); - const nodeId = validateNodeReference( - requiredLifecycleString(rawSource, "node_id", recordIndex, sourceIndex), - "lifecycle source artifact node ID" - ); - const matches = upstream.filter( - (candidate) => candidate.findingId === findingId && candidate.identities.includes(nodeId) - ); - if (matches.length !== 1) { - throw new FindingsValidationError( - `finding lifecycle source artifact does not identify exactly one dependency finding: ${nodeId}:${findingId}` - ); - } - const matched = matches[0]!; - if (covered.has(matched.index)) { - throw new FindingsValidationError( - `dependency finding appears more than once in the lifecycle ledger: ${nodeId}:${findingId}` - ); - } - covered.add(matched.index); - appendUnique(sourceNodes, matched.sourceNodes); - } - expectations.push({ - finding_keys: keys, - source_nodes: sourceNodes, - dedupe_keys: identity.dedupeKeys, - family_ids: identity.familyIds, - finding_ids: identity.findingIds, - lifecycle_record: true - }); - } - if (input.requireLifecycleCoverage === true && covered.size !== upstream.length) { - const missing = upstream - .filter((entry) => !covered.has(entry.index)) - .map((entry) => `${entry.nodeId}:${entry.findingId}`); - throw new FindingsValidationError(`finding lifecycle ledger omitted dependency findings: ${missing.join(", ")}`); - } - return expectations; -} - -/** Resolve a downstream finding to its exact authoritative discovery-source set. */ -export function resolveExpectedFindingSourceNodes( - finding: unknown, - expectations: readonly FindingSourceExpectation[] -): string[] | undefined { - if (!isPlainRecord(finding)) { - throw new FindingsValidationError("finding provenance candidate must be an object"); - } - return expectedSourceNodes(finding, expectations); -} - -function resolveFindingsSource(artifactDir: string, relativePath: string): string { - const findingsPath = safeResolveInside(artifactDir, relativePath, "findings source path"); - if (fs.existsSync(findingsPath) && fs.lstatSync(findingsPath).isSymbolicLink()) { - throw new ArtifactPathError("symlink-escape", `findings source path cannot be a symlink: ${relativePath}`); - } - if (!fs.existsSync(findingsPath)) throw new Error(`missing ${relativePath} in ${artifactDir}`); - return findingsPath; -} - -function normalizeFindingProvenance( - value: unknown, - index: number, - input: NormalizeFindingsInput -): Record { - if (!isPlainRecord(value)) throw new FindingsValidationError(`finding ${index} must be an object`); - const normalized = { ...value }; - const provenance = input.provenance ?? {}; - const nodeId = input.nodeId ?? provenance.nodeId; - const producerNodeId = provenance.producerNodeId ?? provenance.nodeId ?? input.nodeId; - if (producerNodeId !== undefined) { - normalized.producer_node_id = validateNodeReference(producerNodeId, "finding producer node ID"); - } - normalizeSourceNodes( - normalized, - nodeId, - producerNodeId, - input.preserveSourceNodes === true, - input.requireSourceNodes === true, - input.allowedSourceNodes, - input.sourceExpectations, - input.requireSourceExpectation === true - ); - assignIfMissing(normalized, "strategy", provenance.strategy); - assignIfMissing(normalized, "attempt_index", provenance.attemptIndex); - assignIfMissing(normalized, "model_id", provenance.modelId); - assignIfMissing(normalized, "model", provenance.model); - assignIfMissing(normalized, "model_index", provenance.modelIndex); - assignIfMissing(normalized, "loop_index", provenance.loopIndex); - return normalized; -} - -function normalizeSourceNodes( - record: Record, - nodeId: string | undefined, - producerNodeId: string | undefined, - preserveExisting: boolean, - requireSourceNodes: boolean, - allowedSourceNodes: readonly string[] | undefined, - sourceExpectations: readonly FindingSourceExpectation[] | undefined, - requireSourceExpectation: boolean -): void { - const existing = record.source_nodes; - if (existing !== undefined && !Array.isArray(existing)) { - throw new FindingsValidationError("field source_nodes must be an array of non-empty strings"); - } - const values = preserveExisting - ? [ - ...(Array.isArray(existing) ? existing : []), - typeof record.source_node_id === "string" ? record.source_node_id : undefined - ] - : [producerNodeId]; - const normalized: string[] = []; - for (const value of values) { - if (value === undefined) continue; - if (typeof value !== "string" || value.trim().length === 0) { - throw new FindingsValidationError("field source_nodes must be an array of non-empty strings"); - } - const candidate = validateNodeReference(value.trim(), "finding source node ID"); - appendUnique(normalized, [candidate]); - } - if (normalized.length === 0) { - if (requireSourceNodes) { - throw new FindingsValidationError("field source_nodes must retain at least one discovery node ID"); - } - delete record.source_nodes; - delete record.source_node_id; - delete record.producer_attempt_id; - return; - } - if (allowedSourceNodes !== undefined) { - const allowed = new Set( - allowedSourceNodes.map((sourceNode) => validateNodeReference(sourceNode, "allowed finding source node ID")) - ); - const invented = normalized.filter((sourceNode) => !allowed.has(sourceNode)); - if (invented.length > 0) { - throw new FindingsValidationError( - `field source_nodes contains IDs not present in dependency findings: ${invented.join(", ")}` - ); - } - } - const expected = expectedSourceNodes(record, sourceExpectations); - if (expected === undefined && requireSourceExpectation) { - throw new FindingsValidationError("finding does not match any dependency provenance record"); - } - if (expected !== undefined) { - if (!sameStringSet(normalized, expected)) { - throw new FindingsValidationError( - "field source_nodes does not preserve the exact dependency discovery-source union" - ); - } - normalized.splice(0, normalized.length, ...expected); - } - record.source_nodes = normalized; - record.source_node_id = normalized[0]; - if (preserveExisting) { - if (record.producer_attempt_id !== undefined) { - const attemptId = nonEmptyString(record.producer_attempt_id); - if (attemptId === undefined) { - throw new FindingsValidationError("field producer_attempt_id must be a non-empty string"); - } - record.producer_attempt_id = validateNodeReference(attemptId, "finding producer attempt ID"); - } - } else { - delete record.producer_attempt_id; - if (nodeId !== undefined && nodeId !== normalized[0]) { - record.producer_attempt_id = validateNodeReference(nodeId, "finding producer attempt ID"); - } - } -} - -function expectedSourceNodes( - record: Record, - expectations: readonly FindingSourceExpectation[] | undefined -): string[] | undefined { - if (expectations === undefined) return undefined; - const identity = findingIdentity(record); - const tiers: Array<{ - label: string; - values: readonly string[]; - expectationValues: (expectation: FindingSourceExpectation) => readonly string[]; - rejectUnknown: boolean; - }> = [ - { - label: "dedupe key", - values: identity.dedupeKeys, - expectationValues: (expectation) => expectation.dedupe_keys, - rejectUnknown: true - }, - { - label: "family ID", - values: identity.familyIds, - expectationValues: (expectation) => expectation.family_ids, - rejectUnknown: false - }, - { - label: "finding reference", - values: identity.referenceIds, - expectationValues: (expectation) => expectation.finding_ids, - rejectUnknown: true - } - ]; - let authoritative: string[] | undefined; - for (const tier of tiers) { - for (const value of tier.values) { - const matches = expectations.filter((expectation) => tier.expectationValues(expectation).includes(value)); - if (matches.length === 0) { - if (tier.rejectUnknown) { - throw new FindingsValidationError(`finding ${tier.label} does not match any dependency provenance record`); - } - continue; - } - const lifecycleMatches = matches.filter((expectation) => expectation.lifecycle_record === true); - const resolved = exactExpectedSourceNodes(lifecycleMatches.length > 0 ? lifecycleMatches : matches, tier.label); - if (authoritative === undefined) authoritative = resolved; - else if (!sameStringSet(resolved, authoritative)) { - throw new FindingsValidationError(`finding ${tier.label} conflicts with higher-priority dependency provenance`); - } - } - } - if (authoritative !== undefined || identity.ownIds.length === 0) return authoritative; - const resolvedOwnIds = identity.ownIds.flatMap((value) => { - const matches = expectations.filter((expectation) => expectation.finding_ids.includes(value)); - return matches.length === 0 ? [] : [exactExpectedSourceNodes(matches, "finding ID")]; - }); - const expected = resolvedOwnIds[0]; - if (expected === undefined) return undefined; - if (resolvedOwnIds.some((candidate) => !sameStringSet(candidate, expected))) { - throw new FindingsValidationError("finding ID values resolve to conflicting dependency provenance"); - } - return expected; -} - -function exactExpectedSourceNodes(matches: readonly FindingSourceExpectation[], identityLabel: string): string[] { - const candidates = matches.map((expectation) => - uniqueNonEmptyStrings( - expectation.source_nodes.map((sourceNode) => validateNodeReference(sourceNode, "expected finding source node ID")) - ) - ); - const expected = candidates[0] ?? []; - if (candidates.some((candidate) => !sameStringSet(candidate, expected))) { - throw new FindingsValidationError(`finding ${identityLabel} matches conflicting dependency provenance records`); - } - return expected; -} - -function normalizeUpstreamFindingSource(entry: UpstreamFindingSource, index: number) { - if (!isPlainRecord(entry.finding)) { - throw new FindingsValidationError(`dependency finding ${index} must be an object`); - } - const nodeId = validateNodeReference(entry.node_id, "dependency finding node ID"); - const findingId = nonEmptyString(entry.finding.id); - if (findingId === undefined) throw new FindingsValidationError(`finding ${index} missing required field id`); - const identity = findingIdentity(entry.finding); - const keys = findingIdentityKeys(entry.finding); - if (keys.length === 0) throw new FindingsValidationError(`dependency finding ${index} has no stable finding key`); - const sourceNodes = sourceNodesFromFinding(entry.finding, index); - const identities = [nodeId]; - if (typeof entry.finding.producer_node_id === "string") { - appendUnique(identities, [validateNodeReference(entry.finding.producer_node_id, "dependency producer node ID")]); - } - appendUnique(identities, sourceNodes); - return { index, nodeId, findingId, identities, keys, identity, sourceNodes }; -} - -interface FindingIdentity { - dedupeKeys: string[]; - familyIds: string[]; - findingIds: string[]; - referenceIds: string[]; - ownIds: string[]; -} - -function findingIdentity(value: unknown): FindingIdentity { - if (!isPlainRecord(value)) { - return { dedupeKeys: [], familyIds: [], findingIds: [], referenceIds: [], ownIds: [] }; - } - const lifecycle = isPlainRecord(value.lifecycle) ? value.lifecycle : undefined; - const referenceIds = uniqueNonEmptyStrings([value.upstream_id, value.source_finding_id, value.finding_id]); - const ownIds = uniqueNonEmptyStrings([value.id]); - return { - dedupeKeys: uniqueNonEmptyStrings([value.dedupe_key, lifecycle?.dedupe_key]), - familyIds: uniqueNonEmptyStrings([value.family_id]), - findingIds: uniqueNonEmptyStrings([...referenceIds, ...ownIds]), - referenceIds, - ownIds - }; -} - -function sourceNodesFromFinding(finding: Record, index: number): string[] { - const raw = Array.isArray(finding.source_nodes) - ? finding.source_nodes - : typeof finding.source_node_id === "string" - ? [finding.source_node_id] - : []; - if (raw.length === 0) { - throw new FindingsValidationError(`dependency finding ${index} has no discovery source nodes`); - } - const result: string[] = []; - for (const sourceNode of raw) { - const normalized = nonEmptyString(sourceNode); - if (normalized === undefined) { - throw new FindingsValidationError(`dependency finding ${index} has an invalid discovery source node`); - } - appendUnique(result, [validateNodeReference(normalized, "dependency finding source node ID")]); - } - return result; -} - -function requiredLifecycleString( - record: Record, - key: string, - recordIndex: number, - sourceIndex: number -): string { - const value = nonEmptyString(record[key]); - if (value === undefined) { - throw new FindingsValidationError( - `finding lifecycle record ${recordIndex} source_artifact ${sourceIndex} requires ${key}` - ); - } - return value; -} - -function nonEmptyString(value: unknown): string | undefined { - return typeof value === "string" && value.trim() !== "" ? value.trim() : undefined; -} - -function uniqueNonEmptyStrings(values: readonly unknown[]): string[] { - const result: string[] = []; - for (const value of values) { - const normalized = nonEmptyString(value); - if (normalized !== undefined) appendUnique(result, [normalized]); - } - return result; -} - -function appendUnique(target: string[], values: readonly string[]): void { - for (const value of values) if (!target.includes(value)) target.push(value); -} - -function sameStringSet(left: readonly string[], right: readonly string[]): boolean { - return left.length === right.length && left.every((value) => right.includes(value)); -} - -function assignIfMissing(target: Record, key: string, value: unknown): void { - if (target[key] === undefined && value !== undefined) target[key] = value; -} - -function isPlainRecord(value: unknown): value is Record { - return typeof value === "object" && value !== null && !Array.isArray(value); -} diff --git a/packages/artifacts/src/findings-schema.ts b/packages/artifacts/src/findings-schema.ts index bbae66f98..e1c5522ba 100644 --- a/packages/artifacts/src/findings-schema.ts +++ b/packages/artifacts/src/findings-schema.ts @@ -25,7 +25,7 @@ import { import { validateRegisteredJsonSchema } from "./json-schema-validator.js"; import { hasAtMostCodePoints } from "./portable-json-primitives.js"; import { NODE_REFERENCE_PATTERN } from "./safe-paths.js"; -import { schemaErrorMessage, validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; +import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; import { jsonPointerPath } from "./lang-primitives.js"; export const FINDING_JSON_SCHEMA_ID = "urn:ultrafuzz:schema:artifacts:finding:2" as const; @@ -1040,19 +1040,3 @@ function validateRegisteredFindingSchema( } return { ok: true, issues: [], value: value as T }; } - -export function assertFindingSchema(value: unknown): NormalizedFinding { - const result = validateFindingSchema(value); - if (!result.ok || !result.value) { - throw new Error(schemaErrorMessage("finding", result.issues)); - } - return result.value; -} - -export function assertFindingsSchema(value: unknown): NormalizedFinding[] { - const result = validateFindingsSchema(value); - if (!result.ok || !result.value) { - throw new Error(schemaErrorMessage("findings", result.issues)); - } - return result.value; -} diff --git a/packages/artifacts/src/findings.ts b/packages/artifacts/src/findings.ts index 5dc82579b..ca4e36a71 100644 --- a/packages/artifacts/src/findings.ts +++ b/packages/artifacts/src/findings.ts @@ -1,5 +1,4 @@ export const FINDINGS_SCHEMA_VERSION = "ultrafuzz.finding.v2" as const; -export const FINDINGS_SCHEMA_VERSIONS = [FINDINGS_SCHEMA_VERSION] as const; export const FINDINGS_FILE = "findings.json"; export const FINDING_STATUSES = [ @@ -24,12 +23,3 @@ export const TRIAGE_CLASSIFICATIONS = [ "spec-gated", "defensive-hardening" ] as const; - -export type FindingStatus = (typeof FINDING_STATUSES)[number]; -export type FindingSeverity = (typeof FINDING_SEVERITIES)[number]; -export type FindingConfidence = (typeof FINDING_CONFIDENCE_LEVELS)[number]; -export type TriageClassification = (typeof TRIAGE_CLASSIFICATIONS)[number]; - -export function isSupportedFindingsSchemaVersion(value: string): boolean { - return value === FINDINGS_SCHEMA_VERSION; -} diff --git a/packages/artifacts/src/generated-tests.ts b/packages/artifacts/src/generated-tests.ts index 70aba365d..124e1cdfa 100644 --- a/packages/artifacts/src/generated-tests.ts +++ b/packages/artifacts/src/generated-tests.ts @@ -1,39 +1,16 @@ -import fs from "node:fs"; -import path from "node:path"; - import { z } from "zod/v4"; +import { MAX_GENERATED_TEST_BUNDLE_BYTES, MAX_GENERATED_TEST_BUNDLE_ENTRIES } from "./artifact-limits.js"; import { - MAX_GENERATED_TEST_BUNDLE_BYTES, - MAX_GENERATED_TEST_BUNDLE_ENTRIES, - MAX_GENERATED_TEST_COMPANION_BYTES -} from "./artifact-limits.js"; -import { - GENERATED_TESTS_DIR, generatedTestEntriesSchema, generatedTestFrameworkSchema, generatedTestProvenanceSchema } from "./generated-test-schema.js"; import { validateRegisteredJsonSchema } from "./json-schema-validator.js"; -import { normalizeArtifactProvenance, type ArtifactProvenance } from "./manifests.js"; -import { getNodeArtifactDir, type RunLayout } from "./run-layout.js"; -import { - ArtifactPathError, - assertRegularFileInside, - ensureSafeDirectory, - normalizeSafeRelativePath, - prepareSafeFilePath, - readSinglyLinkedRegularFileSnapshotInside, - safeResolveInside, - sha256Bytes, - writeFileDurable, - writeJsonDurable -} from "./safe-paths.js"; +import { type ArtifactProvenance } from "./manifests.js"; import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; -import { parseStrictJsonBytes } from "./strict-json.js"; export const GENERATED_TESTS_SCHEMA_VERSION = "ultrafuzz.generated-tests.v3" as const; -export const GENERATED_TESTS_MANIFEST = "generated-tests.json"; export const GENERATED_TESTS_JSON_SCHEMA_ID = "urn:ultrafuzz:schema:artifacts:generated-tests:3" as const; export { @@ -58,15 +35,6 @@ export { export type GeneratedTestProvenance = Partial>; -export interface GeneratedTestInput { - /** Exact normalized POSIX manifest path beginning with `generated-tests/`. */ - path: string; - content?: string | Uint8Array; - language?: string; - description?: string; - provenance?: GeneratedTestProvenance; -} - export interface GeneratedTestEntry { path: string; size_bytes: number; @@ -180,281 +148,6 @@ export function assertGeneratedTestBundleResourceBounds(manifest: GeneratedTestM } } -export function writeGeneratedTestManifest(input: { - layout: RunLayout; - nodeId: string; - framework: string; - tests: GeneratedTestInput[]; - supportFiles: GeneratedTestInput[]; - provenance?: GeneratedTestProvenance; -}): GeneratedTestManifest { - const provenance = normalizeArtifactProvenance(input.layout, input.nodeId, input.provenance); - preflightGeneratedTestInputShapes(input, provenance); - const nodeDir = getNodeArtifactDir(input.layout, input.nodeId); - const manifestPath = safeResolveInside(nodeDir, GENERATED_TESTS_MANIFEST, "generated tests manifest path"); - assertGeneratedTestDestinationCanBeReplaced(manifestPath); - preflightGeneratedTestFiles(nodeDir, input.tests, input.supportFiles); - getNodeArtifactDir(input.layout, input.nodeId, { create: true }); - ensureSafeDirectory(nodeDir, GENERATED_TESTS_DIR); - const generated_tests = input.tests.map((test) => writeGeneratedTestEntry(nodeDir, test, provenance)); - const support_files = input.supportFiles.map((supportFile) => - writeGeneratedTestEntry(nodeDir, supportFile, provenance) - ); - const manifest: GeneratedTestManifest = { - schema_version: GENERATED_TESTS_SCHEMA_VERSION, - run_id: input.layout.runId, - node_id: input.nodeId, - framework: input.framework, - generated_tests, - support_files, - provenance - }; - assertGeneratedTestManifestSchema(manifest); - writeJsonDurable(manifestPath, manifest); - return manifest; -} - -export function readGeneratedTestManifest(layout: RunLayout, nodeId: string): GeneratedTestManifest { - const nodeDir = getNodeArtifactDir(layout, nodeId); - const manifestPath = path.join(nodeDir, GENERATED_TESTS_MANIFEST); - return assertGeneratedTestManifestSchema( - parseStrictJsonBytes( - readSinglyLinkedRegularFileSnapshotInside(nodeDir, manifestPath, 64 * 1024 * 1024, "generated tests manifest") - ) - ); -} - -function writeGeneratedTestEntry( - nodeDir: string, - input: GeneratedTestInput, - manifestProvenance: GeneratedTestProvenance & Pick -): GeneratedTestEntry { - const safeRelativeTestPath = canonicalGeneratedTestRelativePath(input.path, "generated test path"); - const generatedTestsRoot = path.join(nodeDir, GENERATED_TESTS_DIR); - const absolutePath = prepareSafeFilePath(generatedTestsRoot, safeRelativeTestPath); - assertGeneratedTestDestinationCanBeReplaced(absolutePath); - if (input.content !== undefined) { - writeFileDurable(absolutePath, input.content); - } - assertRegularFileInside(generatedTestsRoot, absolutePath, "generated test file"); - const contents = readSinglyLinkedRegularFileSnapshotInside( - generatedTestsRoot, - absolutePath, - MAX_GENERATED_TEST_COMPANION_BYTES, - "generated test file" - ); - if (contents.length === 0) { - throw new Error(`generated test file must be non-empty: ${absolutePath}`); - } - assertStrictUtf8(contents, "generated test file", absolutePath); - return generatedTestEntryFromSnapshot(safeRelativeTestPath, input, contents, manifestProvenance); -} - -function generatedTestEntryFromSnapshot( - safeRelativeTestPath: string, - input: GeneratedTestInput, - contents: Uint8Array, - manifestProvenance: GeneratedTestProvenance & Pick -): GeneratedTestEntry { - const entry: GeneratedTestEntry = { - path: `${GENERATED_TESTS_DIR}/${safeRelativeTestPath}`, - size_bytes: contents.length, - sha256: sha256Bytes(contents), - provenance: normalizeArtifactProvenance( - { runId: manifestProvenance.run_id ?? "" }, - manifestProvenance.producer_node_id, - { - ...manifestProvenance, - ...input.provenance - } - ) - }; - if (input.language !== undefined) { - entry.language = input.language; - } - if (input.description !== undefined) { - entry.description = input.description; - } - return entry; -} - -function preflightGeneratedTestInputShapes( - input: { - layout: RunLayout; - nodeId: string; - framework: string; - tests: readonly GeneratedTestInput[]; - supportFiles: readonly GeneratedTestInput[]; - }, - provenance: GeneratedTestProvenance & Pick -): void { - const { tests, supportFiles } = input; - if (tests.length + supportFiles.length > MAX_GENERATED_TEST_BUNDLE_ENTRIES) { - throw new Error( - `generated tests manifest exceeds the ${MAX_GENERATED_TEST_BUNDLE_ENTRIES}-entry combined bundle limit` - ); - } - if (tests.length === 0 && supportFiles.length > 0) { - throw new Error("generated tests manifest cannot declare support files without a runnable generated test"); - } - const paths = new Set(); - for (const [label, entries] of [ - ["generated test", tests], - ["generated-test support", supportFiles] - ] as const) { - for (const entry of entries) { - const safeRelativePath = canonicalGeneratedTestRelativePath(entry.path, `${label} path`); - const manifestPath = `${GENERATED_TESTS_DIR}/${safeRelativePath}`; - if (paths.has(manifestPath)) { - throw new Error(`generated tests manifest repeats path ${JSON.stringify(manifestPath)}`); - } - paths.add(manifestPath); - if (entry.content !== undefined && entry.content.length === 0) { - throw new Error(`${label} file must be non-empty: ${manifestPath}`); - } - if (entry.content !== undefined && Buffer.byteLength(entry.content) > MAX_GENERATED_TEST_COMPANION_BYTES) { - throw new Error(`${label} file exceeds the ${MAX_GENERATED_TEST_COMPANION_BYTES}-byte limit: ${manifestPath}`); - } - if (entry.content !== undefined) { - assertStrictUtf8(Buffer.from(entry.content), `${label} file`, manifestPath); - } - } - } - const suppliedBytes = [...tests, ...supportFiles].reduce( - (total, entry) => total + (entry.content === undefined ? 0 : Buffer.byteLength(entry.content)), - 0 - ); - if (suppliedBytes > MAX_GENERATED_TEST_BUNDLE_BYTES) { - throw new Error( - `generated-test bundle supplied content exceeds the ${MAX_GENERATED_TEST_BUNDLE_BYTES}-byte combined limit` - ); - } - assertGeneratedTestPathsAreMaterializable(paths); - const placeholder = Buffer.from("x", "utf8"); - assertGeneratedTestManifestSchema({ - schema_version: GENERATED_TESTS_SCHEMA_VERSION, - run_id: input.layout.runId, - node_id: input.nodeId, - framework: input.framework, - generated_tests: tests.map((entry) => - generatedTestEntryFromSnapshot( - canonicalGeneratedTestRelativePath(entry.path, "generated test path"), - entry, - placeholder, - provenance - ) - ), - support_files: supportFiles.map((entry) => - generatedTestEntryFromSnapshot( - canonicalGeneratedTestRelativePath(entry.path, "generated-test support path"), - entry, - placeholder, - provenance - ) - ), - provenance - }); -} - -function preflightGeneratedTestFiles( - nodeDir: string, - tests: readonly GeneratedTestInput[], - supportFiles: readonly GeneratedTestInput[] -): void { - const generatedTestsRoot = path.join(nodeDir, GENERATED_TESTS_DIR); - const existingFiles: Array<{ label: string; manifestPath: string; absolutePath: string; sizeBytes: number }> = []; - let totalBytes = 0; - for (const [label, entries] of [ - ["generated test", tests], - ["generated-test support", supportFiles] - ] as const) { - for (const entry of entries) { - const safeRelativePath = canonicalGeneratedTestRelativePath(entry.path, `${label} path`); - const manifestPath = `${GENERATED_TESTS_DIR}/${safeRelativePath}`; - const absolutePath = safeResolveInside(generatedTestsRoot, safeRelativePath, `${label} path`); - assertGeneratedTestDestinationCanBeReplaced(absolutePath); - if (entry.content !== undefined) { - totalBytes += Buffer.byteLength(entry.content); - continue; - } - assertRegularFileInside(generatedTestsRoot, absolutePath, `${label} file`); - const sizeBytes = fs.lstatSync(absolutePath).size; - totalBytes += sizeBytes; - existingFiles.push({ label, manifestPath, absolutePath, sizeBytes }); - } - } - if (totalBytes > MAX_GENERATED_TEST_BUNDLE_BYTES) { - throw new Error( - `generated-test bundle exceeds the ${MAX_GENERATED_TEST_BUNDLE_BYTES}-byte combined companion limit` - ); - } - for (const { label, manifestPath, sizeBytes } of existingFiles) { - if (sizeBytes === 0) { - throw new Error(`${label} file must be non-empty: ${manifestPath}`); - } - if (sizeBytes > MAX_GENERATED_TEST_COMPANION_BYTES) { - throw new Error(`${label} file exceeds the ${MAX_GENERATED_TEST_COMPANION_BYTES}-byte limit: ${manifestPath}`); - } - } - for (const { label, manifestPath, absolutePath } of existingFiles) { - const contents = readSinglyLinkedRegularFileSnapshotInside( - generatedTestsRoot, - absolutePath, - MAX_GENERATED_TEST_COMPANION_BYTES, - `${label} file` - ); - if (contents.length === 0) { - throw new Error(`${label} file must be non-empty: ${manifestPath}`); - } - assertStrictUtf8(contents, `${label} file`, manifestPath); - } -} - -function assertStrictUtf8(contents: Uint8Array, label: string, filePath: string): void { - try { - new TextDecoder("utf-8", { fatal: true }).decode(contents); - } catch { - throw new Error(`${label} must be strict UTF-8 text: ${filePath}`); - } -} - -function assertGeneratedTestDestinationCanBeReplaced(filePath: string): void { - let fileStats: fs.Stats; - try { - fileStats = fs.lstatSync(filePath); - } catch (error) { - if ((error as NodeJS.ErrnoException).code === "ENOENT") { - return; - } - throw error; - } - if (fileStats.isSymbolicLink()) { - throw new ArtifactPathError("symlink-escape", `generated-test bundle destination cannot be a symlink: ${filePath}`); - } - if (!fileStats.isFile()) { - throw new ArtifactPathError("not-file", `generated-test bundle destination must be a regular file: ${filePath}`); - } - if (fileStats.nlink !== 1) { - throw new ArtifactPathError("hard-link", `generated-test bundle destination must be singly linked: ${filePath}`); - } -} - -function canonicalGeneratedTestRelativePath(value: string, label: string): string { - const prefix = `${GENERATED_TESTS_DIR}/`; - if (!value.startsWith(prefix)) { - throw new ArtifactPathError("noncanonical-path", `${label} must begin with ${JSON.stringify(prefix)}`); - } - const relativePath = value.slice(prefix.length); - const normalized = normalizeSafeRelativePath(relativePath, label); - if (relativePath !== normalized) { - throw new ArtifactPathError( - "noncanonical-path", - `${label} must already be a normalized relative POSIX path: ${JSON.stringify(value)}` - ); - } - return normalized; -} - function assertGeneratedTestPathsAreMaterializable(paths: ReadonlySet): void { const directoryOrderedPaths = [...paths].sort((left, right) => { const leftDirectory = `${left}/`; diff --git a/packages/artifacts/src/goal-plan.ts b/packages/artifacts/src/goal-plan.ts index b45e679f2..9f1b0ee3d 100644 --- a/packages/artifacts/src/goal-plan.ts +++ b/packages/artifacts/src/goal-plan.ts @@ -332,7 +332,6 @@ const applicabilityDecisionSchema = z }); export const GOAL_LANE_KINDS = ["threat", "class", "roaming"] as const; -export type GoalLaneKind = (typeof GOAL_LANE_KINDS)[number]; /** * One goal lane: a named unit of hunting work and the concrete node IDs it owns. @@ -598,9 +597,6 @@ export const goalPlanJsonSchema = { } as Record; export type GoalPlan = z.infer; -export type ThreatGoalPlanItem = z.infer; -export type ClassGoalPlanItem = z.infer; -export type ApplicabilityDecision = z.infer; export function validateGoalPlan(value: unknown, path = "$"): SchemaValidationResult { return validateWithZod(goalPlanSchema, value, { path, code: "GOAL_PLAN_SCHEMA_INVALID" }); diff --git a/packages/artifacts/src/index.ts b/packages/artifacts/src/index.ts index ed7e2fc5b..7836d9e2e 100644 --- a/packages/artifacts/src/index.ts +++ b/packages/artifacts/src/index.ts @@ -10,7 +10,6 @@ export * from "./artifact-contracts.js"; export * from "./artifact-schema-metadata.js"; export * from "./cloud-selected-task.js"; export * from "./findings.js"; -export * from "./finding-provenance.js"; export * from "./findings-schema.js"; export * from "./finding-note-vocabulary.js"; export * from "./coverage-evidence.js"; diff --git a/packages/artifacts/src/invariant-ledger.ts b/packages/artifacts/src/invariant-ledger.ts index 24b98698a..3ee30fb48 100644 --- a/packages/artifacts/src/invariant-ledger.ts +++ b/packages/artifacts/src/invariant-ledger.ts @@ -1,6 +1,6 @@ import { z } from "zod/v4"; -import { validateWithZod, type SchemaValidationIssue, type SchemaValidationResult } from "./schema-validation.js"; +import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; import { executeSemanticGates } from "./semantic-gates.js"; export const INVARIANT_LEDGER_SCHEMA_VERSION = "ultrafuzz.invariant-evidence-ledger.v1" as const; @@ -142,7 +142,6 @@ export const invariantLedgerSchema = z }); export type InvariantLedgerEntry = z.infer; -export type InvariantInventoryRow = z.infer; export type InvariantLedgerArtifact = z.infer; export function validateInvariantLedgerSchema( @@ -203,10 +202,6 @@ function publicInvariantLedgerSemanticPath(semanticPath: string, message: string : semanticPath; } -export function invariantLedgerSchemaIssues(value: unknown, path = "$"): SchemaValidationIssue[] { - return validateInvariantLedgerSchema(value, path).issues; -} - export const invariantLedgerJsonSchema = { $schema: "https://json-schema.org/draft/2020-12/schema", $id: "urn:ultrafuzz:schema:artifacts:invariant-evidence-ledger:1", diff --git a/packages/artifacts/src/invariant-source-proof.ts b/packages/artifacts/src/invariant-source-proof.ts index 06607eaca..e089b7038 100644 --- a/packages/artifacts/src/invariant-source-proof.ts +++ b/packages/artifacts/src/invariant-source-proof.ts @@ -1,6 +1,6 @@ import { z } from "zod/v4"; -import { validateWithZod, type SchemaValidationIssue, type SchemaValidationResult } from "./schema-validation.js"; +import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; import { executeSemanticGate } from "./semantic-gates.js"; export const INVARIANT_SOURCE_PROOF_SCHEMA_VERSION = "ultrafuzz.invariant-source-proof.v1" as const; @@ -41,7 +41,6 @@ export const invariantSourceProofSchema = z.strictObject({ }); export type InvariantSourceProof = z.infer; -export type InvariantSourceProofFile = z.infer; export function validateInvariantSourceProofSchema( value: unknown, @@ -68,10 +67,6 @@ function prefixedSemanticPath(rootPath: string, semanticPath: string): string { return semanticPath === "$" ? rootPath : `${rootPath}${semanticPath.slice(1)}`; } -export function invariantSourceProofSchemaIssues(value: unknown, path = "$"): SchemaValidationIssue[] { - return validateInvariantSourceProofSchema(value, path).issues; -} - export const invariantSourceProofJsonSchema = { $schema: "https://json-schema.org/draft/2020-12/schema", $id: "urn:ultrafuzz:schema:artifacts:invariant-source-proof:1", diff --git a/packages/artifacts/src/manifests.ts b/packages/artifacts/src/manifests.ts index 022b6a0ff..7a50d1965 100644 --- a/packages/artifacts/src/manifests.ts +++ b/packages/artifacts/src/manifests.ts @@ -8,7 +8,6 @@ import { listSafeFiles, normalizeSafeRelativePath, prepareSafeFilePath, - safeResolveInside, sha256File, validateNodeReference, validateSafeId, @@ -443,49 +442,6 @@ export function readArtifactManifest(layout: RunLayout, nodeId: string): Artifac return manifest as ArtifactManifest; } -export interface RunArtifactIndexEntry extends ArtifactManifestEntry { - node_id: string; -} - -export interface RunArtifactIndex { - schema_version: string; - run_id: string; - artifacts: RunArtifactIndexEntry[]; -} - -export function buildRunArtifactIndex(layout: RunLayout): RunArtifactIndex { - const artifacts: RunArtifactIndexEntry[] = []; - if (!fs.existsSync(layout.artifactsDir)) { - return { schema_version: ARTIFACT_MANIFEST_SCHEMA_VERSION, run_id: layout.runId, artifacts }; - } - - for (const dirent of fs.readdirSync(layout.artifactsDir, { withFileTypes: true })) { - if (!dirent.isDirectory()) { - continue; - } - const nodeId = validateSafeId(dirent.name, "node ID"); - const nodeDir = path.join(layout.artifactsDir, nodeId); - const manifestPath = path.join(nodeDir, ARTIFACT_MANIFEST_FILE); - if (!fs.existsSync(manifestPath)) { - continue; - } - assertRegularFileInside(layout.artifactsDir, manifestPath, "artifact manifest path"); - const manifest = readArtifactManifest(layout, nodeId); - for (const file of manifest.files) { - const artifactPath = safeResolveInside(nodeDir, file.path, "artifact manifest file path"); - assertRegularFileInside(nodeDir, artifactPath, "artifact manifest file path"); - artifacts.push({ - ...file, - node_id: nodeId, - path: `artifacts/${nodeId}/${file.path}` - }); - } - } - - artifacts.sort((left, right) => left.node_id.localeCompare(right.node_id) || left.path.localeCompare(right.path)); - return { schema_version: ARTIFACT_MANIFEST_SCHEMA_VERSION, run_id: layout.runId, artifacts }; -} - function assertValidArtifactManifest(value: unknown): void { const validation = validateArtifactManifest(value); if (validation.ok) return; diff --git a/packages/artifacts/src/property-provenance.ts b/packages/artifacts/src/property-provenance.ts index 25aaf7954..90e9607a6 100644 --- a/packages/artifacts/src/property-provenance.ts +++ b/packages/artifacts/src/property-provenance.ts @@ -3,12 +3,7 @@ import { z } from "zod/v4"; import { canonicalArtifactRelativePathSchema } from "./artifact-path-primitives.js"; import { validateRegisteredJsonSchema } from "./json-schema-validator.js"; import { canonicalTimestampSchema, hasAtMostCodePoints } from "./portable-json-primitives.js"; -import { - schemaErrorMessage, - validateWithZod, - type SchemaValidationIssue, - type SchemaValidationResult -} from "./schema-validation.js"; +import { validateWithZod, type SchemaValidationIssue, type SchemaValidationResult } from "./schema-validation.js"; import { jsonPointerPath } from "./lang-primitives.js"; export const PROPERTIES_SCHEMA_VERSION = "ultrafuzz.properties.v2" as const; @@ -1156,14 +1151,6 @@ export function validatePropertyCampaignSchema( ); } -export function assertPropertiesSchema(value: unknown): PropertiesArtifact { - const result = validatePropertiesSchema(value); - if (!result.ok || result.value === undefined) { - throw new Error(schemaErrorMessage("properties", result.issues)); - } - return result.value; -} - export function validatePropertyReferences( catalog: PropertiesArtifact, references: readonly PropertyReferenceInput[] diff --git a/packages/artifacts/src/report-observation.ts b/packages/artifacts/src/report-observation.ts index ae2a11413..e086ed327 100644 --- a/packages/artifacts/src/report-observation.ts +++ b/packages/artifacts/src/report-observation.ts @@ -48,6 +48,5 @@ export const reportObservedCompletionSchema = z.strictObject({ }); export type ReportVerification = z.infer; -export type ReportVerificationReasonCode = ReportVerification["reason_codes"][number]; export type ReportObservedCompletion = z.infer; export type ObservedReportCompletion = ReportObservedCompletion; diff --git a/packages/artifacts/src/runtime-schemas.ts b/packages/artifacts/src/runtime-schemas.ts index 93bbd4972..8892d3c0d 100644 --- a/packages/artifacts/src/runtime-schemas.ts +++ b/packages/artifacts/src/runtime-schemas.ts @@ -192,29 +192,6 @@ export const agentSourceProofJsonSchema = { } } as const; -export interface AgentSourceProof { - schema_version: typeof AGENT_SOURCE_PROOF_SCHEMA_VERSION; - attempt_id: string; - commit: string; - tree: string; - base_ref: "refs/heads/ultrafuzz-pinned"; - refs: Array<{ name: string; object: string }>; - remotes: []; - revision_count: 1; - commit_object_count: 1; - dependencies: { - schema_version: "ultrafuzz.pinned-submodules-expectation.v1"; - source_commit: string; - source_tree: string; - manifest_sha256: string; - top_level_roots: string[]; - recursive_gitlinks: Array<{ path: string; commit: string; tree: string }>; - entry_count: number; - file_count: number; - total_file_bytes: number; - } | null; -} - export interface ArtifactVerificationEntry { path: string; contract: ArtifactContractId; diff --git a/packages/artifacts/src/safe-paths.ts b/packages/artifacts/src/safe-paths.ts index 395fe012e..8698f7d27 100644 --- a/packages/artifacts/src/safe-paths.ts +++ b/packages/artifacts/src/safe-paths.ts @@ -368,40 +368,6 @@ export function writeJsonDurable(filePath: string, value: unknown): void { writeFileDurable(filePath, `${JSON.stringify(value, null, 2)}\n`); } -export function appendLineDurable(filePath: string, line: string, trustedRoot?: string): void { - const directory = path.dirname(filePath); - if (trustedRoot !== undefined) { - assertNoSymlinkComponents(trustedRoot, directory, "append directory"); - } - fs.mkdirSync(directory, { recursive: true }); - if (trustedRoot !== undefined) { - assertNoSymlinkComponents(trustedRoot, filePath, "append path"); - } - const fd = fs.openSync( - filePath, - fs.constants.O_APPEND | fs.constants.O_CREAT | fs.constants.O_WRONLY | fs.constants.O_NOFOLLOW, - 0o600 - ); - runWithClosedDescriptor(fd, `failed to durably append ${filePath} and close its descriptor`, () => { - if (!fs.fstatSync(fd).isFile()) { - throw new ArtifactPathError("not-file", `append path must be a regular file: ${filePath}`); - } - if (trustedRoot !== undefined) { - assertNoSymlinkComponents(trustedRoot, filePath, "append path"); - } - const bytes = Buffer.from(line.endsWith("\n") ? line : `${line}\n`, "utf8"); - const written = fs.writeSync(fd, bytes); - if (written !== bytes.length) { - throw new ArtifactPathError( - "short-write", - `durable append wrote ${written} of ${bytes.length} bytes: ${filePath}` - ); - } - fs.fsyncSync(fd); - }); - fsyncDirectory(directory); -} - /** * Durably creates a new regular file without accepting an intervening writer. * @@ -482,43 +448,6 @@ export function appendBytesDurableAt( fsyncDirectory(path.dirname(filePath)); } -/** - * Durably discards an unterminated trailing fragment. - * - * The expected size fences the repair decision against a concurrent append; - * hard links are rejected so truncation cannot mutate another named file. - */ -export function truncateDurable( - filePath: string, - length: number, - options: { expectedSize: number; trustedRoot?: string } -): void { - if (options.trustedRoot !== undefined) { - assertNoSymlinkComponents(options.trustedRoot, filePath, "truncate path"); - } - const fd = fs.openSync(filePath, fs.constants.O_WRONLY | fs.constants.O_NOFOLLOW | fs.constants.O_NONBLOCK); - try { - const stat = fs.fstatSync(fd); - if (!stat.isFile()) { - throw new ArtifactPathError("not-file", `truncate path must be a regular file: ${filePath}`); - } - if (stat.nlink !== 1) { - throw new ArtifactPathError("not-file", `truncate path must not be hard-linked: ${filePath}`); - } - if (stat.size !== options.expectedSize) { - throw new ArtifactPathError("not-file", `truncate path changed size before truncation: ${filePath}`); - } - if (length > stat.size) { - throw new ArtifactPathError("not-file", `truncate length exceeds the file size: ${filePath}`); - } - fs.ftruncateSync(fd, length); - fs.fsyncSync(fd); - } finally { - fs.closeSync(fd); - } - fsyncDirectory(path.dirname(filePath)); -} - export function readJsonFile(filePath: string): T { return parseStrictJsonBytes(readRegularFileSnapshot(filePath, 64 * 1024 * 1024), { maxBytes: 64 * 1024 * 1024, diff --git a/packages/artifacts/src/state-schema.ts b/packages/artifacts/src/state-schema.ts index bf1a9aaf4..61714ae13 100644 --- a/packages/artifacts/src/state-schema.ts +++ b/packages/artifacts/src/state-schema.ts @@ -23,8 +23,7 @@ import { TERMINAL_NODE_STATE_STATUSES, isTerminalNodeStatus, type NodeState, - type RunState, - type TerminalDispositionDocument + type RunState } from "./state.js"; import { schemaErrorMessage, validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; @@ -825,16 +824,6 @@ export function validateRunStateSchema(value: unknown, path = "$"): SchemaValida }); } -export function validateTerminalDispositionSchema( - value: unknown, - path = "$" -): SchemaValidationResult { - return validateWithZod(terminalDispositionSchema as z.ZodType, value, { - path, - code: "TERMINAL_DISPOSITION_SCHEMA_INVALID" - }); -} - export function validateNodeStateSchema(value: unknown, path = "$"): SchemaValidationResult { return validateWithZod(nodeStateSchema as z.ZodType, value, { path, @@ -849,11 +838,3 @@ export function assertRunStateSchema(value: unknown): RunState { } return result.value; } - -export function assertNodeStateSchema(value: unknown): NodeState { - const result = validateNodeStateSchema(value); - if (!result.ok || !result.value) { - throw new Error(schemaErrorMessage("node state", result.issues)); - } - return result.value; -} diff --git a/packages/artifacts/src/state.ts b/packages/artifacts/src/state.ts index f79aa0e2d..b370d88b5 100644 --- a/packages/artifacts/src/state.ts +++ b/packages/artifacts/src/state.ts @@ -1,5 +1,3 @@ -import fs from "node:fs"; - import { redactSecretsInText, SENSITIVE_REDACTION_PLACEHOLDER } from "@ultrafuzz/security"; import type { ArtifactContractId } from "./artifact-contract-ids.js"; @@ -557,19 +555,6 @@ function assertCurrentRunState(value: unknown): asserts value is RunState { } } -export function loadOrCreateRunState( - target: RunLayoutStateLike | string, - state: RunState, - options: RunStateWriteOptions = {} -): RunState { - const statePath = resolveStatePath(target); - if (fs.existsSync(statePath)) { - return readRunState(statePath); - } - writeRunState(statePath, state, options); - return state; -} - export function updateRunStatus( target: RunLayoutStateLike | string, status: RunStatus, diff --git a/packages/artifacts/src/threat-model.ts b/packages/artifacts/src/threat-model.ts index ba2d4800d..b1a37033b 100644 --- a/packages/artifacts/src/threat-model.ts +++ b/packages/artifacts/src/threat-model.ts @@ -272,7 +272,6 @@ export const threatModelJsonSchema = { export type ThreatModel = z.infer; export type ThreatModelEvidenceReference = z.infer; -export type CapabilityStatus = (typeof CAPABILITY_STATUSES)[number]; export function validateThreatModel(value: unknown, path = "$"): SchemaValidationResult { return validateWithZod(threatModelSchema, value, { path, code: "THREAT_MODEL_SCHEMA_INVALID" }); diff --git a/packages/artifacts/src/usage-ledger.ts b/packages/artifacts/src/usage-ledger.ts index 034677421..fa59c125d 100644 --- a/packages/artifacts/src/usage-ledger.ts +++ b/packages/artifacts/src/usage-ledger.ts @@ -24,7 +24,6 @@ export const USAGE_FIELDS = [ "cache_write_tokens", "reasoning_tokens" ] as const; -export type UsageField = (typeof USAGE_FIELDS)[number]; export interface NormalizedUsage { model: string; diff --git a/packages/artifacts/src/workflow-contracts.ts b/packages/artifacts/src/workflow-contracts.ts index e9e2a93ee..b103ba3c6 100644 --- a/packages/artifacts/src/workflow-contracts.ts +++ b/packages/artifacts/src/workflow-contracts.ts @@ -34,7 +34,6 @@ import { import { canonicalTimestampSchema } from "./portable-json-primitives.js"; import { PROPERTY_PRIORITIES } from "./property-provenance.js"; import { SAFE_ID_PATTERN } from "./safe-paths.js"; -import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; import { canonicalJsonValueKey } from "./lang-primitives.js"; const nonEmptyString = z.string().min(1); @@ -75,10 +74,6 @@ function withDocumentMetadata(schema: T, slug: string, vers }) as T; } -export const HARNESS_REPAIRS_SCHEMA_VERSION = "ultrafuzz.harness-repairs.v1" as const; -export const STRATEGY_DETECTIONS_SCHEMA_VERSION = "ultrafuzz.strategy-detections.v1" as const; -export const TRIAGED_FINDINGS_SCHEMA_VERSION = "ultrafuzz.triaged-findings.v1" as const; -export const SEVERITY_CLASSIFIED_FINDINGS_SCHEMA_VERSION = "ultrafuzz.severity-classified-findings.v1" as const; export const REFERENCE_MANIFEST_SCHEMA_VERSION = "ultrafuzz.reference-manifest.v1" as const; export const BOUNDARY_RECIPES_SCHEMA_VERSION = "ultrafuzz.boundary-recipes.v1" as const; export const ADMIN_CONFIG_BOUNDARY_MATRIX_SCHEMA_VERSION = "ultrafuzz.admin-config-boundary-matrix.v1" as const; @@ -1814,44 +1809,8 @@ export const findingLifecycleLedgerJsonSchema = workflowContractJsonSchemas["ult export const aggregationManifestJsonSchema = workflowContractJsonSchemas["ultrafuzz/aggregation-manifest@1"]; export const reportJsonSchema = workflowContractJsonSchemas["ultrafuzz/report@3"]; -export function validateWorkflowContract( - contract: WorkflowContractId, - value: unknown, - path = "$" -): SchemaValidationResult { - return validateWithZod(workflowContractSchemas[contract] as z.ZodType, value, { - path, - code: "WORKFLOW_ARTIFACT_SCHEMA_INVALID" - }); -} - -export type HarnessRepairs = z.infer; -export type StrategyDetections = z.infer; export type TriagedFindings = z.infer; export type SeverityClassifiedFindings = z.infer; -export type ReferenceManifest = z.infer; -export type BoundaryRecipes = z.infer; -export type AdminConfigBoundaryMatrix = z.infer; -export type DependencyScopeMatrix = z.infer; -export type ExternalizedStateAccounting = z.infer; -export type CoverageGoal = z.infer; -export type InvariantCampaignPlan = z.infer; -export type CampaignSummary = z.infer; -export type DifferentialPlan = z.infer; -export type ReferenceHarness = z.infer; -export type AuditedDifferentialLanes = z.infer; -export type DifferentialLaneResult = z.infer; -export type SemanticRedRegistry = z.infer; -export type DifferentialRedTriage = z.infer; -export type DifferentialRepairSummary = z.infer; -export type DifferentialGapReview = z.infer; -export type DifferentialReportReview = z.infer; -export type DynamicStrategyPlan = z.infer; -export type DynamicEnumeratorOutputs = z.infer; -export type SelectedStrategies = z.infer; -export type DynamicStrategyProvenance = z.infer; -export type FindingLifecycleLedger = z.infer; -export type AggregationManifest = z.infer; export type TerminalReport = z.infer; export const WORKFLOW_SCHEMA_FILES = { diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 390b021cc..8b287424e 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -6,17 +6,12 @@ import path from "node:path"; import test from "node:test"; import { - ArtifactPathError, ArtifactSecretGateError, ARTIFACT_MANIFEST_FILE, - GENERATED_TESTS_SCHEMA_VERSION, - MAX_GENERATED_TEST_BUNDLE_BYTES, - MAX_GENERATED_TEST_BUNDLE_ENTRIES, appendUsageEvents, assertArtifactPublicationsContainNoSecrets, appendNodeAttempt, appendEvent, - appendLineDurable, assertUsageLedgerEntry, createRunLayout, getNodeArtifactDir, @@ -29,7 +24,6 @@ import { queryEvents, readArtifactManifest, readEventQueryFacade, - readGeneratedTestManifest, readRunState, replayEvents, replayUsageEvents, @@ -42,7 +36,6 @@ import { verifyArtifactManifestPrerequisites, writeArtifact, writeArtifactManifest, - writeGeneratedTestManifest as writeGeneratedTestManifestWithFramework, writeRunState } from "../src/index.js"; @@ -50,18 +43,6 @@ function tempProject(): string { return fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-artifacts-")); } -function writeGeneratedTestManifest( - input: Omit[0], "framework"> -): ReturnType { - return writeGeneratedTestManifestWithFramework({ ...input, framework: "foundry" }); -} - -function errnoError(code: string): NodeJS.ErrnoException { - const error = new Error(`injected ${code}`) as NodeJS.ErrnoException; - error.code = code; - return error; -} - test("createRunLayout persists product-owned run evidence outside checkpoints", () => { const project = tempProject(); const layout = createRunLayout({ @@ -146,7 +127,7 @@ test("validated usage events append idempotently with exact Smithers identities" assert.equal(ledger.entries[0]?.control_generation, "a".repeat(64)); assert.equal(ledger.entries.length, 1); - appendLineDurable(layout.usageLedgerPath, "{malformed", layout.root); + fs.appendFileSync(layout.usageLedgerPath, "{malformed\n"); assert.throws(() => replayUsageEvents(layout), /invalid strict JSON/u); }); @@ -191,11 +172,7 @@ test("usage ledger replay rejects entries copied from another run", () => { }; appendUsageEvents(firstLayout, [input]); appendUsageEvents(secondLayout, [input]); - appendLineDurable( - secondLayout.usageLedgerPath, - fs.readFileSync(firstLayout.usageLedgerPath, "utf8"), - secondLayout.root - ); + fs.appendFileSync(secondLayout.usageLedgerPath, fs.readFileSync(firstLayout.usageLedgerPath)); assert.throws(() => replayUsageEvents(secondLayout), /run_id belongs to/u); }); @@ -1181,649 +1158,3 @@ test("updateNodeState accepts an explicit transition timestamp", () => { assert.equal(running.last_transition_at, "2026-01-01T00:00:05.000Z"); assert.equal(running.nodes["node-a"]?.wait_since, "2026-01-01T00:00:05.000Z"); }); - -test("durable append rejects a symlinked parent before creating outside directories", () => { - const root = tempProject(); - const outside = tempProject(); - fs.symlinkSync(outside, path.join(root, "linked"), "dir"); - - assert.throws( - () => appendLineDurable(path.join(root, "linked", "created", "audit.jsonl"), "entry", root), - /crosses symlink/u - ); - assert.equal(fs.existsSync(path.join(outside, "created")), false); -}); - -test("durable append rejects a half write without retrying the record", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const expected = Buffer.from("0123456789\n", "utf8"); - const partialLength = Math.floor(expected.length / 2); - const realWriteSync = fs.writeSync; - const realCloseSync = fs.closeSync; - const writeSync = t.mock.method(fs, "writeSync", (fd: number, bytes: Uint8Array) => - realWriteSync(fd, Buffer.from(bytes).subarray(0, partialLength)) - ); - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => realCloseSync(fd)); - - assert.throws( - () => appendLineDurable(filePath, "0123456789"), - (error: unknown) => { - assert.ok(error instanceof ArtifactPathError); - assert.equal(error.code, "short-write"); - assert.match(error.message, new RegExp(`wrote ${partialLength} of ${expected.length} bytes`, "u")); - return true; - } - ); - - assert.equal(writeSync.mock.callCount(), 1); - assert.equal(closeSync.mock.callCount(), 1); - assert.deepEqual(fs.readFileSync(filePath), expected.subarray(0, partialLength)); -}); - -test("durable append rejects a zero write without retrying the record", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const realCloseSync = fs.closeSync; - const writeSync = t.mock.method(fs, "writeSync", () => 0); - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => realCloseSync(fd)); - - assert.throws( - () => appendLineDurable(filePath, "entry"), - (error: unknown) => { - assert.ok(error instanceof ArtifactPathError); - assert.equal(error.code, "short-write"); - assert.match(error.message, /wrote 0 of 6 bytes/u); - return true; - } - ); - - assert.equal(writeSync.mock.callCount(), 1); - assert.equal(closeSync.mock.callCount(), 1); - assert.equal(fs.readFileSync(filePath, "utf8"), ""); -}); - -test("durable append propagates a directory fsync I/O failure", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const realFsyncSync = fs.fsyncSync; - const realCloseSync = fs.closeSync; - let fsyncCall = 0; - const fsyncSync = t.mock.method(fs, "fsyncSync", (fd: number) => { - fsyncCall += 1; - if (fsyncCall === 1) { - realFsyncSync(fd); - return; - } - throw errnoError("EIO"); - }); - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => realCloseSync(fd)); - - assert.throws( - () => appendLineDurable(filePath, "entry"), - (error: unknown) => { - assert.ok(error instanceof Error && "code" in error); - assert.equal(error.code, "EIO"); - return true; - } - ); - - assert.equal(fsyncSync.mock.callCount(), 2); - assert.equal(closeSync.mock.callCount(), 2); - assert.equal(fs.readFileSync(filePath, "utf8"), "entry\n"); -}); - -test("durable append tolerates only an unsupported directory fsync operation", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const realFsyncSync = fs.fsyncSync; - const realCloseSync = fs.closeSync; - let fsyncCall = 0; - const fsyncSync = t.mock.method(fs, "fsyncSync", (fd: number) => { - fsyncCall += 1; - if (fsyncCall === 1) { - realFsyncSync(fd); - return; - } - throw errnoError("EINVAL"); - }); - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => realCloseSync(fd)); - - appendLineDurable(filePath, "entry"); - - assert.equal(fsyncSync.mock.callCount(), 2); - assert.equal(closeSync.mock.callCount(), 2); - assert.equal(fs.readFileSync(filePath, "utf8"), "entry\n"); -}); - -test("durable append propagates a close failure without retrying close", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const realCloseSync = fs.closeSync; - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => { - realCloseSync(fd); - throw errnoError("EIO"); - }); - - assert.throws( - () => appendLineDurable(filePath, "entry"), - (error: unknown) => { - assert.ok(error instanceof Error && "code" in error); - assert.equal(error.code, "EIO"); - return true; - } - ); - - assert.equal(closeSync.mock.callCount(), 1); - assert.equal(fs.readFileSync(filePath, "utf8"), "entry\n"); -}); - -test("durable append aggregates an operation failure with its close failure", (t) => { - const root = tempProject(); - const filePath = path.join(root, "audit.jsonl"); - const realCloseSync = fs.closeSync; - const writeSync = t.mock.method(fs, "writeSync", () => { - throw errnoError("EIO"); - }); - const closeSync = t.mock.method(fs, "closeSync", (fd: number) => { - realCloseSync(fd); - throw errnoError("EBADF"); - }); - - assert.throws( - () => appendLineDurable(filePath, "entry"), - (error: unknown) => { - assert.ok(error instanceof AggregateError); - assert.deepEqual( - error.errors.map((entry: unknown) => (entry instanceof Error && "code" in entry ? entry.code : undefined)), - ["EIO", "EBADF"] - ); - return true; - } - ); - - assert.equal(writeSync.mock.callCount(), 1); - assert.equal(closeSync.mock.callCount(), 1); - assert.equal(fs.readFileSync(filePath, "utf8"), ""); -}); - -test("generated-test manifests persist explicit generated files with provenance", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-1" }); - const generatedContents = "contract InvariantTest {} // π\n"; - const supportContents = "library InvariantFixture {} // café\n"; - const manifest = writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - provenance: { agent_ref: "CodexAgent", workflow_task_id: "node:strategy-a", attempt_index: 0 }, - tests: [ - { - path: "generated-tests/Invariant.t.sol", - content: generatedContents, - language: "solidity" - } - ], - supportFiles: [ - { - path: "generated-tests/helpers/InvariantFixture.sol", - content: supportContents, - language: "solidity" - } - ] - }); - - assert.equal(manifest.schema_version, GENERATED_TESTS_SCHEMA_VERSION); - assert.equal(manifest.framework, "foundry"); - assert.equal(manifest.generated_tests.length, 1); - assert.equal(manifest.generated_tests[0]!.path, "generated-tests/Invariant.t.sol"); - assert.equal(manifest.generated_tests[0]!.size_bytes, Buffer.byteLength(generatedContents)); - assert.equal( - manifest.generated_tests[0]!.sha256, - crypto.createHash("sha256").update(generatedContents).digest("hex") - ); - assert.equal(manifest.generated_tests[0]!.provenance!.agent_ref, "CodexAgent"); - assert.equal(manifest.support_files[0]!.path, "generated-tests/helpers/InvariantFixture.sol"); - assert.equal(manifest.support_files[0]!.size_bytes, Buffer.byteLength(supportContents)); - assert.equal(manifest.support_files[0]!.sha256, crypto.createHash("sha256").update(supportContents).digest("hex")); - assert.equal( - fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests", "Invariant.t.sol")), - true - ); -}); - -test("generated-test manifest writer rejects zero-byte companion files", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-empty-generated-test" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - supportFiles: [], - tests: [{ path: "generated-tests/Empty.t.sol", content: "" }] - }), - /generated test file must be non-empty/u - ); - assert.equal(fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests.json")), false); -}); - -test("generated-test manifest writer rejects support-only bundles before writing files", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-support-only-generated-test" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [], - supportFiles: [{ path: "generated-tests/Helper.sol", content: "library Helper {}\n" }] - }), - /cannot declare support files without a runnable generated test/u - ); - assert.equal(fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests.json")), false); - assert.equal( - fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests", "Helper.sol")), - false - ); -}); - -test("generated-test manifest writer rejects excess combined entries before creating bundle paths", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-excess-generated-test-entries" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: Array.from({ length: MAX_GENERATED_TEST_BUNDLE_ENTRIES + 1 }, (_, index) => ({ - path: `generated-tests/Test-${index}.sol`, - content: "x" - })), - supportFiles: [] - }), - /1024-entry combined bundle limit/u - ); - assert.equal(fs.existsSync(path.join(layout.artifactsDir, "strategy-a")), false); -}); - -test("generated-test manifest writer rejects excess cumulative existing bytes before reading companions", (t) => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-excess-generated-test-bytes" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir); - const tests = Array.from({ length: 5 }, (_, index) => { - const relativePath = `generated-tests/Test-${index}.sol`; - const absolutePath = path.join(nodeDir, relativePath); - fs.writeFileSync(absolutePath, "x", "utf8"); - fs.truncateSync(absolutePath, MAX_GENERATED_TEST_BUNDLE_BYTES / 4); - return { path: relativePath }; - }); - const readSync = t.mock.method(fs, "readSync", () => { - throw new Error("companion content was read before cumulative resource preflight completed"); - }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests, - supportFiles: [] - }), - /67108864-byte combined companion limit/u - ); - assert.equal(readSync.mock.callCount(), 0); - assert.equal(fs.existsSync(path.join(nodeDir, "generated-tests.json")), false); -}); - -test("generated-test manifest writer rejects cross-array duplicate paths before changing bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-duplicate-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const companionPath = path.join(generatedTestsDir, "Shared.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - fs.writeFileSync(companionPath, "sentinel companion\n", "utf8"); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Shared.sol", content: "replacement test\n" }], - supportFiles: [{ path: "generated-tests/Shared.sol", content: "replacement support\n" }] - }), - /repeats path "generated-tests\/Shared\.sol"/u - ); - assert.equal(fs.readFileSync(companionPath, "utf8"), "sentinel companion\n"); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); -}); - -test("generated-test manifest writer preflights missing existing companions before overwriting earlier files", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-missing-support-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [{ path: "generated-tests/MissingHelper.sol" }] - }), - /generated-test support file does not exist/u - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); -}); - -test("generated-test manifest writer preflights non-file destinations before overwriting earlier files", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-directory-support-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - const supportDirectory = path.join(generatedTestsDir, "InvariantFixture.sol"); - fs.mkdirSync(supportDirectory, { recursive: true }); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [{ path: "generated-tests/InvariantFixture.sol", content: "library InvariantFixture {}\n" }] - }), - /generated-test bundle destination must be a regular file/u - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); - assert.deepEqual(fs.readdirSync(supportDirectory), []); -}); - -test("generated-test manifest writer validates entry metadata before changing bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-invalid-metadata-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [ - { - path: "generated-tests/InvariantFixture.sol", - content: "library InvariantFixture {}\n", - language: "" - } - ] - }), - /generated tests manifest is schema-invalid/u - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.equal(fs.existsSync(path.join(generatedTestsDir, "InvariantFixture.sol")), false); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); -}); - -test("generated-test manifest writer preflights its manifest destination before changing companion bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-directory-manifest-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - const manifestDirectory = path.join(nodeDir, "generated-tests.json"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - fs.mkdirSync(manifestDirectory); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [] - }), - /generated-test bundle destination must be a regular file/u - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.deepEqual(fs.readdirSync(manifestDirectory), []); -}); - -test("generated-test manifest writer rejects a hard-linked manifest destination before changing bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-hardlinked-manifest-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - const manifestAlias = path.join(tempProject(), "generated-tests-alias.json"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - fs.linkSync(manifestPath, manifestAlias); - const manifestInode = fs.lstatSync(manifestPath).ino; - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [] - }), - (error: unknown) => { - assert.ok(error instanceof ArtifactPathError); - assert.equal(error.code, "hard-link"); - assert.match(error.message, /bundle destination must be singly linked/u); - return true; - } - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); - assert.equal(fs.readFileSync(manifestAlias, "utf8"), "sentinel manifest\n"); - assert.equal(fs.lstatSync(manifestPath).ino, manifestInode); - assert.equal(fs.lstatSync(manifestPath).nlink, 2); -}); - -test("generated-test manifest writer rejects a hard-linked companion before changing bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-hardlinked-companion-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const testPath = path.join(generatedTestsDir, "Replay.t.sol"); - const supportPath = path.join(generatedTestsDir, "InvariantFixture.sol"); - const supportAlias = path.join(tempProject(), "InvariantFixture-alias.sol"); - const manifestPath = path.join(nodeDir, "generated-tests.json"); - fs.writeFileSync(testPath, "sentinel test\n", "utf8"); - fs.writeFileSync(supportPath, "sentinel support\n", "utf8"); - fs.linkSync(supportPath, supportAlias); - fs.writeFileSync(manifestPath, "sentinel manifest\n", "utf8"); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "replacement test\n" }], - supportFiles: [{ path: "generated-tests/InvariantFixture.sol", content: "replacement support\n" }] - }), - (error: unknown) => { - assert.ok(error instanceof ArtifactPathError); - assert.equal(error.code, "hard-link"); - assert.match(error.message, /bundle destination must be singly linked/u); - return true; - } - ); - assert.equal(fs.readFileSync(testPath, "utf8"), "sentinel test\n"); - assert.equal(fs.readFileSync(supportPath, "utf8"), "sentinel support\n"); - assert.equal(fs.readFileSync(supportAlias, "utf8"), "sentinel support\n"); - assert.equal(fs.readFileSync(manifestPath, "utf8"), "sentinel manifest\n"); - assert.equal(fs.lstatSync(supportPath).nlink, 2); -}); - -test("generated-test manifest reader rejects a hard-linked manifest", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-read-hardlinked-generated-test" }); - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "contract Replay {}\n" }], - supportFiles: [] - }); - const manifestPath = path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests.json"); - const manifestAlias = path.join(tempProject(), "generated-tests-alias.json"); - fs.linkSync(manifestPath, manifestAlias); - const manifestBytes = fs.readFileSync(manifestPath); - - assert.throws( - () => readGeneratedTestManifest(layout, "strategy-a"), - (error: unknown) => { - assert.ok(error instanceof ArtifactPathError); - assert.equal(error.code, "hard-link"); - assert.match(error.message, /manifest must be a singly linked regular file/u); - return true; - } - ); - assert.deepEqual(fs.readFileSync(manifestPath), manifestBytes); - assert.deepEqual(fs.readFileSync(manifestAlias), manifestBytes); -}); - -test("generated-test manifest writer rejects file-directory path collisions before creating bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-prefix-collision-generated-test" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "contract Replay {}\n" }], - supportFiles: [ - { - path: "generated-tests/Replay.t.sol/InvariantFixture.sol", - content: "library InvariantFixture {}\n" - } - ] - }), - /conflicts with file path/u - ); - assert.equal(fs.existsSync(path.join(layout.artifactsDir, "strategy-a")), false); -}); - -test("generated-test manifest writer rejects noncanonical path aliases without rewriting them", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-noncanonical-path-generated-test" }); - - for (const candidate of [ - "Replay.t.sol", - "generated-tests/sub/../Replay.t.sol", - "generated-tests/./Replay.t.sol", - "generated-tests/sub//Replay.t.sol" - ]) { - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: candidate, content: "contract Replay {}\n" }], - supportFiles: [] - }), - /must (?:begin with|already be a normalized relative POSIX path)/u, - candidate - ); - assert.equal(fs.existsSync(path.join(layout.artifactsDir, "strategy-a")), false, candidate); - } -}); - -test("generated-test manifest writer rejects non-UTF-8 supplied support before creating bundle bytes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-binary-support-generated-test" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - tests: [{ path: "generated-tests/Replay.t.sol", content: "contract Replay {}\n" }], - supportFiles: [{ path: "generated-tests/fixture.dat", content: Buffer.from([0xff]) }] - }), - /generated-test support file must be strict UTF-8 text/u - ); - assert.equal(fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests.json")), false); - assert.equal(fs.existsSync(path.join(getNodeArtifactDir(layout, "strategy-a"), "generated-tests")), false); -}); - -test("generated-test manifest writer rejects final symlinks without touching outside files", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-symlinked-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const outsideDir = tempProject(); - const outsideFile = path.join(outsideDir, "Outside.t.sol"); - fs.writeFileSync(outsideFile, "outside sentinel\n"); - const symlinkPath = path.join(generatedTestsDir, "Linked.t.sol"); - fs.symlinkSync(outsideFile, symlinkPath); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - supportFiles: [], - tests: [{ path: "generated-tests/Linked.t.sol", content: "replacement\n" }] - }), - /symlink/u - ); - assert.equal(fs.lstatSync(symlinkPath).isSymbolicLink(), true); - assert.equal(fs.readFileSync(outsideFile, "utf8"), "outside sentinel\n"); - assert.equal(fs.existsSync(path.join(nodeDir, "generated-tests.json")), false); -}); - -test("generated-test manifest writer rejects broken final symlinks before writing content", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-broken-symlink-generated-test" }); - const nodeDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); - const generatedTestsDir = path.join(nodeDir, "generated-tests"); - fs.mkdirSync(generatedTestsDir, { recursive: true }); - const missingOutsideFile = path.join(tempProject(), "Missing.t.sol"); - const symlinkPath = path.join(generatedTestsDir, "Broken.t.sol"); - fs.symlinkSync(missingOutsideFile, symlinkPath); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - supportFiles: [], - tests: [{ path: "generated-tests/Broken.t.sol", content: "replacement\n" }] - }), - /symlink/u - ); - assert.equal(fs.lstatSync(symlinkPath).isSymbolicLink(), true); - assert.equal(fs.existsSync(missingOutsideFile), false); -}); - -test("generated-test manifest writer rejects paths outside the generated-tests root", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-outside-generated-test" }); - - assert.throws( - () => - writeGeneratedTestManifest({ - layout, - nodeId: "strategy-a", - supportFiles: [], - tests: [{ path: "generated-tests/../../Outside.t.sol", content: "outside\n" }] - }), - /cannot traverse outside/u - ); - assert.equal(fs.existsSync(path.join(layout.artifactsDir, "Outside.t.sol")), false); -}); diff --git a/packages/artifacts/test/events-tail-repair.test.ts b/packages/artifacts/test/events-tail-repair.test.ts index e4565cdde..0ea0dde9d 100644 --- a/packages/artifacts/test/events-tail-repair.test.ts +++ b/packages/artifacts/test/events-tail-repair.test.ts @@ -11,8 +11,7 @@ import { appendEventRecord, createEventRecord, createRunLayout, - replayEvents, - truncateDurable + replayEvents } from "../src/index.js"; function tempProject(): string { @@ -94,13 +93,12 @@ test("a torn event index rejects the whole append before the canonical journal c assert.equal(replayEvents(layout).records[0]?.event_id, first.event_id); }); -test("durable repair mutations reject stale sizes and hard-linked files", () => { +test("durable appends reject stale sizes and hard-linked files", () => { const filePath = path.join(tempProject(), "events.jsonl"); fs.writeFileSync(filePath, "first\nsecond", "utf8"); const observedSize = fs.statSync(filePath).size; fs.appendFileSync(filePath, "-raced\n", "utf8"); - assert.throws(() => truncateDurable(filePath, 6, { expectedSize: observedSize }), /changed size/u); assert.throws( () => appendBytesDurableAt(filePath, Buffer.from("\n"), { expectedSize: observedSize }), /changed size/u @@ -109,7 +107,6 @@ test("durable repair mutations reject stale sizes and hard-linked files", () => const currentSize = fs.statSync(filePath).size; const linkPath = path.join(path.dirname(filePath), "events-link.jsonl"); fs.linkSync(filePath, linkPath); - assert.throws(() => truncateDurable(filePath, 0, { expectedSize: currentSize }), /must not be hard-linked/u); assert.throws( () => appendBytesDurableAt(filePath, Buffer.from("x"), { expectedSize: currentSize }), /must not be hard-linked/u diff --git a/packages/artifacts/test/threat-goal-artifacts.test.ts b/packages/artifacts/test/threat-goal-artifacts.test.ts index 1c0df73cb..c4b1b5ec9 100644 --- a/packages/artifacts/test/threat-goal-artifacts.test.ts +++ b/packages/artifacts/test/threat-goal-artifacts.test.ts @@ -9,11 +9,9 @@ import { GOAL_PLAN_POLICY, GOAL_PLAN_SCHEMA_VERSION, THREAT_MODEL_SCHEMA_VERSION, - buildFindingSourceExpectations, goalPlanExpansionFacts, materializeCanonicalThreatModelMarkdown, renderThreatModelMarkdown, - normalizeFindings, validateArtifactContract, validateGoalPlan, validateThreatModel, @@ -724,383 +722,6 @@ test("an applicable class remains planned when no explicit threat maps to it", ( assert.equal(result.value?.class_goals[0]?.coverage_gap, true); }); -test("initial finding normalization ignores agent provenance and seeds the runtime-controlled producer", () => { - const artifactDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-goal-provenance-")); - fs.writeFileSync( - path.join(artifactDir, "findings.json"), - JSON.stringify([ - { - title: "Overdue liquidation bypass", - status: "candidate", - severity_guess: "high", - confidence: "high", - summary: "The position can be liquidated before its required overdue boundary.", - source_node_id: "legacy-node", - source_nodes: ["dynamic:class:liquidation:fixed-term-before-overdue", "legacy-node"] - } - ]) - ); - - const result = normalizeFindings({ - artifactDir, - nodeId: "filesystem-safe-attempt", - provenance: { producerNodeId: "dynamic:threat:liquidation:overdue" } - }); - - assert.equal(result.findings[0]?.source_node_id, "dynamic:threat:liquidation:overdue"); - assert.equal(result.findings[0]?.producer_node_id, "dynamic:threat:liquidation:overdue"); - assert.deepEqual(result.findings[0]?.source_nodes, ["dynamic:threat:liquidation:overdue"]); -}); - -test("downstream normalization preserves discovery sources without treating the transformer as a discoverer", () => { - const artifactDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-goal-provenance-transform-")); - fs.writeFileSync( - path.join(artifactDir, "deduped-findings.json"), - JSON.stringify([ - { - title: "Overdue liquidation bypass", - status: "candidate", - severity_guess: "high", - confidence: "high", - summary: "Two focused hunters corroborated the same root cause.", - source_nodes: ["dynamic:threat:liquidation:overdue", "dynamic:class:liquidation:fixed-term-before-overdue"] - } - ]) - ); - - const result = normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: ["dynamic:threat:liquidation:overdue", "dynamic:class:liquidation:fixed-term-before-overdue"] - }); - assert.equal(result.findings[0]?.producer_node_id, "dedupe-findings"); - assert.equal(result.findings[0]?.source_node_id, "dynamic:threat:liquidation:overdue"); - assert.deepEqual(result.findings[0]?.source_nodes, [ - "dynamic:threat:liquidation:overdue", - "dynamic:class:liquidation:fixed-term-before-overdue" - ]); - - const tamperedDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-goal-provenance-tamper-")); - fs.writeFileSync( - path.join(tamperedDir, "deduped-findings.json"), - JSON.stringify([ - { - title: "Invented source", - status: "candidate", - severity_guess: "high", - confidence: "high", - summary: "The review node tried to invent discovery provenance.", - source_nodes: ["dynamic:threat:liquidation:invented"] - } - ]) - ); - assert.throws( - () => - normalizeFindings({ - artifactDir: tamperedDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: ["dynamic:threat:liquidation:overdue"] - }), - /not present in dependency findings/u - ); -}); - -test("dedupe provenance rejects dropped corroborating sources and incomplete lifecycle coverage", () => { - const upstream = [ - { - node_id: "dynamic:threat:liquidation:overdue", - artifact_path: "artifacts/threat/findings.json", - finding: { - id: "threat-finding", - dedupe_key: "raw:threat", - producer_node_id: "dynamic:threat:liquidation:overdue", - source_nodes: ["dynamic:threat:liquidation:overdue"] - } - }, - { - node_id: "dynamic:class:liquidation:fixed-term-before-overdue", - artifact_path: "artifacts/class/findings.json", - finding: { - id: "class-finding", - dedupe_key: "raw:class", - producer_node_id: "dynamic:class:liquidation:fixed-term-before-overdue", - source_nodes: ["dynamic:class:liquidation:fixed-term-before-overdue"] - } - } - ]; - const lifecycleLedger = { - schema_version: "1.0", - records: [ - { - dedupe_key: "root:fixed-term-overdue", - source_artifacts: [ - { - node_id: "dynamic:threat:liquidation:overdue", - finding_id: "threat-finding" - }, - { - node_id: "dynamic:class:liquidation:fixed-term-before-overdue", - finding_id: "class-finding" - } - ] - } - ] - }; - const expectations = buildFindingSourceExpectations({ - upstream, - lifecycleLedger, - requireLifecycleCoverage: true - }); - const artifactDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-goal-provenance-drop-")); - fs.writeFileSync( - path.join(artifactDir, "deduped-findings.json"), - JSON.stringify([ - { - id: "kept-finding", - dedupe_key: "root:fixed-term-overdue", - title: "Fixed-term liquidation before overdue", - status: "candidate", - severity_guess: "high", - confidence: "high", - summary: "Two focused hunters found the same lifecycle boundary bug.", - source_nodes: ["dynamic:threat:liquidation:overdue"] - } - ]) - ); - assert.throws( - () => - normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }), - /does not preserve the exact dependency discovery-source union/u - ); - - const complete = JSON.parse(fs.readFileSync(path.join(artifactDir, "deduped-findings.json"), "utf8")) as Array< - Record - >; - complete[0]!.source_nodes = [ - "dynamic:class:liquidation:fixed-term-before-overdue", - "dynamic:threat:liquidation:overdue" - ]; - fs.writeFileSync(path.join(artifactDir, "deduped-findings.json"), JSON.stringify(complete)); - const normalized = normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }); - assert.deepEqual(normalized.findings[0]?.source_nodes, [ - "dynamic:threat:liquidation:overdue", - "dynamic:class:liquidation:fixed-term-before-overdue" - ]); - - const incompleteLedger = structuredClone(lifecycleLedger); - incompleteLedger.records[0]!.source_artifacts.pop(); - assert.throws( - () => - buildFindingSourceExpectations({ - upstream, - lifecycleLedger: incompleteLedger, - requireLifecycleCoverage: true - }), - /omitted dependency findings/u - ); -}); - -test("generic finding ID collisions cannot union unrelated discovery lanes", () => { - const upstream = [ - { - node_id: "dynamic:threat:accounting:rounding", - artifact_path: "artifacts/threat/findings.json", - finding: { - id: "finding-1", - upstream_id: "threat-original", - dedupe_key: "raw:threat", - source_nodes: ["dynamic:threat:accounting:rounding"] - } - }, - { - node_id: "dynamic:class:authorization:roles", - artifact_path: "artifacts/class/findings.json", - finding: { - id: "finding-1", - upstream_id: "class-original", - dedupe_key: "raw:class", - source_nodes: ["dynamic:class:authorization:roles"] - } - } - ]; - const lifecycleLedger = { - schema_version: "1.0", - records: [ - { - dedupe_key: "root:rounding", - family_id: "family:rounding", - source_artifacts: [ - { - node_id: "dynamic:threat:accounting:rounding", - finding_id: "finding-1" - } - ] - }, - { - dedupe_key: "root:roles", - family_id: "family:roles", - source_artifacts: [ - { - node_id: "dynamic:class:authorization:roles", - finding_id: "finding-1" - } - ] - } - ] - }; - const expectations = buildFindingSourceExpectations({ - upstream, - lifecycleLedger, - requireLifecycleCoverage: true - }); - const artifactDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-provenance-id-collision-")); - const finding = { - id: "finding-1", - dedupe_key: "root:rounding", - title: "Rounding drift", - status: "candidate", - severity_guess: "high", - confidence: "high", - summary: "The rounding lane found a loss of accounting precision.", - source_nodes: ["dynamic:threat:accounting:rounding"] - }; - fs.writeFileSync(path.join(artifactDir, "deduped-findings.json"), JSON.stringify([finding])); - const normalized = normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }); - assert.deepEqual(normalized.findings[0]?.source_nodes, ["dynamic:threat:accounting:rounding"]); - - for (const [label, contradictoryIdentity] of [ - ["family ID", { family_id: "family:roles" }], - ["finding reference", { upstream_id: "class-original" }] - ] as const) { - fs.writeFileSync( - path.join(artifactDir, "deduped-findings.json"), - JSON.stringify([{ ...finding, ...contradictoryIdentity }]) - ); - assert.throws( - () => - normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }), - new RegExp(`${label} conflicts with higher-priority dependency provenance`, "u") - ); - } - - const freshOutputIdentity = { - ...finding, - id: "kept-finding-new", - family_id: "family:new", - source_nodes: ["dynamic:threat:accounting:rounding"] - }; - fs.writeFileSync(path.join(artifactDir, "deduped-findings.json"), JSON.stringify([freshOutputIdentity])); - assert.deepEqual( - normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }).findings[0]?.source_nodes, - ["dynamic:threat:accounting:rounding"] - ); - - const ambiguous = { - ...finding, - dedupe_key: undefined, - source_nodes: upstream.flatMap((entry) => entry.finding.source_nodes) - }; - fs.writeFileSync(path.join(artifactDir, "deduped-findings.json"), JSON.stringify([ambiguous])); - assert.throws( - () => - normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }), - /finding ID matches conflicting dependency provenance records/u - ); - - const forgedStrongIdentity = { - ...finding, - dedupe_key: "root:unknown", - source_nodes: upstream.flatMap((entry) => entry.finding.source_nodes) - }; - fs.writeFileSync(path.join(artifactDir, "deduped-findings.json"), JSON.stringify([forgedStrongIdentity])); - assert.throws( - () => - normalizeFindings({ - artifactDir, - relativePath: "deduped-findings.json", - provenance: { producerNodeId: "dedupe-findings" }, - preserveSourceNodes: true, - requireSourceNodes: true, - allowedSourceNodes: upstream.flatMap((entry) => entry.finding.source_nodes), - sourceExpectations: expectations, - requireSourceExpectation: true - }), - /dedupe key does not match any dependency provenance record/u - ); - - const duplicateKeyLedger = structuredClone(lifecycleLedger); - duplicateKeyLedger.records[1]!.dedupe_key = "root:rounding"; - assert.throws( - () => - buildFindingSourceExpectations({ - upstream, - lifecycleLedger: duplicateKeyLedger, - requireLifecycleCoverage: true - }), - /duplicate dedupe_key/u - ); -}); - test("selected vulnerability-class snapshots are manifest-backed exact artifacts", () => { const artifactDir = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-selected-classes-")); const contents = Buffer.from("# Share inflation\n\nFocused hunter instructions.\n"); @@ -1412,56 +1033,3 @@ function classGoalPlanFixture(selectedPath: string): Record { sealGoalPlanCardinality(plan); return plan; } - -test("a cross-node dedupe root only resolves through its lifecycle ledger record", () => { - // Regression for the workflow-sync path, which rebuilt these expectations - // without the ledger. Without it the union expectation does not exist at all, - // so a legitimate merge matches nothing and a succeeded node is reported as a - // task-output-validation-failure. - const upstream = [ - { - node_id: "dynamic:threat:liquidation:overdue", - artifact_path: "/runs/r/artifacts/a/findings.json", - finding: { id: "f-threat", source_nodes: ["dynamic:threat:liquidation:overdue"] } - }, - { - node_id: "dynamic:class:liquidation:fixed-term-before-overdue", - artifact_path: "/runs/r/artifacts/b/findings.json", - finding: { id: "f-class", source_nodes: ["dynamic:class:liquidation:fixed-term-before-overdue"] } - } - ]; - const lifecycleLedger = { - schema_version: "1.0", - records: [ - { - dedupe_key: "root:fixed-term-overdue", - source_artifacts: [ - { path: "a/findings.json", node_id: "dynamic:threat:liquidation:overdue", finding_id: "f-threat" }, - { - path: "b/findings.json", - node_id: "dynamic:class:liquidation:fixed-term-before-overdue", - finding_id: "f-class" - } - ] - } - ] - }; - const unionOf = (expectations: ReturnType): string[][] => - expectations.filter((expectation) => expectation.source_nodes.length > 1).map((e) => [...e.source_nodes].sort()); - - assert.deepEqual(unionOf(buildFindingSourceExpectations({ upstream, requireLifecycleCoverage: true })), []); - assert.deepEqual( - unionOf(buildFindingSourceExpectations({ upstream, lifecycleLedger, requireLifecycleCoverage: true })), - [["dynamic:class:liquidation:fixed-term-before-overdue", "dynamic:threat:liquidation:overdue"]] - ); - - // The ledger keys the union expectation, so the retained finding must carry - // that same key. Both prompts now say so explicitly. - const withLedger = buildFindingSourceExpectations({ upstream, lifecycleLedger, requireLifecycleCoverage: true }); - assert.ok( - withLedger.some( - (expectation) => - expectation.finding_keys.includes("root:fixed-term-overdue") && expectation.source_nodes.length === 2 - ) - ); -}); diff --git a/packages/security/src/path-policy.ts b/packages/security/src/path-policy.ts index c61c53cfa..9d642cbb4 100644 --- a/packages/security/src/path-policy.ts +++ b/packages/security/src/path-policy.ts @@ -1,18 +1,6 @@ -import { realpathSync } from "node:fs"; import path from "node:path"; import { type PolicyDiagnostic, type PolicyResult, policyError, policyResult } from "./types.js"; -export interface ResolveInsideOptions { - allowAbsolute?: boolean; - mustExist?: boolean; -} - -export interface ResolvedPathPolicy { - root: string; - path: string; - relativePath: string; -} - export function validateSafeId(label: string, id: string): PolicyResult { const diagnostics: PolicyDiagnostic[] = []; if (id.length === 0) { @@ -62,63 +50,6 @@ export function validateSafeRelativePath(relativePath: string): PolicyResult { - const diagnostics: PolicyDiagnostic[] = []; - if (requestedPath.includes("\\")) { - diagnostics.push(policyError("PATH_BACKSLASH", `path \`${requestedPath}\` must use forward slashes`)); - } - if (isAbsoluteLike(requestedPath) && !options.allowAbsolute) { - diagnostics.push(policyError("PATH_ABSOLUTE", `path \`${requestedPath}\` must be relative`)); - } - - let rootReal = ""; - let candidateReal = ""; - try { - rootReal = realpathSync.native(root); - } catch (error) { - diagnostics.push( - policyError("PATH_ROOT_MISSING", `root \`${root}\` is not accessible`, { - details: { error: String(error) } - }) - ); - } - - const candidate = isAbsoluteLike(requestedPath) ? requestedPath : path.join(root, requestedPath); - if (!isAbsoluteLike(requestedPath)) { - diagnostics.push(...validateSafeRelativePath(requestedPath).diagnostics); - } - try { - candidateReal = canonicalExistingOrParent(candidate, options.mustExist ?? false); - } catch (error) { - diagnostics.push( - policyError("PATH_CANDIDATE_MISSING", `path \`${requestedPath}\` is not accessible`, { - path: requestedPath, - details: { error: String(error) } - }) - ); - } - - if (rootReal && candidateReal && !isPathInside(rootReal, candidateReal)) { - diagnostics.push( - policyError("PATH_ESCAPE", `path \`${requestedPath}\` resolves outside \`${rootReal}\``, { - path: requestedPath, - details: { root: rootReal, resolved: candidateReal } - }) - ); - } - - const relative = rootReal && candidateReal ? toPosixRelative(rootReal, candidateReal) : ""; - return policyResult(diagnostics, { - root: rootReal, - path: candidateReal, - relativePath: relative - }); -} - export function normalizeRelativePath(value: string): string { return value.replaceAll("\\", "/").replace(/^\.\/+/, ""); } @@ -141,27 +72,6 @@ export function isPathInside(root: string, candidate: string): boolean { return relative === "" || (!relative.startsWith("..") && !path.isAbsolute(relative)); } -function canonicalExistingOrParent(candidate: string, mustExist: boolean): string { - if (mustExist) { - return realpathSync.native(candidate); - } - let current = candidate; - const missingParts: string[] = []; - while (current !== path.dirname(current)) { - try { - return path.join(realpathSync.native(current), ...missingParts.reverse()); - } catch { - missingParts.push(path.basename(current)); - current = path.dirname(current); - } - } - return path.join(realpathSync.native(current), ...missingParts.reverse()); -} - -function toPosixRelative(root: string, candidate: string): string { - return normalizeRelativePath(path.relative(root, candidate)); -} - function isReservedWindowsName(component: string): boolean { return /^(?:con|prn|aux|nul|com[1-9]|lpt[1-9])(?:\..*)?$/i.test(component); } diff --git a/packages/security/src/types.ts b/packages/security/src/types.ts index b5e399b6b..f5b335d96 100644 --- a/packages/security/src/types.ts +++ b/packages/security/src/types.ts @@ -9,24 +9,10 @@ export interface PolicyDiagnostic { details?: Record; } -export interface UnsafeModeAuditInput { - mode: string; - scope: string; - reason: string; - operator: string; - acknowledgement: string; - approvedAt?: string; -} - -export interface UnsafeModeAuditRecord extends UnsafeModeAuditInput { - approvedAt: string; -} - export interface PolicyResult { ok: boolean; diagnostics: PolicyDiagnostic[]; value?: T; - audit: UnsafeModeAuditRecord[]; } export function policyError( @@ -37,15 +23,10 @@ export function policyError( return { code, message, severity: "error", ...fields }; } -export function policyResult( - diagnostics: PolicyDiagnostic[], - value?: T, - audit: UnsafeModeAuditRecord[] = [] -): PolicyResult { +export function policyResult(diagnostics: PolicyDiagnostic[], value?: T): PolicyResult { const result: PolicyResult = { ok: !diagnostics.some((diagnostic) => diagnostic.severity === "error"), - diagnostics, - audit + diagnostics }; if (value !== undefined) { result.value = value; diff --git a/packages/security/test/path-policy.test.ts b/packages/security/test/path-policy.test.ts index fb0ae1983..24c0c56ee 100644 --- a/packages/security/test/path-policy.test.ts +++ b/packages/security/test/path-policy.test.ts @@ -1,9 +1,6 @@ import assert from "node:assert/strict"; -import { mkdtempSync, realpathSync, symlinkSync, writeFileSync } from "node:fs"; -import { tmpdir } from "node:os"; -import path from "node:path"; import { test } from "node:test"; -import { resolvePathInside, validateSafeId, validateSafeRelativePath } from "../src/index.js"; +import { validateSafeId, validateSafeRelativePath } from "../src/index.js"; test("safe relative path policy rejects traversal and absolute injection", () => { assert.equal(validateSafeRelativePath("src/Test.sol").ok, true); @@ -19,11 +16,6 @@ test("shared path policy rejects backslash paths before normalization", () => { const relative = validateSafeRelativePath("runs\\run-1\\report.json"); assert.equal(relative.ok, false); assert.ok(relative.diagnostics.some((diagnostic) => diagnostic.code === "PATH_BACKSLASH")); - - const root = mkdtempSync(path.join(realpathSync(tmpdir()), "ufz-path-backslash-")); - const resolved = resolvePathInside(root, "runs\\run-1\\report.json"); - assert.equal(resolved.ok, false); - assert.ok(resolved.diagnostics.some((diagnostic) => diagnostic.code === "PATH_BACKSLASH")); }); test("safe IDs reject traversal, slashes, and dot edges", () => { @@ -32,25 +24,3 @@ test("safe IDs reject traversal, slashes, and dot edges", () => { assert.equal(validateSafeId("node id", "bad/node").ok, false); assert.equal(validateSafeId("node id", ".hidden").ok, false); }); - -test("resolvePathInside rejects symlink escapes after canonicalization", () => { - const root = mkdtempSync(path.join(realpathSync(tmpdir()), "ufz-path-root-")); - const outside = mkdtempSync(path.join(realpathSync(tmpdir()), "ufz-path-outside-")); - writeFileSync(path.join(outside, "secret.txt"), "secret"); - symlinkSync(outside, path.join(root, "linked-outside")); - - const result = resolvePathInside(root, "linked-outside/secret.txt"); - assert.equal(result.ok, false); - assert.ok(result.diagnostics.some((diagnostic) => diagnostic.code === "PATH_ESCAPE")); -}); - -test("resolvePathInside rejects traversal and preserves non-existing leaf paths", () => { - const root = mkdtempSync(path.join(realpathSync(tmpdir()), "ufz-path-root-")); - const safe = resolvePathInside(root, "nested/new-file.txt"); - assert.equal(safe.ok, true, JSON.stringify(safe.diagnostics)); - assert.equal(safe.value?.relativePath, "nested/new-file.txt"); - - const traversal = resolvePathInside(root, "../outside.txt"); - assert.equal(traversal.ok, false); - assert.ok(traversal.diagnostics.some((diagnostic) => diagnostic.code === "PATH_TRAVERSAL")); -}); From 5ae10bebb4fd12b93931b1d2e3394b34647df654 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:48:10 +0000 Subject: [PATCH 055/206] fix(artifacts): event appends stop re-reading the journal and writing an unread index appendEvent validated the whole journal several times per append: it pre-read events.jsonl and five events.index/* fan-out files, then each append re-read its file in full, and appendStrictJsonlRecords re-read the file once more after writing. Nothing reads events.index/* or its query-inputs.json facade (queryEvents filters events.jsonl; the report bundle only copies the directory). The 100,000-record cap on the event codec also failed every later append and replay once reached, and the sync, lifecycle and cancel paths all append events. - appendEvent writes only events.jsonl. It checks the new record against the final record and the trailing records that share its timestamp. The event ID hashes the timestamp and timestamps never decrease, so a record outside that window can repeat the ID only through a 96-bit hash collision. Readers still validate the whole journal, and the byte length read before the write still fences it. - appendStrictJsonlRecords returns the snapshot it validated plus the fenced append instead of re-reading the file. - The event codec has no record cap; its 64 MiB byte limit still applies. - createRunLayout no longer creates events.index/ or query-inputs.json, and RunLayout drops eventsIndexDir. The facade schema stays registered because removing a schema file would change the schema-bundle digest. - The report bundle validates its journal snapshot with parseEventJournalBytes instead of its own copy of the parser, which still enforced the 100,000-record cap. Measured CPU per append with fsync disabled (eatmydata, loaded 32-core host): 17.8 -> 4.8 ms at 250 existing events, 478 -> 12.5 ms at 10,000, 2,183 -> 39 ms at 50,000. The remaining growth is the byte read. "event appends and replays keep working past 100,000 records" fails on main ("event journal exceeds the record limit") and passes here. The same-millisecond duplicate test passes on both, and fails if the window is shrunk to the final record only. Refs #462 Co-Authored-By: Claude Opus 5.5 --- docs/SPECS.md | 1 - docs/reference/artifacts-reports.md | 12 +- packages/artifacts/src/events.ts | 171 +++--------------- packages/artifacts/src/run-layout.ts | 18 +- packages/artifacts/src/strict-jsonl.ts | 164 +++++++++++++---- packages/artifacts/test/artifacts.test.ts | 132 +++++--------- .../artifacts/test/events-tail-repair.test.ts | 55 ++---- packages/cli/src/commands/report/bundle.ts | 87 +-------- 8 files changed, 220 insertions(+), 420 deletions(-) diff --git a/docs/SPECS.md b/docs/SPECS.md index 0fe0c4cea..a73223290 100644 --- a/docs/SPECS.md +++ b/docs/SPECS.md @@ -470,7 +470,6 @@ Before or at launch, each run MUST persist: - immutable rendered prompt snapshots under `prompt-snapshots/` - per-node artifacts under `artifacts/` - review artifacts under `review/` -- event query indexes under `events.index/` - workspace metadata under `workspaces/` Reporting is agentic and lives in final-report artifacts. diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index 2683bb70e..678908da7 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -30,14 +30,18 @@ trusted-cli.json trusted-bin/ artifacts/ review/ -events.index/ workspaces/ workspaces.json ``` -`source-run.json` is present when the run derives from another run. Event query -indexes are JSONL files derived from `events.jsonl`; SQLite events are not part -of the artifact contract. +`source-run.json` is present when the run derives from another run. +`events.jsonl` is the only event journal: event queries filter it, and SQLite +events are not part of the artifact contract. An append checks the new event +against the final event and any trailing events with the same timestamp; +`replayEvents` and `queryEvents` validate the whole journal. The journal has no +record-count limit; its 64 MiB byte limit still applies. Runs created before +this change may also have an `events.index/` directory. Nothing reads it, and +report bundles still copy it. `usage.jsonl` is an append-only ledger of normalized workflow usage events. Each entry's immutable identity is the exact Smithers pair diff --git a/packages/artifacts/src/events.ts b/packages/artifacts/src/events.ts index 194a22787..fccb79277 100644 --- a/packages/artifacts/src/events.ts +++ b/packages/artifacts/src/events.ts @@ -1,7 +1,4 @@ import crypto from "node:crypto"; -import fs from "node:fs"; -import path from "node:path"; -import { isDeepStrictEqual } from "node:util"; import { redactSecretsInValue, type SecretScanMode } from "@ultrafuzz/security"; import { z } from "zod/v4"; @@ -14,19 +11,12 @@ import { canonicalUuidSchema } from "./portable-json-primitives.js"; import { type RunLayout } from "./run-layout.js"; -import { - SAFE_ID_PATTERN, - createFileDurableExclusive, - prepareSafeFilePath, - readJsonFile, - safeResolveInside, - validateSafeId -} from "./safe-paths.js"; +import { SAFE_ID_PATTERN, validateSafeId } from "./safe-paths.js"; import { schemaErrorMessage, validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; import { - appendStrictJsonlRecords, + appendStrictJsonlRecordsAfterTail, + parseStrictJsonlBytes, readStrictJsonlSnapshot, - validateStrictJsonlHistory, type StrictJsonlCodec } from "./strict-jsonl.js"; import { @@ -47,7 +37,6 @@ export const DEFAULT_EVENT_REPLAY_LIMIT = 10_000; const MAX_EVENT_INDEX_FILENAME_LENGTH = 128; const EVENT_INDEX_EXTENSION = ".jsonl"; const EVENT_INDEX_DIRECT_MAX_ID_LENGTH = MAX_EVENT_INDEX_FILENAME_LENGTH - EVENT_INDEX_EXTENSION.length; -const EVENT_INDEX_LONG_DIRECTORY = "sha256"; const EVENT_INDEX_KEY_SCHEMA_VERSION = "ultrafuzz.event-index-key.v1" as const; export const EVENT_RECORD_TYPES = [ @@ -96,36 +85,6 @@ export interface EventQuery { limit?: number; } -export interface EventQueryFacade { - schema_version: typeof EVENT_QUERY_FACADE_SCHEMA_VERSION; - run_id: string; - append_log: string; - index_root: string; - indexes: ["run", "node", "type", "status", "timestamp"]; - filters: { - run_id: "events.index/run/.jsonl"; - node_id: "events.index/node/.jsonl"; - event_type: "events.index/type/.jsonl"; - status: "events.index/status/.jsonl"; - timestamp: "events.index/timestamp/.jsonl"; - }; - long_filters: { - run_id: "events.index/run/sha256/.jsonl"; - node_id: "events.index/node/sha256/.jsonl"; - event_type: "events.index/type/sha256/.jsonl"; - status: "events.index/status/sha256/.jsonl"; - }; - index_key_encoding: { - version: typeof EVENT_INDEX_KEY_SCHEMA_VERSION; - direct_max_id_length: number; - direct_id_path: "/.jsonl"; - long_id_path: "/sha256/.jsonl"; - digest: "sha256"; - hash_input_encoding: "utf8"; - digest_encoding: "hex"; - }; -} - const eventIdSchema = z.string().regex(/^evt-[a-f0-9]{24}$/u); const timestampSchema = canonicalTimestampSchema; const safeIdSchema = z.string().regex(SAFE_ID_PATTERN); @@ -1216,36 +1175,18 @@ export function assertEventRecord(value: unknown, recordPath = "$"): EventRecord return result.value; } -export function validateEventQueryFacade(value: unknown, recordPath = "$"): SchemaValidationResult { - return validateWithZod(eventQueryFacadeSchema as z.ZodType, value, { - path: recordPath, - code: "EVENT_QUERY_FACADE_SCHEMA_INVALID" - }); -} - -export function assertEventQueryFacade(value: unknown, recordPath = "$"): EventQueryFacade { - const result = validateEventQueryFacade(value, recordPath); - if (!result.ok || result.value === undefined) { - throw new Error(schemaErrorMessage("event query facade", result.issues)); - } - return result.value; -} - export function appendEvent(layout: RunLayout, input: AppendEventInput): EventRecord { const record = createEventRecord(layout, input); - const queryFacadePath = safeResolveInside(layout.eventsIndexDir, "query-inputs.json", "event query facade"); - assertExistingQueryFacade(layout, queryFacadePath); - const targets = [layout.eventsPath, ...eventIndexPaths(layout, record)]; - for (const target of targets) { - const codec = eventRecordCodec(layout.runId); - const existing = readStrictJsonlSnapshot(target, codec).records; - validateStrictJsonlHistory([...existing, record], codec); - } - appendEventRecord(layout.eventsPath, record, layout.root, layout.runId); - for (const target of targets.slice(1)) { - appendEventRecord(target, record, layout.root, layout.runId); - } - writeQueryFacadeInputs(layout, queryFacadePath); + // The event ID hashes the timestamp and timestamps never decrease, so only the + // trailing records that share this timestamp can repeat the ID. Checking that + // window keeps an append from costing a parse of the whole journal. + appendStrictJsonlRecordsAfterTail( + layout.eventsPath, + [record], + eventRecordCodec(layout.runId), + (existing) => existing.timestamp === record.timestamp, + layout.root + ); return record; } @@ -1279,16 +1220,6 @@ export function createEventRecord(layout: Pick, input: Appen }); } -export function appendEventRecord( - eventsPath: string, - record: EventRecord, - trustedRoot?: string, - expectedRunId?: string -): void { - const canonical = assertEventRecord(record); - appendStrictJsonlRecords(eventsPath, [canonical], eventRecordCodec(expectedRunId), trustedRoot); -} - export function replayEvents(layoutOrPath: RunLayout | string, limit = DEFAULT_EVENT_REPLAY_LIMIT): EventReplay { if (!Number.isSafeInteger(limit) || limit < 0) throw new Error("event replay limit must be a non-negative safe integer"); @@ -1302,6 +1233,11 @@ export function replayEvents(layoutOrPath: RunLayout | string, limit = DEFAULT_E }; } +/** Parse one captured event-journal byte snapshot with the rules replayEvents applies. */ +export function parseEventJournalBytes(bytes: Uint8Array, expectedRunId?: string): EventRecord[] { + return parseStrictJsonlBytes(bytes, eventRecordCodec(expectedRunId)).records; +} + export function queryEvents(layout: RunLayout, query: EventQuery = {}): EventRecord[] { const normalizedQuery = normalizeEventQuery(query); const limit = normalizedQuery.limit ?? DEFAULT_EVENT_REPLAY_LIMIT; @@ -1326,42 +1262,6 @@ export function normalizeEventQuery(query: EventQuery = {}): EventQuery { return parsed.data; } -export function readEventQueryFacade(layout: RunLayout): EventQueryFacade { - return assertEventQueryFacade(readJsonFile(path.join(layout.eventsIndexDir, "query-inputs.json"))); -} - -export function createEventQueryFacadeInputs(layout: RunLayout): EventQueryFacade { - return assertEventQueryFacade({ - schema_version: EVENT_QUERY_FACADE_SCHEMA_VERSION, - run_id: layout.runId, - append_log: path.relative(layout.root, layout.eventsPath).split(path.sep).join("/"), - index_root: path.relative(layout.root, layout.eventsIndexDir).split(path.sep).join("/"), - indexes: ["run", "node", "type", "status", "timestamp"], - filters: { - run_id: "events.index/run/.jsonl", - node_id: "events.index/node/.jsonl", - event_type: "events.index/type/.jsonl", - status: "events.index/status/.jsonl", - timestamp: "events.index/timestamp/.jsonl" - }, - long_filters: { - run_id: "events.index/run/sha256/.jsonl", - node_id: "events.index/node/sha256/.jsonl", - event_type: "events.index/type/sha256/.jsonl", - status: "events.index/status/sha256/.jsonl" - }, - index_key_encoding: { - version: EVENT_INDEX_KEY_SCHEMA_VERSION, - direct_max_id_length: EVENT_INDEX_DIRECT_MAX_ID_LENGTH, - direct_id_path: "/.jsonl", - long_id_path: "/sha256/.jsonl", - digest: "sha256", - hash_input_encoding: "utf8", - digest_encoding: "hex" - } - }); -} - export function redactValue( value: unknown, forbiddenSecretValues: readonly string[] = [], @@ -1370,29 +1270,10 @@ export function redactValue( return redactSecretsInValue(value, undefined, forbiddenSecretValues, mode); } -function eventIndexPaths(layout: RunLayout, record: EventRecord): string[] { - const nodeId = eventRecordNodeId(record); - const targets = [ - ["run", ...eventIndexPath(record.run_id)], - ["type", ...eventIndexPath(record.event_type)], - ["timestamp", ...eventIndexPath(record.timestamp.slice(0, 10))] - ]; - if (nodeId !== undefined) targets.push(["node", ...eventIndexPath(nodeId)]); - targets.push(["status", ...eventIndexPath(record.status)]); - return targets.map((segments) => prepareSafeFilePath(layout.eventsIndexDir, segments.join("/"))); -} - function eventRecordNodeId(record: EventRecord): string | undefined { return "node_id" in record ? record.node_id : undefined; } -function eventIndexPath(value: string): string[] { - const direct = `${value}${EVENT_INDEX_EXTENSION}`; - if (direct.length <= MAX_EVENT_INDEX_FILENAME_LENGTH) return [direct]; - const digest = crypto.createHash("sha256").update(value, "utf8").digest("hex"); - return [EVENT_INDEX_LONG_DIRECTORY, `${digest}${EVENT_INDEX_EXTENSION}`]; -} - function eventRecordIdentity(record: EventRecord): string { return record.event_id; } @@ -1400,6 +1281,9 @@ function eventRecordIdentity(record: EventRecord): string { function eventRecordCodec(expectedRunId?: string): StrictJsonlCodec { return { label: "event journal", + // Only the byte limit bounds this journal: a record cap, once reached, would + // fail every later sync, lifecycle and cancel call for the rest of the run. + maxRecords: Number.MAX_SAFE_INTEGER, parseRecord: (value, recordPath) => { const record = assertEventRecord(value, recordPath); if (expectedRunId !== undefined && record.run_id !== expectedRunId) { @@ -1425,16 +1309,3 @@ function eventRecordCodec(expectedRunId?: string): StrictJsonlCodec } }; } - -function assertExistingQueryFacade(layout: RunLayout, facadePath: string): void { - if (!fs.existsSync(facadePath)) return; - const actual = assertEventQueryFacade(readJsonFile(facadePath)); - const expected = createEventQueryFacadeInputs(layout); - if (!isDeepStrictEqual(actual, expected)) throw new Error("event query facade conflicts with the current run layout"); -} - -function writeQueryFacadeInputs(layout: RunLayout, facadePath: string): void { - if (fs.existsSync(facadePath)) return; - const bytes = Buffer.from(`${JSON.stringify(createEventQueryFacadeInputs(layout), null, 2)}\n`, "utf8"); - createFileDurableExclusive(facadePath, bytes, layout.root); -} diff --git a/packages/artifacts/src/run-layout.ts b/packages/artifacts/src/run-layout.ts index 48c59ba56..6ce66e84c 100644 --- a/packages/artifacts/src/run-layout.ts +++ b/packages/artifacts/src/run-layout.ts @@ -2,7 +2,6 @@ import crypto from "node:crypto"; import fs from "node:fs"; import path from "node:path"; -import { createEventQueryFacadeInputs } from "./events.js"; import { assertPlannedGraph, PLANNED_GRAPH_SCHEMA_VERSION, type PlannedGraphDocument } from "./planned-graph.js"; import { CONFIG_REDACTIONS_SCHEMA_VERSION, @@ -44,7 +43,6 @@ export interface RunLayout { eventsPath: string; usageLedgerPath: string; attemptLedgerPath: string; - eventsIndexDir: string; } export interface CreateRunLayoutInput { @@ -80,13 +78,7 @@ export function createRunLayout(input: CreateRunLayoutInput): RunLayout { assertNoSymlinkComponents(guardRoot, root, "run root"); const layout = layoutForRunRoot(root, runId); - for (const directory of [ - layout.root, - layout.artifactsDir, - layout.workspacesDir, - layout.eventsIndexDir, - layout.reviewDir - ]) { + for (const directory of [layout.root, layout.artifactsDir, layout.workspacesDir, layout.reviewDir]) { fs.mkdirSync(directory, { recursive: true }); } @@ -165,11 +157,6 @@ export function createRunLayout(input: CreateRunLayoutInput): RunLayout { if ((input.overwrite ?? false) || !fs.existsSync(layout.attemptLedgerPath)) { writeFileDurable(layout.attemptLedgerPath, ""); } - writeJsonIfNeeded( - path.join(layout.eventsIndexDir, "query-inputs.json"), - createEventQueryFacadeInputs(layout), - input.overwrite ?? false - ); return layout; } @@ -193,8 +180,7 @@ export function layoutForRunRoot(root: string, runId = path.basename(root)): Run statePath: path.join(absoluteRoot, "state.json"), eventsPath: path.join(absoluteRoot, "events.jsonl"), usageLedgerPath: path.join(absoluteRoot, "usage.jsonl"), - attemptLedgerPath: path.join(absoluteRoot, "attempts.jsonl"), - eventsIndexDir: path.join(absoluteRoot, "events.index") + attemptLedgerPath: path.join(absoluteRoot, "attempts.jsonl") }; } diff --git a/packages/artifacts/src/strict-jsonl.ts b/packages/artifacts/src/strict-jsonl.ts index 3afea1ce4..255df935f 100644 --- a/packages/artifacts/src/strict-jsonl.ts +++ b/packages/artifacts/src/strict-jsonl.ts @@ -33,15 +33,18 @@ export function readStrictJsonlSnapshot( filePath: string, codec: StrictJsonlCodec ): StrictJsonlSnapshot { + const bytes = readStrictJsonlBytes(filePath, codec); + return bytes === undefined ? { records: [], byteLength: 0, exists: false } : parseStrictJsonlBytes(bytes, codec); +} + +function readStrictJsonlBytes(filePath: string, codec: StrictJsonlCodec): Buffer | undefined { try { fs.lstatSync(filePath); } catch (error) { - if (isErrnoException(error, "ENOENT")) return { records: [], byteLength: 0, exists: false }; + if (isErrnoException(error, "ENOENT")) return undefined; throw new Error(`failed to inspect ${codec.label} ${filePath}`, { cause: error }); } - const maxBytes = codec.maxBytes ?? DEFAULT_STRICT_JSONL_MAX_BYTES; - const bytes = readRegularFileSnapshot(filePath, maxBytes); - return parseStrictJsonlBytes(bytes, codec); + return readRegularFileSnapshot(filePath, codec.maxBytes ?? DEFAULT_STRICT_JSONL_MAX_BYTES); } /** Parse one already-captured immutable JSONL byte snapshot. */ @@ -72,35 +75,42 @@ export function parseStrictJsonlBytes( throw new Error(`${codec.label} exceeds the ${maxRecords}-record limit`); } - const maxRecordBytes = codec.maxRecordBytes ?? DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES; - const records: RecordType[] = []; - for (const [index, line] of lines.entries()) { - const lineNumber = index + 1; - if (line.trim().length === 0) throw new Error(`${codec.label} contains a blank record at line ${lineNumber}`); - const lineBytes = Buffer.from(line, "utf8"); - if (lineBytes.byteLength > maxRecordBytes) { - throw new Error(`${codec.label} record ${lineNumber} exceeds the ${maxRecordBytes}-byte limit`); - } - let parsed: unknown; - try { - parsed = parseStrictJsonBytes(lineBytes, { - maxBytes: maxRecordBytes, - maxDepth: 128, - maxItems: 100_000, - maxProperties: 100_000 - }); - } catch (error) { - throw new Error( - `${codec.label} record ${lineNumber} is invalid strict JSON: ${error instanceof Error ? error.message : String(error)}`, - { cause: error } - ); - } - records.push(codec.parseRecord(parsed, `$[${index}]`)); - } + const records = lines.map((line, index) => + parseStrictJsonlLine(line, Buffer.from(line, "utf8"), String(index + 1), `$[${String(index)}]`, codec) + ); validateStrictJsonlHistory(records, codec); return { records, byteLength: bytes.byteLength, exists: true }; } +function parseStrictJsonlLine( + line: string, + lineBytes: Buffer, + lineLabel: string, + recordPath: string, + codec: StrictJsonlCodec +): RecordType { + if (line.trim().length === 0) throw new Error(`${codec.label} contains a blank record at line ${lineLabel}`); + const maxRecordBytes = codec.maxRecordBytes ?? DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES; + if (lineBytes.byteLength > maxRecordBytes) { + throw new Error(`${codec.label} record ${lineLabel} exceeds the ${String(maxRecordBytes)}-byte limit`); + } + let parsed: unknown; + try { + parsed = parseStrictJsonBytes(lineBytes, { + maxBytes: maxRecordBytes, + maxDepth: 128, + maxItems: 100_000, + maxProperties: 100_000 + }); + } catch (error) { + throw new Error( + `${codec.label} record ${lineLabel} is invalid strict JSON: ${error instanceof Error ? error.message : String(error)}`, + { cause: error } + ); + } + return codec.parseRecord(parsed, recordPath); +} + function isErrnoException(error: unknown, code: string): error is NodeJS.ErrnoException { return error instanceof Error && "code" in error && (error as NodeJS.ErrnoException).code === code; } @@ -138,7 +148,84 @@ export function appendStrictJsonlRecords( const existing = readStrictJsonlSnapshot(filePath, codec); if (records.length === 0) return existing; - const canonicalRecords = records.map((record, index) => { + const canonicalRecords = canonicalStrictJsonlRecords( + records, + codec, + (index) => `$[${String(existing.records.length + index)}]` + ); + const combined = [...existing.records, ...canonicalRecords]; + if (combined.length > (codec.maxRecords ?? DEFAULT_STRICT_JSONL_MAX_RECORDS)) { + throw new Error(`${codec.label} exceeds the record limit`); + } + validateStrictJsonlHistory(combined, codec); + const byteLength = writeStrictJsonlRecords(filePath, existing, canonicalRecords, codec, trustedRoot); + // The write is fenced at the snapshot length, so the file now holds exactly these records. + return { records: combined, byteLength, exists: true }; +} + +/** + * Append records after validating them against only the journal's trailing + * records: the final record, then earlier ones while `inWindow` holds for + * them. This suits journals whose history rules relate a record only to the + * records inside that window. It does not count records, so it is for + * journals bounded by bytes alone. Readers still validate the whole journal, + * and the byte length read here still fences the write. + */ +export function appendStrictJsonlRecordsAfterTail( + filePath: string, + records: readonly RecordType[], + codec: StrictJsonlCodec, + inWindow: (existing: RecordType) => boolean, + trustedRoot?: string +): void { + if (records.length === 0) return; + const bytes = readStrictJsonlBytes(filePath, codec); + const tail = bytes === undefined ? [] : parseStrictJsonlTail(bytes, codec, inWindow); + const canonicalRecords = canonicalStrictJsonlRecords(records, codec, (index) => `$[new ${String(index)}]`); + validateStrictJsonlHistory([...tail, ...canonicalRecords], codec); + writeStrictJsonlRecords( + filePath, + { exists: bytes !== undefined, byteLength: bytes?.byteLength ?? 0 }, + canonicalRecords, + codec, + trustedRoot + ); +} + +function parseStrictJsonlTail( + bytes: Buffer, + codec: StrictJsonlCodec, + inWindow: (existing: RecordType) => boolean +): RecordType[] { + if (bytes.byteLength === 0) return []; + if (bytes[bytes.byteLength - 1] !== 0x0a) { + throw new Error(`${codec.label} has a torn or unterminated final record`); + } + const tail: RecordType[] = []; + for (let end = bytes.byteLength - 1; end >= 0;) { + const start = end === 0 ? 0 : bytes.lastIndexOf(0x0a, end - 1) + 1; + const lineBytes = bytes.subarray(start, end); + const fromEnd = tail.length + 1; + const record = parseStrictJsonlLine( + lineBytes.toString("utf8"), + lineBytes, + `${String(fromEnd)} from the end`, + `$[-${String(fromEnd)}]`, + codec + ); + tail.unshift(record); + if (!inWindow(record)) break; + end = start - 1; + } + return tail; +} + +function canonicalStrictJsonlRecords( + records: readonly RecordType[], + codec: StrictJsonlCodec, + recordPath: (index: number) => string +): RecordType[] { + return records.map((record, index) => { const serialized = JSON.stringify(record); const parsed = parseStrictJsonBytes(Buffer.from(serialized, "utf8"), { maxBytes: codec.maxRecordBytes ?? DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES, @@ -146,14 +233,17 @@ export function appendStrictJsonlRecords( maxItems: 100_000, maxProperties: 100_000 }); - return codec.parseRecord(parsed, `$[${existing.records.length + index}]`); + return codec.parseRecord(parsed, recordPath(index)); }); - const combined = [...existing.records, ...canonicalRecords]; - if (combined.length > (codec.maxRecords ?? DEFAULT_STRICT_JSONL_MAX_RECORDS)) { - throw new Error(`${codec.label} exceeds the record limit`); - } - validateStrictJsonlHistory(combined, codec); +} +function writeStrictJsonlRecords( + filePath: string, + existing: { exists: boolean; byteLength: number }, + canonicalRecords: readonly RecordType[], + codec: StrictJsonlCodec, + trustedRoot: string | undefined +): number { const payload = Buffer.from(`${canonicalRecords.map((record) => JSON.stringify(record)).join("\n")}\n`, "utf8"); const nextSize = existing.byteLength + payload.byteLength; if (nextSize > (codec.maxBytes ?? DEFAULT_STRICT_JSONL_MAX_BYTES)) { @@ -167,5 +257,5 @@ export function appendStrictJsonlRecords( } else { createFileDurableExclusive(filePath, payload, trustedRoot); } - return readStrictJsonlSnapshot(filePath, codec); + return nextSize; } diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 8b287424e..07956ce2b 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -12,6 +12,7 @@ import { assertArtifactPublicationsContainNoSecrets, appendNodeAttempt, appendEvent, + createEventRecord, assertUsageLedgerEntry, createRunLayout, getNodeArtifactDir, @@ -23,7 +24,6 @@ import { queryNodeAttempts, queryEvents, readArtifactManifest, - readEventQueryFacade, readRunState, replayEvents, replayUsageEvents, @@ -92,12 +92,11 @@ test("createRunLayout persists product-owned run evidence outside checkpoints", layout.statePath, layout.eventsPath, layout.usageLedgerPath, - layout.attemptLedgerPath, - path.join(layout.eventsIndexDir, "query-inputs.json") + layout.attemptLedgerPath ]) { assert.equal(fs.existsSync(expected), true, expected); } - for (const expected of [layout.artifactsDir, layout.reviewDir, layout.eventsIndexDir]) { + for (const expected of [layout.artifactsDir, layout.reviewDir]) { assert.equal(fs.statSync(expected).isDirectory(), true, expected); } assert.equal(path.basename(getNodeArtifactDir(layout, "node-a", { create: true })), "node-a"); @@ -773,7 +772,7 @@ test("artifact manifest reuse checks the complete prerequisite chain", () => { }); }); -test("events append to JSONL, replay, and expose query indexes", () => { +test("events append to JSONL, replay, and filter by node and status", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-1" }); appendEvent(layout, { eventType: "node-synced", @@ -796,90 +795,53 @@ test("events append to JSONL, replay, and expose query indexes", () => { workflow_state: "in-progress" }); assert.equal(queryEvents(layout, { nodeId: "node-a", status: "succeeded" }).length, 1); - assert.equal(fs.existsSync(path.join(layout.eventsIndexDir, "node", "node-a.jsonl")), true); - assert.equal(fs.existsSync(path.join(layout.eventsIndexDir, "status", "succeeded.jsonl")), true); }); -test("event indexes encode long IDs in a collision-free hash namespace", () => { - const maximumRunId = "r".repeat(128); - const layout = createRunLayout({ projectRoot: tempProject(), runId: maximumRunId }); - const facadeBeforeAppend = readEventQueryFacade(layout); - const maximumNodeId = "n".repeat(128); - const directBoundaryNodeId = "d".repeat(122); - const longBoundaryNodeId = "e".repeat(123); - const legacyCollisionNodeId = `${maximumNodeId.slice(0, 97)}-${crypto - .createHash("sha256") - .update(maximumNodeId, "utf8") - .digest("hex") - .slice(0, 24)}`; - assert.equal(legacyCollisionNodeId.length, 122); - - const appendNodeSynced = (nodeId: string, workflowTaskId: string): void => { - appendEvent(layout, { - eventType: "node-synced", - nodeId, - status: "running", - payload: { workflow_run_id: "workflow-1", workflow_task_id: workflowTaskId } - }); - }; - - appendNodeSynced(maximumNodeId, "task-maximum"); - appendNodeSynced(longBoundaryNodeId, "task-long-boundary"); - appendNodeSynced(legacyCollisionNodeId, "task-legacy-collision"); - appendNodeSynced(directBoundaryNodeId, "task-direct-boundary"); - - const hashedIndexPath = (dimension: string, value: string): string => - path.join( - layout.eventsIndexDir, - dimension, - "sha256", - `${crypto.createHash("sha256").update(value, "utf8").digest("hex")}.jsonl` - ); - const maximumRunIndex = hashedIndexPath("run", maximumRunId); - assert.equal(fs.existsSync(maximumRunIndex), true); - assert.equal(fs.existsSync(path.join(layout.eventsIndexDir, "node", `${legacyCollisionNodeId}.jsonl`)), true); - assert.equal(fs.existsSync(path.join(layout.eventsIndexDir, "node", `${directBoundaryNodeId}.jsonl`)), true); - assert.equal(fs.existsSync(hashedIndexPath("node", longBoundaryNodeId)), true); - assert.equal(fs.existsSync(hashedIndexPath("node", maximumNodeId)), true); - - const maximumRecords = fs - .readFileSync(maximumRunIndex, "utf8") - .trimEnd() - .split("\n") - .map((line) => JSON.parse(line) as { run_id: string }); - assert.equal(maximumRecords.length, 4); - assert.equal( - maximumRecords.every((record) => record.run_id === maximumRunId), - true - ); - const facadeAfterAppend = readEventQueryFacade(layout) as { - filters?: unknown; - long_filters?: unknown; - index_key_encoding?: unknown; - }; - assert.deepEqual(facadeAfterAppend, facadeBeforeAppend); - assert.deepEqual(facadeAfterAppend.filters, { - run_id: "events.index/run/.jsonl", - node_id: "events.index/node/.jsonl", - event_type: "events.index/type/.jsonl", - status: "events.index/status/.jsonl", - timestamp: "events.index/timestamp/.jsonl" - }); - assert.deepEqual(facadeAfterAppend.long_filters, { - run_id: "events.index/run/sha256/.jsonl", - node_id: "events.index/node/sha256/.jsonl", - event_type: "events.index/type/sha256/.jsonl", - status: "events.index/status/sha256/.jsonl" +test("event appends and replays keep working past 100,000 records", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-long-journal" }); + const template = createEventRecord(layout, { + eventType: "node-synced", + nodeId: "node-a", + status: "running", + timestamp: "2026-08-05T00:00:00.000Z", + payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-1" } }); - assert.deepEqual(facadeAfterAppend.index_key_encoding, { - version: "ultrafuzz.event-index-key.v1", - direct_max_id_length: 122, - direct_id_path: "/.jsonl", - long_id_path: "/sha256/.jsonl", - digest: "sha256", - hash_input_encoding: "utf8", - digest_encoding: "hex" + const lines = Array.from({ length: 100_000 }, (_, index) => + JSON.stringify({ ...template, event_id: `evt-${index.toString(16).padStart(24, "0")}` }) + ); + fs.writeFileSync(layout.eventsPath, `${lines.join("\n")}\n`); + + const appended = appendEvent(layout, { + eventType: "node-synced", + nodeId: "node-a", + status: "succeeded", + timestamp: "2026-08-05T00:00:01.000Z", + payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-1" } }); + const replay = replayEvents(layout, Number.MAX_SAFE_INTEGER); + assert.equal(replay.records.length, 100_001); + assert.equal(replay.records.at(-1)?.event_id, appended.event_id); +}); + +test("an event append refuses a repeated or out-of-order event without changing the journal", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-event-window" }); + const event = (nodeId: string, timestamp: string) => + ({ + eventType: "node-synced", + nodeId, + status: "succeeded", + timestamp, + payload: { workflow_run_id: "workflow-1", workflow_task_id: `task-${nodeId}` } + }) as const; + appendEvent(layout, event("node-a", "2026-08-05T00:00:01.000Z")); + appendEvent(layout, event("node-b", "2026-08-05T00:00:01.000Z")); + const before = fs.readFileSync(layout.eventsPath); + + // The same event again in the same millisecond, behind a different one, is still a repeat. + assert.throws(() => appendEvent(layout, event("node-a", "2026-08-05T00:00:01.000Z")), /duplicate identity/u); + assert.throws(() => appendEvent(layout, event("node-c", "2026-08-05T00:00:00.000Z")), /not ordered/u); + assert.deepEqual(fs.readFileSync(layout.eventsPath), before); + assert.equal(replayEvents(layout).records.length, 2); }); test("event redaction covers token families, AWS keys, URL credentials, and private keys", () => { diff --git a/packages/artifacts/test/events-tail-repair.test.ts b/packages/artifacts/test/events-tail-repair.test.ts index 0ea0dde9d..b0a38abdb 100644 --- a/packages/artifacts/test/events-tail-repair.test.ts +++ b/packages/artifacts/test/events-tail-repair.test.ts @@ -5,14 +5,7 @@ import os from "node:os"; import path from "node:path"; import test from "node:test"; -import { - appendBytesDurableAt, - appendEvent, - appendEventRecord, - createEventRecord, - createRunLayout, - replayEvents -} from "../src/index.js"; +import { appendBytesDurableAt, appendEvent, createEventRecord, createRunLayout, replayEvents } from "../src/index.js"; function tempProject(): string { return fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-events-tail-")); @@ -28,16 +21,19 @@ test("a torn event journal is rejected without repair or append", () => { payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-1", attempt: 1 } }); fs.writeFileSync(layout.eventsPath, `${JSON.stringify(first)}\n{"event_id":"evt-tor`, "utf8"); - const second = createEventRecord(layout, { - eventType: "node-synced", - nodeId: "node-b", - status: "succeeded", - timestamp: "2026-08-05T00:00:01.000Z", - payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-2", attempt: 2 } - }); const before = fs.readFileSync(layout.eventsPath); - assert.throws(() => appendEventRecord(layout.eventsPath, second), /torn or unterminated/u); + assert.throws( + () => + appendEvent(layout, { + eventType: "node-synced", + nodeId: "node-b", + status: "succeeded", + timestamp: "2026-08-05T00:00:01.000Z", + payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-2", attempt: 2 } + }), + /torn or unterminated/u + ); assert.deepEqual(fs.readFileSync(layout.eventsPath), before); assert.throws(() => replayEvents(layout), /torn or unterminated/u); }); @@ -51,33 +47,9 @@ test("a complete but unterminated trailing object is rejected without repair", ( timestamp: "2026-08-05T00:00:00.000Z", payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-1" } }); - const second = createEventRecord(layout, { - eventType: "node-synced", - nodeId: "node-b", - status: "succeeded", - timestamp: "2026-08-05T00:00:01.000Z", - payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-2" } - }); fs.writeFileSync(layout.eventsPath, JSON.stringify(first), "utf8"); const before = fs.readFileSync(layout.eventsPath); - assert.throws(() => appendEventRecord(layout.eventsPath, second), /torn or unterminated/u); - assert.deepEqual(fs.readFileSync(layout.eventsPath), before); -}); - -test("a torn event index rejects the whole append before the canonical journal changes", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-index-torn" }); - const first = appendEvent(layout, { - eventType: "node-synced", - nodeId: "node-a", - status: "succeeded", - timestamp: "2026-08-05T00:00:00.000Z", - payload: { workflow_run_id: "workflow-1", workflow_task_id: "task-1" } - }); - const indexPath = path.join(layout.eventsIndexDir, "run", `${layout.runId}.jsonl`); - fs.appendFileSync(indexPath, '{"event_id":"evt-tor', "utf8"); - - const canonicalBefore = fs.readFileSync(layout.eventsPath); assert.throws( () => appendEvent(layout, { @@ -89,8 +61,7 @@ test("a torn event index rejects the whole append before the canonical journal c }), /torn or unterminated/u ); - assert.deepEqual(fs.readFileSync(layout.eventsPath), canonicalBefore); - assert.equal(replayEvents(layout).records[0]?.event_id, first.event_id); + assert.deepEqual(fs.readFileSync(layout.eventsPath), before); }); test("durable appends reject stale sizes and hard-linked files", () => { diff --git a/packages/cli/src/commands/report/bundle.ts b/packages/cli/src/commands/report/bundle.ts index a76fdf005..e0cff4858 100644 --- a/packages/cli/src/commands/report/bundle.ts +++ b/packages/cli/src/commands/report/bundle.ts @@ -5,18 +5,12 @@ import path from "node:path"; import { Args, Command, Flags } from "@oclif/core"; import { DEFAULT_STRICT_JSONL_MAX_BYTES, - DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES, - DEFAULT_STRICT_JSONL_MAX_RECORDS, - assertEventRecord, assertNoSymlinkComponents, assertPathInside, assertRegularFileInside, layoutForRunRoot, - parseStrictJsonBytes, + parseEventJournalBytes, readRegularFileSnapshot, - validateStrictJsonlHistory, - type EventRecord, - type StrictJsonlCodec, validateSafeId } from "@ultrafuzz/artifacts"; import { @@ -385,7 +379,7 @@ function loadValidatedEventJournalSnapshot( assertRegularFileInside(runRoot, eventsPath, "event journal"); const contents = readRegularFileSnapshot(eventsPath, DEFAULT_STRICT_JSONL_MAX_BYTES); - validateEventJournalSnapshot(contents, expectedRunId); + parseEventJournalBytes(contents, expectedRunId); return { absolutePath: eventsPath, archivePath: "events.jsonl", @@ -393,83 +387,6 @@ function loadValidatedEventJournalSnapshot( }; } -function validateEventJournalSnapshot(contents: Buffer, expectedRunId: string): void { - if (contents.byteLength === 0) return; - if (contents[contents.byteLength - 1] !== 0x0a) { - throw new Error("event journal has a torn or unterminated final record"); - } - - let text: string; - try { - text = new TextDecoder("utf-8", { fatal: true }).decode(contents); - } catch (error) { - throw new Error("event journal is not valid UTF-8", { cause: error }); - } - const lines = text.split("\n"); - lines.pop(); - if (lines.length > DEFAULT_STRICT_JSONL_MAX_RECORDS) { - throw new Error(`event journal exceeds the ${DEFAULT_STRICT_JSONL_MAX_RECORDS}-record limit`); - } - - const codec = eventJournalCodec(expectedRunId); - const records: EventRecord[] = []; - for (const [index, line] of lines.entries()) { - const lineNumber = index + 1; - if (line.trim().length === 0) throw new Error(`event journal contains a blank record at line ${lineNumber}`); - const lineBytes = Buffer.from(line, "utf8"); - if (lineBytes.byteLength > DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES) { - throw new Error( - `event journal record ${lineNumber} exceeds the ${DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES}-byte limit` - ); - } - let parsed: unknown; - try { - parsed = parseStrictJsonBytes(lineBytes, { - maxBytes: DEFAULT_STRICT_JSONL_MAX_RECORD_BYTES, - maxDepth: 128, - maxItems: 100_000, - maxProperties: 100_000 - }); - } catch (error) { - throw new Error( - `event journal record ${lineNumber} is invalid strict JSON: ${error instanceof Error ? error.message : String(error)}`, - { cause: error } - ); - } - records.push(codec.parseRecord(parsed, `$[${index}]`)); - } - validateStrictJsonlHistory(records, codec); -} - -function eventJournalCodec(expectedRunId: string): StrictJsonlCodec { - return { - label: "event journal", - parseRecord: (value, recordPath) => { - const record = assertEventRecord(value, recordPath); - if (record.run_id !== expectedRunId) { - throw new Error( - `${recordPath}.run_id belongs to ${JSON.stringify(record.run_id)}, expected ${JSON.stringify(expectedRunId)}` - ); - } - return record; - }, - identity: (record) => record.event_id, - validateHistory: (records) => { - const firstRunId = records[0]?.run_id; - let priorTimestamp = records[0]?.timestamp; - for (const [index, record] of records.entries()) { - if (firstRunId !== undefined && record.run_id !== firstRunId) { - throw new Error(`event journal changes run_id at record ${index + 1}`); - } - if (priorTimestamp !== undefined && record.timestamp < priorTimestamp) { - throw new Error(`event journal timestamps are not ordered at record ${index + 1}`); - } - priorTimestamp = record.timestamp; - } - } - }; -} - function collectBundleFiles( runRoot: string, diagnostics: RuntimeDiagnostic[], From eaa1ed5987eb51dfb5f144818c3bc0fcc1e3ed7f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:48:26 +0000 Subject: [PATCH 056/206] refactor(artifacts): findings and property validators return the Ajv verdict validateFindingSchema, validateFindingsSchema and the five property validators ran the registered JSON Schema (documented as authoritative), then parsed the same document again with its Zod schema and threw an internal "schema parity invariant violated" error if Zod rejected or transformed it. The Zod result was discarded either way. Those validators back workflow-sync and several artifact gates, so a disagreement would have aborted sync or verification with an internal error instead of returning a validation result. Return the Ajv result directly. The Zod schemas still generate the JSON Schema and the TypeScript types, and contract-fixtures.test.ts still checks that Ajv and Zod agree on every fixture, so a drift fails CI instead of a campaign. No test here discriminates: no document is known on which the two validators disagree. The contract fixtures pin their agreement, and a probe of timestamp, integer and string edge cases found none. On such documents the change only removes the second parse. Refs #462 Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/src/findings-schema.ts | 32 ++------- packages/artifacts/src/property-provenance.ts | 68 ++++++------------- 2 files changed, 26 insertions(+), 74 deletions(-) diff --git a/packages/artifacts/src/findings-schema.ts b/packages/artifacts/src/findings-schema.ts index e1c5522ba..e4f45991a 100644 --- a/packages/artifacts/src/findings-schema.ts +++ b/packages/artifacts/src/findings-schema.ts @@ -1,5 +1,3 @@ -import { isDeepStrictEqual } from "node:util"; - import { z } from "zod/v4"; import { @@ -25,7 +23,7 @@ import { import { validateRegisteredJsonSchema } from "./json-schema-validator.js"; import { hasAtMostCodePoints } from "./portable-json-primitives.js"; import { NODE_REFERENCE_PATTERN } from "./safe-paths.js"; -import { validateWithZod, type SchemaValidationResult } from "./schema-validation.js"; +import { type SchemaValidationResult } from "./schema-validation.js"; import { jsonPointerPath } from "./lang-primitives.js"; export const FINDING_JSON_SCHEMA_ID = "urn:ultrafuzz:schema:artifacts:finding:2" as const; @@ -992,24 +990,22 @@ export const findingsJsonSchema = { } as const; export function validateFindingSchema(value: unknown, path = "$"): SchemaValidationResult { - return validateRegisteredFindingSchema(FINDING_JSON_SCHEMA_ID, findingSchema as z.ZodType, value, { + return validateRegisteredFindingSchema(FINDING_JSON_SCHEMA_ID, value, { path, code: "FINDING_SCHEMA_INVALID" }); } export function validateFindingsSchema(value: unknown, path = "$"): SchemaValidationResult { - return validateRegisteredFindingSchema( - FINDINGS_JSON_SCHEMA_ID, - findingsSchema as z.ZodType, - value, - { path, code: "FINDINGS_SCHEMA_INVALID" } - ); + return validateRegisteredFindingSchema(FINDINGS_JSON_SCHEMA_ID, value, { + path, + code: "FINDINGS_SCHEMA_INVALID" + }); } +/** The registered JSON Schema is authoritative, so its verdict is returned as is. */ function validateRegisteredFindingSchema( schemaId: string, - zodSchema: z.ZodType, value: unknown, options: { path: string; code: string } ): SchemaValidationResult { @@ -1024,19 +1020,5 @@ function validateRegisteredFindingSchema( })) }; } - - // The checked-in JSON Schema is authoritative. Zod remains only as a - // non-transforming parity assertion for typed access by existing callers. - const parity = validateWithZod(zodSchema, value, options); - if (!parity.ok) { - throw new Error( - `internal schema parity invariant violated: registered JSON Schema ${schemaId} accepted a document rejected by its retained Zod parser` - ); - } - if (!isDeepStrictEqual(parity.value, value)) { - throw new Error( - `internal schema parity invariant violated: retained Zod parser for registered JSON Schema ${schemaId} transformed its input` - ); - } return { ok: true, issues: [], value: value as T }; } diff --git a/packages/artifacts/src/property-provenance.ts b/packages/artifacts/src/property-provenance.ts index 90e9607a6..5a7ee1dab 100644 --- a/packages/artifacts/src/property-provenance.ts +++ b/packages/artifacts/src/property-provenance.ts @@ -3,7 +3,7 @@ import { z } from "zod/v4"; import { canonicalArtifactRelativePathSchema } from "./artifact-path-primitives.js"; import { validateRegisteredJsonSchema } from "./json-schema-validator.js"; import { canonicalTimestampSchema, hasAtMostCodePoints } from "./portable-json-primitives.js"; -import { validateWithZod, type SchemaValidationIssue, type SchemaValidationResult } from "./schema-validation.js"; +import { type SchemaValidationIssue, type SchemaValidationResult } from "./schema-validation.js"; import { jsonPointerPath } from "./lang-primitives.js"; export const PROPERTIES_SCHEMA_VERSION = "ultrafuzz.properties.v2" as const; @@ -1040,20 +1040,15 @@ export function validateLensPropertiesSchema( value: unknown, path = "$" ): SchemaValidationResult { - return validateRegisteredPropertySchema( - PROPERTY_LENS_JSON_SCHEMA_ID, - lensPropertiesSchema as z.ZodType, - value, - { - path, - code: "PROPERTY_LENS_SCHEMA_INVALID" - } - ); + return validateRegisteredPropertySchema(PROPERTY_LENS_JSON_SCHEMA_ID, value, { + path, + code: "PROPERTY_LENS_SCHEMA_INVALID" + }); } +/** The registered JSON Schema is authoritative, so its verdict is returned as is. */ function validateRegisteredPropertySchema( schemaId: string, - zodSchema: z.ZodType, value: unknown, options: { path: string; code: string } ): SchemaValidationResult { @@ -1068,15 +1063,6 @@ function validateRegisteredPropertySchema( })) }; } - - // The checked-in JSON Schema is authoritative. Zod remains only as a - // non-transforming parity assertion for typed access by existing callers. - const parity = validateWithZod(zodSchema, value, options); - if (!parity.ok) { - throw new Error( - `internal schema parity invariant violated: registered JSON Schema ${schemaId} accepted a document rejected by its retained Zod parser` - ); - } return { ok: true, issues: [], value: value as T }; } @@ -1084,27 +1070,17 @@ export function validateReferenceExpectationsSchema( value: unknown, path = "$" ): SchemaValidationResult { - return validateRegisteredPropertySchema( - REFERENCE_EXPECTATIONS_JSON_SCHEMA_ID, - referenceExpectationsSchema as z.ZodType, - value, - { - path, - code: "REFERENCE_EXPECTATIONS_SCHEMA_INVALID" - } - ); + return validateRegisteredPropertySchema(REFERENCE_EXPECTATIONS_JSON_SCHEMA_ID, value, { + path, + code: "REFERENCE_EXPECTATIONS_SCHEMA_INVALID" + }); } export function validatePropertiesSchema(value: unknown, path = "$"): SchemaValidationResult { - return validateRegisteredPropertySchema( - PROPERTIES_JSON_SCHEMA_ID, - propertiesSchema as z.ZodType, - value, - { - path, - code: "PROPERTIES_SCHEMA_INVALID" - } - ); + return validateRegisteredPropertySchema(PROPERTIES_JSON_SCHEMA_ID, value, { + path, + code: "PROPERTIES_SCHEMA_INVALID" + }); } export function validateImplementedPropertiesSchema( @@ -1112,9 +1088,8 @@ export function validateImplementedPropertiesSchema( path = "$", options: { requireSelection?: boolean } = {} ): SchemaValidationResult { - const result = validateRegisteredPropertySchema( + const result = validateRegisteredPropertySchema( IMPLEMENTED_PROPERTIES_JSON_SCHEMA_ID, - implementedPropertiesSchema as z.ZodType, value, { path, @@ -1140,15 +1115,10 @@ export function validatePropertyCampaignSchema( value: unknown, path = "$" ): SchemaValidationResult { - return validateRegisteredPropertySchema( - PROPERTY_CAMPAIGN_JSON_SCHEMA_ID, - propertyCampaignSchema as z.ZodType, - value, - { - path, - code: "PROPERTY_CAMPAIGN_SCHEMA_INVALID" - } - ); + return validateRegisteredPropertySchema(PROPERTY_CAMPAIGN_JSON_SCHEMA_ID, value, { + path, + code: "PROPERTY_CAMPAIGN_SCHEMA_INVALID" + }); } export function validatePropertyReferences( From 15944d038e78f7d3b8887bee05dbb3633fe8564a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:50:00 +0000 Subject: [PATCH 057/206] docs(changelog): note the artifacts dead-code and event-journal changes Refs #462 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..c0d40d502 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[artifacts] [security] [cli] [docs]** Deletes the 48 artifact semantic gates that no production code dispatches (the planned-graph ones duplicated `assertPlannedGraphSemantics`), and production-dead code such as `finding-provenance.ts`, the generated-test manifest writer and `resolvePathInside`. Schema bytes, contract digests and the validator build identity are unchanged. Event appends now write only `events.jsonl` and check the new event against the final event and any trailing events with the same timestamp, instead of re-parsing the whole journal and up to five unread `events.index/` files; the journal no longer stops accepting events at 100,000 records (its 64 MiB byte limit remains), and new runs have no `events.index/` directory. The findings and property validators return the registered JSON Schema verdict instead of throwing when their Zod parser disagrees (#462). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From ebe5ef6321995796774581859a7468f3d76bc0a8 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:55:11 +0000 Subject: [PATCH 058/206] test(runtime): pin host aggregation attribution for model-fanout producers The sealed-authority aggregation context attributes each source bundle to its producer's planned loop attempt index, which is what the in-workflow verifier uses (dependencyTask.metadata.loop.attemptIndex) and what the task manifest validator binds to the planned node. The previous host walk used the model-fanout attempt index instead, so for a two-profile producer the host expected attempt_index 1 for the second model while the verifier required 0, and any manifest one gate accepted the other rejected. On origin/main this test fails with "Aggregation source bundle attribution does not match authority" for strategy-a__model_1__attempt_1; it passes with the sealed-authority context. Destination paths in the fixture manifest now include the attempt ID so two attempts of one logical node do not collide. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/test/artifact-gates.test.ts | 66 +++++++++++++++++++- 1 file changed, 65 insertions(+), 1 deletion(-) diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 67cbf706c..226522f22 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -3822,7 +3822,7 @@ function writeAggregationManifestCopying( source_manifest_sha256: source.manifestSha256 }); const files = sources.map((source) => { - const destinationRelativePath = `test/foundry/${source.logicalNodeId}/attempt-0/${path.basename(source.testPath)}`; + const destinationRelativePath = `test/foundry/${source.attemptId}/${path.basename(source.testPath)}`; const destinationPath = path.join(workspace, ...destinationRelativePath.split("/")); fs.mkdirSync(path.dirname(destinationPath), { recursive: true }); fs.writeFileSync(destinationPath, source.testBytes); @@ -3951,6 +3951,70 @@ test("host aggregation intake keys a directly consumed dynamic producer by its s assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); }); +test("host aggregation intake attributes model-fanout bundles to the loop attempt index the verifier uses", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-aggregation-model-fanout" }); + const fanout: PlannedGraphNode = { + ...aggregationFixtureNode("strategy-a", { group: "strategies" }), + model_fanout: [0, 1].map((modelIndex) => ({ + model_profile_id: `model-${modelIndex}`, + agent_ref: "CodexAgent", + model_name: "gpt-test", + reasoning_effort: "high", + model_index: modelIndex, + loop_index: 0, + attempt_index: modelIndex + })) + }; + const aggregationNode = aggregationFixtureNode("aggregate-test-files", { + group: "review", + dependsOn: [fanout.id], + outputs: [boundOutput("aggregation.json", "ultrafuzz/aggregation-manifest@1", true)] + }); + const nodes = [fanout, aggregationNode]; + writePlannedGraph(layout, nodes, AGGREGATION_FIXTURE_GROUPS); + const fanoutTasks = [0, 1].map((modelIndex) => + smithersTaskForNode({ + layout, + node: fanout, + attemptId: `${fanout.id}__model_${modelIndex}__attempt_${modelIndex}`, + modelIndex + }) + ); + const aggregationTask: SmithersTaskManifestTask = { + ...sealedTaskForNode(layout, aggregationNode, fanoutTasks), + optionalDependencyArtifactDirs: fanoutTasks.map((task) => task.artifactDir) + }; + const tasks = [...fanoutTasks, aggregationTask]; + const sources = fanoutTasks.map((task) => { + const source = finalizeGeneratedTestsProducer(layout, fanout, task.attemptId); + // The controller records the model attempt index (here 0 and 1) in each + // producer manifest; the verifier attributes both bundles to loop attempt 0. + const manifestPath = path.join(task.artifactDir, "artifact-manifest.json"); + const manifest = JSON.parse(fs.readFileSync(manifestPath, "utf8")) as { provenance: { attempt_index?: number } }; + manifest.provenance.attempt_index = task.metadata.model.attemptIndex; + fs.writeFileSync(manifestPath, JSON.stringify(manifest)); + const state = readRunState(layout); + const provenance = state.nodes[task.attemptId]?.provenance; + assert.ok(provenance !== undefined && "output_contracts" in provenance && provenance.output_contracts); + provenance.output_contracts.artifact_manifest_sha256 = createHash("sha256") + .update(fs.readFileSync(manifestPath)) + .digest("hex"); + writeRunState(layout, state); + return source; + }); + writeSealedFixtureTaskAuthority(layout, nodes, tasks); + + writeAggregationManifestCopying(layout, aggregationNode, sources); + const result = verifyAggregationAttempt( + layout, + aggregationTask, + tasks, + fanoutTasks.map((task) => task.attemptId) + ); + + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); +}); + test("host aggregation intake follows the sealed closure below a dynamic group's direct dependent", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-aggregation-transitive-dynamic" }); const source = aggregationFixtureNode("goal-plan", { From d0091aba21e0b0b5db4d5303ebc4de5d69ed787c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:55:12 +0000 Subject: [PATCH 059/206] docs(changelog): record the host artifact gate false-positive fixes Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..b0c691aed 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [docs]** Host artifact gates stop rejecting verifier-accepted output in several cases. Aggregation builds its sources from the attempt's sealed ancestor closure and verifier admission, so a failed `failure_policy: continue` generated-tests producer is skipped, a directly consumed dynamic goal producer is read under its storage attempt ID, and model-fanout bundles carry the loop attempt index the verifier expects; nodes below a dynamic group's direct dependent (triage and the later review nodes in the default topology) no longer fail the sealed-closure check for omitting generated attempts that dynamic lowering never adds to their closure. The host-only severity-matrix gate, which rejected schema-valid non-promoted records that omit `severity`, `impact`, and `likelihood`, is deleted; the `severity-classification-matrix` registry gate still checks records that carry them, but report rows authored in bounded classification mode (for example in the smoke topology) are no longer matrix-checked. Coverage percentages and fractions in report and coverage prose that name no exact declaration-completeness scope are now warnings instead of errors. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From a037bc9f4f5318e7b70d06e6c6a71e3bbebd1433 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:01:27 +0000 Subject: [PATCH 060/206] fix(runtime): a no-op resume attach leaves run state alone and says so `resume` treats a run whose derived Smithers state is still active as an idempotent attach and starts no controller. That includes a run still draining a pause and, for up to the 30 s heartbeat window, a run whose controller just died. submitSmithersContinuation nevertheless called recordNativeContinuationState, which rewrote state.json and moved workflow_deadline_at forward on every such no-op, and the CLI printed "Submitted resume" regardless of `submitted`. Only record continuation state when a controller was actually started, and drop the now-unused alreadyRunning parameter and branch. The CLI now prints "Run already active: ; no new controller was started" when nothing was submitted, with the hint to resume again once a draining pause has parked. JSON output is unchanged. #1153 asked for an Ultrafuzz-side controller-generation fence. The pinned engine already provides the guarantee, so this pins it instead of re-implementing it. A real-Smithers integration test shows that: - a graceful pause lets held in-flight tasks finish; - while the owner drains, `up --resume --force` and `timetravel --force` are refused with RUN_OWNER_ALIVE, and inspect still reports the run state as `running`, so `ultrafuzz resume` only attaches; - a resume after the park runs only the remaining task, in a new controller. Swapping --force for --steal-ownership in that test makes it fail, so it detects a takeover. The state test fails on main (workflow_deadline_at moves on a submitted:false resume) and the CLI test fails on main ("Submitted resume: ..."); both pass with this change. The pin test passes on both by design. Co-Authored-By: Claude Opus 5.5 --- docs/reference/cli.md | 17 +- docs/reference/configuration.md | 3 +- packages/cli/src/commands/resume.ts | 6 +- packages/cli/test/cli.test.ts | 20 +++ packages/runtime/src/start-run.ts | 36 ++--- packages/runtime/test/runtime.test.ts | 6 + ...thers-preparation-race.integration.test.ts | 148 +++++++++++++++++- 7 files changed, 204 insertions(+), 32 deletions(-) diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 68bd3bd33..a9ac9e2dc 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -459,13 +459,16 @@ an available agent-written report. `pause` requests a graceful stop: no new tasks are scheduled, in-flight tasks finish, and the run settles in the resumable `paused` state. `resume` reports -`submitted: false` instead of launching a duplicate continuation when the linked -workflow is still in an active state (running, in-progress, started, queued, -retrying, or waiting). `resume --reset-node` retries one failed workflow node and -its dependents in the same linked run; the applied reset is recorded so retrying -the command after a failed continuation resumes the already-reset run instead of -repeating the reset. `fork` may start from a checkpoint frame and may reset one -workflow node before starting the fork. +`submitted: false` (text output `Run already active`) instead of launching a +duplicate continuation when the linked workflow is still active (its Smithers +run state is `running`, `recovering`, or one of the `waiting-*` states), and +leaves the run's recorded state and workflow deadline unchanged. A run still +finishing its in-flight tasks after `pause` is still active; resume it again +once `status` reports `paused`. `resume --reset-node` retries one failed +workflow node and its dependents in the same linked run; the applied reset is +recorded so retrying the command after a failed continuation resumes the +already-reset run instead of repeating the reset. `fork` may start from a +checkpoint frame and may reset one workflow node before starting the fork. Every command in this section takes an Ultrafuzz run ID and resolves the linked workflow run from existing product evidence; none of them require the diff --git a/docs/reference/configuration.md b/docs/reference/configuration.md index 1095c7155..1f5f87da9 100644 --- a/docs/reference/configuration.md +++ b/docs/reference/configuration.md @@ -139,7 +139,8 @@ resume. `workflow_deadline_seconds` is not a guaranteed wall-clock limit. Ultrafuzz records `workflow_deadline_at` in run state when the run is created, and again -from each `resume`, `replay`, or `fork`, but nothing enforces it on a timer: an +from each `replay`, `fork`, or `resume` that starts a controller (not one that +finds the run still active), but nothing enforces it on a timer: an unattended run keeps executing, and incurring provider cost, past its deadline. The deadline is checked only when a command synchronizes the run: `ultrafuzz status` (including each `--watch` poll), `inspect`, `why`, and `stats`, plus diff --git a/packages/cli/src/commands/resume.ts b/packages/cli/src/commands/resume.ts index 2f3bb79da..e172870b7 100644 --- a/packages/cli/src/commands/resume.ts +++ b/packages/cli/src/commands/resume.ts @@ -40,7 +40,11 @@ export default class Resume extends Command { emitCommandResult( this, "resume", - commandFromRuntime("resume", result, (value) => `Submitted ${value.action}: ${value.workflow_run_id}\n`), + commandFromRuntime("resume", result, (value) => + value.submitted + ? `Submitted ${value.action}: ${value.workflow_run_id}\n` + : `Run already active: ${value.workflow_run_id}; no new controller was started. If a pause is still draining, resume again once status reports paused.\n` + ), flags.json === true ); } diff --git a/packages/cli/test/cli.test.ts b/packages/cli/test/cli.test.ts index 18676ca3a..750066e23 100644 --- a/packages/cli/test/cli.test.ts +++ b/packages/cli/test/cli.test.ts @@ -1918,6 +1918,26 @@ test("status surfaces a terminal product and live workflow lifecycle divergence" ); }); +test("resume of an already-active run says no controller was started instead of claiming a submission", async () => { + const project = tempProject(); + const env = fakeSmithersEnv(project); + assert.equal((await cli(project, ["init", "--json"], env)).code, 0); + writeSmallTopology(project); + const runId = "resume-already-active"; + const run = await cli(project, ["run", "--run-id", runId, "--json"], env); + assert.equal(run.code, 0, run.stderr); + + // The fake runner still reports the run as running, so resume only attaches. + const resumed = await cli(project, ["resume", runId], env); + + assert.equal(resumed.code, 0, `${resumed.stderr}\n${resumed.stdout}`); + assert.match( + resumed.stdout, + /^Run already active: ultrafuzz-resume-already-active; no new controller was started\./mu + ); + assert.doesNotMatch(resumed.stdout, /Submitted/u); +}); + test("status --watch stops immediately on a degraded verdict even while product state is nonterminal", async () => { const project = tempProject(); const env = fakeSmithersEnv(project); diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index 2e8f44da9..9d92cfcc5 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -672,12 +672,11 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { trustedCli.environmentVariableNames ) }); - recordNativeContinuationState({ - layout, - config, - requestedConcurrency: input.maxConcurrency, - alreadyRunning: result.alreadyRunning ?? false - }); + // Attaching to an already-active run started no controller. Its state + // (status, lease, deadline) belongs to the live owner, so leave it alone. + if (result.alreadyRunning !== true) { + recordNativeContinuationState({ layout, config, requestedConcurrency: input.maxConcurrency }); + } return runtimeResult(true, { run_id: runId, workflow_run_id: smithersRunId, @@ -743,7 +742,6 @@ function recordNativeContinuationState(input: { layout: RunLayout; config: ResolvedConfig | undefined; requestedConcurrency: number | undefined; - alreadyRunning: boolean; }): void { try { const submittedAt = new Date().toISOString(); @@ -753,19 +751,17 @@ function recordNativeContinuationState(input: { state.status = "running"; state.started_at ??= submittedAt; delete state.finished_at; - if (!input.alreadyRunning) { - const leaseDurationMs = - (input.config?.run.controllerLeaseSeconds ?? Math.max(1, state.controller_lease.duration_ms / 1_000)) * 1_000; - state.controller_lease = { - ...state.controller_lease, - status: "active", - duration_ms: leaseDurationMs, - renewed_at: submittedAt, - expires_at: new Date(submittedAtMs + leaseDurationMs).toISOString() - }; - state.concurrency.requested_concurrency = - input.requestedConcurrency ?? input.config?.run.maxParallelAgents ?? state.concurrency.requested_concurrency; - } + const leaseDurationMs = + (input.config?.run.controllerLeaseSeconds ?? Math.max(1, state.controller_lease.duration_ms / 1_000)) * 1_000; + state.controller_lease = { + ...state.controller_lease, + status: "active", + duration_ms: leaseDurationMs, + renewed_at: submittedAt, + expires_at: new Date(submittedAtMs + leaseDurationMs).toISOString() + }; + state.concurrency.requested_concurrency = + input.requestedConcurrency ?? input.config?.run.maxParallelAgents ?? state.concurrency.requested_concurrency; if (input.config !== undefined) { state.workflow_deadline_at = new Date( submittedAtMs + input.config.run.workflowDeadlineSeconds * 1_000 diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 5fd3a3f9f..2394621f2 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -26463,6 +26463,11 @@ test("ordinary resume checks active-run ownership before detached preflight", as }); const run = await startRun({ projectRoot: project, runId: "active-lifecycle-run", env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + // An attach starts no controller, so the run's state (status, lease and + // workflow deadline) must stay exactly as its live owner left it (#1153). + assert.ok(run.value); + const statePath = path.join(run.value.run_root, "state.json"); + const stateBefore = fs.readFileSync(statePath, "utf8"); fs.writeFileSync(env.SMITHERS_FAKE_LOG!, "", "utf8"); // A duplicate `up --resume --detach` renders the workflow before Smithers // checks ownership. Keep that path fatal so this regression proves active @@ -26492,6 +26497,7 @@ test("ordinary resume checks active-run ownership before detached preflight", as const forcedCommands = fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"); assert.match(forcedCommands, /inspect ultrafuzz-active-lifecycle-run --format json --full-output/u); assert.doesNotMatch(forcedCommands, /^up /mu); + assert.equal(fs.readFileSync(statePath, "utf8"), stateBefore); }); test("resume derives reset identities from the canonical nodes of a failed workflow", async () => { diff --git a/packages/runtime/test/smithers-preparation-race.integration.test.ts b/packages/runtime/test/smithers-preparation-race.integration.test.ts index 204af09e9..ad778d442 100644 --- a/packages/runtime/test/smithers-preparation-race.integration.test.ts +++ b/packages/runtime/test/smithers-preparation-race.integration.test.ts @@ -1,6 +1,6 @@ import assert from "node:assert/strict"; import { temporaryRoot } from "./temporary-root.js"; -import { execFileSync } from "node:child_process"; +import { execFileSync, spawnSync } from "node:child_process"; import fs from "node:fs"; import path from "node:path"; import test from "node:test"; @@ -226,6 +226,96 @@ process.stdout.write(JSON.stringify(result));` } }); +// #1153 asked Ultrafuzz for its own controller-generation fence around pause +// and continuation handoff. The pinned engine already provides the guarantee: +// a graceful pause lets in-flight tasks finish instead of aborting them, a +// second controller is refused while the owner is alive (`--force` does not +// take ownership), and a resume after the park runs only the remaining work. +// Pin those facts so an engine bump that regresses them fails here. +test("graceful pause drains in-flight tasks and refuses a second controller until the run parks", async () => { + const root = temporaryRoot("ultrafuzz-smithers-pause-handoff-"); + const workflowDir = path.join(root, ".smithers", "workflows"); + const workflowPath = path.join(workflowDir, "pause-handoff.tsx"); + const traceLog = path.join(root, "trace.log"); + const releasePath = path.join(root, "release-in-flight"); + const runId = `pause-handoff-${process.pid}-${Date.now()}`; + const runner = (args: string[]) => + spawnSync(smithersBinary(), args, { + cwd: root, + encoding: "utf8", + env: { ...process.env, SMITHERS_POST_FAILURE: "0" } + }); + const trace = (): string[] => + fs.existsSync(traceLog) ? fs.readFileSync(traceLog, "utf8").trim().split("\n").filter(Boolean) : []; + const pidOf = (lines: readonly string[], event: string) => + lines.find((line) => line.endsWith(` ${event}`))?.split(" ")[0]; + + try { + fs.mkdirSync(workflowDir, { recursive: true }); + initFixtureRepository(root); + const smithersPackageRoot = fs.realpathSync(path.join(runtimePackageRoot(), "node_modules", "smthrs")); + fs.symlinkSync(path.dirname(smithersPackageRoot), path.join(root, ".smithers", "node_modules"), "dir"); + fs.writeFileSync(workflowPath, pauseHandoffWorkflowSource({ traceLog, releasePath }), "utf8"); + + const launched = runner([ + "up", + workflowPath, + "--detach", + "--run-id", + runId, + "--root", + root, + "--input", + "{}", + "--format", + "json" + ]); + assert.equal(launched.status, 0, launched.stderr); + await waitUntil(() => trace().filter((line) => line.endsWith(" start")).length === 2, 60_000, "a and b start"); + + const pause = runner(["pause", runId, "--format", "json"]); + assert.match(pause.stdout, /"pause-requested"/u, pause.stderr); + // Both in-flight tasks are held open, so the owner is still draining: a + // replacement controller and a node reset are refused despite `--force`. + for (const takeover of [ + ["up", workflowPath, "--resume", runId, "--run-id", runId, "--force", "--detach", "--format", "json"], + ["timetravel", workflowPath, "--run-id", runId, "--node-id", "a", "--no-vcs", "--force", "--format", "json"] + ]) { + const refused = runner(takeover); + assert.notEqual(refused.status, 0, takeover.join(" ")); + assert.match(refused.stdout + refused.stderr, /RUN_OWNER_ALIVE/u, takeover.join(" ")); + } + // The draining run still reports an active state, so `ultrafuzz resume` + // only attaches to it instead of starting a controller. + const draining = JSON.parse(runner(["inspect", runId, "--format", "json", "--full-output"]).stdout) as { + data?: { runState?: { state?: string } }; + }; + assert.equal(draining.data?.runState?.state, "running"); + assert.equal(trace().length, 2, "no task ended or started while the pause drained"); + + fs.writeFileSync(releasePath, "", "utf8"); + await waitForStatus(root, runId, "paused", 60_000); + const parked = trace(); + const ownerPid = pidOf(parked, "a start"); + assert.equal(pidOf(parked, "a end"), ownerPid, "the draining owner finished a"); + assert.equal(pidOf(parked, "b end"), ownerPid, "the draining owner finished b"); + assert.equal(pidOf(parked, "c start"), undefined, "the pause stopped new scheduling"); + + const resumed = runner(["up", workflowPath, "--resume", runId, "--run-id", runId, "--detach", "--format", "json"]); + assert.equal(resumed.status, 0, resumed.stderr); + await waitForSuccessfulCompletion(root, runId, 60_000); + const finished = trace(); + for (const event of ["a start", "b start", "c start"]) { + const runs = finished.filter((line) => line.endsWith(` ${event}`)).length; + assert.equal(runs, 1, `${event} ran ${runs} times: ${JSON.stringify(finished)}`); + } + assert.notEqual(pidOf(finished, "c start"), ownerPid, "only the replacement controller ran c"); + } finally { + fs.writeFileSync(releasePath, "", "utf8"); + fs.rmSync(root, { recursive: true, force: true }); + } +}); + function initFixtureRepository(root: string): void { execGit(root, ["init", "--quiet", "--initial-branch=main"]); execGit(root, ["config", "user.name", "Ultrafuzz Synthetic Test"]); @@ -368,6 +458,45 @@ export default smithers((ctx) => ( `; } +function pauseHandoffWorkflowSource(input: { traceLog: string; releasePath: string }): string { + return `/** @jsxImportSource smthrs */ +import fs from "node:fs"; +import { createSmithers } from "smthrs"; +import { z } from "zod/v4"; + +const traceLog = ${JSON.stringify(input.traceLog)}; +const releasePath = ${JSON.stringify(input.releasePath)}; +const trace = (event) => fs.appendFileSync(traceLog, process.pid + " " + event + "\\n", "utf8"); +const { Workflow, Task, Parallel, Sequence, smithers, outputs } = createSmithers({ + input: z.object({}), + step: z.object({ done: z.literal(true) }) +}); +// Hold each in-flight task until the test releases it, bounded so a failed +// test cannot leave the detached engine polling forever. +const held = (id) => async () => { + trace(id + " start"); + const deadline = Date.now() + 120000; + while (!fs.existsSync(releasePath) && Date.now() < deadline) { + await new Promise((resolve) => setTimeout(resolve, 50)); + } + trace(id + " end"); + return { done: true }; +}; + +export default smithers(() => ( + + + + {held("a")} + {held("b")} + + {() => (trace("c start"), { done: true })} + + +)); +`; +} + async function waitForFile(filePath: string, timeoutMs: number): Promise { const deadline = Date.now() + timeoutMs; while (Date.now() < deadline) { @@ -377,7 +506,20 @@ async function waitForFile(filePath: string, timeoutMs: number): Promise { throw new Error(`timed out waiting for ${filePath}`); } +async function waitUntil(condition: () => boolean, timeoutMs: number, label: string): Promise { + const deadline = Date.now() + timeoutMs; + while (Date.now() < deadline) { + if (condition()) return; + await new Promise((resolve) => setTimeout(resolve, 100)); + } + throw new Error(`timed out waiting until ${label}`); +} + async function waitForSuccessfulCompletion(root: string, runId: string, timeoutMs: number): Promise { + await waitForStatus(root, runId, "finished", timeoutMs); +} + +async function waitForStatus(root: string, runId: string, expected: string, timeoutMs: number): Promise { const deadline = Date.now() + timeoutMs; let status = "unknown"; while (Date.now() < deadline) { @@ -393,13 +535,13 @@ async function waitForSuccessfulCompletion(root: string, runId: string, timeoutM await new Promise((resolve) => setTimeout(resolve, 100)); continue; } - if (status === "finished") return; + if (status === expected) return; if (["failed", "cancelled", "canceled"].includes(status)) { throw new Error(`synthetic Smithers workflow ended with status ${status}`); } await new Promise((resolve) => setTimeout(resolve, 100)); } - throw new Error(`synthetic Smithers workflow did not finish; final status ${status}`); + throw new Error(`synthetic Smithers workflow did not reach ${expected}; final status ${status}`); } function smithersNodeOutput(root: string, runId: string, nodeId: string): Buffer { From 8298212f4ccebeefe176c1be5bdf2319efddaf40 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:01:44 +0000 Subject: [PATCH 061/206] fix(runtime): resume warns instead of dropping the trusted launcher or failing on worktree cleanup Two resume-time preparation steps were wrong in opposite directions. When prepareTrustedCliEnvironment threw on resume, for example because the run was launched under a different Node binary and the launcher bytes no longer match a fresh render, the error was swallowed and the trusted CLI variables were cleared. composeSmithersCommandPath then left /trusted-bin off the engine PATH, so every task's validator preflight ran whatever `ultrafuzz` the operator's PATH held (or hit ENOENT). Keep ULTRAFUZZ_TRUSTED_BIN pointed at the run-owned launcher when it exists: the launcher re-verifies its closure on every dispatch. The failure is now reported as a WORKFLOW_TRUSTED_CLI_UNVERIFIED warning. repairPrunableRunWorktreeRegistrations is cleanup, but a failed `git worktree remove` (or an unreadable workspace path) threw out of resume and failed the whole continuation. It now yields a WORKFLOW_WORKTREE_REPAIR_FAILED warning and the continuation proceeds. Resume returns these warnings as diagnostics (also appended after the error on a failed resume), and the CLI prints them in text mode. Both tests fail on main (the runner PATH starts with the caller's PATH instead of trusted-bin; the resume fails with WORKFLOW_LIFECYCLE_FAILED) and pass with this change. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/cli.md | 4 +- docs/schemas.md | 8 ++-- packages/cli/src/commands/resume.ts | 18 ++++---- packages/runtime/src/start-run.ts | 56 +++++++++++++++++++---- packages/runtime/test/runtime.test.ts | 66 +++++++++++++++++++++++++++ 6 files changed, 130 insertions(+), 23 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..66f2072d1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [cli] [docs]** `resume` now points agent adapters at the run's TOML agent config (`smithers/execution-config.toml`) instead of `resolved-config.json`, which their TOML reader parsed as empty, so a resumed run keeps each agent's `auth`, `api_key_env`, and `config_dir` (the default `CodexAgent` API-key auth had fallen back to subscription). A resume that only attaches to an already-active run no longer rewrites `state.json` or pushes back `workflow_deadline_at`, and prints `Run already active` instead of `Submitted resume`. A failed stale task-worktree cleanup no longer fails the resume and a failed trusted-CLI re-verification is no longer silent: both are now warnings, and in the latter case the run-owned launcher stays first on `PATH` instead of leaving tasks to whatever `ultrafuzz` the operator's `PATH` holds (#1153, #1143, #1110). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index a9ac9e2dc..3c16047e1 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -324,7 +324,9 @@ they are not resume authorization. Smithers decides which finished rows can be reused and which newly rendered or unfinished tasks run. Ultrafuzz does not rewrite historical artifacts or automatically reset, replay, timetravel, or fork completed work. Agent adapters in the continued workflow read the same -agent config as at launch, from the run's `smithers/execution-config.toml`. +agent config as at launch, from the run's `smithers/execution-config.toml`. If +resume cannot prune stale task-worktree registrations, it reports a +`WORKFLOW_WORKTREE_REPAIR_FAILED` warning and continues. `resume --refresh-controller` first renders the currently installed Ultrafuzz controller and stock adapters beside the historical source, then delegates to diff --git a/docs/schemas.md b/docs/schemas.md index c8e8caedc..da37ec8e9 100644 --- a/docs/schemas.md +++ b/docs/schemas.md @@ -84,9 +84,11 @@ a working-tree rebuild cannot change an active run. Ambient Node loader/search variables are removed and both ESM and CommonJS module resolution must stay inside that snapshot; document reads are unaffected. A path lookup alone is not a preflight. Ordinary resume now delegates continuation to Smithers instead of -using the historical launcher or closure as an authorization gate. A current -controller refresh publishes a new controller path without rewriting the -historical closure. +using the historical launcher or closure as an authorization gate. When resume +cannot re-verify the launcher, it reports a `WORKFLOW_TRUSTED_CLI_UNVERIFIED` +warning and keeps the run-owned launcher first on `PATH`; that launcher still +verifies its closure before every dispatch. A current controller refresh +publishes a new controller path without rewriting the historical closure. Exit `0` establishes portable document-shape conformance only. Cross-file joins, projected-key uniqueness, filesystem and Git facts, digest relationships, diff --git a/packages/cli/src/commands/resume.ts b/packages/cli/src/commands/resume.ts index e172870b7..09f2fc96d 100644 --- a/packages/cli/src/commands/resume.ts +++ b/packages/cli/src/commands/resume.ts @@ -5,6 +5,7 @@ import { cliEntrypoint, cliIo, commandFromRuntime, + diagnosticsText, emitCommandResult, globalFlags, projectRoot @@ -37,15 +38,14 @@ export default class Resume extends Command { resetNode: flags["reset-node"], env: cliIo().env }); - emitCommandResult( - this, - "resume", - commandFromRuntime("resume", result, (value) => - value.submitted - ? `Submitted ${value.action}: ${value.workflow_run_id}\n` - : `Run already active: ${value.workflow_run_id}; no new controller was started. If a pause is still draining, resume again once status reports paused.\n` - ), - flags.json === true + const commandResult = commandFromRuntime("resume", result, (value) => + value.submitted + ? `Submitted ${value.action}: ${value.workflow_run_id}\n` + : `Run already active: ${value.workflow_run_id}; no new controller was started. If a pause is still draining, resume again once status reports paused.\n` ); + if (result.ok && result.diagnostics.length > 0) { + commandResult.text = `${commandResult.text ?? ""}${diagnosticsText(result.diagnostics)}`; + } + emitCommandResult(this, "resume", commandResult, flags.json); } } diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index 9d92cfcc5..eb03c2139 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -54,6 +54,7 @@ import { prepareTrustedCliEnvironment, runTrustedJsonValidatorPreflight, TRUSTED_CLI_ENVIRONMENT_VARIABLES, + ULTRAFUZZ_TRUSTED_BIN_ENV, type TrustedCliEnvironment } from "./trusted-cli.js"; import { hasRuntimeErrors, runtimeFailure, runtimeResult } from "./utils.js"; @@ -478,6 +479,7 @@ export async function resumeRun(input: WorkflowLifecycleInput) { async function submitSmithersContinuation(input: WorkflowLifecycleInput) { let releaseLifecycleLock: (() => Promise) | undefined; + const diagnostics: RuntimeDiagnostic[] = []; try { const projectRoot = path.resolve(input.projectRoot); const runsRoot = await runsRootForProject(projectRoot); @@ -597,9 +599,23 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { }); if (prepared.active) runTrustedJsonValidatorPreflight({ layout, trusted: prepared }); trustedCli = prepared; - } catch { + } catch (error) { // Historical validator identity is task setup provenance, not authority - // to prevent Smithers from continuing the workflow. + // to prevent Smithers from continuing the workflow. Keep the run-owned + // launcher first on PATH anyway: it re-verifies its closure on every + // call, while dropping it lets tasks run whatever `ultrafuzz` is on PATH. + const trustedBin = path.join(layout.root, "trusted-bin"); + const launcherKept = fs.existsSync( + path.join(trustedBin, process.platform === "win32" ? "ultrafuzz.cmd" : "ultrafuzz") + ); + if (launcherKept) trustedCli.env[ULTRAFUZZ_TRUSTED_BIN_ENV] = trustedBin; + diagnostics.push( + resumeWarning( + "WORKFLOW_TRUSTED_CLI_UNVERIFIED", + `resume could not re-verify the run's trusted Ultrafuzz CLI${launcherKept ? "; tasks keep calling the run-owned launcher, which checks itself on every call" : ""}`, + error + ) + ); } } const agentRefs = tasks.flatMap((task) => task.agentChain.map((profile) => profile.agentRef)); @@ -615,7 +631,15 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { assertCurrentCloudAgentCredentialEnvironment(config, tasks, lifecycleEnvironment); } if (typeof metadata.source_revision === "string") { - repairPrunableRunWorktreeRegistrations({ projectRoot, runRoot: layout.root, runId }); + try { + repairPrunableRunWorktreeRegistrations({ projectRoot, runRoot: layout.root, runId }); + } catch (error) { + // Pruning is cleanup: a stale registration it leaves behind surfaces + // when Smithers recreates that task's worktree, so do not stop here. + diagnostics.push( + resumeWarning("WORKFLOW_WORKTREE_REPAIR_FAILED", "resume could not prune stale task worktrees", error) + ); + } } const result = await runSmithersLifecycleCommand({ action: "resume", @@ -677,14 +701,21 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { if (result.alreadyRunning !== true) { recordNativeContinuationState({ layout, config, requestedConcurrency: input.maxConcurrency }); } - return runtimeResult(true, { - run_id: runId, - workflow_run_id: smithersRunId, - action: "resume" as const, - submitted: !result.alreadyRunning - }); + return runtimeResult( + true, + { + run_id: runId, + workflow_run_id: smithersRunId, + action: "resume" as const, + submitted: !result.alreadyRunning + }, + diagnostics + ); } catch (error) { - return runtimeFailure([smithersDiagnostic(error, "WORKFLOW_LIFECYCLE_FAILED")]); + return runtimeFailure([ + smithersDiagnostic(error, "WORKFLOW_LIFECYCLE_FAILED"), + ...diagnostics + ]); } finally { await releaseLifecycleLock?.(); } @@ -779,6 +810,11 @@ function recordNativeContinuationState(input: { } } +function resumeWarning(code: string, context: string, error: unknown): RuntimeDiagnostic { + const diagnostic = smithersDiagnostic(error, code); + return { ...diagnostic, message: `${context}: ${diagnostic.message}`, severity: "warning", source: "runtime" }; +} + export async function replayRun(input: WorkflowLifecycleInput) { return submitLifecycleAction(input, "replay"); } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 2394621f2..ca30a97d7 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -25046,11 +25046,14 @@ test("native continuation does not use historical trusted CLI identity as an aut writeSmallTopology(project); const runId = "controller-refresh-trusted-cli"; const env = controllerRefreshTerminalEnv(project, runId); + const upPathLog = path.join(project, "fake-smithers-up-path.log"); + logFakeRunnerUpVariable(env, "PATH", upPathLog); const launched = await startRun({ projectRoot: project, runId, env }); assert.equal(launched.ok, true, JSON.stringify(launched.diagnostics)); const trustedMetadataPath = path.join(launched.value!.run_root, "trusted-cli.json"); fs.writeFileSync(trustedMetadataPath, "{}\n", "utf8"); fs.writeFileSync(env.SMITHERS_FAKE_LOG!, "", "utf8"); + fs.writeFileSync(upPathLog, "", "utf8"); const ordinary = await resumeRun({ projectRoot: project, runId, env }); assert.equal(ordinary.ok, true, JSON.stringify(ordinary.diagnostics)); @@ -25060,6 +25063,23 @@ test("native continuation does not use historical trusted CLI identity as an aut assert.equal(refreshed.ok, true, JSON.stringify(refreshed.diagnostics)); assert.equal(refreshed.value?.submitted, true); assert.equal(fs.readFileSync(trustedMetadataPath, "utf8"), "{}\n"); + // The run-owned launcher, which re-verifies itself on every call, stays + // first on the runner's PATH instead of leaving tasks to whatever + // `ultrafuzz` the operator's PATH holds, and the failure is reported (#1143). + const upPaths = fs.readFileSync(upPathLog, "utf8").trim().split("\n"); + assert.equal(upPaths.length, 2); + for (const upPath of upPaths) { + assert.equal(upPath.split(path.delimiter)[0], path.join(path.dirname(trustedMetadataPath), "trusted-bin")); + } + for (const resumed of [ordinary, refreshed]) { + assert.equal( + resumed.diagnostics.some( + (diagnostic) => diagnostic.code === "WORKFLOW_TRUSTED_CLI_UNVERIFIED" && diagnostic.severity === "warning" + ), + true, + JSON.stringify(resumed.diagnostics) + ); + } assert.equal( fs .readFileSync(env.SMITHERS_FAKE_LOG!, "utf8") @@ -26500,6 +26520,52 @@ test("ordinary resume checks active-run ownership before detached preflight", as assert.equal(fs.readFileSync(statePath, "utf8"), stateBefore); }); +test("resume continues with a warning when stale task-worktree cleanup fails", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "worktree-cleanup-failure"; + const env = controllerRefreshTerminalEnv(project, runId); + const commandLog = env.SMITHERS_FAKE_LOG; + assert.ok(commandLog); + const launched = await startRun({ projectRoot: project, runId, env }); + assert.equal(launched.ok, true, JSON.stringify(launched.diagnostics)); + assert.ok(launched.value); + const runRoot = launched.value.run_root; + // Cleanup runs only for runs launched from a Git revision. + const metadataPath = path.join(runRoot, "run.json"); + const metadata = JSON.parse(fs.readFileSync(metadataPath, "utf8")) as Record; + fs.writeFileSync(metadataPath, `${JSON.stringify({ ...metadata, source_revision: "0".repeat(40) }, null, 2)}\n`); + // Git lists a prunable registration owned by this run, then cannot remove it. + const previousPath = process.env.PATH ?? ""; + const gitBin = temporaryRoot("ufz-failing-git-"); + fs.writeFileSync( + path.join(gitBin, "git"), + [ + "#!/bin/sh", + 'case "$*" in', + ` *"worktree list --porcelain"*) printf 'worktree %s\\nbranch refs/heads/ultrafuzz/%s/stale\\nprunable\\n\\n' ${shellQuote(path.join(runRoot, "workspaces", "stale"))} ${runId} ;;`, + ' *"worktree remove"*) echo "fatal: synthetic removal failure" >&2; exit 1 ;;', + ` *) PATH=${shellQuote(previousPath)} exec git "$@" ;;`, + "esac", + "" + ].join("\n"), + { mode: 0o755 } + ); + process.env.PATH = [gitBin, previousPath].join(path.delimiter); + const resumed = await resumeRun({ projectRoot: project, runId, env }).finally(() => { + process.env.PATH = previousPath; + }); + + assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); + assert.equal(resumed.value?.submitted, true); + const warning = resumed.diagnostics.find((diagnostic) => diagnostic.code === "WORKFLOW_WORKTREE_REPAIR_FAILED"); + assert.ok(warning, JSON.stringify(resumed.diagnostics)); + assert.equal(warning.severity, "warning"); + assert.match(warning.message, /synthetic removal failure/u); + assert.match(fs.readFileSync(commandLog, "utf8"), /^up .*--resume ultrafuzz-worktree-cleanup-failure/mu); +}); + test("resume derives reset identities from the canonical nodes of a failed workflow", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); From 7643f7a38198311be7901f065a10ad36ee05c167 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:59:27 +0000 Subject: [PATCH 062/206] style(runtime): satisfy strict lint on the changed gate code The diff-limited strict lint CI runs flags the lines this branch touched: non-null assertions in the un-indented lens authority check and two tests, a number in a template literal, and a test body over the function length limit. Replace the assertions with explicit narrowing, stringify the sequence length once, and hoist the forged timeout-evidence cases to module scope. No behaviour change. Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/src/semantic-gates.ts | 10 +- packages/runtime/src/artifact-gates.ts | 9 +- packages/runtime/test/artifact-gates.test.ts | 304 ++++++++++--------- packages/runtime/test/runtime.test.ts | 4 +- 4 files changed, 165 insertions(+), 162 deletions(-) diff --git a/packages/artifacts/src/semantic-gates.ts b/packages/artifacts/src/semantic-gates.ts index 04df4f516..75eca0e68 100644 --- a/packages/artifacts/src/semantic-gates.ts +++ b/packages/artifacts/src/semantic-gates.ts @@ -6748,13 +6748,9 @@ function propertyCampaignTimeoutEvidenceIssues(document: unknown, context: Seman issue("$.exact_command", `Recon command must contain exactly one --test-limit ${RECON_MAX_TEST_LIMIT} flag`) ); } - if (sequenceLengthValues.length !== 1 || sequenceLengthValues[0] !== String(RECON_STATEFUL_SEQUENCE_LENGTH)) { - issues.push( - issue( - "$.exact_command", - `Recon command must contain exactly one --seq-len ${RECON_STATEFUL_SEQUENCE_LENGTH} flag` - ) - ); + const sequenceLength = String(RECON_STATEFUL_SEQUENCE_LENGTH); + if (sequenceLengthValues.length !== 1 || sequenceLengthValues[0] !== sequenceLength) { + issues.push(issue("$.exact_command", `Recon command must contain exactly one --seq-len ${sequenceLength} flag`)); } if (!hasExactCampaignHostTimeoutWrapper(resultCommand, configuredTimeoutSeconds)) { issues.push( diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index d3998f9bf..ef2a2e1f6 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -4966,10 +4966,12 @@ function verifyLensReferenceExpectationAuthority( (output) => output.contract === PROPERTY_LENS_CONTRACT ); const plannedOutputs = node.outputs.filter((output) => output.contract === PROPERTY_LENS_CONTRACT); + const declaration = sealedOutputs.length === 1 ? sealedOutputs[0] : undefined; + const plannedOutput = plannedOutputs.length === 1 ? plannedOutputs[0] : undefined; if ( - sealedOutputs.length !== 1 || - plannedOutputs.length !== 1 || - !smithersOutputMatchesPlanned(sealedOutputs[0]!, plannedOutputs[0]!) + declaration === undefined || + plannedOutput === undefined || + !smithersOutputMatchesPlanned(declaration, plannedOutput) ) { return [ { @@ -4981,7 +4983,6 @@ function verifyLensReferenceExpectationAuthority( } ]; } - const declaration = sealedOutputs[0]!; const lensPath = safeResolveInside(artifactDir, declaration.path, "property lens output"); const lensDocument = parseCurrentArtifactJson(artifactDir, lensPath, authenticated); if (lensDocument === undefined) { diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index e7172f540..52becf417 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -10503,7 +10503,7 @@ function timeoutEvidenceIssuePaths(result: ReturnType { @@ -10558,157 +10558,161 @@ test("current campaign timeout gate accepts a sealed model-profile or run-defaul assert.deepEqual(timeoutEvidenceIssuePaths(result), [], JSON.stringify(result.diagnostics)); }); -test("current campaign timeout gate rejects forged budgets, Recon flags, and sequence lengths", () => { - const replaceCommand = (search: string, replacement: string) => (fixture: CampaignTimeoutFixture) => - withCampaignCommand(fixture, fixture.backend.exact_command.replace(search, replacement)); - const cases: Array<{ name: string; path: string; mutate: (fixture: CampaignTimeoutFixture) => void }> = [ - { - name: "plan configured timeout", - path: "$.campaign_plan_ref#configured_fuzzer_timeout_seconds", - mutate: (fixture) => { - fixture.plan.configured_fuzzer_timeout_seconds = 3300; - } - }, - { - name: "Recon internal timeout", - path: "$.campaign_plan_ref#recon_internal_timeout_seconds", - mutate: (fixture) => { - fixture.plan.recon_internal_timeout_seconds = 3300; - } - }, - { - name: "host soft timeout", - path: "$.campaign_plan_ref#host_soft_timeout_seconds", - mutate: (fixture) => { - fixture.plan.host_soft_timeout_seconds = 3300; - } - }, - { - name: "backend configured timeout", - path: "$.configured_timeout_seconds", - mutate: (fixture) => { - fixture.backend.configured_timeout_seconds = 3300; - } - }, - { - name: "wrong host force-kill grace", - path: "$.campaign_plan_ref#host_force_kill_grace_seconds", - mutate: (fixture) => { - fixture.plan.host_force_kill_grace_seconds = 30; - } - }, - { - name: "forged artifact finalization reserve", - path: "$.campaign_plan_ref#artifact_finalization_reserve_seconds", - mutate: (fixture) => { - fixture.plan.artifact_finalization_reserve_seconds = 299; - } - }, - { - name: "plan test limit", - path: "$.campaign_plan_ref#recon_test_limit", - mutate: (fixture) => { - fixture.plan.recon_test_limit = "50000"; - } - }, - { - name: "reserve-subtracted command timeout", - path: "$.exact_command", - mutate: replaceCommand("--timeout 3600", "--timeout 3300") - }, - { - name: "duplicate timeout flag", - path: "$.exact_command", - mutate: replaceCommand("--timeout 3600", "--timeout 3600 --timeout 3600") - }, - { name: "missing timeout flag", path: "$.exact_command", mutate: replaceCommand("--timeout 3600 ", "") }, - { - name: "missing GNU timeout wrapper", - path: "$.exact_command", - mutate: replaceCommand("timeout --preserve-status --signal=INT --kill-after=300s 3600s ", "") - }, - { - name: "wrong GNU timeout soft deadline", - path: "$.exact_command", - mutate: replaceCommand("300s 3600s recon", "300s 3300s recon") - }, - { - name: "foreground wrapper", - path: "$.exact_command", - mutate: replaceCommand("--preserve-status", "--preserve-status --foreground") - }, - { - name: "bounded default test limit", - path: "$.exact_command", - mutate: replaceCommand(`--test-limit ${RECON_TIMEOUT_TEST_LIMIT}`, "--test-limit 50000") - }, - { - name: "one-step stateful sequence command", - path: "$.exact_command", - mutate: replaceCommand("--seq-len 100", "--seq-len 1") - }, - { - name: "missing stateful sequence command flag", - path: "$.exact_command", - mutate: replaceCommand(" --seq-len 100", "") - }, - { - name: "duplicate stateful sequence flags", - path: "$.exact_command", - mutate: replaceCommand("--seq-len 100", "--seq-len 100 --seq-len 100") - }, - { - name: "one-step stateful sequence in plan", - path: "$.campaign_plan_ref#recon_sequence_length", - mutate: (fixture) => { - fixture.plan.recon_sequence_length = 1; - } - }, - { - name: "one-step stateful sequence in result", - path: "$.sequence_length", - mutate: (fixture) => { - fixture.backend.sequence_length = 1; - } - }, - { - name: "one-step stateful sequence in summary", - path: "$.campaign_summary_ref#sequence_length", - mutate: (fixture) => { - fixture.summary.sequence_length = 1; - } - }, - { - name: "backend command differs from plan", - path: "$.exact_command", - mutate: (fixture) => { - fixture.backend.exact_command = `${fixture.backend.exact_command} --quiet`; - } - }, - { - name: "backend start differs from plan", - path: "$.start_timestamp", - mutate: (fixture) => { - fixture.backend.start_timestamp = "2026-08-11T00:00:01.000Z"; - } - }, - ...(["fuzzing_deadline_utc", "force_kill_deadline_utc", "final_artifact_deadline_utc"] as const).map((field) => ({ - name: `deadline arithmetic ${field}`, - path: `$.campaign_plan_ref#${field}`, - mutate: (fixture: CampaignTimeoutFixture) => { - fixture.plan[field] = "2026-08-11T01:00:02.000Z"; - } - })), - { - name: "summary outcome", - path: "$.campaign_summary_ref#outcome", - mutate: (fixture) => { - fixture.summary.outcome = "partial"; - } +const withReplacedCampaignCommand = (search: string, replacement: string) => (fixture: CampaignTimeoutFixture) => + withCampaignCommand(fixture, fixture.backend.exact_command.replace(search, replacement)); +const forgedCampaignTimeoutCases: Array<{ + name: string; + path: string; + mutate: (fixture: CampaignTimeoutFixture) => void; +}> = [ + { + name: "plan configured timeout", + path: "$.campaign_plan_ref#configured_fuzzer_timeout_seconds", + mutate: (fixture) => { + fixture.plan.configured_fuzzer_timeout_seconds = 3300; } - ]; + }, + { + name: "Recon internal timeout", + path: "$.campaign_plan_ref#recon_internal_timeout_seconds", + mutate: (fixture) => { + fixture.plan.recon_internal_timeout_seconds = 3300; + } + }, + { + name: "host soft timeout", + path: "$.campaign_plan_ref#host_soft_timeout_seconds", + mutate: (fixture) => { + fixture.plan.host_soft_timeout_seconds = 3300; + } + }, + { + name: "backend configured timeout", + path: "$.configured_timeout_seconds", + mutate: (fixture) => { + fixture.backend.configured_timeout_seconds = 3300; + } + }, + { + name: "wrong host force-kill grace", + path: "$.campaign_plan_ref#host_force_kill_grace_seconds", + mutate: (fixture) => { + fixture.plan.host_force_kill_grace_seconds = 30; + } + }, + { + name: "forged artifact finalization reserve", + path: "$.campaign_plan_ref#artifact_finalization_reserve_seconds", + mutate: (fixture) => { + fixture.plan.artifact_finalization_reserve_seconds = 299; + } + }, + { + name: "plan test limit", + path: "$.campaign_plan_ref#recon_test_limit", + mutate: (fixture) => { + fixture.plan.recon_test_limit = "50000"; + } + }, + { + name: "reserve-subtracted command timeout", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("--timeout 3600", "--timeout 3300") + }, + { + name: "duplicate timeout flag", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("--timeout 3600", "--timeout 3600 --timeout 3600") + }, + { name: "missing timeout flag", path: "$.exact_command", mutate: withReplacedCampaignCommand("--timeout 3600 ", "") }, + { + name: "missing GNU timeout wrapper", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("timeout --preserve-status --signal=INT --kill-after=300s 3600s ", "") + }, + { + name: "wrong GNU timeout soft deadline", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("300s 3600s recon", "300s 3300s recon") + }, + { + name: "foreground wrapper", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("--preserve-status", "--preserve-status --foreground") + }, + { + name: "bounded default test limit", + path: "$.exact_command", + mutate: withReplacedCampaignCommand(`--test-limit ${RECON_TIMEOUT_TEST_LIMIT}`, "--test-limit 50000") + }, + { + name: "one-step stateful sequence command", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("--seq-len 100", "--seq-len 1") + }, + { + name: "missing stateful sequence command flag", + path: "$.exact_command", + mutate: withReplacedCampaignCommand(" --seq-len 100", "") + }, + { + name: "duplicate stateful sequence flags", + path: "$.exact_command", + mutate: withReplacedCampaignCommand("--seq-len 100", "--seq-len 100 --seq-len 100") + }, + { + name: "one-step stateful sequence in plan", + path: "$.campaign_plan_ref#recon_sequence_length", + mutate: (fixture) => { + fixture.plan.recon_sequence_length = 1; + } + }, + { + name: "one-step stateful sequence in result", + path: "$.sequence_length", + mutate: (fixture) => { + fixture.backend.sequence_length = 1; + } + }, + { + name: "one-step stateful sequence in summary", + path: "$.campaign_summary_ref#sequence_length", + mutate: (fixture) => { + fixture.summary.sequence_length = 1; + } + }, + { + name: "backend command differs from plan", + path: "$.exact_command", + mutate: (fixture) => { + fixture.backend.exact_command = `${fixture.backend.exact_command} --quiet`; + } + }, + { + name: "backend start differs from plan", + path: "$.start_timestamp", + mutate: (fixture) => { + fixture.backend.start_timestamp = "2026-08-11T00:00:01.000Z"; + } + }, + ...(["fuzzing_deadline_utc", "force_kill_deadline_utc", "final_artifact_deadline_utc"] as const).map((field) => ({ + name: `deadline arithmetic ${field}`, + path: `$.campaign_plan_ref#${field}`, + mutate: (fixture: CampaignTimeoutFixture) => { + fixture.plan[field] = "2026-08-11T01:00:02.000Z"; + } + })), + { + name: "summary outcome", + path: "$.campaign_summary_ref#outcome", + mutate: (fixture) => { + fixture.summary.outcome = "partial"; + } + } +]; - for (const entry of cases) { +test("current campaign timeout gate rejects forged budgets, Recon flags, and sequence lengths", () => { + for (const entry of forgedCampaignTimeoutCases) { const result = runCampaignTimeoutGate(entry.mutate); assert.equal(result.ok, false, entry.name); assert.ok( diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 0785e34d8..03478f046 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -24561,7 +24561,9 @@ test("artifact gates validate a historical bundle through its active sealed sche fs.readFileSync(path.join(layout.root, "smithers", "tasks.json"), "utf8") ) as SmithersTaskManifestDocument ).tasks; - const attemptAuthority = { task: sealedTasks.find((task) => task.attemptId === node.id)!, tasks: sealedTasks }; + const sealedTask = sealedTasks.find((task) => task.attemptId === node.id); + assert.ok(sealedTask); + const attemptAuthority = { task: sealedTask, tasks: sealedTasks }; const verified = verifyRequiredArtifactsForAttempt(layout, node, node.id, attemptAuthority); assert.equal(verified.ok, true, JSON.stringify(verified.diagnostics)); From 3803dff7f688d70edd285b000173984115e576b4 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:02:18 +0000 Subject: [PATCH 063/206] fix(config): accept smol-toml 1.9 tables when loading project config smol-toml 1.9.0 (released 2026-09-22) builds every parsed table with Object.create(null). The config loader treats only objects whose prototype is Object.prototype as tables, and it sorts Object.entries(table) with the default comparator, which stringifies each value. A packed install has no lockfile and resolves the ^1.7.1 range to 1.9.0, so `ultrafuzz init` failed while loading the shipped defaults ("project must be a table; run must be a table; ..."). This is the packed-install failure that turned the package-gates lane red in main run 36495460966. Rebuild the parsed document with structuredClone, which returns the same data as ordinary objects, so the loader sees what it saw with 1.7. `__proto__` keys stay own data properties. Co-Authored-By: Claude Opus 5.5 --- packages/config/src/loader.ts | 5 ++- .../config/test/toml-null-prototype.test.ts | 42 +++++++++++++++++++ 2 files changed, 46 insertions(+), 1 deletion(-) create mode 100644 packages/config/test/toml-null-prototype.test.ts diff --git a/packages/config/src/loader.ts b/packages/config/src/loader.ts index 3deeb87b7..cfdb4af0e 100644 --- a/packages/config/src/loader.ts +++ b/packages/config/src/loader.ts @@ -121,7 +121,10 @@ export async function loadProjectConfig( export function parseProjectConfigToml(text: string, file = CONFIG_FILE_NAME): ConfigResult { let table: TomlTable; try { - table = parse(text); + // smol-toml 1.9 builds every table with Object.create(null), which this + // loader's plain-object checks and entry sorts do not accept. structuredClone + // returns the same data as ordinary objects. + table = structuredClone(parse(text)); } catch (error) { return fail([ diagnostic("CONFIG_TOML_PARSE_FAILED", `${file} is not valid TOML`, [], "project-toml", tomlLocation(error, file)) diff --git a/packages/config/test/toml-null-prototype.test.ts b/packages/config/test/toml-null-prototype.test.ts new file mode 100644 index 000000000..0d9e544f8 --- /dev/null +++ b/packages/config/test/toml-null-prototype.test.ts @@ -0,0 +1,42 @@ +import fs from "node:fs"; + +import type * as SmolToml from "smol-toml"; +import { describe, expect, it, vi } from "vitest"; + +// smol-toml 1.9 builds every table with Object.create(null). The workspace +// lockfile resolves 1.7, but a packed install resolves the newest 1.x, so +// reproduce the 1.9 shape here. +vi.mock("smol-toml", async (importOriginal) => { + const actual = await importOriginal(); + return { ...actual, parse: (text: string) => withNullPrototypes(actual.parse(text)) }; +}); + +import { createDefaultResolvedConfig, parseProjectConfigToml } from "../src/index.js"; + +function withNullPrototypes(value: unknown): unknown { + if (Array.isArray(value)) return value.map(withNullPrototypes); + if (typeof value !== "object" || value === null || value instanceof Date) return value; + const table = Object.create(null) as Record; + for (const [key, entry] of Object.entries(value)) table[key] = withNullPrototypes(entry); + return table; +} + +describe("TOML tables without a prototype", () => { + it("load the shipped defaults and a project config", () => { + const defaults = parseProjectConfigToml(fs.readFileSync(new URL("../defaults.toml", import.meta.url), "utf8")); + expect(defaults.ok, JSON.stringify(defaults.diagnostics)).toBe(true); + expect(createDefaultResolvedConfig().models.default).toBe("default"); + + const project = parseProjectConfigToml(` +[models.fast] +agent = "CodexAgent" +model = "gpt-5.5" + +[agents.CodexAgent] +auth = "subscription" +`); + expect(project.ok, JSON.stringify(project.diagnostics)).toBe(true); + if (!project.ok) return; + expect(project.value.models?.profiles?.fast).toMatchObject({ agent: "CodexAgent", model: "gpt-5.5" }); + }); +}); From 99d227eb96d2ded14c4269f0c7cfcd48e20ce447 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:50:39 +0000 Subject: [PATCH 064/206] fix(runtime): DeepSeek tasks no longer hang on Claude Code result lines DeepSeekAgent overrode the Claude Code output interpreter with a strict usage parser that required DeepSeek's OpenAI-compatible field names (prompt_cache_miss_tokens/prompt_cache_hit_tokens) and rejected input_tokens and cache_read_input_tokens as "legacy aliases". Claude Code always reports Anthropic field names, so every finished DeepSeek task threw from onStdoutLine. That hook runs inside the child process 'data' listener, so the throw escaped the listener instead of failing the invocation, and the task sat until the idle or node timeout. Delete the usage overrides (createOutputInterpreter, generate, stream and their parsing/attachment helpers). Smithers' ClaudeCodeAgent already reads the Anthropic names. Failed attempts no longer carry adapter-attached usage, which matches ClaudeAgent. The old tests fed a result line with DeepSeek field names that Claude Code never prints, so they passed while real runs hung. The replacement runs a Claude-Code-shaped result line through generate(): on main it fails with "DeepSeek result usage contains unsupported legacy alias input_tokens" thrown from the stdout listener; now it resolves with the reported usage. Co-Authored-By: Claude Opus 5.5 --- docs/config.md | 7 +- docs/reference/agent-adapter-boundaries.md | 2 +- .../runtime/scripts/run-pr-smoke-tests.mjs | 5 +- .../templates/smithers/agents/deepseek.tsx | 220 +---------------- .../test/agent-adapter-boundaries.test.ts | 6 +- packages/runtime/test/runtime.test.ts | 225 +++--------------- 6 files changed, 51 insertions(+), 414 deletions(-) diff --git a/docs/config.md b/docs/config.md index 4ccc55e3e..a8d974eb8 100644 --- a/docs/config.md +++ b/docs/config.md @@ -247,9 +247,10 @@ value is rejected before execution. See DeepSeek's and [Anthropic API guide](https://api-docs.deepseek.com/guides/anthropic_api). DeepSeek's automatic disk cache reports cache misses and hits independently. -Ultrafuzz records those as uncached input and cache-read tokens, records no -cache-write charge, and treats the provider's output count as already including -thinking tokens rather than publishing a second reasoning component. Pricing +Claude Code reports them under Anthropic field names (`input_tokens`, +`cache_read_input_tokens`), which the pinned Smithers Claude Code adapter +already reads, so the DeepSeek adapter does no usage parsing of its own. The +output count already includes thinking tokens. Pricing is pinned to the first-party `deepseek` catalog entry so a same-named hosted or subscription plan cannot supply a zero or unrelated rate. The current [DeepSeek price table](https://api-docs.deepseek.com/quick_start/pricing) lists diff --git a/docs/reference/agent-adapter-boundaries.md b/docs/reference/agent-adapter-boundaries.md index 902d4f73d..5f97ebdc8 100644 --- a/docs/reference/agent-adapter-boundaries.md +++ b/docs/reference/agent-adapter-boundaries.md @@ -20,7 +20,7 @@ without a matching re-export. | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | `claude.tsx` | Thin mapping. Its override applies Ultrafuzz's child-environment policy but does not rebuild an orchestrator responsibility. | Uses `model`, `extraArgs`, `addDir`, `permissionMode`, `settingSources`, `apiKey`, `configDir`, and `env`. | None. | | `codex.tsx` | Partly avoidable. The local argv rewrite and resume awareness exist because a working constructor option is serialized incorrectly upstream. Provider-home inspection and child-environment filtering are local policy. | `addDir` exists, but multiple values become one `--add-dir` occurrence. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622) | -| `deepseek.tsx` | Inherent after removing its avoidable reasoning-effort argv mapping. Route, auth, and effort are now thin mappings; result parsing and token normalization have no typed upstream surface. | Uses `model`, first-class `effort`, `addDir`, `permissionMode`, `settingSources`, `env`, and `configDir`; a custom-provider usage normalizer is missing. | [smithers#1624](https://github.com/smithersai/smithers/issues/1624) | +| `deepseek.tsx` | Thin mapping. Route, auth, and effort map onto constructor options; Claude Code reports usage under Anthropic field names that Smithers already reads, so the adapter parses no output. | Uses `model`, first-class `effort`, `addDir`, `permissionMode`, `settingSources`, `env`, and `configDir`. | None. | | `kimi.tsx` | Inherent with the current dependency except for Ultrafuzz-specific bounded-I/O and credential-governance checks. Usage discovery, actual-session recovery, argv compatibility, and runtime-home isolation cannot be expressed by constructor options. | `model`, `extraArgs`, `env`, `configDir`, and `session` exist; invocation-local usage, actual-session resolution, separate credential/runtime homes, and a Kimi Code 0.29.x command dialect are missing. | [smithers#1623](https://github.com/smithersai/smithers/issues/1623), [smithers#1626](https://github.com/smithersai/smithers/issues/1626) | | `openrouter.tsx` | Inherent with the current dependency except for local credential/config materialization. Provider-output quarantine and exact-session retry are orchestration responsibilities; the inherited Codex argv workaround is separately avoidable. | `config`, `configDir`, `env`, `model`, and `addDir` cover the route; a bounded provider-recovery policy is missing. | [smithers#1622](https://github.com/smithersai/smithers/issues/1622), [smithers#1625](https://github.com/smithersai/smithers/issues/1625) | diff --git a/packages/runtime/scripts/run-pr-smoke-tests.mjs b/packages/runtime/scripts/run-pr-smoke-tests.mjs index 5b6c249df..f8458d974 100644 --- a/packages/runtime/scripts/run-pr-smoke-tests.mjs +++ b/packages/runtime/scripts/run-pr-smoke-tests.mjs @@ -37,10 +37,9 @@ const namedTests = new Map([ const bunTestFile = "test/runtime.test.ts"; const bunTestNamePrefix = "Bun adapter contract: "; const bunTestNames = [ - "generated DeepSeek adapter uses the official endpoint and preserves independent usage components", + "generated DeepSeek adapter uses the official endpoint and isolates Claude routing", "generated DeepSeek adapter cleans an upstream command when environment policy rejects it", - "generated DeepSeek adapter corrects Smithers result and failed-attempt telemetry", - "generated DeepSeek adapter rejects ambiguous or noncanonical result telemetry", + "generated DeepSeek adapter completes on a Claude Code result line and reports its usage", // Needs bun:sqlite, so it can only run in this lane. Gates the claim that the // pinned runner's schema migrations are additive over a stopped 0.34.0 store. "pinned store migrations are additive over a 0.34.0 database" diff --git a/packages/runtime/src/templates/smithers/agents/deepseek.tsx b/packages/runtime/src/templates/smithers/agents/deepseek.tsx index b61f11424..4fe9c08b6 100644 --- a/packages/runtime/src/templates/smithers/agents/deepseek.tsx +++ b/packages/runtime/src/templates/smithers/agents/deepseek.tsx @@ -3,7 +3,6 @@ import path from "node:path"; import { ClaudeCodeAgent as SmithersClaudeCodeAgent } from "smthrs"; import { workflowControlChildEnvironment, workflowControlCredentialValue } from "./environment"; import { resolveProviderHome } from "./provider-home"; -import { parseStrictJson } from "./strict-json"; import { readStringTable, stringField } from "./toml"; type DeepSeekAuthConfig = { auth?: string; api_key_env?: string; config_dir?: string }; @@ -12,45 +11,18 @@ export type DeepSeekTaskOptions = { model?: string; reasoningEffort?: string; ad type DeepSeekAgentOptions = ConstructorParameters[0] & DeepSeekAuthOptions; type DeepSeekCommandParams = Parameters[0]; type DeepSeekCommand = Awaited>; -type DeepSeekOutputInterpreter = ReturnType; type DeepSeekReasoningEffort = "low" | "high" | "max"; -type DeepSeekUsage = { - inputTokens: number; - outputTokens: number; - cacheReadTokens: number; - cacheWriteTokens: 0; - totalTokens: number; -}; -type DeepSeekSmithersUsage = { - inputTokens: number; - inputTokenDetails: { noCacheTokens: number; cacheReadTokens: number; cacheWriteTokens: 0 }; - outputTokens: number; - outputTokenDetails: { textTokens: undefined; reasoningTokens: undefined }; - totalTokens: number; -}; const DEEPSEEK_ANTHROPIC_BASE_URL = "https://api.deepseek.com/anthropic"; const DEEPSEEK_REASONING_EFFORTS = ["low", "high", "max"] as const; -const DEEPSEEK_RESULT_MAX_BYTES = 1024 * 1024; -const DEEPSEEK_RESULT_MAX_DEPTH = 32; -const DEEPSEEK_RESULT_MAX_ITEMS = 10_000; -const DEEPSEEK_RESULT_MAX_PROPERTIES = 10_000; -const DEEPSEEK_LEGACY_USAGE_FIELDS = [ - "input_tokens", - "inputTokens", - "outputTokens", - "completion_tokens", - "cache_read_input_tokens", - "cacheReadTokens", - "cached_input_tokens" -] as const; /** * Runs the Claude Code harness against DeepSeek's Anthropic-compatible * endpoint: a compatibility pairing, not DeepSeek's first-party coding agent. - * Keep it as a distinct factory so credentials, model selection, telemetry - * semantics, and pricing provenance never inherit Anthropic defaults - * accidentally. + * Keep it as a distinct factory so credentials, model selection, and pricing + * provenance never inherit Anthropic defaults accidentally. Token usage needs + * no adapter code: Claude Code reports it under Anthropic field names, which + * Smithers' ClaudeCodeAgent already reads. */ export function createDeepSeekAgent(options: DeepSeekTaskOptions = {}): SmithersClaudeCodeAgent { const reasoningEffort = deepSeekReasoningEffort(options.reasoningEffort); @@ -67,7 +39,6 @@ export function createDeepSeekAgent(options: DeepSeekTaskOptions = {}): Smithers } export class DeepSeekClaudeCodeAgent extends SmithersClaudeCodeAgent { - private pendingUsage: DeepSeekSmithersUsage | undefined; private readonly ultrafuzzApiKey: string; // Smithers 0.35.0's BaseCliAgent rejects unknown constructor options with a @@ -79,35 +50,7 @@ export class DeepSeekClaudeCodeAgent extends SmithersClaudeCodeAgent { this.ultrafuzzApiKey = ultrafuzzApiKey; } - override generate( - ...args: Parameters - ): ReturnType { - return this.withDeepSeekUsage(super.generate(...args)) as ReturnType; - } - - override stream( - ...args: Parameters - ): ReturnType { - return this.withDeepSeekStreamUsage(super.stream(...args)) as ReturnType; - } - - override createOutputInterpreter(): DeepSeekOutputInterpreter { - const base = super.createOutputInterpreter(); - return { - ...base, - onStdoutLine: (line) => { - const usage = deepSeekUsageFromResultLine(line); - if (usage !== undefined) this.pendingUsage = deepSeekSmithersUsage(usage); - const events = base.onStdoutLine?.(line) ?? []; - if (usage === undefined) return events; - const completedUsage = deepSeekCompletedUsage(usage); - return events.map((event) => (event.type === "completed" ? { ...event, usage: completedUsage } : event)); - } - }; - } - override async buildCommand(params: DeepSeekCommandParams): Promise { - this.pendingUsage = undefined; this.opts.settingSources = ""; const command = await super.buildCommand(params); try { @@ -161,28 +104,6 @@ export class DeepSeekClaudeCodeAgent extends SmithersClaudeCodeAgent { throw error; } } - - private withDeepSeekUsage(promise: Promise): Promise { - return promise - .then((result) => attachDeepSeekResultUsage(result, this.pendingUsage)) - .catch((error: unknown) => { - throw attachDeepSeekFailureUsage(error, this.pendingUsage); - }) - .finally(() => { - this.pendingUsage = undefined; - }); - } - - private withDeepSeekStreamUsage(promise: Promise): Promise { - return promise - .then((result) => attachDeepSeekStreamUsage(result, this.pendingUsage)) - .catch((error: unknown) => { - throw attachDeepSeekFailureUsage(error, this.pendingUsage); - }) - .finally(() => { - this.pendingUsage = undefined; - }); - } } function deepSeekAuthOptions(): DeepSeekAuthOptions { @@ -222,136 +143,3 @@ function deepSeekReasoningEffort(value: string | undefined): DeepSeekReasoningEf } throw new Error(`DeepSeekAgent reasoning effort must be one of ${DEEPSEEK_REASONING_EFFORTS.join(", ")}: ${value}`); } - -/** - * DeepSeek bills cache misses and cache hits independently. Its completion - * token count already includes thinking tokens, so exposing a separate - * reasoning count would double-count both tokens and spend. - */ -function deepSeekUsageFromResultLine(line: string): DeepSeekUsage | undefined { - const first = firstNonJsonWhitespace(line); - if (first === undefined) return undefined; - const objectCandidate = first === "{"; - let payload: unknown; - try { - payload = parseStrictJson(line, { - maxBytes: DEEPSEEK_RESULT_MAX_BYTES, - maxDepth: DEEPSEEK_RESULT_MAX_DEPTH, - maxItems: DEEPSEEK_RESULT_MAX_ITEMS, - maxProperties: DEEPSEEK_RESULT_MAX_PROPERTIES - }); - } catch (error) { - if (objectCandidate) { - throw new Error( - `DeepSeek result output is invalid strict JSON: ${error instanceof Error ? error.message : String(error)}`, - { cause: error } - ); - } - return undefined; - } - if (!isRecord(payload) || payload.type !== "result") return undefined; - if (!isRecord(payload.usage)) throw new Error("DeepSeek result usage must be an object"); - const usage = payload.usage; - for (const legacyField of DEEPSEEK_LEGACY_USAGE_FIELDS) { - if (Object.prototype.hasOwnProperty.call(usage, legacyField)) { - throw new Error(`DeepSeek result usage contains unsupported legacy alias ${legacyField}`); - } - } - const inputTokens = requiredDeepSeekTokenCount(usage, "prompt_cache_miss_tokens"); - const outputTokens = requiredDeepSeekTokenCount(usage, "output_tokens"); - const cacheReadTokens = requiredDeepSeekTokenCount(usage, "prompt_cache_hit_tokens"); - const normalized = { - inputTokens, - outputTokens, - cacheReadTokens, - cacheWriteTokens: 0 as const - }; - const totalTokens = normalized.inputTokens + normalized.cacheReadTokens + normalized.outputTokens; - if (!Number.isSafeInteger(totalTokens)) throw new Error("DeepSeek result usage exceeds the safe integer range"); - return { ...normalized, totalTokens }; -} - -function firstNonJsonWhitespace(value: string): string | undefined { - for (const character of value) { - if (character !== " " && character !== "\t" && character !== "\n" && character !== "\r") return character; - } - return undefined; -} - -function requiredDeepSeekTokenCount(value: Record, field: string): number { - const candidate = value[field]; - if (typeof candidate !== "number" || !Number.isSafeInteger(candidate) || candidate < 0) { - throw new Error(`DeepSeek result usage.${field} must be a non-negative safe integer`); - } - return candidate; -} - -function deepSeekCompletedUsage(usage: DeepSeekUsage): Record { - return { - input_tokens: deepSeekProviderInputTokens(usage), - fresh_input_tokens: usage.inputTokens, - output_tokens: usage.outputTokens, - cache_read_input_tokens: usage.cacheReadTokens, - cache_creation_input_tokens: usage.cacheWriteTokens, - total_tokens: usage.totalTokens - }; -} - -function deepSeekSmithersUsage(usage: DeepSeekUsage): DeepSeekSmithersUsage { - return { - inputTokens: deepSeekProviderInputTokens(usage), - inputTokenDetails: { - noCacheTokens: usage.inputTokens, - cacheReadTokens: usage.cacheReadTokens, - cacheWriteTokens: usage.cacheWriteTokens - }, - outputTokens: usage.outputTokens, - outputTokenDetails: { - textTokens: undefined, - reasoningTokens: undefined - }, - totalTokens: usage.totalTokens - }; -} - -function deepSeekProviderInputTokens(usage: DeepSeekUsage): number { - return usage.totalTokens - usage.outputTokens; -} - -function attachDeepSeekResultUsage(result: T, usage: DeepSeekSmithersUsage | undefined): T { - if (usage === undefined || !isRecord(result)) return result; - try { - result.usage = usage; - result.totalUsage = usage; - } catch { - // Telemetry must never turn a successful provider invocation into a model - // failure if an exotic Smithers result becomes immutable. - } - return result; -} - -function attachDeepSeekStreamUsage(result: T, usage: DeepSeekSmithersUsage | undefined): T { - if (usage === undefined || !isRecord(result)) return result; - try { - result.usage = Promise.resolve(usage); - result.totalUsage = Promise.resolve(usage); - } catch { - // Telemetry must never turn a successful provider invocation into a model - // failure if an exotic Smithers stream result becomes immutable. - } - return result; -} - -function attachDeepSeekFailureUsage(error: unknown, usage: DeepSeekSmithersUsage | undefined): unknown { - if (usage === undefined || !isRecord(error)) return error; - try { - error.usage = usage; - } catch { - // Preserve the original failure if an exotic error object is immutable. - } - return error; -} - -function isRecord(value: unknown): value is Record { - return typeof value === "object" && value !== null && !Array.isArray(value); -} diff --git a/packages/runtime/test/agent-adapter-boundaries.test.ts b/packages/runtime/test/agent-adapter-boundaries.test.ts index daf8c2d7a..2bb3fbd5f 100644 --- a/packages/runtime/test/agent-adapter-boundaries.test.ts +++ b/packages/runtime/test/agent-adapter-boundaries.test.ts @@ -84,11 +84,7 @@ const adapterPolicies: Record = { responsibilities: ["argv-construction", "session-handling"], upstreamIssues: ["https://github.com/smithersai/smithers/issues/1622"] }, - "deepseek.tsx": { - purpose: "adapter", - responsibilities: ["output-interpretation", "token-accounting"], - upstreamIssues: ["https://github.com/smithersai/smithers/issues/1624"] - }, + "deepseek.tsx": { purpose: "adapter", responsibilities: [], upstreamIssues: [] }, "environment.tsx": { purpose: "data-governance", responsibilities: [], upstreamIssues: [] }, "index.tsx": { purpose: "registry", responsibilities: [], upstreamIssues: [] }, "kimi.tsx": { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..6508cdc92 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -3532,9 +3532,6 @@ test("init preserves existing project-owned files and validate exposes launch po assert.match(deepSeekAgentText, /settingSources:\s*""/); assert.match(deepSeekAgentText, /effort:\s*reasoningEffort/); assert.doesNotMatch(deepSeekAgentText, /extraArgs:\s*\["--effort"/); - assert.match(deepSeekAgentText, /cacheReadTokens/); - assert.match(deepSeekAgentText, /reasoningTokens: undefined/); - assert.match(deepSeekAgentText, /import \{ parseStrictJson \} from "\.\/strict-json";/u); assert.doesNotMatch(deepSeekAgentText, /=\s*createDeepSeekAgent\(\)/); const kimiAgentText = fs.readFileSync(path.join(project, ".smithers/agents/kimi.ts"), "utf8"); assert.match(kimiAgentText, /KimiAgent/); @@ -6984,7 +6981,7 @@ bunAdapterTest( ); bunAdapterTest( - "generated DeepSeek adapter uses the official endpoint and preserves independent usage components", + "generated DeepSeek adapter uses the official endpoint and isolates Claude routing", { timeout: 30_000 }, async () => { const project = tempProject(); @@ -7066,33 +7063,6 @@ bunAdapterTest( ]) { assert.equal(command.env?.[name], "", `${name} must not leak into DeepSeek Claude Code invocations`); } - - const resultLine = JSON.stringify({ - type: "result", - subtype: "success", - is_error: false, - result: "done", - usage: { - prompt_cache_miss_tokens: 120, - output_tokens: 30, - prompt_cache_hit_tokens: 400, - cache_creation_input_tokens: 999, - reasoning_tokens: 20 - } - }); - const events = agent.createOutputInterpreter().onStdoutLine?.(resultLine) as Array<{ - type?: string; - usage?: Record; - }>; - const completed = events.find((event) => event.type === "completed"); - assert.deepEqual(completed?.usage, { - input_tokens: 520, - fresh_input_tokens: 120, - output_tokens: 30, - cache_read_input_tokens: 400, - cache_creation_input_tokens: 0, - total_tokens: 550 - }); } finally { await command.cleanup?.(); } @@ -7147,173 +7117,56 @@ bunAdapterTest( ); bunAdapterTest( - "generated DeepSeek adapter corrects Smithers result and failed-attempt telemetry", + "generated DeepSeek adapter completes on a Claude Code result line and reports its usage", { timeout: 30_000 }, async () => { const project = tempProject(); const init = initProject({ projectRoot: project, force: true }); assert.equal(init.ok, true, JSON.stringify(init.diagnostics)); const { DeepSeekClaudeCodeAgent } = await loadGeneratedDeepSeekAgent(project); - const providerUsage = { - prompt_cache_miss_tokens: 101, - prompt_cache_hit_tokens: 400, - output_tokens: 23, - reasoning_tokens: 17 - }; - const normalizedUsage = { - inputTokens: 501, - inputTokenDetails: { noCacheTokens: 101, cacheReadTokens: 400, cacheWriteTokens: 0 }, - outputTokens: 23, - outputTokenDetails: { textTokens: undefined, reasoningTokens: undefined }, - totalTokens: 524 - }; - - const successful = new DeepSeekClaudeCodeAgent({ model: "deepseek-v4-pro", ultrafuzzApiKey: "test-key" }); - successful.buildCommand = async () => ({ - command: process.execPath, - args: [ - "-e", - `console.log(${JSON.stringify( - JSON.stringify({ - type: "result", - subtype: "success", - is_error: false, - result: "done", - session_id: "deepseek-session", - usage: providerUsage - }) - )})` - ], - outputFormat: "stream-json" + // The shape Claude Code prints for a DeepSeek-routed session: Anthropic + // usage field names. (The Claude Code 2.1.284 binary contains no + // prompt_cache_hit_tokens or prompt_cache_miss_tokens string at all.) + const resultLine = JSON.stringify({ + type: "result", + subtype: "success", + is_error: false, + duration_ms: 1200, + num_turns: 1, + result: "done", + session_id: "deepseek-session", + total_cost_usd: 0.01, + usage: { + input_tokens: 120, + cache_creation_input_tokens: 7, + cache_read_input_tokens: 400, + output_tokens: 30, + server_tool_use: { web_search_requests: 0 }, + service_tier: "standard" + } }); - const result = await successful.generate({ prompt: "Telemetry", rootDir: project }); - assert.deepEqual(result.usage, normalizedUsage); - - const streamed = await successful.stream({ prompt: "Stream telemetry", rootDir: project }); - assert.deepEqual(await streamed.usage, normalizedUsage); - assert.deepEqual(await streamed.totalUsage, normalizedUsage); - - const failed = new DeepSeekClaudeCodeAgent({ model: "deepseek-v4-pro", ultrafuzzApiKey: "test-key" }); - failed.buildCommand = async () => ({ + // A short idle timeout turns a stalled invocation into a prompt failure + // instead of waiting for the node timeout. + const agent = new DeepSeekClaudeCodeAgent({ + model: "deepseek-v4-pro", + ultrafuzzApiKey: "test-key", + idleTimeoutMs: 5_000 + }); + agent.buildCommand = async () => ({ command: process.execPath, - args: [ - "-e", - `console.log(${JSON.stringify( - JSON.stringify({ - type: "result", - subtype: "error", - is_error: true, - error: "provider failed", - usage: providerUsage - }) - )}); process.exit(17)` - ], + args: ["-e", `console.log(${JSON.stringify(resultLine)})`], outputFormat: "stream-json" }); - let failure: unknown; - try { - await failed.generate({ prompt: "Failed telemetry", rootDir: project }); - } catch (error) { - failure = error; - } - assert.ok(failure instanceof Error); - assert.deepEqual((failure as Error & { usage?: unknown }).usage, normalizedUsage); - } -); - -bunAdapterTest( - "generated DeepSeek adapter rejects ambiguous or noncanonical result telemetry", - { timeout: 30_000 }, - async () => { - const project = tempProject(); - const init = initProject({ projectRoot: project, force: true }); - assert.equal(init.ok, true, JSON.stringify(init.diagnostics)); - const { DeepSeekClaudeCodeAgent } = await loadGeneratedDeepSeekAgent(project); - const agent = new DeepSeekClaudeCodeAgent({ model: "deepseek-v4-pro", ultrafuzzApiKey: "test-key" }); - const interpreter = agent.createOutputInterpreter(); - assert.doesNotThrow(() => - interpreter.onStdoutLine?.(JSON.stringify({ type: "assistant", message: { content: "working" } })) - ); - assert.doesNotThrow(() => interpreter.onStdoutLine?.("provider banner: still starting")); + const result = (await agent.generate({ prompt: "Telemetry", rootDir: project })) as { + text?: string; + usage?: { inputTokens?: number; outputTokens?: number; inputTokenDetails?: { cacheReadTokens?: number } }; + }; - const tooDeep = `${"[".repeat(34)}null${"]".repeat(34)}`; - const invalid = [ - { - label: "duplicate key", - line: '{"type":"result","type":"result","usage":{"prompt_cache_miss_tokens":1,"prompt_cache_hit_tokens":2,"output_tokens":3}}', - expected: /duplicate/iu - }, - { - label: "malformed candidate", - line: '{"type":"result","usage":', - expected: /invalid strict JSON/iu - }, - { - label: "malformed object without result marker", - line: '{"provider_status":', - expected: /invalid strict JSON/iu - }, - { - label: "legacy aliases", - line: JSON.stringify({ - type: "result", - usage: { input_tokens: 1, cache_read_input_tokens: 2, completion_tokens: 3 } - }), - expected: /legacy alias/iu - }, - { - label: "legacy alias alongside canonical fields", - line: JSON.stringify({ - type: "result", - usage: { - prompt_cache_miss_tokens: 1, - prompt_cache_hit_tokens: 2, - output_tokens: 3, - input_tokens: 1 - } - }), - expected: /legacy alias input_tokens/iu - }, - { - label: "missing exact field", - line: JSON.stringify({ - type: "result", - usage: { prompt_cache_miss_tokens: 1, output_tokens: 3 } - }), - expected: /prompt_cache_hit_tokens/iu - }, - { - label: "oversize raw line whitespace", - line: - " ".repeat(1024 * 1024) + - JSON.stringify({ - type: "result", - usage: { prompt_cache_miss_tokens: 1, prompt_cache_hit_tokens: 2, output_tokens: 3 } - }), - expected: /1048576-byte limit/iu - }, - { - label: "excessive depth", - line: `{"type":"result","future":${tooDeep},"usage":{"prompt_cache_miss_tokens":1,"prompt_cache_hit_tokens":2,"output_tokens":3}}`, - expected: /nesting-depth limit of 32/iu - }, - { - label: "unsafe aggregate", - line: JSON.stringify({ - type: "result", - usage: { - prompt_cache_miss_tokens: Number.MAX_SAFE_INTEGER, - prompt_cache_hit_tokens: 1, - output_tokens: 0 - } - }), - expected: /safe integer range/iu - } - ]; - for (const fixture of invalid) { - assert.throws(() => interpreter.onStdoutLine?.(fixture.line), fixture.expected, fixture.label); - } + assert.equal(result.text, "done"); + assert.equal(result.usage?.inputTokens, 120); + assert.equal(result.usage?.outputTokens, 30); + assert.equal(result.usage?.inputTokenDetails?.cacheReadTokens, 400); } ); From 45a708c8b9b00a81b227b0e0884df71f5e38543d Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:57:05 +0000 Subject: [PATCH 065/206] fix(runtime): Kimi and Pi telemetry parsing fails open instead of hanging tasks The Kimi and Pi adapters parse CLI output inside onStdoutLine and onStderrLine, which Smithers calls from the child process 'data' listeners. A throw there escapes the listener instead of failing the invocation, so the task waits for its idle or node timeout. Kimi threw on any stdout/stderr line starting with "{" that was not strict JSON (a Node util.inspect dump such as "{ code: 'ECONNRESET' }", a duplicate key, a line over 1 MiB) and on a malformed resume hint; Pi threw when a provider's totalTokens did not equal its components or its cost breakdown was missing or inconsistent. Kimi also read session recovery and wire usage from files in onExit and at buildCommand, where any torn, oversized or replaced record failed the invocation. A wire torn by a killed attempt made every later resume of that session fail at buildCommand. Make these reads total: an ambiguous resume hint carries no session, an unreadable wire or baseline leaves the invocation's usage absent, and an inconsistent Pi usage event stays uncounted. Usage is still never partially counted or fabricated. Tests that asserted these throws are flipped. New end-to-end tests run the adapters on malformed output; against main they fail with the throws quoted above (e.g. "Kimi output JSON is invalid", "Pi assistant usage totalTokens does not equal its token component sum", "Kimi wire has a torn or unterminated final record"). Co-Authored-By: Claude Opus 5.5 --- docs/config.md | 4 +- .../src/templates/smithers/agents/kimi.tsx | 47 +- .../src/templates/smithers/agents/pi.tsx | 23 +- packages/runtime/test/runtime.test.ts | 473 ++++++++++-------- 4 files changed, 311 insertions(+), 236 deletions(-) diff --git a/docs/config.md b/docs/config.md index a8d974eb8..857b2c0bf 100644 --- a/docs/config.md +++ b/docs/config.md @@ -198,7 +198,9 @@ Kimi's four components — uncached input, output, cache reads, and cache creation — are reported independently; Kimi already folds thinking tokens into output, so no separate reasoning total is published. Malformed or absent usage stays absent rather than becoming zeros, which keeps accounting honest about -what it does not know. Kimi model pricing resolves against the Moonshot +what it does not know. An unreadable wire, including inherited history torn by +a killed attempt, leaves that invocation's usage absent instead of failing the +invocation. Kimi model pricing resolves against the Moonshot provider entry in the pricing catalog, so the configured alias must match a Moonshot catalog model id such as `kimi-k3`; anything else is reported as an unresolved model instead of being priced from a same-named third-party entry. diff --git a/packages/runtime/src/templates/smithers/agents/kimi.tsx b/packages/runtime/src/templates/smithers/agents/kimi.tsx index ddca25b66..08c59c366 100644 --- a/packages/runtime/src/templates/smithers/agents/kimi.tsx +++ b/packages/runtime/src/templates/smithers/agents/kimi.tsx @@ -161,11 +161,14 @@ export class KimiCode029Agent extends SmithersKimiAgent { return base.onStderrLine?.(line) ?? []; }, onExit: (result) => { - this.issuedSessionId ??= sessionIdFromIndex(this.activeRuntimeHome); + // Session recovery and usage are read back from files the CLI wrote. + // An unreadable or unexpected record leaves them unknown; it must never + // replace the invocation's own result with a telemetry failure. + this.issuedSessionId ??= failOpen(() => sessionIdFromIndex(this.activeRuntimeHome)); // Kimi Code 0.29.1 prints no usage on stdout, so the pinned Smithers // BaseCliAgent falls back to the completed event. Attach this // invocation's own wire-record delta there. - const delta = this.invocationUsage(); + const delta = failOpen(() => this.invocationUsage()); this.pendingFailureUsage = delta === undefined ? undefined : kimiSmithersUsage(delta); const events = base.onExit?.(result) ?? []; if (delta === undefined) return events; @@ -234,11 +237,14 @@ export class KimiCode029Agent extends SmithersKimiAgent { ); } if (executionConfigDir !== undefined && sessionStoreDir !== undefined) { - runtimeHome = createKimiRuntimeHome(executionConfigDir, sessionStoreDir); - seedKimiSessionState(sessionStoreDir, runtimeHome, knownSession); + const seededHome = createKimiRuntimeHome(executionConfigDir, sessionStoreDir); + runtimeHome = seededHome; + seedKimiSessionState(sessionStoreDir, seededHome, knownSession); // Baseline AFTER seeding so a resumed session reports only the tokens - // this invocation adds, never the history it inherited. - usageBaseline = kimiUsageBaseline(runtimeHome); + // this invocation adds, never the history it inherited. Unreadable + // inherited history (for example a wire torn by a killed attempt) + // leaves this invocation's usage unknown instead of failing it. + usageBaseline = failOpen(() => kimiUsageBaseline(seededHome)); } } catch (error) { await command.cleanup?.(); @@ -1508,27 +1514,22 @@ function combineCleanup( }; } +// Runs inside the child's stdout/stderr listeners, where a throw escapes the +// invocation instead of failing it and leaves the task waiting for its timeout. +// Any line that is not an unambiguous resume hint therefore carries no session. function sessionIdFromJsonLine(line: string): string | undefined { const first = firstNonJsonWhitespace(line); if (first === undefined || first !== "{") return undefined; - let value: unknown; - try { - value = parseStrictJson(line, { + const value = failOpen(() => + parseStrictJson(line, { maxBytes: KIMI_WIRE_MAX_LINE_BYTES, maxDepth: KIMI_JSON_MAX_DEPTH, maxItems: KIMI_JSON_MAX_ITEMS, maxProperties: KIMI_JSON_MAX_PROPERTIES - }); - } catch (error) { - throw new Error(`Kimi output JSON is invalid: ${error instanceof Error ? error.message : String(error)}`, { - cause: error - }); - } + }) + ); if (!isRecord(value) || value.type !== KIMI_RESUME_HINT_TYPE) return undefined; - if (typeof value.session_id !== "string") throw new Error("Kimi session.resume_hint session_id must be a string"); - const sessionId = validSessionId(value.session_id); - if (sessionId === undefined) throw new Error("Kimi session.resume_hint session_id is invalid"); - return sessionId; + return typeof value.session_id === "string" ? validSessionId(value.session_id) : undefined; } function firstNonJsonWhitespace(value: string): string | undefined { @@ -1559,6 +1560,14 @@ function isRecord(value: unknown): value is Record { return typeof value === "object" && value !== null && !Array.isArray(value); } +function failOpen(read: () => T): T | undefined { + try { + return read(); + } catch { + return undefined; + } +} + function pathEntryExists(filePath: string): boolean { try { lstatSync(filePath); diff --git a/packages/runtime/src/templates/smithers/agents/pi.tsx b/packages/runtime/src/templates/smithers/agents/pi.tsx index 48ff75343..38c9a7448 100644 --- a/packages/runtime/src/templates/smithers/agents/pi.tsx +++ b/packages/runtime/src/templates/smithers/agents/pi.tsx @@ -99,7 +99,10 @@ export class CompatiblePiAgent extends SmithersPiAgent { return { ...interpreter, onStdoutLine: (line) => { - const observation = observePiLine(usage, line); + // This runs inside the child's stdout listener, where a throw escapes + // the invocation instead of failing it and leaves the task waiting for + // its timeout. A usage event Pi reports inconsistently stays uncounted. + const observation = failOpen(() => observePiLine(usage, line)); sessionId = observation?.sessionId ?? sessionId; // A fresh Smithers interpreter sees only the latest authoritative // assistant message and subsequent deltas. Reusing it as an oracle @@ -274,10 +277,10 @@ function safePiUsageSum(left: number, right: number): number { function reportedPiUsage(totals: PiInvocationUsage): PiReportedUsage | undefined { if (totals.messageCount === 0) return undefined; - const inputTokens = safePiUsageSum( - safePiUsageSum(totals.freshInputTokens, totals.cacheReadTokens), - totals.cacheWriteTokens - ); + const inputTokens = totals.freshInputTokens + totals.cacheReadTokens + totals.cacheWriteTokens; + const totalTokens = inputTokens + totals.outputTokens; + // An aggregate outside the safe-integer range is unknown usage, not a failure. + if (!Number.isSafeInteger(totalTokens)) return undefined; return { inputTokens, freshInputTokens: totals.freshInputTokens, @@ -285,7 +288,7 @@ function reportedPiUsage(totals: PiInvocationUsage): PiReportedUsage | undefined cacheReadTokens: totals.cacheReadTokens, cacheWriteTokens: totals.cacheWriteTokens, reasoningTokens: totals.reasoningTokens, - totalTokens: safePiUsageSum(inputTokens, totals.outputTokens), + totalTokens, reportedCostUsd: totals.reportedCostUsd }; } @@ -362,6 +365,14 @@ function objectRecord(value: unknown): Record | undefined { return value && typeof value === "object" && !Array.isArray(value) ? (value as Record) : undefined; } +function failOpen(read: () => T): T | undefined { + try { + return read(); + } catch { + return undefined; + } +} + function applyPiTerminalAnswer( events: T, terminalEvents: unknown, diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 6508cdc92..d01b8ee12 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -7850,6 +7850,86 @@ bunAdapterTest( } ); +bunAdapterTest( + "generated Pi adapter leaves inconsistent usage uncounted instead of aborting the invocation", + { timeout: 30_000 }, + async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const { createPiAgent } = await loadGeneratedPiAgent(project); + const previous = { + config: process.env.ULTRAFUZZ_CONFIG_PATH, + openRouter: process.env.OPENROUTER_API_KEY + }; + process.env.ULTRAFUZZ_CONFIG_PATH = path.join(project, "ultrafuzz.toml"); + process.env.OPENROUTER_API_KEY = "sk-or-v1-not-a-real-openrouter-credential"; + try { + const cost = { input: 0.001, output: 0.002, cacheRead: 0, cacheWrite: 0, total: 0.003 }; + const assistant = (responseId: string, text: string, usage: Record) => ({ + type: "message_end", + message: { role: "assistant", responseId, content: [{ type: "text", text }], usage } + }); + const lines = [ + { type: "session", id: "inconsistent-usage-session" }, + // A provider that folds reasoning into totalTokens, one that omits the + // cost breakdown, and one whose cost total disagrees with its parts. + assistant("reasoning-in-total", "first", { + input: 1, + output: 2, + cacheRead: 0, + cacheWrite: 0, + totalTokens: 9, + cost + }), + assistant("no-cost", "second", { input: 1, output: 2, cacheRead: 0, cacheWrite: 0, totalTokens: 3 }), + assistant("cost-mismatch", "third", { + input: 1, + output: 2, + cacheRead: 0, + cacheWrite: 0, + totalTokens: 3, + cost: { ...cost, total: 1 } + }), + assistant("consistent", "final answer", { + input: 4, + output: 5, + cacheRead: 0, + cacheWrite: 0, + totalTokens: 9, + cost + }), + { type: "agent_end", messages: [{ role: "assistant", content: [{ type: "text", text: "final answer" }] }] } + ]; + const agent = createPiAgent({ model: "openai/gpt-mini-latest" }); + agent.buildCommand = async () => ({ + command: process.execPath, + args: ["-e", lines.map((line) => `console.log(${JSON.stringify(JSON.stringify(line))});`).join("")], + outputFormat: "stream-json" + }); + + // A short idle timeout turns a stalled invocation into a prompt failure + // instead of waiting for the node timeout. + const result = await agent.generate({ prompt: "inconsistent usage", timeout: { idleMs: 5_000 } }); + + assert.equal(result.text, "final answer"); + // Only the consistent response is counted. + assert.deepEqual(result.usage, { + inputTokens: 4, + inputTokenDetails: { noCacheTokens: 4, cacheReadTokens: 0, cacheWriteTokens: 0 }, + outputTokens: 5, + outputTokenDetails: { textTokens: undefined, reasoningTokens: 0 }, + totalTokens: 9, + reportedCostUsd: 0.003 + }); + } finally { + if (previous.config === undefined) delete process.env.ULTRAFUZZ_CONFIG_PATH; + else process.env.ULTRAFUZZ_CONFIG_PATH = previous.config; + if (previous.openRouter === undefined) delete process.env.OPENROUTER_API_KEY; + else process.env.OPENROUTER_API_KEY = previous.openRouter; + } + } +); + bunAdapterTest( "generated OpenCode adapter preserves its isolated environment and keeps credentials out of argv", { timeout: 30_000 }, @@ -8725,7 +8805,7 @@ function kimiCompletedEvent(events: unknown): KimiInterpreterEvent { } bunAdapterTest( - "generated Kimi adapter strictly parses credentials and resume hints without normalization", + "generated Kimi adapter strictly parses credentials and ignores ambiguous resume hints", { timeout: 30_000 }, async () => { const project = tempProject(); @@ -8837,26 +8917,23 @@ bunAdapterTest( const hintAgent = new KimiCode029Agent(kimiSubscriptionOptions(sourceConfig)); const hintInterpreter = hintAgent.createOutputInterpreter(); const session = "00000000-0000-0000-0000-000000000301"; - assert.throws( - () => - hintInterpreter.onStdoutLine?.( - `{"type":"session.resume_hint","type":"session.resume_hint","session_id":${JSON.stringify(session)}}` - ), - /duplicate/iu - ); - assert.throws( - () => hintInterpreter.onStdoutLine?.(JSON.stringify({ type: "session.resume_hint", session_id: ` ${session}` })), - /session_id is invalid/iu - ); - assert.throws( - () => - hintInterpreter.onStdoutLine?.( - " ".repeat(1024 * 1024) + JSON.stringify({ type: "session.resume_hint", session_id: session }) - ), - /1048576-byte limit/iu - ); - assert.throws(() => hintInterpreter.onStdoutLine?.('{"provider_status":'), /Kimi output JSON is invalid/iu); - assert.doesNotThrow(() => hintInterpreter.onStdoutLine?.("Kimi Code provider banner")); + // These hooks run inside the child's stdout/stderr listeners, where a throw + // escapes the invocation, so an ambiguous line carries no session instead. + for (const line of [ + `{"type":"session.resume_hint","type":"session.resume_hint","session_id":${JSON.stringify(session)}}`, + JSON.stringify({ type: "session.resume_hint", session_id: ` ${session}` }), + JSON.stringify({ type: "session.resume_hint", session_id: 301 }), + " ".repeat(1024 * 1024) + JSON.stringify({ type: "session.resume_hint", session_id: session }), + '{"provider_status":', + "{ code: 'ECONNRESET', errno: -104 }", + "Kimi Code provider banner" + ]) { + assert.doesNotThrow(() => hintInterpreter.onStdoutLine?.(line), line.slice(0, 60)); + assert.doesNotThrow(() => hintInterpreter.onStderrLine?.(line), line.slice(0, 60)); + } + assert.equal(hintAgent.issuedSessionId, undefined); + hintInterpreter.onStdoutLine?.(JSON.stringify({ type: "session.resume_hint", session_id: session })); + assert.equal(hintAgent.issuedSessionId, session); } ); @@ -9222,112 +9299,193 @@ bunAdapterTest( ); bunAdapterTest( - "generated Kimi adapter rejects malformed wire records instead of fabricating tokens", - { timeout: 30_000 }, + "generated Kimi adapter reports no usage for an unreadable wire instead of failing the invocation", + { timeout: 60_000 }, async () => { const project = tempProject(); const init = initProject({ projectRoot: project, force: true }); assert.equal(init.ok, true, JSON.stringify(init.diagnostics)); const { KimiCode029Agent } = await loadGeneratedKimiAgent(project); - const sourceConfig = writeKimiSourceConfig(project, "kimi-usage-malformed"); + const sourceConfig = writeKimiSourceConfig(project, "kimi-usage-unreadable"); const options = kimiSubscriptionOptions(sourceConfig); + const tooDeep = `${"[".repeat(34)}null${"]".repeat(34)}`; + const wire = (home: string, name: string) => + path.join(home, "sessions", `wd_${name}`, "session-303", "agents", "main", "wire.jsonl"); + const anomalies: Array<{ label: string; write: (home: string) => void }> = [ + { + label: "malformed records beside a valid one", + write: (home) => + writeKimiWire(home, "wd_malformed/session-203", "main", [ + "not json at all", + "{", + JSON.stringify({ + type: "usage.record", + usage: { inputOther: "12", output: 1, inputCacheRead: 0, inputCacheCreation: 0 } + }), + JSON.stringify({ type: "usage.record", usage: { inputOther: 4, output: 2, inputCacheRead: 1 } }), + kimiUsageRecordLine(33, 7, 2, 1) + ]) + }, + ...[ + [ + "duplicate key", + '{"type":"usage.record","type":"usage.record","usage":{"inputOther":1,"output":2,"inputCacheRead":3,"inputCacheCreation":4}}\n' + ], + ["invalid UTF-8", Buffer.from([0x7b, 0xff, 0x7d, 0x0a])], + ["oversize line", `${JSON.stringify({ type: "message.appended", padding: "x".repeat(1024 * 1024) })}\n`], + ["excessive depth", `{"type":"message.appended","future":${tooDeep}}\n`], + ["torn final record", '{"type":"message.appended"'], + [ + "unsafe aggregate", + `${kimiUsageRecordLine(Number.MAX_SAFE_INTEGER, 0, 0, 0)}\n${kimiUsageRecordLine(0, 1, 0, 0)}\n` + ] + ].map(([label, bytes]) => ({ + label: label as string, + write: (home: string) => { + const target = wire(home, (label as string).replaceAll(" ", "_")); + fs.mkdirSync(path.dirname(target), { recursive: true }); + fs.writeFileSync(target, bytes as string | Buffer); + } + })), + { + label: "wire file budget", + write: (home) => { + for (let index = 0; index < 513; index += 1) { + writeKimiWire(home, "wd_budget/session-211", `agent-${index.toString().padStart(3, "0")}`, [ + kimiUsageRecordLine(1, 1, 0, 0) + ]); + } + } + } + ]; + for (const anomaly of anomalies) { + const agent = new KimiCode029Agent(options); + const command = await agent.buildCommand({ prompt: anomaly.label, cwd: project, options: {} }); + const home = command.env?.KIMI_CODE_HOME; + assert.ok(home); + anomaly.write(home); + let events: unknown; + assert.doesNotThrow(() => { + events = agent.createOutputInterpreter().onExit?.(kimiExitResult(command.args)); + }, anomaly.label); + const completed = kimiCompletedEvent(events); + assert.equal(completed.ok, true, anomaly.label); + // Unknown usage stays absent; it is never partially counted or fabricated. + assert.equal(Object.prototype.hasOwnProperty.call(completed, "usage"), false, anomaly.label); + await command.cleanup?.(); + } - const malformed = new KimiCode029Agent(options); - const malformedCommand = await malformed.buildCommand({ - prompt: "Malformed usage", + // A resumed session whose stored wire was torn by an earlier, killed + // attempt still launches; its usage is unknown rather than miscounted. + const session = "00000000-0000-0000-0000-000000000210"; + const bucket = "wd_target_000000000210"; + const storedSessionDir = path.join(sourceConfig, "sessions", bucket, session); + const storedWire = writeKimiWire(sourceConfig, `${bucket}/${session}`, "main", [ + kimiUsageRecordLine(1_000, 200, 3_000, 40) + ]); + fs.appendFileSync(storedWire, '{"type":"usage.record","usage":{"inputOther":5', "utf8"); + fs.writeFileSync( + path.join(storedSessionDir, "state.json"), + `${JSON.stringify({ + workDir: "/workspace/target", + agents: { + main: { homedir: path.join(storedSessionDir, "agents", "main"), type: "agent", parentAgentId: null } + } + })}\n`, + "utf8" + ); + fs.writeFileSync( + path.join(sourceConfig, "session_index.jsonl"), + `${JSON.stringify({ sessionId: session, sessionDir: storedSessionDir, workDir: "/workspace/target" })}\n`, + "utf8" + ); + const torn = new KimiCode029Agent(options); + const tornCommand = await torn.buildCommand({ + prompt: "Resume a torn wire", cwd: "/workspace/target", - options: {} + options: { resumeSession: session } }); - const malformedHome = malformedCommand.env?.KIMI_CODE_HOME; - assert.ok(malformedHome); - writeKimiWire(malformedHome, "wd_target_000000000203/session-203", "main", [ - "not json at all", - "{", - JSON.stringify({ - type: "usage.record", - usage: { inputOther: "12", output: 1, inputCacheRead: 0, inputCacheCreation: 0 } - }), - JSON.stringify({ - type: "usage.record", - usage: { inputOther: -5, output: 1, inputCacheRead: 0, inputCacheCreation: 0 } - }), - JSON.stringify({ type: "usage.record", usage: { inputOther: 4, output: 2, inputCacheRead: 1 } }), - JSON.stringify({ type: "usage.record", model: "kimi-k3" }), - JSON.stringify({ type: "message.appended", usage: { inputOther: 999, output: 999 } }), - kimiUsageRecordLine(33, 7, 2, 1) - ]); - assert.throws( - () => malformed.createOutputInterpreter().onExit?.(kimiExitResult(malformedCommand.args)), - /invalid strict JSON/iu + const tornHome = tornCommand.env?.KIMI_CODE_HOME; + assert.ok(tornHome); + fs.appendFileSync( + path.join(tornHome, "sessions", bucket, session, "agents", "main", "wire.jsonl"), + `\n${kimiUsageRecordLine(99, 88, 77, 66)}\n`, + "utf8" ); - await malformedCommand.cleanup?.(); - - const absent = new KimiCode029Agent(options); - const absentCommand = await absent.buildCommand({ prompt: "No usage", cwd: "/workspace/target", options: {} }); - const absentHome = absentCommand.env?.KIMI_CODE_HOME; - assert.ok(absentHome); - writeKimiWire(absentHome, "wd_target_000000000204/session-204", "main", [ - kimiWireHeaderLine("session-204"), - "still not json", - JSON.stringify({ type: "usage.record", usage: { inputOther: Number.NaN } }) - ]); - assert.throws( - () => absent.createOutputInterpreter().onExit?.(kimiExitResult(absentCommand.args)), - /invalid strict JSON/iu + const tornCompleted = kimiCompletedEvent(torn.createOutputInterpreter().onExit?.(kimiExitResult(tornCommand.args))); + assert.equal(Object.prototype.hasOwnProperty.call(tornCompleted, "usage"), false); + assert.equal(tornCompleted.resume, session); + await tornCommand.cleanup?.(); + + // A resumed wire replaced during the invocation is no longer comparable + // with its baseline, so its usage is unknown too. + fs.writeFileSync(storedWire, `${kimiUsageRecordLine(1_000, 200, 3_000, 40)}\n`, "utf8"); + const replaced = new KimiCode029Agent(options); + const replacedCommand = await replaced.buildCommand({ + prompt: "Replaced resumed wire", + cwd: "/workspace/target", + options: { resumeSession: session } + }); + const replacedHome = replacedCommand.env?.KIMI_CODE_HOME; + assert.ok(replacedHome); + const runtimeWire = path.join(replacedHome, "sessions", bucket, session, "agents", "main", "wire.jsonl"); + fs.rmSync(runtimeWire); + fs.writeFileSync(runtimeWire, `${kimiUsageRecordLine(99, 88, 77, 66)}\n`, "utf8"); + const replacedCompleted = kimiCompletedEvent( + replaced.createOutputInterpreter().onExit?.(kimiExitResult(replacedCommand.args)) ); - await absentCommand.cleanup?.(); + assert.equal(Object.prototype.hasOwnProperty.call(replacedCompleted, "usage"), false); + await replacedCommand.cleanup?.(); } ); bunAdapterTest( - "generated Kimi adapter rejects ambiguous, invalidly encoded, oversized, or deep wire JSON", + "generated Kimi adapter completes when its output and telemetry are malformed", { timeout: 30_000 }, async () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); const { KimiCode029Agent } = await loadGeneratedKimiAgent(project); - const sourceConfig = writeKimiSourceConfig(project, "kimi-wire-strict-json"); - const tooDeep = `${"[".repeat(34)}null${"]".repeat(34)}`; - const invalidWires: Array<{ label: string; bytes: Buffer; expected: RegExp }> = [ - { - label: "duplicate key", - bytes: Buffer.from( - '{"type":"usage.record","type":"usage.record","usage":{"inputOther":1,"output":2,"inputCacheRead":3,"inputCacheCreation":4}}\n' - ), - expected: /duplicate/iu - }, - { label: "invalid UTF-8", bytes: Buffer.from([0x7b, 0xff, 0x7d, 0x0a]), expected: /UTF-8/iu }, - { - label: "oversize line", - bytes: Buffer.from(`${JSON.stringify({ type: "message.appended", padding: "x".repeat(1024 * 1024) })}\n`), - expected: /line exceeded its byte budget/iu - }, - { - label: "excessive depth", - bytes: Buffer.from(`{"type":"message.appended","future":${tooDeep}}\n`), - expected: /nesting-depth limit of 32/iu - }, - { - label: "torn final record", - bytes: Buffer.from('{"type":"message.appended"'), - expected: /torn or unterminated/iu - } - ]; - for (const [index, fixture] of invalidWires.entries()) { - const agent = new KimiCode029Agent(kimiSubscriptionOptions(sourceConfig)); - const command = await agent.buildCommand({ prompt: fixture.label, cwd: project, options: {} }); + const sourceConfig = writeKimiSourceConfig(project, "kimi-malformed-output"); + const agent = new KimiCode029Agent(kimiSubscriptionOptions(sourceConfig)); + const originalBuildCommand = agent.buildCommand.bind(agent); + agent.buildCommand = async (params) => { + const command = await originalBuildCommand(params); const home = command.env?.KIMI_CODE_HOME; assert.ok(home); - const wire = path.join(home, "sessions", `bucket-${index}`, "session-303", "agents", "main", "wire.jsonl"); - fs.mkdirSync(path.dirname(wire), { recursive: true }); - fs.writeFileSync(wire, fixture.bytes); - assert.throws( - () => agent.createOutputInterpreter().onExit?.(kimiExitResult(command.args)), - fixture.expected, - fixture.label - ); - await command.cleanup?.(); - } + const wire = path.join(home, "sessions", "wd_target_000000000401", "session-401", "agents", "main", "wire.jsonl"); + const script = [ + 'const fs = require("node:fs");', + 'const path = require("node:path");', + `fs.mkdirSync(path.dirname(${JSON.stringify(wire)}), { recursive: true });`, + // A usage record torn mid-write, as a killed CLI leaves it. + `fs.writeFileSync(${JSON.stringify(wire)}, ${JSON.stringify('{"type":"usage.record"')});`, + // A Node util.inspect dump on stderr and a truncated JSON line on stdout. + `console.error(${JSON.stringify("{ code: 'ECONNRESET', errno: -104 }")});`, + `console.log(${JSON.stringify('{"provider_status":')});`, + `console.log(${JSON.stringify(JSON.stringify({ role: "assistant", content: "Done." }))});` + ].join("\n"); + return { + ...command, + command: process.execPath, + args: ["-e", script], + stdin: undefined, + outputFormat: "stream-json" + }; + }; + + // A short idle timeout turns a stalled invocation into a prompt failure + // instead of waiting for the node timeout. + const result = (await agent.generate({ + prompt: "Malformed output", + rootDir: project, + timeout: { idleMs: 5_000 } + })) as { text?: string; usage?: { inputTokens?: number; totalTokens?: number } }; + + assert.equal(result.text, "Done."); + // The torn wire record leaves this invocation's usage unknown. + assert.equal(result.usage?.inputTokens, undefined); + assert.equal(result.usage?.totalTokens, undefined); } ); @@ -9375,111 +9533,6 @@ bunAdapterTest( } ); -bunAdapterTest( - "generated Kimi adapter fails telemetry closed on unsafe wire bounds and replacement", - { timeout: 30_000 }, - async () => { - const project = tempProject(); - const init = initProject({ projectRoot: project, force: true }); - assert.equal(init.ok, true, JSON.stringify(init.diagnostics)); - const { KimiCode029Agent } = await loadGeneratedKimiAgent(project); - const sourceConfig = writeKimiSourceConfig(project, "kimi-usage-bounds"); - const options = kimiSubscriptionOptions(sourceConfig); - - const oversized = new KimiCode029Agent(options); - const oversizedCommand = await oversized.buildCommand({ - prompt: "Oversized usage wire", - cwd: "/workspace/target", - options: {} - }); - const oversizedHome = oversizedCommand.env?.KIMI_CODE_HOME; - assert.ok(oversizedHome); - const oversizedWire = writeKimiWire(oversizedHome, "wd_target_000000000208/session-208", "main", [ - kimiUsageRecordLine(9, 8, 7, 6) - ]); - fs.appendFileSync(oversizedWire, "x".repeat(1024 * 1024 + 1), "utf8"); - assert.throws( - () => oversized.createOutputInterpreter().onExit?.(kimiExitResult(oversizedCommand.args)), - /torn or unterminated|line exceeded/iu - ); - await oversizedCommand.cleanup?.(); - - const overflowing = new KimiCode029Agent(options); - const overflowingCommand = await overflowing.buildCommand({ - prompt: "Overflowing usage values", - cwd: "/workspace/target", - options: {} - }); - const overflowingHome = overflowingCommand.env?.KIMI_CODE_HOME; - assert.ok(overflowingHome); - writeKimiWire(overflowingHome, "wd_target_000000000209/session-209", "main", [ - kimiUsageRecordLine(Number.MAX_SAFE_INTEGER, 0, 0, 0), - // Each component total is independently safe, but the combined total is not. - kimiUsageRecordLine(0, 1, 0, 0) - ]); - assert.throws( - () => overflowing.createOutputInterpreter().onExit?.(kimiExitResult(overflowingCommand.args)), - /safe integer range/iu - ); - await overflowingCommand.cleanup?.(); - - const tooMany = new KimiCode029Agent(options); - const tooManyCommand = await tooMany.buildCommand({ - prompt: "Too many usage wires", - cwd: "/workspace/target", - options: {} - }); - const tooManyHome = tooManyCommand.env?.KIMI_CODE_HOME; - assert.ok(tooManyHome); - for (let index = 0; index < 513; index += 1) { - writeKimiWire(tooManyHome, "wd_target_000000000211/session-211", `agent-${index.toString().padStart(3, "0")}`, [ - kimiUsageRecordLine(1, 1, 0, 0) - ]); - } - assert.throws( - () => tooMany.createOutputInterpreter().onExit?.(kimiExitResult(tooManyCommand.args)), - /file budget/iu - ); - await tooManyCommand.cleanup?.(); - - const session = "00000000-0000-0000-0000-000000000210"; - const bucket = "wd_target_000000000210"; - const storedSessionDir = path.join(sourceConfig, "sessions", bucket, session); - writeKimiWire(sourceConfig, `${bucket}/${session}`, "main", [kimiUsageRecordLine(1_000, 200, 3_000, 40)]); - fs.writeFileSync( - path.join(storedSessionDir, "state.json"), - `${JSON.stringify({ - workDir: "/workspace/target", - agents: { - main: { homedir: path.join(storedSessionDir, "agents", "main"), type: "agent", parentAgentId: null } - } - })}\n`, - "utf8" - ); - fs.writeFileSync( - path.join(sourceConfig, "session_index.jsonl"), - `${JSON.stringify({ sessionId: session, sessionDir: storedSessionDir, workDir: "/workspace/target" })}\n`, - "utf8" - ); - const replaced = new KimiCode029Agent(options); - const replacedCommand = await replaced.buildCommand({ - prompt: "Replaced resumed wire", - cwd: "/workspace/target", - options: { resumeSession: session } - }); - const replacedHome = replacedCommand.env?.KIMI_CODE_HOME; - assert.ok(replacedHome); - const runtimeWire = path.join(replacedHome, "sessions", bucket, session, "agents", "main", "wire.jsonl"); - fs.rmSync(runtimeWire); - fs.writeFileSync(runtimeWire, `${kimiUsageRecordLine(99, 88, 77, 66)}\n`, "utf8"); - assert.throws( - () => replaced.createOutputInterpreter().onExit?.(kimiExitResult(replacedCommand.args)), - /replaced or truncated/iu - ); - await replacedCommand.cleanup?.(); - } -); - bunAdapterTest( "generated Kimi completed-event usage is what pinned Smithers 0.35.0 consumes", { timeout: 30_000 }, From 71c6007b099a064f5f74b43d1470c60d2f4cd825 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:00:44 +0000 Subject: [PATCH 066/206] fix(runtime): native continuations stop blanking adapter-owned paths under the target workflowControlChildEnvironment blanks every child value that contains a controller root. It derived those roots from ULTRAFUZZ_WORKFLOW_ PERSISTED_PATH and from "/modules/", "/controls/", "/dependencies/" and "/.smithers/workflows/" markers in controller variables. A native continuation (`ultrafuzz resume`) persists the target's own .smithers/workflows/*.tsx, so the whole project root became a "controller root" and adapter-supplied paths under it were set to "": OpenCode's run-scoped XDG_* and OPENCODE_DB fell back to the operator's real directories and Kimi's API-key home disappeared. Only Pi had been special-cased (#1035). Derive the roots only from the snapshot names the process anchor advertises (ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT/PROCESS_ROOT/SOURCE_ROOT), which every sealed launch sets and a native continuation does not, and drop the Pi special case. The marker heuristics also matched ordinary install paths (a checkout under any ".../modules/..." directory made its parent a root). ULTRAFUZZ_BUN_MODULE_CONFINEMENT, previously blanked only through those heuristics, is now blanked by name like the other controller-only variables. The continuation test asserted the old over-blanking; it now checks that inherited and adapter-supplied homes under the target survive while paths under an advertised snapshot are still blanked. A new OpenCode test fails on main with XDG_CONFIG_HOME === "". Co-Authored-By: Claude Opus 5.5 --- .../templates/smithers/agents/environment.tsx | 45 +++----- packages/runtime/test/runtime.test.ts | 104 +++++++++++++----- 2 files changed, 93 insertions(+), 56 deletions(-) diff --git a/packages/runtime/src/templates/smithers/agents/environment.tsx b/packages/runtime/src/templates/smithers/agents/environment.tsx index 1b44d8573..f25934279 100644 --- a/packages/runtime/src/templates/smithers/agents/environment.tsx +++ b/packages/runtime/src/templates/smithers/agents/environment.tsx @@ -12,6 +12,7 @@ const CONTROLLER_ONLY_ENVIRONMENT_VARIABLES = [ "SMITHERS_CLI_SRC_DIR", "ULTRAFUZZ_AGENT_ENV_ALLOWLIST", "ULTRAFUZZ_ARTIFACTS_MODULE", + "ULTRAFUZZ_BUN_MODULE_CONFINEMENT", "ULTRAFUZZ_CONFIG_PATH", "ULTRAFUZZ_DATA_DISCLOSURE_ACKNOWLEDGEMENTS", "ULTRAFUZZ_MODAL_PUBLIC_BENCHMARK", @@ -93,22 +94,12 @@ export function workflowControlChildEnvironment( ].map((name) => [name, ""]) ); const roots = workflowExecutionSnapshotRoots(source); - // Native continuations name the target's workflow, not a sealed execution - // tree. Pi's configured home may live under that target. Continue rejecting - // homes that alias a real advertised snapshot or another controller root. - const piHomeRoots = - source.ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT || - source.ULTRAFUZZ_SNAPSHOT_PROCESS_ROOT || - source.ULTRAFUZZ_SNAPSHOT_SOURCE_ROOT - ? roots - : workflowExecutionSnapshotRoots({ ...source, ULTRAFUZZ_WORKFLOW_PERSISTED_PATH: undefined }); const aliasesControl = (name: string, value: string): boolean => { // PATH has already passed the controller's command-path admission. Its // trusted-bin/forge guard intentionally lives under the target; blanking // the whole list also removes external CLIs admitted by the controller. if (name === "PATH") return false; - const selectedRoots = name === "PI_CODING_AGENT_DIR" ? piHomeRoots : roots; - return selectedRoots.some((root) => environmentPath(value).includes(root)); + return roots.some((root) => environmentPath(value).includes(root)); }; for (const [name, value] of Object.entries(source)) { if (value !== undefined && aliasesControl(name, value)) child[name] = ""; @@ -391,26 +382,24 @@ export function workflowControlCredentialValue( return value; } +/** + * The sealed execution snapshot this process runs from, under every name the + * process anchor advertises for it. A native continuation runs the target's + * own workflow and advertises none: treating that project as controller-only + * state would blank every adapter-owned path under it, such as OpenCode's + * run-scoped XDG roots and Kimi's API-key home. + */ function workflowExecutionSnapshotRoots(source: Record): string[] { const roots = new Set(); - const persistedWorkflow = source.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH; - if (persistedWorkflow !== undefined && path.isAbsolute(persistedWorkflow)) { - const workflowsDirectory = path.dirname(persistedWorkflow); - const smithersDirectory = path.dirname(workflowsDirectory); - if (path.basename(workflowsDirectory) === "workflows" && path.basename(smithersDirectory) === ".smithers") { - roots.add(path.dirname(smithersDirectory)); - } - } - for (const name of CONTROLLER_ONLY_ENVIRONMENT_VARIABLES) { - const value = source[name]; - if (value === undefined) continue; - const candidate = environmentPath(value); - for (const marker of ["/dependencies/", "/modules/", "/controls/", "/.smithers/workflows/"]) { - const index = candidate.indexOf(marker); - if (index > 0) roots.add(candidate.slice(0, index)); - } + for (const name of [ + "ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT", + "ULTRAFUZZ_SNAPSHOT_PROCESS_ROOT", + "ULTRAFUZZ_SNAPSHOT_SOURCE_ROOT" + ]) { + const root = source[name]?.trim(); + if (root && path.isAbsolute(root) && root !== path.parse(root).root) roots.add(root); } - return [...roots].filter((root) => root !== path.parse(root).root); + return [...roots]; } function environmentPath(value: string): string { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index d01b8ee12..bc85851d1 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -6961,19 +6961,19 @@ bunAdapterTest( const previous = { config: process.env.ULTRAFUZZ_CONFIG_PATH, - workflow: process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH, + snapshot: process.env.ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT, alias: process.env.MY_ALIAS }; process.env.ULTRAFUZZ_CONFIG_PATH = configPath; - process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH = persistedWorkflow; + process.env.ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT = snapshotRoot; process.env.MY_ALIAS = aliasedControlPath; try { assert.throws(() => createCodexAgent(), /credential MY_ALIAS resolves inside controller-only execution state/u); } finally { if (previous.config === undefined) delete process.env.ULTRAFUZZ_CONFIG_PATH; else process.env.ULTRAFUZZ_CONFIG_PATH = previous.config; - if (previous.workflow === undefined) delete process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH; - else process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH = previous.workflow; + if (previous.snapshot === undefined) delete process.env.ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT; + else process.env.ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT = previous.snapshot; if (previous.alias === undefined) delete process.env.MY_ALIAS; else process.env.MY_ALIAS = previous.alias; } @@ -7178,47 +7178,95 @@ bunAdapterTest( initProject({ projectRoot: project, force: true }); const { workflowControlChildEnvironment } = await loadGeneratedCodexAgent(project); const externalBin = temporaryRoot("ultrafuzz-continuation-external-bin-"); - const trustedBin = path.join(project, ".ultrafuzz", "runs", "continued", "trusted-bin"); + const runRoot = path.join(project, ".ultrafuzz", "runs", "continued"); + const trustedBin = path.join(runRoot, "trusted-bin"); const admittedPath = [trustedBin, externalBin].join(path.delimiter); - const piHome = path.join(project, ".ultrafuzz", "pi-coding-agent"); + // Homes adapters place under the target: Pi's configured home, OpenCode's + // run-scoped XDG roots, and Kimi's API-key home. + const adapterOwned = { + PI_CODING_AGENT_DIR: path.join(project, ".ultrafuzz", "pi-coding-agent"), + XDG_CONFIG_HOME: path.join(runRoot, "opencode", "config"), + OPENCODE_DB: path.join(runRoot, "opencode", "data", "opencode", "opencode.db"), + KIMI_CODE_HOME: path.join(project, ".ultrafuzz", "kimi-code") + }; const source = { PATH: admittedPath, - PI_CODING_AGENT_DIR: piHome, ULTRAFUZZ_WORKFLOW_PERSISTED_PATH: path.join(project, ".smithers", "workflows", "continued.tsx"), - ULTRAFUZZ_CONFIG_PATH: path.join(project, ".ultrafuzz", "runs", "continued", "smithers", "resolved-config.json"), - CONTROL_ALIAS: path.join(project, ".smithers", "workflows", "continued.tsx"), + ULTRAFUZZ_CONFIG_PATH: path.join(runRoot, "smithers", "resolved-config.json"), OPENAI_API_KEY: "unrelated-provider-key" }; - for (const additions of [{}, { PATH: admittedPath, PI_CODING_AGENT_DIR: piHome }]) { - const child = { ...source, ...workflowControlChildEnvironment(additions, source) }; + // A native continuation runs the target's own workflow and advertises no + // sealed snapshot, so nothing under the target is controller-only state: + // neither inherited values nor the homes an adapter supplies are blanked. + // (An inherited KIMI_CODE_HOME is still dropped as a provider home.) + const { KIMI_CODE_HOME: _kimiHome, ...inherited } = adapterOwned; + const continuation = { ...source, ...inherited }; + for (const [additions, expected] of [ + [{}, inherited], + [{ PATH: admittedPath, ...adapterOwned }, adapterOwned] + ] as const) { + const child: Record = { + ...continuation, + ...workflowControlChildEnvironment(additions, continuation) + }; assert.equal(child.PATH, admittedPath); - assert.equal(child.PI_CODING_AGENT_DIR, piHome); - assert.equal(child.CONTROL_ALIAS, ""); + for (const [name, value] of Object.entries(expected)) assert.equal(child[name], value, name); assert.equal(child.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH, ""); + assert.equal(child.ULTRAFUZZ_CONFIG_PATH, ""); assert.equal(child.OPENAI_API_KEY, ""); } - const snapshotRoot = path.join( - project, - ".ultrafuzz", - "runs", - "continued", - "smithers", - "execution-snapshots", - "a".repeat(64) - ); + const snapshotRoot = path.join(runRoot, "smithers", "execution-snapshots", "a".repeat(64)); const snapshotHome = path.join(snapshotRoot, "controls"); const snapshotSource = { ...source, ULTRAFUZZ_WORKFLOW_PERSISTED_PATH: path.join(snapshotRoot, ".smithers", "workflows", "continued.tsx"), ULTRAFUZZ_CONFIG_PATH: path.join(snapshotRoot, "controls", "ultrafuzz.toml"), ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT: snapshotRoot, - PI_CODING_AGENT_DIR: snapshotHome + PI_CODING_AGENT_DIR: snapshotHome, + XDG_CONFIG_HOME: snapshotHome }; - assert.equal(workflowControlChildEnvironment({}, snapshotSource).PI_CODING_AGENT_DIR, ""); - assert.equal( - workflowControlChildEnvironment({ PI_CODING_AGENT_DIR: snapshotHome }, snapshotSource).PI_CODING_AGENT_DIR, - "" - ); + for (const additions of [{}, { PI_CODING_AGENT_DIR: snapshotHome, XDG_CONFIG_HOME: snapshotHome }]) { + const child = workflowControlChildEnvironment(additions, snapshotSource); + assert.equal(child.PI_CODING_AGENT_DIR, ""); + assert.equal(child.XDG_CONFIG_HOME, ""); + } + } +); + +bunAdapterTest( + "generated OpenCode adapter keeps its run-scoped state roots in a native continuation", + { timeout: 30_000 }, + async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const { createOpenCodeAgent } = await loadGeneratedOpenCodeAgent(project); + const runRoot = path.join(project, ".ultrafuzz", "runs", "opencode-continued"); + const names = ["ULTRAFUZZ_CONFIG_PATH", "ULTRAFUZZ_WORKFLOW_PERSISTED_PATH", "OPENROUTER_API_KEY"] as const; + const saved = Object.fromEntries(names.map((name) => [name, process.env[name]])); + process.env.ULTRAFUZZ_CONFIG_PATH = path.join(project, "ultrafuzz.toml"); + // What `ultrafuzz resume` hands a native continuation: the target's own + // persisted workflow and no advertised execution snapshot. + process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH = path.join(project, ".smithers", "workflows", "continued.tsx"); + process.env.OPENROUTER_API_KEY = "sk-or-v1-not-a-real-opencode-credential"; + try { + const agent = createOpenCodeAgent({ + model: "openrouter/test-model", + addDir: [path.join(runRoot, "artifacts", "attempt")] + }); + const command = await agent.buildCommand({ prompt: "inspect", cwd: project, options: {} }); + const childEnv = { ...process.env, ...agent.opts.env, ...command.env }; + const stateRoot = path.join(runRoot, "opencode"); + assert.equal(childEnv.XDG_CONFIG_HOME, path.join(stateRoot, "config")); + assert.equal(childEnv.XDG_DATA_HOME, path.join(stateRoot, "data")); + assert.equal(childEnv.OPENCODE_DB, path.join(stateRoot, "data", "opencode", "opencode.db")); + assert.equal(childEnv.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH, ""); + } finally { + for (const name of names) { + const value = saved[name]; + if (value === undefined) Reflect.deleteProperty(process.env, name); + else process.env[name] = value; + } + } } ); From 427cba98dce79a1700c0c5db3651e2c1ac722a94 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:28:43 +0000 Subject: [PATCH 067/206] fix(runtime): route digests cover only route-bearing input, with one implementation The provider route ID an adapter re-verifies on every invocation was computed by two hand-maintained copies (environment.tsx for the adapters, data-governance.ts at plan time) and both digested input that does not select a destination: - a Codex config.toml with any provider line was hashed whole, so the CLI's own rewrites ([marketplaces.*] last_updated, [projects.*] trust_level, #908) failed every later task; - proxy variables counted for every agent, and AWS_/GOOGLE_/AZURE_/ FOUNDRY_ variables counted for Claude even without a cloud platform flag (FOUNDRY_PROFILE is forge's profile variable), so resuming from another shell failed every task with "provider route changed after disclosure acknowledgement". providerRouteDestination in data-governance.ts is now the only implementation. Plan-time modelDestination calls it, and the generated environment.tsx loads it from the runtime module the rendered workflow already imports (ULTRAFUZZ_RUNTIME_MODULE, with the same snapshot fallback as workflow.tsx). It digests the agent's endpoint and platform variables, with Claude cloud prefixes counted only while a matching CLAUDE_CODE_USE_* flag is set; the Codex provider that model_provider (or a profile) selects, limited to its id, base_url, wire_api and env_key and parsed with smol-toml; Claude settings helpers and routing env entries; and, as before, the whole Kimi config. A config file the reader cannot parse is digested by its bytes, as before. Proxies no longer count. effectiveRouteEnvironment keeps its wider set for cloud credential forwarding. Route IDs that included the dropped inputs change, so affected policies must be re-acknowledged and affected paused runs re-planned. New tests fail on main with "provider route changed" for a proxy change, a Codex rewrite of unrelated sections, and changed ambient Claude cloud settings. Co-Authored-By: Claude Opus 5.5 --- docs/security.md | 14 + packages/runtime/package.json | 1 + packages/runtime/src/data-governance.ts | 250 ++++++++++++------ .../templates/smithers/agents/environment.tsx | 160 +++-------- packages/runtime/test/data-governance.test.ts | 67 ++++- packages/runtime/test/runtime.test.ts | 185 +++++++++++-- pnpm-lock.yaml | 3 + 7 files changed, 441 insertions(+), 239 deletions(-) diff --git a/docs/security.md b/docs/security.md index 9b5c134dc..e2d9cd2e5 100644 --- a/docs/security.md +++ b/docs/security.md @@ -93,6 +93,20 @@ remain fail-closed pending the separate R-26 disclosure authorization. Public Modal runs record `cloud:modal`. These controls are not a sandbox or egress filter: YOLO agents remain unrestricted. +A model destination is `model:`, or `model:-route-` +when any route-bearing input is present. The digest covers only that input: the +agent's endpoint and platform environment variables, the provider that a Codex +`config.toml` selects through `model_provider` (its id, `base_url`, `wire_api`, +and `env_key`), Claude `settings.json` credential helpers and routing `env` +entries, and, for Kimi subscription auth, the whole Kimi `config.toml`. +Claude's `AWS_`, `GOOGLE_`/`CLOUD_ML_`, and `AZURE_`/`FOUNDRY_` variables count +only while a `CLAUDE_CODE_USE_*` flag for that platform is set in the environment +or in `settings.json` `env` (for example `CLAUDE_CODE_USE_BEDROCK`, +`CLAUDE_CODE_USE_VERTEX`, or `CLAUDE_CODE_USE_FOUNDRY`). Proxy variables and a CLI's own rewrites of +unrelated config sections do not change the route. Every agent +invocation recomputes the destination with the same runtime function and fails +if it no longer matches an acknowledged one. + ## Production dependency advisories CI and release validation run `pnpm security:dependency-advisories`. The gate diff --git a/packages/runtime/package.json b/packages/runtime/package.json index e80034ba5..787e9e0a8 100644 --- a/packages/runtime/package.json +++ b/packages/runtime/package.json @@ -37,6 +37,7 @@ "mdast-util-from-markdown": "2.0.3", "npm": "11.19.0", "proper-lockfile": "4.1.2", + "smol-toml": "1.7.1", "typescript": "6.0.3", "zod": "4.4.3" }, diff --git a/packages/runtime/src/data-governance.ts b/packages/runtime/src/data-governance.ts index 72f6f1d2a..520e11101 100644 --- a/packages/runtime/src/data-governance.ts +++ b/packages/runtime/src/data-governance.ts @@ -6,6 +6,7 @@ import path from "node:path"; import { parseStrictJsonBytes, readSinglyLinkedRegularFileSnapshotInside } from "@ultrafuzz/artifacts"; import type { ResolvedConfig } from "@ultrafuzz/config"; import { isSensitiveEnvironmentName } from "@ultrafuzz/security"; +import { parse as parseToml } from "smol-toml"; import { retryFallbackProfileIds } from "./retry-chain.js"; import { DATA_DISCLOSURE_ACKNOWLEDGEMENTS_JSON_SCHEMA_ID, @@ -43,6 +44,27 @@ const ROUTE_PROXY_ENV = [ "https_proxy", "no_proxy" ] as const; +// Claude Code reads cloud-provider settings only for the platform a +// CLAUDE_CODE_USE_* flag selects (the flags Claude Code 2.1.284 checks). +// Without one an ambient AWS_PROFILE or GOOGLE_CLOUD_PROJECT routes nothing, +// yet pinning it made the acknowledged route depend on which shell resumed. +const CLAUDE_CLOUD_ROUTE_PREFIXES: Readonly> = { + CLAUDE_CODE_USE_ANTHROPIC_AWS: ["AWS_"], + CLAUDE_CODE_USE_ANTHROPIC_GOOGLE_CLOUD: ["CLOUD_ML_", "GOOGLE_"], + CLAUDE_CODE_USE_BEDROCK: ["AWS_"], + CLAUDE_CODE_USE_FOUNDRY: ["AZURE_", "FOUNDRY_"], + CLAUDE_CODE_USE_MANTLE: ["AWS_"], + CLAUDE_CODE_USE_VERTEX: ["CLOUD_ML_", "GOOGLE_"] +}; +const PROVIDER_DESTINATIONS: Readonly> = { + ClaudeAgent: "anthropic", + CodexAgent: "openai", + DeepSeekAgent: "deepseek", + KimiAgent: "moonshot", + OpenCodeAgent: "openrouter", + OpenRouterAgent: "openrouter", + PiAgent: "openrouter" +}; export interface DataGovernanceDestinationPolicy { destination: string; processor: string; @@ -301,102 +323,162 @@ function requiredDestinations(config: ResolvedConfig, graph: PlannedGraph, env: }; } export function modelDestination(agent: string, config: ResolvedConfig, env: NodeJS.ProcessEnv): string { - const builtins: Record = { - CodexAgent: "openai", - ClaudeAgent: "anthropic", - KimiAgent: "moonshot", - DeepSeekAgent: "deepseek", - OpenCodeAgent: "openrouter", - OpenRouterAgent: "openrouter", - PiAgent: "openrouter" - }, - route = effectiveRoute(agent, config, env); - if (route !== undefined) return `model:${agent.toLowerCase().replace("agent", "")}-route-${route}`; - if (builtins[agent] === undefined) throw new Error(`cannot derive a data destination for ${agent}`); - return `model:${builtins[agent]}`; + const routeConfig = providerHomeRouteConfig(agent, config, env); + if ( + routeConfig !== undefined && + config.execution.mode === "cloud" && + routeBearingConfig(agent, routeConfig, env) !== undefined + ) + throw new Error( + "cloud execution cannot use host provider-home routing; select and acknowledge the route through environment variables" + ); + return providerRouteDestination(agent, env, routeConfig); } -function effectiveRoute(agent: string, config: ResolvedConfig, env: NodeJS.ProcessEnv): string | undefined { - const routeEnvironment = effectiveRouteEnvironment(agent, env), - provider = { CodexAgent: "codex", KimiAgent: "kimi", ClaudeAgent: "claude" }[agent]; +/** + * The data destination an agent's model traffic reaches. Plan-time disclosure + * acknowledgement calls this, and every generated adapter re-verifies its + * invocation with this same function (loaded through the runtime module), so + * there is no second implementation to keep in step. Only route-bearing input + * participates: proxies, unused cloud-provider settings, and a provider CLI's + * rewrites of unrelated config sections change nothing. + */ +export function providerRouteDestination( + agent: string, + env: Record, + routeConfig?: Uint8Array +): string { + // A CLAUDE_CODE_USE_* flag in Claude's settings env selects the platform for + // the process environment's cloud variables too, and vice versa. + const settingsEnv = + agent === "ClaudeAgent" && routeConfig !== undefined ? claudeSettings(routeConfig)?.env : undefined, + route = routeBearingEnvironment(agent, env, [env, settingsEnv ?? {}]), + config = routeConfig === undefined ? undefined : routeBearingConfig(agent, routeConfig, env); + if (route.length === 0 && config === undefined) { + const destination = PROVIDER_DESTINATIONS[agent]; + if (destination === undefined) throw new Error(`cannot derive a data destination for ${agent}`); + return `model:${destination}`; + } + return `model:${agent.toLowerCase().replace("agent", "")}-route-${sha256Stable({ agent, config: config ?? null, route })}`; +} +/** The provider CLI's own config file, when that file can select this agent's route. */ +function providerHomeRouteConfig(agent: string, config: ResolvedConfig, env: NodeJS.ProcessEnv): Buffer | undefined { + const provider = { CodexAgent: "codex", KimiAgent: "kimi", ClaudeAgent: "claude" }[agent]; if (provider === undefined || (agent === "KimiAgent" && config.agents?.KimiAgent?.auth === "api-key")) - return routeEnvironment.length > 0 ? sha256Stable({ agent, config: null, route: routeEnvironment }) : undefined; + return undefined; const configured = config.agents?.[agent]?.configDir, - selectedRoot = env.ULTRAFUZZ_PROVIDER_HOME_ROOT?.trim(), - userHome = env.HOME?.trim() || os.homedir(), - defaultRoot = path.join( + home = providerHome(agent, provider, configured, env), + routeConfig = path.join(home, agent === "ClaudeAgent" ? "settings.json" : "config.toml"); + if (!fs.existsSync(routeConfig)) return undefined; + return readSinglyLinkedRegularFileSnapshotInside(home, routeConfig, 1024 * 1024, "provider route config"); +} +/** The home directory the stock adapter gives this provider's CLI. */ +function providerHome(agent: string, provider: string, configured: string | undefined, env: NodeJS.ProcessEnv): string { + const selectedRoot = env.ULTRAFUZZ_PROVIDER_HOME_ROOT?.trim(), + userHome = env.HOME?.trim() || os.homedir(); + if (configured) { + const defaultRoot = path.join( env.XDG_STATE_HOME?.trim() || path.join(userHome, ".local", "state"), "ultrafuzz", "provider-homes" - ), - home = configured - ? path.join(selectedRoot || defaultRoot, provider, configured) - : selectedRoot - ? path.join(selectedRoot, provider) - : agent === "CodexAgent" - ? env.CODEX_HOME?.trim() || path.join(userHome, ".codex") - : agent === "KimiAgent" - ? env.KIMI_CODE_HOME?.trim() || env.KIMI_SHARE_DIR?.trim() || path.join(userHome, ".kimi-code") - : env.CLAUDE_CONFIG_DIR?.trim() || path.join(userHome, ".claude"), - routeConfig = path.join(home, agent === "ClaudeAgent" ? "settings.json" : "config.toml"); - let configDigest: string | undefined; - if (fs.existsSync(routeConfig)) { - const bytes = readSinglyLinkedRegularFileSnapshotInside(home, routeConfig, 1024 * 1024, "provider route config"); - const affectsRoute = - agent === "ClaudeAgent" - ? claudeSettingsAffectRoute(bytes) - : agent !== "CodexAgent" || codexConfigAffectsRoute(bytes.toString("utf8")); - if (affectsRoute) configDigest = hash(bytes); - if (configDigest !== undefined && config.execution.mode === "cloud") - throw new Error( - "cloud execution cannot use host provider-home routing; select and acknowledge the route through environment variables" - ); + ); + return path.join(selectedRoot || defaultRoot, provider, configured); } - return routeEnvironment.length > 0 - ? sha256Stable({ agent, config: configDigest ?? null, route: routeEnvironment }) - : configDigest; + if (selectedRoot) return path.join(selectedRoot, provider); + if (agent === "CodexAgent") return env.CODEX_HOME?.trim() || path.join(userHome, ".codex"); + if (agent === "KimiAgent") + return env.KIMI_CODE_HOME?.trim() || env.KIMI_SHARE_DIR?.trim() || path.join(userHome, ".kimi-code"); + return env.CLAUDE_CONFIG_DIR?.trim() || path.join(userHome, ".claude"); } /** - * The Codex CLI rewrites its own config.toml on invocation — marketplace - * `last_updated` timestamps, plugin toggles, and project trust levels — so - * digesting the whole file makes the acknowledged route change the moment the - * CLI first runs in a fresh HOME, which failed every sandbox agent task after - * disclosure (#908). Mirror claudeSettingsAffectRoute: only content that can - * actually redirect traffic — a `model_provider` selection, a - * `[model_providers…]` table, or a `base_url` assignment, the same fields - * codexProviderRouting reads — participates in the route digest. A config - * that gains any of these after acknowledgement still fails closed. + * Environment entries that select where an agent's model traffic goes. A + * Claude cloud prefix counts while any of `selectors` sets its platform flag. */ -function codexConfigAffectsRoute(text: string): boolean { - return ( - /(?:^|\n)\s*(?:model_provider|"model_provider"|'model_provider')\s*=/u.test(text) || - /(?:^|\n)\s*\[[^\]\n]*model_providers[^\]\n]*\]/u.test(text) || - /(?:^|\n)\s*(?:base_url|"base_url"|'base_url')\s*=/u.test(text) +function routeBearingEnvironment( + agent: string, + env: Record, + selectors: ReadonlyArray> +): Array<[string, string]> { + const selected = new Set( + Object.entries(CLAUDE_CLOUD_ROUTE_PREFIXES).flatMap(([flag, prefixes]) => + selectors.some((source) => source[flag]?.trim()) ? prefixes : [] + ) + ), + inactivePrefixes = + agent === "ClaudeAgent" + ? Object.values(CLAUDE_CLOUD_ROUTE_PREFIXES) + .flat() + .filter((prefix) => !selected.has(prefix)) + : []; + return effectiveRouteEnvironment(agent, env).filter( + ([name]) => !ROUTE_PROXY_ENV.includes(name as never) && !inactivePrefixes.some((prefix) => name.startsWith(prefix)) ); } -function claudeSettingsAffectRoute(bytes: Buffer): boolean { - const parsed = parseStrictJsonBytes(bytes, { - maxBytes: 1024 * 1024, - maxDepth: 32, - maxItems: 4096, - maxProperties: 4096 - }); - if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) - throw new Error("Claude settings must be a JSON object"); - if (Object.keys(parsed).some((name) => /(?:helper|refresh|credentialexport|processwrapper|proxyauth)$/iu.test(name))) - return true; - const configuredEnv = (parsed as Record).env; - if (configuredEnv === undefined) return false; - if (configuredEnv === null || typeof configuredEnv !== "object" || Array.isArray(configuredEnv)) - throw new Error("Claude settings env must be a JSON object"); - return Object.keys(configuredEnv).some((name) => { - const upper = name.toUpperCase(); - return ( - ["ALL_PROXY", "HTTP_PROXY", "HTTPS_PROXY", "NO_PROXY"].includes(upper) || - (!NON_ROUTING_PROVIDER_ENVIRONMENT_NAMES.has(upper) && - !isCredentialLikeEnvironmentVariableName(upper) && - ROUTE_ENV_PREFIXES.ClaudeAgent!.some((prefix) => upper.startsWith(prefix))) - ); - }); +/** + * The part of a provider CLI's own config file that can redirect traffic, or + * undefined when it selects no route. The CLIs rewrite unrelated sections of + * these files themselves (Codex refreshes marketplace timestamps and project + * trust levels, #908), so only these fields participate: + * - Codex `config.toml`: the selected `model_provider` (a `profile` may select + * it) and that provider's `base_url`, `wire_api`, and `env_key`; + * - Claude `settings.json`: credential/process helper keys and routing `env`; + * - Kimi `config.toml`: the whole file. + * A file these readers cannot parse routes by its exact bytes, as before. + */ +function routeBearingConfig(agent: string, bytes: Uint8Array, env: Record): unknown { + if (agent === "CodexAgent") return codexRouteConfig(bytes); + if (agent === "ClaudeAgent") return claudeRouteConfig(bytes, env); + return hash(bytes); +} +function codexRouteConfig(bytes: Uint8Array): unknown { + let config: Record; + try { + config = parseToml(Buffer.from(bytes).toString("utf8")); + } catch { + return { unparsed: hash(bytes) }; + } + const profile = typeof config.profile === "string" ? record(record(config.profiles)?.[config.profile]) : undefined, + selected = profile?.model_provider ?? config.model_provider; + if (typeof selected !== "string") return undefined; + const provider = record(record(config.model_providers)?.[selected]) ?? {}; + return { + model_provider: selected, + base_url: provider.base_url ?? null, + wire_api: provider.wire_api ?? null, + env_key: provider.env_key ?? null + }; +} +function claudeRouteConfig(bytes: Uint8Array, env: Record): unknown { + const parsed = claudeSettings(bytes); + if (parsed === undefined) return { unparsed: hash(bytes) }; + const helpers = Object.entries(parsed.settings) + .filter(([name]) => /(?:helper|refresh|credentialexport|processwrapper|proxyauth)$/iu.test(name)) + .sort(([left], [right]) => (left < right ? -1 : left > right ? 1 : 0)), + routeEnv = routeBearingEnvironment("ClaudeAgent", parsed.env, [env, parsed.env]); + return helpers.length === 0 && routeEnv.length === 0 ? undefined : { helpers, env: routeEnv }; +} +/** Claude settings.json, or undefined when it is not an object with an object `env`. */ +function claudeSettings( + bytes: Uint8Array +): { settings: Record; env: Record } | undefined { + let settings: Record | undefined; + try { + settings = record(JSON.parse(Buffer.from(bytes).toString("utf8"))); + } catch { + return undefined; + } + const env = settings === undefined ? undefined : record(settings.env ?? {}); + if (settings === undefined || env === undefined) return undefined; + return { + settings, + env: Object.fromEntries( + Object.entries(env).filter((entry): entry is [string, string] => typeof entry[1] === "string") + ) + }; +} +function record(value: unknown): Record | undefined { + return value !== null && typeof value === "object" && !Array.isArray(value) + ? (value as Record) + : undefined; } export function effectiveRouteEnvironment(agent: string, env: NodeJS.ProcessEnv): Array<[string, string]> { const names = new Set( diff --git a/packages/runtime/src/templates/smithers/agents/environment.tsx b/packages/runtime/src/templates/smithers/agents/environment.tsx index f25934279..98e5faabc 100644 --- a/packages/runtime/src/templates/smithers/agents/environment.tsx +++ b/packages/runtime/src/templates/smithers/agents/environment.tsx @@ -1,4 +1,3 @@ -import { createHash } from "node:crypto"; import { existsSync } from "node:fs"; import path from "node:path"; import { fileURLToPath } from "node:url"; @@ -7,6 +6,20 @@ import { parseStrictJsonBytes, readRegularFileSnapshot } from "./strict-json"; export const PROVIDER_SCOPED_SENSITIVE_ENVIRONMENT_CAPABILITY = "ultrafuzz.provider-scoped-sensitive-environment.v1" as const; +type ProviderRouteDestination = ( + agent: string, + env: Record, + routeConfig?: Uint8Array +) => string; +// Plan-time disclosure acknowledgement computes route IDs with +// @ultrafuzz/runtime's providerRouteDestination. Re-verify with that same +// function, loaded from the module the rendered workflow imports, rather than +// with a second copy that has to be kept in step with it. +const { providerRouteDestination } = (await import( + process.env.ULTRAFUZZ_RUNTIME_MODULE ?? + new URL("../../modules/@ultrafuzz/runtime/dist/index.js", import.meta.url).href +)) as { providerRouteDestination: ProviderRouteDestination }; + const CONTROLLER_ONLY_ENVIRONMENT_VARIABLES = [ "SMITHERS_BIN", "SMITHERS_CLI_SRC_DIR", @@ -60,21 +73,10 @@ const ROUTE_ENV_PREFIXES: Readonly> = { CodexAgent: ["AZURE_OPENAI_", "OPENAI_"], KimiAgent: ["KIMI_", "MOONSHOT_"] }; -const NON_ROUTING_PROVIDER_ENVIRONMENT_NAMES = new Set(["AZURE_EXTENSION_DIR"]); // Keep this generated, dependency-free boundary in parity with // @ultrafuzz/security's isSensitiveEnvironmentName contract. const SENSITIVE_ENVIRONMENT_NAME_PATTERN = /(?:^|_)(?:API_?KEY|TOKEN|SECRET|PASSWORD|PASSWD|PRIVATE_?KEY|ACCESS_?KEY|CLIENT_?SECRET|CREDENTIALS?|AUTH(?:ORIZATION)?)(?:_|$)/iu; -const ROUTE_PROXY_ENV = [ - "ALL_PROXY", - "HTTP_PROXY", - "HTTPS_PROXY", - "NO_PROXY", - "all_proxy", - "http_proxy", - "https_proxy", - "no_proxy" -] as const; /** * Smithers agents inherit the controller environment by default. Remove every @@ -249,127 +251,27 @@ function assertWorkflowDataRoute( required = (record as { required_source_destinations?: unknown }).required_source_destinations; if (!Array.isArray(required) || required.some((entry) => typeof entry !== "string")) throw new Error("sealed data-governance authority is invalid"); - const effective = effectiveWorkflowDataRoute(route, effectiveEnvironment, authorityEnvironment); + const configPath = + route.configDir === undefined || route.agent === "OpenRouterAgent" + ? undefined + : path.join(route.configDir, route.agent === "ClaudeAgent" ? "settings.json" : "config.toml"); + const effective = providerRouteDestination( + route.agent, + { + ...effectiveEnvironment, + ULTRAFUZZ_AGENT_ENV_ALLOWLIST: authorityEnvironment.ULTRAFUZZ_AGENT_ENV_ALLOWLIST, + // The Codex adapter derives OPENAI_BASE_URL from the provider config, + // which the config part of the route already covers. + ...(route.agent === "CodexAgent" && !authorityEnvironment.OPENAI_BASE_URL?.trim() + ? { OPENAI_BASE_URL: undefined } + : {}) + }, + configPath !== undefined && existsSync(configPath) ? readRegularFileSnapshot(configPath, 1024 * 1024) : undefined + ); if (!required.includes(effective)) throw new Error(`effective ${route.agent} provider route changed after disclosure acknowledgement`); } -function effectiveWorkflowDataRoute( - route: WorkflowDataRoute, - source: Record, - authority: Record -): string { - const routeSource = { ...source, ULTRAFUZZ_AGENT_ENV_ALLOWLIST: authority.ULTRAFUZZ_AGENT_ENV_ALLOWLIST }, - routeEnvironment = effectiveRouteEnvironment( - route.agent, - route.agent === "CodexAgent" && !authority.OPENAI_BASE_URL?.trim() - ? { ...routeSource, OPENAI_BASE_URL: undefined } - : routeSource - ); - let configDigest: string | undefined; - if (route.configDir !== undefined && route.agent !== "OpenRouterAgent") { - const configPath = path.join(route.configDir, route.agent === "ClaudeAgent" ? "settings.json" : "config.toml"); - if (existsSync(configPath)) { - const bytes = readRegularFileSnapshot(configPath, 1024 * 1024); - const affectsRoute = - route.agent === "ClaudeAgent" - ? claudeSettingsAffectRoute(bytes) - : route.agent !== "CodexAgent" || codexConfigAffectsRoute(bytes.toString("utf8")); - if (affectsRoute) configDigest = sha256(bytes); - } - } - const digest = - routeEnvironment.length > 0 - ? sha256(JSON.stringify({ agent: route.agent, config: configDigest ?? null, route: routeEnvironment })) - : configDigest, - provider = { - ClaudeAgent: "anthropic", - CodexAgent: "openai", - DeepSeekAgent: "deepseek", - KimiAgent: "moonshot", - OpenRouterAgent: "openrouter" - }[route.agent]; - return digest === undefined - ? `model:${provider}` - : `model:${route.agent.toLowerCase().replace("agent", "")}-route-${digest}`; -} - -/** - * The Codex CLI rewrites its own config.toml on invocation — marketplace - * `last_updated` timestamps, plugin toggles, and project trust levels — so - * digesting the whole file makes the acknowledged route change the moment the - * CLI first runs in a fresh HOME, which failed every sandbox agent task after - * disclosure (#908). Keep this generated copy in exact parity with - * @ultrafuzz/runtime data-governance.ts: only content that can actually - * redirect traffic — a `model_provider` selection, a `[model_providers…]` - * table, or a `base_url` assignment — participates in the route digest. - */ -function codexConfigAffectsRoute(text: string): boolean { - return ( - /(?:^|\n)\s*(?:model_provider|"model_provider"|'model_provider')\s*=/u.test(text) || - /(?:^|\n)\s*\[[^\]\n]*model_providers[^\]\n]*\]/u.test(text) || - /(?:^|\n)\s*(?:base_url|"base_url"|'base_url')\s*=/u.test(text) - ); -} - -function claudeSettingsAffectRoute(bytes: Buffer): boolean { - const parsed = parseStrictJsonBytes(bytes, { - maxBytes: 1024 * 1024, - maxDepth: 32, - maxItems: 4096, - maxProperties: 4096 - }); - if (parsed === null || typeof parsed !== "object" || Array.isArray(parsed)) - throw new Error("Claude settings must be a JSON object"); - if (Object.keys(parsed).some((name) => /(?:helper|refresh|credentialexport|processwrapper|proxyauth)$/iu.test(name))) - return true; - const configuredEnv = (parsed as Record).env; - if (configuredEnv === undefined) return false; - if (configuredEnv === null || typeof configuredEnv !== "object" || Array.isArray(configuredEnv)) - throw new Error("Claude settings env must be a JSON object"); - return Object.keys(configuredEnv).some((name) => { - const upper = name.toUpperCase(); - return ( - ["ALL_PROXY", "HTTP_PROXY", "HTTPS_PROXY", "NO_PROXY"].includes(upper) || - (!NON_ROUTING_PROVIDER_ENVIRONMENT_NAMES.has(upper) && - !isCredentialLikeEnvironmentVariableName(upper) && - ROUTE_ENV_PREFIXES.ClaudeAgent!.some((prefix) => upper.startsWith(prefix))) - ); - }); -} - -function effectiveRouteEnvironment( - agent: WorkflowRouteAgent, - env: Record -): Array<[string, string]> { - const names = new Set( - (env.ULTRAFUZZ_AGENT_ENV_ALLOWLIST ?? "").split(",").map((entry) => entry.trim().toUpperCase()) - ); - for (const name of ROUTE_PROXY_ENV) names.add(name); - for (const name of Object.keys(env)) - if ( - !isCredentialLikeEnvironmentVariableName(name) && - ROUTE_ENV_PREFIXES[agent]?.some((prefix) => name.startsWith(prefix)) - ) - names.add(name); - if (agent === "CodexAgent") names.add("OPENAI_BASE_URL"); - if (agent === "KimiAgent") names.add("KIMI_BASE_URL"); - for (const name of NON_ROUTING_PROVIDER_ENVIRONMENT_NAMES) names.delete(name); - names.delete("KIMI_CODE_HOME"); - names.delete("KIMI_SHARE_DIR"); - return [...names].sort().flatMap((name): Array<[string, string]> => { - const value = env[name]; - return value !== undefined && - value.trim() !== "" && - !isCredentialLikeEnvironmentVariableName(name) && - (ROUTE_PROXY_ENV.includes(name as never) || ROUTE_ENV_PREFIXES[agent]?.some((prefix) => name.startsWith(prefix))) - ? [[name, value]] - : []; - }); -} - -const sha256 = (value: string | Buffer): string => createHash("sha256").update(value).digest("hex"); - export function workflowControlCredentialValue( value: string, name: string, diff --git a/packages/runtime/test/data-governance.test.ts b/packages/runtime/test/data-governance.test.ts index 63213c6f0..54e97bc80 100644 --- a/packages/runtime/test/data-governance.test.ts +++ b/packages/runtime/test/data-governance.test.ts @@ -355,7 +355,7 @@ test("policy pins explicit and home routes, Modal, and OpenRouter models", () => { apiKeyHelper: "/operator/helper" }, { processWrapper: "/operator/wrapper" }, { proxyAuth: "operator" }, - { env: { HTTPS_PROXY: "https://proxy.example" } } + { env: { ANTHROPIC_BASE_URL: "https://home-route" } } ]) { fs.writeFileSync(claudeSettings, JSON.stringify(value)); assert.throws( @@ -363,6 +363,13 @@ test("policy pins explicit and home routes, Modal, and OpenRouter models", () => /cloud execution cannot use host provider-home routing/u ); } + // A proxy does not select a destination, in settings or in the environment. + fs.writeFileSync(claudeSettings, JSON.stringify({ env: { HTTPS_PROXY: "https://proxy.example" } })); + assert.equal(modelDestination("ClaudeAgent", cloud, { HOME: homes }), "model:anthropic"); + assert.equal( + modelDestination("ClaudeAgent", routed, { HOME: homes, HTTPS_PROXY: "https://proxy.example" }), + "model:anthropic" + ); const mixedGraph = { nodes: [{ model_fanout: [{ agent_ref: "OpenRouterAgent", model_name: "vendor/review-model" }] }] } as unknown as PlannedGraph, @@ -523,6 +530,19 @@ test("target identity rejects Git content filters before execution", () => { assert.equal(fs.existsSync(marker), false); }); +test("a Claude platform flag in settings pins that platform's process variables", () => { + const homes = temporaryRoot("ufz-claude-settings-platform-"), + settingsPath = path.join(homes, ".claude", "settings.json"); + fs.mkdirSync(path.dirname(settingsPath)); + fs.writeFileSync(settingsPath, JSON.stringify({ env: { CLAUDE_CODE_USE_BEDROCK: "1" } })); + const east = modelDestination("ClaudeAgent", config, { HOME: homes, AWS_REGION: "us-east-1" }); + assert.match(east, /^model:claude-route-/u); + assert.notEqual(modelDestination("ClaudeAgent", config, { HOME: homes, AWS_REGION: "eu-west-1" }), east); + // Without a platform flag anywhere, the same variable routes nothing. + fs.writeFileSync(settingsPath, "{}"); + assert.equal(modelDestination("ClaudeAgent", config, { HOME: homes, AWS_REGION: "us-east-1" }), "model:anthropic"); +}); + test("Codex CLI bookkeeping in config.toml does not change the acknowledged route", () => { const homes = temporaryRoot("ufz-codex-route-"); fs.mkdirSync(path.join(homes, ".codex")); @@ -551,14 +571,39 @@ test("Codex CLI bookkeeping in config.toml does not change the acknowledged rout ); assert.equal(modelDestination("CodexAgent", config, { HOME: homes }), "model:openai"); - // Routing declarations still pin (and, after acknowledgement, still fail - // closed): a provider selection or endpoint makes the digest reappear. - for (const routing of [ - 'model_provider = "private"\n', - '[model_providers.private]\nbase_url = "https://gateway.example/v1"\n', - 'base_url = "https://gateway.example/v1"\n' - ]) { - fs.writeFileSync(configPath, routing); - assert.match(modelDestination("CodexAgent", config, { HOME: homes }), /^model:codex-route-/u); - } + // A provider table routes nothing until model_provider selects it. + fs.writeFileSync(configPath, '[model_providers.private]\nbase_url = "https://gateway.example/v1"\n'); + assert.equal(modelDestination("CodexAgent", config, { HOME: homes }), "model:openai"); + + // A selected provider pins its id, endpoint, wire API, and credential name, + // and nothing else in the file. + const selected = (lines: string[]) => + ['model_provider = "private"', "", "[model_providers.private]", ...lines, ""].join("\n"); + const route = (text: string) => { + fs.writeFileSync(configPath, text); + return modelDestination("CodexAgent", config, { HOME: homes }); + }; + const pinned = route(selected(['base_url = "https://gateway.example/v1"', 'env_key = "PRIVATE_KEY_ENV"'])); + assert.match(pinned, /^model:codex-route-/u); + assert.equal( + route( + `model = "gpt-5.5"\n${selected([ + 'name = "Private gateway"', + 'base_url = "https://gateway.example/v1"', + 'env_key = "PRIVATE_KEY_ENV"', + "request_max_retries = 4" + ])}\n[marketplaces.openai-bundled]\nlast_updated = "2026-08-25T13:35:06Z"\n\n[model_providers.unused]\nbase_url = "https://unused.example/v1"\n` + ), + pinned + ); + for (const drift of [ + selected(['base_url = "https://other.example/v1"', 'env_key = "PRIVATE_KEY_ENV"']), + selected(['base_url = "https://gateway.example/v1"', 'env_key = "OTHER_KEY_ENV"']), + selected(['base_url = "https://gateway.example/v1"', 'env_key = "PRIVATE_KEY_ENV"', 'wire_api = "chat"']), + `profile = "work"\n\n[profiles.work]\nmodel_provider = "other"\n\n${selected([ + 'base_url = "https://gateway.example/v1"', + 'env_key = "PRIVATE_KEY_ENV"' + ])}` + ]) + assert.notEqual(route(drift), pinned, drift); }); diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index bc85851d1..e66ff05ac 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -134,6 +134,9 @@ import { const runningUnderBun = typeof process.versions.bun === "string"; const BUN_ADAPTER_TEST_PREFIX = "Bun adapter contract: "; const bunAdapterTest = prefixTestNames(testWhen(runningUnderBun, { timeout: 30_000 }), BUN_ADAPTER_TEST_PREFIX); +// Generated adapters load their route helper from the runtime module the +// rendered workflow names; point them at this build. +if (runningUnderBun) process.env.ULTRAFUZZ_RUNTIME_MODULE ??= new URL("../src/index.js", import.meta.url).href; const SMITHERS_TEST_ENVIRONMENT_ALLOWLIST = [ "SMITHERS_FAKE_ADMISSION_TIMEOUT_LOG", "SMITHERS_FAKE_CLOUD_ENV_LOG", @@ -3899,7 +3902,7 @@ test("init does not modify a regular file swapped after the anchored open", { co }); bunAdapterTest( - "generated Codex commands accept a sealed ambient proxy and reject drift", + "generated Codex commands ignore ambient proxies and reject endpoint drift", { timeout: 30_000 }, async () => { const project = tempProject(); @@ -3917,6 +3920,7 @@ bunAdapterTest( "http_proxy", "https_proxy", "no_proxy", + "OPENAI_BASE_URL", "ULTRAFUZZ_DATA_GOVERNANCE_PATH", "ULTRAFUZZ_WORKFLOW_PERSISTED_PATH" ], @@ -3925,23 +3929,33 @@ bunAdapterTest( process.env.ULTRAFUZZ_DATA_GOVERNANCE_PATH = authority; process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH = path.join(snapshot, ".smithers/workflows/test.tsx"); try { + const config = { execution: { mode: "local" }, agents: {} } as never; process.env.HTTPS_PROXY = "https://proxy.a.invalid"; - const hash = crypto - .createHash("sha256") - .update( - JSON.stringify({ agent: "CodexAgent", config: null, route: [["HTTPS_PROXY", process.env.HTTPS_PROXY]] }) - ) - .digest("hex"); - fs.writeFileSync(authority, `{"required_source_destinations":["model:codex-route-${hash}"]}`, "utf8"); - const build = () => - new CompatibleCodexAgent().buildCommand({ prompt: "Contract only", cwd: project, options: {} }); - const accepted = await build(); - await accepted.cleanup?.(); + process.env.OPENAI_BASE_URL = "https://gateway.a.invalid/v1"; + fs.writeFileSync( + authority, + JSON.stringify({ required_source_destinations: [modelDestination("CodexAgent", config, process.env)] }) + ); + const build = async () => { + const command = await new CompatibleCodexAgent().buildCommand({ + prompt: "Contract only", + cwd: project, + options: {} + }); + await command.cleanup?.(); + }; + await build(); + // A proxy does not change which provider receives the traffic, so a run + // resumed from a shell with another proxy, or none, keeps its route. + process.env.HTTPS_PROXY = "https://proxy.b.invalid"; + await build(); + delete process.env.HTTPS_PROXY; + await build(); assert.throws( () => workflowControlChildEnvironment( { - HTTPS_PROXY: "https://proxy.b.invalid", + OPENAI_BASE_URL: "https://gateway.b.invalid/v1", ULTRAFUZZ_DATA_GOVERNANCE_PATH: path.join(project, "forged.json") }, process.env, @@ -3949,7 +3963,7 @@ bunAdapterTest( ), /provider route changed/u ); - process.env.HTTPS_PROXY = "https://proxy.b.invalid"; + process.env.OPENAI_BASE_URL = "https://gateway.b.invalid/v1"; await assert.rejects(build(), /provider route changed/u); } finally { for (const name of names) { @@ -4031,9 +4045,21 @@ bunAdapterTest("planned routes equal final generated-adapter validation", { time workflowControlChildEnvironment({}, { ...claudeEnv, CLAUDE_CODE_USE_BEDROCK: "1" }, { agent: "ClaudeAgent" }), /provider route changed/u ); + // Once Bedrock is selected, its AWS settings are part of the route. + const bedrockEnv = { ...claudeEnv, CLAUDE_CODE_USE_BEDROCK: "1" }, + bedrockDestination = modelDestination("ClaudeAgent", config, bedrockEnv); + fs.writeFileSync(authority, JSON.stringify({ required_source_destinations: [bedrockDestination] })); + assert.doesNotThrow(() => workflowControlChildEnvironment({}, bedrockEnv, { agent: "ClaudeAgent" })); + assert.throws( + () => workflowControlChildEnvironment({}, { ...bedrockEnv, AWS_REGION: "eu-west-1" }, { agent: "ClaudeAgent" }), + /provider route changed/u + ); assert.equal(codexDestination.startsWith("model:codex-route-"), true); assert.equal(openRouterDestination, "model:openrouter"); - assert.equal(claudeDestination.startsWith("model:claude-route-"), true); + // AWS_REGION alone routes nothing: Claude Code reads it only for an + // AWS-hosted platform such as Bedrock. + assert.equal(claudeDestination, "model:anthropic"); + assert.equal(bedrockDestination.startsWith("model:claude-route-"), true); } finally { for (const name of names) { const value = saved[name]; @@ -4075,6 +4101,135 @@ bunAdapterTest( } ); +bunAdapterTest( + "generated Codex adapter keeps its acknowledged route when the CLI rewrites unrelated config", + { timeout: 30_000 }, + async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const homes = temporaryRoot("ufz-codex-rewrite-"), + codexHome = path.join(homes, "codex", "configured"), + configPath = path.join(codexHome, "config.toml"), + snapshot = path.join(project, "route-snapshot"), + authority = path.join(snapshot, "controls/data-governance.json"), + names = [ + "ALL_PROXY", + "HTTP_PROXY", + "HTTPS_PROXY", + "NO_PROXY", + "all_proxy", + "http_proxy", + "https_proxy", + "no_proxy", + "OPENAI_BASE_URL", + "ULTRAFUZZ_AGENT_ENV_ALLOWLIST", + "ULTRAFUZZ_DATA_GOVERNANCE_PATH", + "ULTRAFUZZ_WORKFLOW_PERSISTED_PATH" + ], + saved = Object.fromEntries(names.map((name) => [name, process.env[name]])); + fs.mkdirSync(codexHome, { recursive: true }); + fs.mkdirSync(path.dirname(authority), { recursive: true }); + const providerConfig = [ + 'model = "gpt-5.5"', + 'model_provider = "gateway"', + "", + "[model_providers.gateway]", + 'name = "Gateway"', + 'base_url = "https://gateway.invalid/v1"', + 'env_key = "GATEWAY_API_KEY"', + 'wire_api = "responses"', + "" + ].join("\n"); + fs.writeFileSync( + configPath, + `${providerConfig}\n[marketplaces.openai-bundled]\nlast_updated = "2026-08-25T13:35:06Z"\n` + ); + for (const name of names) Reflect.deleteProperty(process.env, name); + process.env.ULTRAFUZZ_DATA_GOVERNANCE_PATH = authority; + process.env.ULTRAFUZZ_WORKFLOW_PERSISTED_PATH = path.join(snapshot, ".smithers/workflows/test.tsx"); + try { + const config = { execution: { mode: "local" }, agents: { CodexAgent: { configDir: "configured" } } } as never, + destination = modelDestination("CodexAgent", config, { ULTRAFUZZ_PROVIDER_HOME_ROOT: homes }); + assert.match(destination, /^model:codex-route-/u); + fs.writeFileSync(authority, JSON.stringify({ required_source_destinations: [destination] })); + const { CompatibleCodexAgent } = await loadGeneratedCodexAgent(project); + const build = async () => { + const command = await new CompatibleCodexAgent({ configDir: codexHome }).buildCommand({ + prompt: "route", + cwd: project, + options: {} + }); + await command.cleanup?.(); + }; + await build(); + // What the Codex CLI itself writes while running: marketplace refresh + // timestamps and project trust levels (#908). + fs.writeFileSync( + configPath, + `${providerConfig}\n[marketplaces.openai-bundled]\nlast_updated = "2026-09-28T00:00:00Z"\n\n` + + `[projects."/workspace/target"]\ntrust_level = "trusted"\n` + ); + await build(); + fs.writeFileSync(configPath, providerConfig.replace("gateway.invalid", "other-gateway.invalid")); + await assert.rejects(build(), /provider route changed/u); + } finally { + for (const name of names) { + const value = saved[name]; + if (value === undefined) Reflect.deleteProperty(process.env, name); + else process.env[name] = value; + } + } + } +); + +bunAdapterTest( + "generated Claude route ignores ambient cloud settings until a cloud platform is selected", + { timeout: 30_000 }, + async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const { workflowControlChildEnvironment } = await loadGeneratedCodexAgent(project); + const snapshot = path.join(project, "route-snapshot"), + authority = path.join(snapshot, "controls", "data-governance.json"); + fs.mkdirSync(path.dirname(authority), { recursive: true }); + const config = { execution: { mode: "local" }, agents: {} } as never, + env = { + ULTRAFUZZ_PROVIDER_HOME_ROOT: temporaryRoot("ufz-claude-ambient-"), + AWS_PROFILE: "operator-a", + GOOGLE_CLOUD_PROJECT: "project-a", + // Foundry's forge profile, not Microsoft Foundry. + FOUNDRY_PROFILE: "ci", + ULTRAFUZZ_DATA_GOVERNANCE_PATH: authority, + ULTRAFUZZ_WORKFLOW_PERSISTED_PATH: path.join(snapshot, ".smithers", "workflows", "test.tsx") + }; + fs.writeFileSync( + authority, + JSON.stringify({ required_source_destinations: [modelDestination("ClaudeAgent", config, env)] }) + ); + // A run resumed from another shell sees different, unused cloud settings. + assert.doesNotThrow(() => + workflowControlChildEnvironment( + {}, + { ...env, AWS_PROFILE: "operator-b", GOOGLE_CLOUD_PROJECT: "project-b", FOUNDRY_PROFILE: "default" }, + { agent: "ClaudeAgent" } + ) + ); + const vertex = { ...env, CLAUDE_CODE_USE_VERTEX: "1" }; + fs.writeFileSync( + authority, + JSON.stringify({ required_source_destinations: [modelDestination("ClaudeAgent", config, vertex)] }) + ); + assert.doesNotThrow(() => + workflowControlChildEnvironment({}, { ...vertex, AWS_PROFILE: "operator-b" }, { agent: "ClaudeAgent" }) + ); + assert.throws( + () => + workflowControlChildEnvironment({}, { ...vertex, GOOGLE_CLOUD_PROJECT: "project-b" }, { agent: "ClaudeAgent" }), + /provider route changed/u + ); + } +); + bunAdapterTest("quoted TOML provider routes are bound and drift fails closed", { timeout: 30_000 }, async () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 35fd8765f..bb97ddff9 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -329,6 +329,9 @@ importers: proper-lockfile: specifier: 4.1.2 version: 4.1.2 + smol-toml: + specifier: 1.7.1 + version: 1.7.1 typescript: specifier: 6.0.3 version: 6.0.3 From 5093390533832a8ce62243f3a9470fd5ef929bde Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:28:54 +0000 Subject: [PATCH 068/206] fix(runtime): plain init refreshes the stock adapter closure Planning accepts only the byte-exact packaged .smithers/agents closure (controller-source.ts), but init treated those files as project-owned: without --force it preserved stale or customized copies, emitted review warnings, and upgraded adapters only during a recognized 0.32 manifest migration. Every adapter fix therefore left existing projects unplannable until the operator ran `init --force`, which also resets ultrafuzz.toml, the topology and the prompts. Every init now rewrites the files in STOCK_CONTROLLER_SOURCE_TEMPLATES, the same list planning checks, and the preserve/review/0.32-adapter upgrade code and its INIT_AGENT_* diagnostics are deleted. Project-owned files are preserved as before. A symlink, hard link or special file at an adapter path now fails init with INIT_PATH_UNSAFE, as `init --force` already did, instead of being preserved with a warning; the replacement open adds O_NONBLOCK so a FIFO fails rather than blocking init. controller-source now tells the operator to rerun `ultrafuzz init`. The tests that asserted preservation and review warnings are rewritten to assert the refresh and that linked paths are never followed or written through; against main they fail (6 of 6). Co-Authored-By: Claude Opus 5.5 --- docs/config.md | 2 +- docs/reference/cli.md | 17 +- packages/cli/test/cli.test.ts | 15 +- packages/runtime/src/controller-source.ts | 4 +- packages/runtime/src/init.ts | 289 ++-------------------- packages/runtime/test/runtime.test.ts | 275 ++++---------------- 6 files changed, 95 insertions(+), 507 deletions(-) diff --git a/docs/config.md b/docs/config.md index 857b2c0bf..9d05a7ebc 100644 --- a/docs/config.md +++ b/docs/config.md @@ -65,7 +65,7 @@ Agent configuration and model profiles may name only `ClaudeAgent`, `CodexAgent` `DeepSeekAgent`, `KimiAgent`, `OpenCodeAgent`, `OpenRouterAgent`, or `PiAgent`. The complete `.smithers/agents` tree must byte-match the packaged stock closure; custom adapters and registries -are unsupported, and `ultrafuzz init --force` restores the authenticated copy. The stock closure always uses YOLO/bypass-permissions; stricter per-project adapters are unsupported. +are unsupported, and `ultrafuzz init` restores the authenticated copy. The stock closure always uses YOLO/bypass-permissions; stricter per-project adapters are unsupported. Each stock agent's `api_key_env` must use its canonical provider credential name. Custom environment variable names fail config validation. diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..1074dfab8 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -104,16 +104,17 @@ ultrafuzz.toml .ultrafuzz/cache/ ``` -Without `--force`, existing config, topology, prompts, reference catalogs, and -agent adapters are preserved except for recognized stock migrations. `init` +Without `--force`, existing config, topology, prompts, and reference catalogs +are preserved. The `.smithers/agents` adapter closure is not project-owned: +planning accepts only the byte-exact packaged copy, so every `init` rewrites it +from the current templates, and an upgrade needs no `--force`. A symbolic link, +hard link, or special file at an adapter path makes `init` fail with +`INIT_PATH_UNSAFE` instead of being followed or written through. `init` also migrates a recognized generated `smithers-orchestrator@0.32.0` manifest or a `smthrs@0.34.0` manifest to the current `smthrs@0.35.0` pin. The 0.34.0 migration -requires the rest of the manifest to match the current generated document. -During a recognized manifest migration, known byte-identical prior stock -adapters are also upgraded; customized or unrecognized adapters remain -preserved. Use `--force` to replace existing generated files with the current -templates. `init` emits an actionable diagnostic when a preserved adapter -requires manual review. +requires the rest of the manifest to match the current generated document. Use +`--force` to replace the other existing generated files with the current +templates. ## Validate diff --git a/packages/cli/test/cli.test.ts b/packages/cli/test/cli.test.ts index 18676ca3a..76f8fced3 100644 --- a/packages/cli/test/cli.test.ts +++ b/packages/cli/test/cli.test.ts @@ -1280,21 +1280,24 @@ test("init and validate emit schema-versioned launch JSON", async () => { assert.match(tampered.stdout + tampered.stderr, /absent\.database/u); }); -test("plain init surfaces a customized stale agent adapter diagnostic", async () => { +test("plain init restores a customized agent adapter without touching project config", async () => { const project = tempProject(); const initial = await cli(project, ["init", "--force"]); assert.equal(initial.code, 0, initial.stderr); const adapterPath = path.join(project, ".smithers", "agents", "codex.ts"); - const customAdapter = 'export const customConfigPath = "ultrafuzz.toml";\n'; - fs.writeFileSync(adapterPath, customAdapter, "utf8"); + const stockAdapter = fs.readFileSync(adapterPath, "utf8"); + fs.writeFileSync(adapterPath, 'export const customConfigPath = "ultrafuzz.toml";\n', "utf8"); + const configPath = path.join(project, "ultrafuzz.toml"); + const customConfig = `${fs.readFileSync(configPath, "utf8")}\n# operator customization\n`; + fs.writeFileSync(configPath, customConfig, "utf8"); const result = await cli(project, ["init"]); assert.equal(result.code, 0, result.stderr); assert.equal(result.stderr, ""); - assert.match(result.stdout, /warning: INIT_AGENT_ADAPTER_UPDATE_REQUIRED:/u); - assert.match(result.stdout, /ULTRAFUZZ_CONFIG_PATH/u); - assert.equal(fs.readFileSync(adapterPath, "utf8"), customAdapter); + assert.doesNotMatch(result.stdout, /warning:/u); + assert.equal(fs.readFileSync(adapterPath, "utf8"), stockAdapter); + assert.equal(fs.readFileSync(configPath, "utf8"), customConfig); }); test("run exposes the trusted reference expectation catalog option", async () => { diff --git a/packages/runtime/src/controller-source.ts b/packages/runtime/src/controller-source.ts index b696b1c20..d38c47d07 100644 --- a/packages/runtime/src/controller-source.ts +++ b/packages/runtime/src/controller-source.ts @@ -12,7 +12,7 @@ const PROVIDER_SCOPED_SENSITIVE_ENVIRONMENT_DECLARATION = Buffer.from( "utf8" ); const UNTRUSTED_SOURCE = - "controller adapter source must exactly match the packaged stock closure; rerun ultrafuzz init --force"; + "controller adapter source must exactly match the packaged stock closure; rerun ultrafuzz init"; const CONTROLLER_NAMES = "claude codex deepseek environment index kimi opencode openrouter pi provider-home strict-json toml".split(" "); export const STOCK_CONTROLLER_SOURCE_TEMPLATES: Readonly> = Object.freeze( @@ -80,7 +80,7 @@ export function assertProviderScopedSensitiveEnvironmentCapability( if (environment === undefined || !environment.contents.includes(PROVIDER_SCOPED_SENSITIVE_ENVIRONMENT_DECLARATION)) { throw new Error( "sealed controller predates provider-scoped sensitive allowlisted environment handling; " + - "rerun ultrafuzz init --force and start a new run" + "rerun ultrafuzz init and start a new run" ); } } diff --git a/packages/runtime/src/init.ts b/packages/runtime/src/init.ts index 425960466..715f6018b 100644 --- a/packages/runtime/src/init.ts +++ b/packages/runtime/src/init.ts @@ -1,4 +1,3 @@ -import crypto from "node:crypto"; import fs from "node:fs"; import path from "node:path"; @@ -16,74 +15,14 @@ import { } from "@ultrafuzz/config"; import { builtInPromptRelativePaths, scaffoldPrompts } from "@ultrafuzz/prompts"; import { defaultReferenceCatalogYaml } from "@ultrafuzz/references"; -import { AGENT_REGISTRY_RELATIVE_PATH, agentRegistryRegisters, inspectAgentRegistry } from "./agent-registry.js"; +import { STOCK_CONTROLLER_SOURCE_TEMPLATES } from "./controller-source.js"; import { loadRuntimeTemplate } from "./runtime-template.js"; import { migrateStockSmithers032PackageManifest, renderSmithersPackageJson } from "./smithers-package.js"; -import type { InitProjectInput, InitProjectResult, RuntimeDiagnostic } from "./types.js"; +import type { InitProjectInput, InitProjectResult } from "./types.js"; import { configDiagnostics, runtimeFailure, runtimeResult, toProjectRelative } from "./utils.js"; const DEFAULT_TOPOLOGY = fs.readFileSync(packagedTopology("default").path, "utf8"); -const MAX_AGENT_ADAPTER_REVIEW_BYTES = 256 * 1024; -const AGENT_TEMPLATES = [ - { - file: "claude.ts", - template: "smithers/agents/claude.tsx", - ref: "ClaudeAgent", - stock032Sha256: ["f2b97c9b57aa45bdc3b42be20d7a3baddc0086a2bae95c161b841599f3262232"] - }, - { - file: "codex.ts", - template: "smithers/agents/codex.tsx", - ref: "CodexAgent", - stock032Sha256: [ - "b932fb7da3c05fdc662f60359e8a751aaabd236ca4072dfeaade1a7bb25a01b5", - "e7e845b2bccf5b7d41a3f0cedfaec7a1457580126513ff96f282a36307c7da98", - "7865f1be1715d36d016c7b2814081b70e70a9aca7e30d5b41f5d91bf2337f681" - ] - }, - { - file: "deepseek.ts", - template: "smithers/agents/deepseek.tsx", - ref: "DeepSeekAgent", - stockSha256: new Set([ - "c23a03c84e2f62d2e6b23ee7b27b1464a633fe20bb5b91c34d2c93d37dcf7e35", - "65bf43f333cbced8ff0157e942d3c78267d8245d6c0041e053c8577a463e7407" - ]) - }, - { - file: "kimi.ts", - template: "smithers/agents/kimi.tsx", - ref: "KimiAgent", - stockSha256: new Set([ - "fdbeaad6ea55122da50e9c9d86ac58a8b6377f419f8e78df23fb1fb401924e0b", - "6de5f4b00b54b533f8fc1aa0628f5a402dbb6521a4584852d83786ee59dcb2fd", - "25c499f8631db6e2529b046b5d2243119b456696c7a4a729345baa6f21f4c4f5", - "f790a3f121da84049032cfc5bf5d300f7bd56e9e3b15d0a51df5ea1b275de75f" - ]) - }, - { - file: "opencode.ts", - template: "smithers/agents/opencode.tsx", - ref: "OpenCodeAgent", - stockSha256: new Set() - }, - { - file: "openrouter.ts", - template: "smithers/agents/openrouter.tsx", - ref: "OpenRouterAgent", - stock032Sha256: [] - }, - { - file: "pi.ts", - template: "smithers/agents/pi.tsx", - ref: "PiAgent", - // No stock digest yet: this adapter has never shipped in a released - // scaffold, so any pi.ts already on disk is the operator's and is preserved. - stockSha256: new Set() - } -] as const; - /** * Canonical artifact JSON Schemas scaffolded into the project. * @@ -154,10 +93,6 @@ export function initProject(input: InitProjectInput) { try { const stockSmithersPackageMigration = input.force === true ? undefined : prepareStockSmithers032PackageMigration(projectRoot); - const upgradedStockAdapters = - stockSmithersPackageMigration === undefined - ? new Set() - : upgradeStockSmithers032Adapters(projectRoot, created, preserved, overwritten); writeProjectFile( projectRoot, "ultrafuzz.toml", @@ -205,59 +140,16 @@ export function initProject(input: InitProjectInput) { preserved, overwritten ); - writeProjectFile( - projectRoot, - ".smithers/agents/index.ts", - loadRuntimeTemplate("smithers/agents/index.tsx"), - input.force === true, - created, - preserved, - overwritten - ); - writeProjectFile( - projectRoot, - ".smithers/agents/toml.ts", - loadRuntimeTemplate("smithers/agents/toml.tsx"), - input.force === true, - created, - preserved, - overwritten - ); - writeProjectFile( - projectRoot, - ".smithers/agents/environment.ts", - loadRuntimeTemplate("smithers/agents/environment.tsx"), - input.force === true, - created, - preserved, - overwritten - ); - writeProjectFile( - projectRoot, - ".smithers/agents/provider-home.ts", - loadRuntimeTemplate("smithers/agents/provider-home.tsx"), - input.force === true, - created, - preserved, - overwritten - ); - writeProjectFile( - projectRoot, - ".smithers/agents/strict-json.ts", - loadRuntimeTemplate("smithers/agents/strict-json.tsx"), - input.force === true, - created, - preserved, - overwritten - ); - for (const agent of AGENT_TEMPLATES) { - const relativePath = `.smithers/agents/${agent.file}`; - if (upgradedStockAdapters.has(relativePath)) continue; + // Planning admits only the byte-exact packaged adapter closure, so these + // files are never project-owned. Refresh them on every init: otherwise an + // upgrade needs `init --force`, which also resets ultrafuzz.toml, the + // topology, and the prompts. + for (const [file, template] of Object.entries(STOCK_CONTROLLER_SOURCE_TEMPLATES)) { writeProjectFile( projectRoot, - relativePath, - loadRuntimeTemplate(agent.template), - input.force === true, + `.smithers/agents/${file}`, + loadRuntimeTemplate(template), + true, created, preserved, overwritten @@ -304,16 +196,12 @@ export function initProject(input: InitProjectInput) { preserved.push(toProjectRelative(projectRoot, absolutePath)); } - return runtimeResult( - true, - { - project_root: projectRoot, - created: publicInitPaths(created), - preserved: publicInitPaths(preserved), - overwritten: publicInitPaths(overwritten) - }, - [...staleAgentRegistryDiagnostics(projectRoot), ...staleAgentAdapterDiagnostics(projectRoot)] - ); + return runtimeResult(true, { + project_root: projectRoot, + created: publicInitPaths(created), + preserved: publicInitPaths(preserved), + overwritten: publicInitPaths(overwritten) + }); } function prepareStockSmithers032PackageMigration(projectRoot: string): string | undefined { @@ -340,143 +228,6 @@ function prepareStockSmithers032PackageMigration(projectRoot: string): string | } } -function upgradeStockSmithers032Adapters( - projectRoot: string, - created: string[], - preserved: string[], - overwritten: string[] -): ReadonlySet { - const upgraded = new Set(); - for (const agent of AGENT_TEMPLATES) { - const relativePath = `.smithers/agents/${agent.file}`; - const filePath = path.join(projectRoot, relativePath); - let bytes: Buffer; - try { - bytes = readStableInitReviewFile( - projectRoot, - filePath, - MAX_AGENT_ADAPTER_REVIEW_BYTES, - "generated 0.32 agent adapter" - ); - } catch { - continue; - } - const digest = crypto.createHash("sha256").update(bytes).digest("hex"); - if (!("stock032Sha256" in agent)) continue; - const stock032Sha256 = agent.stock032Sha256 as readonly string[]; - if (!stock032Sha256.includes(digest)) continue; - writeProjectFile( - projectRoot, - relativePath, - loadRuntimeTemplate(agent.template), - true, - created, - preserved, - overwritten - ); - upgraded.add(relativePath); - } - return upgraded; -} - -function staleAgentAdapterDiagnostics(projectRoot: string): RuntimeDiagnostic[] { - const diagnostics: RuntimeDiagnostic[] = []; - for (const agent of AGENT_TEMPLATES) { - const relativePath = `.smithers/agents/${agent.file}`; - const filePath = path.join(projectRoot, relativePath); - try { - const lexical = fs.lstatSync(filePath, { bigint: true }); - if (lexical.isSymbolicLink() || !lexical.isFile() || lexical.nlink !== 1n) { - diagnostics.push( - manualAgentAdapterReviewDiagnostic( - relativePath, - "is not a physical single-link file, so init preserved it without inspection; replace it with an ordinary file" - ) - ); - continue; - } - if (lexical.size > BigInt(MAX_AGENT_ADAPTER_REVIEW_BYTES)) { - diagnostics.push( - manualAgentAdapterReviewDiagnostic( - relativePath, - "is too large to inspect as a generated adapter and was preserved" - ) - ); - continue; - } - const source = readStableInitReviewFile( - projectRoot, - filePath, - MAX_AGENT_ADAPTER_REVIEW_BYTES, - "generated agent adapter" - ).toString("utf8"); - if (source.includes("ultrafuzz.toml") && !source.includes("ULTRAFUZZ_CONFIG_PATH")) { - diagnostics.push({ - code: "INIT_AGENT_ADAPTER_UPDATE_REQUIRED", - message: `${relativePath} was preserved and still reads mutable project ultrafuzz.toml; update it to read process.env.ULTRAFUZZ_CONFIG_PATH and use workflowControlChildEnvironment before spawning a model process`, - severity: "warning", - source: "runtime", - path: relativePath - }); - } - } catch (error) { - if (isNodeError(error) && error.code === "ENOENT") continue; - diagnostics.push( - manualAgentAdapterReviewDiagnostic(relativePath, "could not be safely inspected during post-init review") - ); - } - } - return diagnostics; -} - -function manualAgentAdapterReviewDiagnostic(relativePath: string, reason: string): RuntimeDiagnostic { - return { - code: "INIT_AGENT_ADAPTER_UPDATE_REQUIRED", - message: `${relativePath} ${reason}; verify manually that it reads process.env.ULTRAFUZZ_CONFIG_PATH and removes controller-only variables before spawning a model process`, - severity: "warning", - source: "runtime", - path: relativePath - }; -} - -// init preserves project-owned files, so a project scaffolded before an agent -// was added keeps its old registry: the new adapter lands on disk but nothing -// exports it, and the agent is only rejected later, at launch. Report it here -// instead of leaving the mismatch silent. -function staleAgentRegistryDiagnostics(projectRoot: string): RuntimeDiagnostic[] { - const registry = inspectAgentRegistry(projectRoot); - if (!registry.exists) return []; - if (registry.error !== undefined) { - // This runs after init may already have written other project files. Keep - // the warning actionable without reflecting an OS/parser error that can - // contain sensitive path or injected error details. - return [ - manualAgentRegistryReviewDiagnostic("was preserved without inspection because it could not be safely inspected") - ]; - } - return AGENT_TEMPLATES.filter( - (agent) => - lstatIfPresent(path.join(projectRoot, ".smithers", "agents", agent.file)) !== undefined && - !agentRegistryRegisters(registry, agent.ref) - ).map((agent) => ({ - code: "INIT_AGENT_REGISTRY_STALE", - message: `${AGENT_REGISTRY_RELATIVE_PATH} does not register ${agent.ref} in agentFactories, so runs cannot select it; rerun ultrafuzz init --force to regenerate the registry, or add the entry by hand`, - severity: "warning" as const, - source: "runtime", - path: AGENT_REGISTRY_RELATIVE_PATH - })); -} - -function manualAgentRegistryReviewDiagnostic(reason: string): RuntimeDiagnostic { - return { - code: "INIT_AGENT_REGISTRY_REVIEW_REQUIRED", - message: `${AGENT_REGISTRY_RELATIVE_PATH} ${reason}; verify manually that agentFactories registers every generated agent before starting a run`, - severity: "warning", - source: "runtime", - path: AGENT_REGISTRY_RELATIVE_PATH - }; -} - function writeProjectFile( projectRoot: string, relativePath: string, @@ -547,7 +298,9 @@ function writeProjectFileNoFollow( expected === undefined ? fs.constants.O_WRONLY | fs.constants.O_CREAT | fs.constants.O_EXCL | (fs.constants.O_NOFOLLOW ?? 0) : fs.constants.O_WRONLY | (fs.constants.O_NOFOLLOW ?? 0); - fileDescriptor = fs.openSync(accessPath, flags, 0o666); + // O_NONBLOCK makes a FIFO planted at a generated path fail the open (or the + // regular-file check below) instead of blocking init on a missing reader. + fileDescriptor = fs.openSync(accessPath, flags | fs.constants.O_NONBLOCK, 0o666); const opened = fs.fstatSync(fileDescriptor, { bigint: true }); // Modal's virtual filesystem can report one device for an opened // directory and another for stable children created through that dirfd. @@ -714,10 +467,6 @@ function writeDescriptorContents(descriptor: number, contents: Buffer): void { } } -function isNodeError(error: unknown): error is NodeJS.ErrnoException { - return error instanceof Error && "code" in error; -} - function uniqueSorted(values: string[]): string[] { return Array.from(new Set(values)).sort(); } diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index e66ff05ac..474c7afc0 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -122,6 +122,8 @@ import { verifyRequiredArtifactsForAttempt } from "../src/artifact-gates.js"; import { projectCanonicalFinalReport } from "../src/final-report-markdown.js"; import { loadReportSnapshot } from "../src/unverified-report.js"; import { projectArtifactSchemaDir, projectArtifactSchemaJson } from "../src/init.js"; +import { inspectControllerSource } from "../src/controller-source.js"; +import { loadRuntimeTemplate } from "../src/runtime-template.js"; import { assertRenderedPromptValidatorCommands } from "../src/prompt-validator-command.js"; import { writeFakeNpmInstaller } from "./fake-npm-installer.js"; import { addOpenRouterProfile } from "./openrouter-profile-fixture.js"; @@ -3403,9 +3405,9 @@ testWhen(process.platform !== "win32" && fs.existsSync("/proc/self/fd"))( const codexPath = path.join(project, ".smithers", "agents", "codex.ts"); fs.writeFileSync(codexPath, V0_0_2_STOCK_CODEX_ADAPTER, "utf8"); - const preserved = initProject({ projectRoot: project }); - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); - assert.equal(fs.readFileSync(codexPath, "utf8"), V0_0_2_STOCK_CODEX_ADAPTER); + const refreshed = initProject({ projectRoot: project }); + assert.equal(refreshed.ok, true, JSON.stringify(refreshed.diagnostics)); + assert.equal(fs.readFileSync(codexPath, "utf8"), loadRuntimeTemplate("smithers/agents/codex.tsx")); const attackedProject = tempProject(); mismatchedPath = path.join(attackedProject, "ultrafuzz.toml"); @@ -3626,26 +3628,30 @@ test("init preserves existing project-owned files and validate exposes launch po assert.equal(validate.value?.resolved_config?.default_reasoning, "xhigh"); }); -test("non-force init preserves historical stock agent adapters and force replaces them", () => { - assert.equal( - crypto.createHash("sha256").update(V0_0_2_STOCK_CODEX_ADAPTER).digest("hex"), - "26dae14e43c09dbe7901aa731cd552b282d502d86cea8cc6726e4a8579cd3236" - ); +test("non-force init refreshes historical and customized adapters to the packaged closure", () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); - const codexPath = path.join(project, ".smithers", "agents", "codex.ts"); - const currentAdapter = fs.readFileSync(codexPath, "utf8"); - fs.writeFileSync(codexPath, V0_0_2_STOCK_CODEX_ADAPTER, "utf8"); - const historicalStats = fs.statSync(codexPath, { bigint: true }); + const configPath = path.join(project, "ultrafuzz.toml"); + const customConfig = `${fs.readFileSync(configPath, "utf8")}\n# operator customization\n`; + fs.writeFileSync(configPath, customConfig, "utf8"); + const agents = path.join(project, ".smithers", "agents"); + fs.writeFileSync(path.join(agents, "codex.ts"), V0_0_2_STOCK_CODEX_ADAPTER, "utf8"); + fs.writeFileSync(path.join(agents, "environment.ts"), "export const customized = true;\n", "utf8"); + fs.writeFileSync( + path.join(agents, "index.ts"), + 'import { createCodexAgent } from "./codex";\nexport const agentFactories = { CodexAgent: createCodexAgent };\n', + "utf8" + ); + fs.rmSync(path.join(agents, "pi.ts")); - const preserved = initProject({ projectRoot: project }); + const refreshed = initProject({ projectRoot: project }); - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); - assert.equal(fs.readFileSync(codexPath, "utf8"), V0_0_2_STOCK_CODEX_ADAPTER); - assert.equal(fs.statSync(codexPath, { bigint: true }).ino, historicalStats.ino); - const forced = initProject({ projectRoot: project, force: true }); - assert.equal(forced.ok, true, JSON.stringify(forced.diagnostics)); - assert.equal(fs.readFileSync(codexPath, "utf8"), currentAdapter); + assert.equal(refreshed.ok, true, JSON.stringify(refreshed.diagnostics)); + assert.deepEqual(refreshed.diagnostics, []); + // Plan admits only the packaged closure, so plain init restores all of it + // while every project-owned file keeps its customization. + assert.doesNotThrow(() => inspectControllerSource(project)); + assert.equal(fs.readFileSync(configPath, "utf8"), customConfig); }); test("non-force init migrates a superseded generated manifest so the project keeps its durable run", () => { @@ -3732,81 +3738,6 @@ test("non-force init migrates the exact generated 0.32 package and immediately p assert.doesNotMatch(source, /smithers-orchestrator/u); }); -test( - "init converts an adapter inspection failure into a preserved manual-review warning", - { concurrency: false }, - () => { - const project = tempProject(); - assert.equal(initProject({ projectRoot: project, force: true }).ok, true); - const codexPath = path.join(project, ".smithers", "agents", "codex.ts"); - const customized = 'export const customConfig = "ultrafuzz.toml"; // project-owned adapter\n'; - fs.writeFileSync(codexPath, customized, "utf8"); - const originalDescriptor = Object.getOwnPropertyDescriptor(fs, "openSync")!; - const originalOpenSync = fs.openSync; - let codexOpenCount = 0; - - Object.defineProperty(fs, "openSync", { - ...originalDescriptor, - value: (...args: unknown[]) => { - if (String(args[0]) === codexPath) { - codexOpenCount += 1; - if (codexOpenCount === 1) { - throw Object.assign(new Error("induced sensitive adapter inspection failure"), { code: "EACCES" }); - } - } - return Reflect.apply(originalOpenSync, fs, args) as number; - } - }); - try { - const preserved = initProject({ projectRoot: project }); - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); - assert.equal(codexOpenCount, 1); - assert.equal(fs.readFileSync(codexPath, "utf8"), customized); - const warning = preserved.diagnostics.find( - (diagnostic) => - diagnostic.code === "INIT_AGENT_ADAPTER_UPDATE_REQUIRED" && diagnostic.path === ".smithers/agents/codex.ts" - ); - assert.equal(warning?.severity, "warning"); - assert.match(warning?.message ?? "", /could not be safely inspected/u); - assert.match(warning?.message ?? "", /verify manually/u); - assert.doesNotMatch(JSON.stringify(preserved.diagnostics), /induced sensitive|EACCES/u); - } finally { - Object.defineProperty(fs, "openSync", originalDescriptor); - } - } -); - -test("non-force init preserves a customized stale adapter and force remains explicit", () => { - const project = tempProject(); - assert.equal(initProject({ projectRoot: project, force: true }).ok, true); - const codexPath = path.join(project, ".smithers", "agents", "codex.ts"); - const customized = [ - 'import { readFileSync } from "node:fs";', - 'import path from "node:path";', - 'export const customConfig = readFileSync(path.join(process.cwd(), "ultrafuzz.toml"), "utf8");', - "// project-owned customization", - "" - ].join("\n"); - fs.writeFileSync(codexPath, customized, "utf8"); - - const preserved = initProject({ projectRoot: project }); - - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); - assert.equal(fs.readFileSync(codexPath, "utf8"), customized); - const warning = preserved.diagnostics.find( - (diagnostic) => - diagnostic.code === "INIT_AGENT_ADAPTER_UPDATE_REQUIRED" && diagnostic.path === ".smithers/agents/codex.ts" - ); - assert.equal(warning?.severity, "warning"); - assert.match(warning?.message ?? "", /process\.env\.ULTRAFUZZ_CONFIG_PATH/u); - assert.match(warning?.message ?? "", /workflowControlChildEnvironment/u); - - const forced = initProject({ projectRoot: project, force: true }); - assert.equal(forced.ok, true, JSON.stringify(forced.diagnostics)); - assert.notEqual(fs.readFileSync(codexPath, "utf8"), customized); - assert.match(fs.readFileSync(codexPath, "utf8"), /process\.env\.ULTRAFUZZ_CONFIG_PATH/u); -}); - test("non-force init never follows or overwrites linked adapter paths", () => { for (const linkKind of ["symbolic", "hard"] as const) { const project = tempProject(); @@ -3822,18 +3753,13 @@ test("non-force init never follows or overwrites linked adapter paths", () => { fs.linkSync(outsidePath, codexPath); } - const preserved = initProject({ projectRoot: project }); + const rejected = initProject({ projectRoot: project }); - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); + assert.equal(rejected.ok, false); + assert.equal(rejected.diagnostics[0]?.code, "INIT_PATH_UNSAFE"); assert.equal(fs.readFileSync(outsidePath, "utf8"), V0_0_2_STOCK_CODEX_ADAPTER); - assert.equal(fs.readFileSync(codexPath, "utf8"), V0_0_2_STOCK_CODEX_ADAPTER); assert.equal(fs.lstatSync(codexPath).isSymbolicLink(), linkKind === "symbolic"); if (linkKind === "hard") assert.equal(fs.statSync(codexPath).nlink, 2); - const warning = preserved.diagnostics.find( - (diagnostic) => - diagnostic.code === "INIT_AGENT_ADAPTER_UPDATE_REQUIRED" && diagnostic.path === ".smithers/agents/codex.ts" - ); - assert.match(warning?.message ?? "", /preserved it without inspection/u); } }); @@ -3849,11 +3775,11 @@ test("startRun rejects a customized controller adapter before submission", async assert.equal(run.ok, false); assert.equal(run.diagnostics[0]?.code, "CONTROLLER_SOURCE_UNTRUSTED"); assert.match(run.diagnostics[0]?.message ?? "", /must exactly match the packaged stock closure/u); - assert.match(run.diagnostics[0]?.message ?? "", /ultrafuzz init --force/u); + assert.match(run.diagnostics[0]?.message ?? "", /rerun ultrafuzz init$/u); assert.equal(fs.existsSync(path.join(project, "smithers-commands.log")), false); }); -test("init preserves a dangling adapter symlink without writing through it", () => { +test("init never writes through a dangling adapter symlink", () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); const codexPath = path.join(project, ".smithers", "agents", "codex.ts"); @@ -3864,7 +3790,8 @@ test("init preserves a dangling adapter symlink without writing through it", () const result = initProject({ projectRoot: project }); - assert.equal(result.ok, true); + assert.equal(result.ok, false); + assert.equal(result.diagnostics[0]?.code, "INIT_PATH_UNSAFE"); assert.equal(fs.existsSync(outsidePath), false); assert.equal(fs.readlinkSync(codexPath), outsidePath); }); @@ -9934,14 +9861,8 @@ test("validate accepts a typed aliased registry composed from static spreads", a ); const validate = await validateProject({ projectRoot: project, env: {} }); - const preserved = initProject({ projectRoot: project }); assert.equal(validate.ok, true, JSON.stringify(validate.diagnostics)); - assert.equal(preserved.ok, true, JSON.stringify(preserved.diagnostics)); - assert.equal( - preserved.diagnostics.some((diagnostic) => diagnostic.code === "INIT_AGENT_REGISTRY_STALE"), - false - ); }); test("validate applies registry overwrite order and rejects nullish or shadowed factories", async () => { @@ -12256,127 +12177,41 @@ test("startRun --agent does not carry the previous agent's model onto the new ag assert.equal(pinnedTasks.tasks[0]?.reasoningEffort ?? null, null); }); -test("init reports an agent registry that does not export a generated agent", async () => { - const project = tempProject(); - initProject({ projectRoot: project, force: true }); - - // Simulate a project scaffolded before ClaudeAgent, DeepSeekAgent, KimiAgent, - // OpenCodeAgent, OpenRouterAgent, and PiAgent existed: the registry predates the adapters, and init preserves - // project-owned files. - const registryPath = path.join(project, ".smithers/agents/index.ts"); - fs.writeFileSync( - registryPath, - 'import { createCodexAgent } from "./codex";\n' + - 'export { createCodexAgent } from "./codex";\n' + - "export const agentFactories = { CodexAgent: createCodexAgent };\n", - "utf8" - ); - - const upgraded = initProject({ projectRoot: project }); - assert.equal(upgraded.ok, true); - const stale = upgraded.diagnostics.filter((entry) => entry.code === "INIT_AGENT_REGISTRY_STALE"); - assert.equal(stale.length, 6, JSON.stringify(upgraded.diagnostics)); - assert.equal(stale[0]?.severity, "warning"); - assert.match(stale.map((entry) => entry.message).join("\n"), /ClaudeAgent/); - assert.match(stale.map((entry) => entry.message).join("\n"), /DeepSeekAgent/); - assert.match(stale.map((entry) => entry.message).join("\n"), /KimiAgent/); - assert.match(stale.map((entry) => entry.message).join("\n"), /OpenCodeAgent/); - assert.match(stale.map((entry) => entry.message).join("\n"), /OpenRouterAgent/); - assert.match(stale.map((entry) => entry.message).join("\n"), /PiAgent/); - - // A registry that names Claude, DeepSeek, Kimi, OpenCode, OpenRouter, and Pi without registering - // their factories is still stale: nothing resolves it, since generated - // adapters export only factories. - fs.writeFileSync( - registryPath, - 'import { createCodexAgent } from "./codex";\n' + - 'export { CodexAgent, createCodexAgent } from "./codex";\n' + - 'export { ClaudeAgent } from "./claude";\n' + - "export const agentFactories = { CodexAgent: createCodexAgent };\n", - "utf8" - ); - const named = initProject({ projectRoot: project }); - const namedStale = named.diagnostics.filter((entry) => entry.code === "INIT_AGENT_REGISTRY_STALE"); - assert.equal(namedStale.length, 6, JSON.stringify(named.diagnostics)); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /ClaudeAgent/); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /DeepSeekAgent/); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /KimiAgent/); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /OpenCodeAgent/); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /OpenRouterAgent/); - assert.match(namedStale.map((entry) => entry.message).join("\n"), /PiAgent/); - - // A registry that exports every generated agent stays quiet. - const regenerated = initProject({ projectRoot: project, force: true }); - assert.equal( - regenerated.diagnostics.filter((entry) => entry.code === "INIT_AGENT_REGISTRY_STALE").length, - 0, - JSON.stringify(regenerated.diagnostics) - ); -}); - -test( - "post-init registry inspection sanitizes access failures instead of failing after mutation", - { concurrency: false }, - () => { - const project = tempProject(); - assert.equal(initProject({ projectRoot: project, force: true }).ok, true); - const registryPath = path.join(project, ".smithers", "agents", "index.ts"); - const originalDescriptor = Object.getOwnPropertyDescriptor(fs, "openSync")!; - const originalOpenSync = fs.openSync; - - Object.defineProperty(fs, "openSync", { - ...originalDescriptor, - value: (...args: unknown[]) => { - if (String(args[0]) === registryPath) { - throw Object.assign(new Error("sensitive registry access detail"), { code: "EACCES" }); - } - return Reflect.apply(originalOpenSync, fs, args) as number; - } - }); - try { - const inspected = initProject({ projectRoot: project }); - assert.equal(inspected.ok, true, JSON.stringify(inspected.diagnostics)); - const warning = inspected.diagnostics.find( - (diagnostic) => diagnostic.code === "INIT_AGENT_REGISTRY_REVIEW_REQUIRED" - ); - assert.equal(warning?.severity, "warning"); - assert.match(warning?.message ?? "", /could not be safely inspected/u); - assert.match(warning?.message ?? "", /verify manually/u); - assert.doesNotMatch(JSON.stringify(inspected.diagnostics), /sensitive registry|EACCES/u); - } finally { - Object.defineProperty(fs, "openSync", originalDescriptor); - } - } -); - testWhen(process.platform !== "win32")( - "post-init registry inspection rejects symlinks, FIFOs, and oversized files without reading them", + "init replaces a stale registry but never writes through a linked or special registry path", () => { - const cases = ["symlink", "fifo", "oversized"] as const; - for (const kind of cases) { + const stock = loadRuntimeTemplate("smithers/agents/index.tsx"); + const staleProject = tempProject(); + assert.equal(initProject({ projectRoot: staleProject, force: true }).ok, true); + const staleRegistry = path.join(staleProject, ".smithers", "agents", "index.ts"); + // A registry scaffolded before the other adapters existed. + fs.writeFileSync( + staleRegistry, + 'import { createCodexAgent } from "./codex";\nexport const agentFactories = { CodexAgent: createCodexAgent };\n', + "utf8" + ); + assert.equal(initProject({ projectRoot: staleProject }).ok, true); + assert.equal(fs.readFileSync(staleRegistry, "utf8"), stock); + + for (const kind of ["symlink", "fifo"] as const) { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); const registryPath = path.join(project, ".smithers", "agents", "index.ts"); fs.unlinkSync(registryPath); + const outside = path.join(tempProject(), "outside-index.ts"); if (kind === "symlink") { - const outside = path.join(tempProject(), "outside-index.ts"); - fs.writeFileSync(outside, "outside registry must not be read\n", "utf8"); + fs.writeFileSync(outside, "outside registry must not be written\n", "utf8"); fs.symlinkSync(outside, registryPath); - } else if (kind === "fifo") { - execFileSync("mkfifo", [registryPath]); } else { - fs.writeFileSync(registryPath, Buffer.alloc(256 * 1024 + 1, 0x61)); + execFileSync("mkfifo", [registryPath]); } - const inspected = initProject({ projectRoot: project }); + const rejected = initProject({ projectRoot: project }); - assert.equal(inspected.ok, true, `${kind}: ${JSON.stringify(inspected.diagnostics)}`); - const warning = inspected.diagnostics.find( - (diagnostic) => diagnostic.code === "INIT_AGENT_REGISTRY_REVIEW_REQUIRED" - ); - assert.equal(warning?.severity, "warning", kind); - assert.match(warning?.message ?? "", /preserved (?:it )?without inspection|too large to inspect/u, kind); - assert.match(warning?.message ?? "", /verify manually/u, kind); + assert.equal(rejected.ok, false, kind); + assert.equal(rejected.diagnostics[0]?.code, "INIT_PATH_UNSAFE", kind); + if (kind === "symlink") assert.equal(fs.readFileSync(outside, "utf8"), "outside registry must not be written\n"); + else assert.equal(fs.lstatSync(registryPath).isFIFO(), true); } } ); From 4dacfd71b860350feb52a508732bbecf17378e5b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:28:58 +0000 Subject: [PATCH 069/206] docs(changelog): record the adapter fail-open and route digest changes Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..3ad6f6f92 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,7 @@ ### Breaking changes +- **[runtime] [docs]** Provider route IDs (`model:-route-`) now digest only route-bearing input, through one runtime function that plan-time acknowledgement and every adapter invocation share. Proxy variables no longer count, Claude's `AWS_`, `GOOGLE_`/`CLOUD_ML_`, and `AZURE_`/`FOUNDRY_` variables count only while a matching `CLAUDE_CODE_USE_*` platform flag is set, and a Codex `config.toml` or Claude `settings.json` contributes only its selected provider (id, `base_url`, `wire_api`, `env_key`) or its credential helpers and routing `env`, so a CLI rewriting its own config, or a resume from a shell with different proxy or unused cloud settings, no longer fails every later task with "provider route changed after disclosure acknowledgement". Route IDs that included the dropped inputs change: update `ULTRAFUZZ_DATA_GOVERNANCE_POLICY` and re-acknowledge, and re-plan rather than resume a run whose acknowledged ID changed (#908). - **[runtime] [cli] [docs]** Restores native Smithers continuation as the normal `resume` behavior. Ordinary resume now sends the persisted workflow and same Smithers run ID directly to Smithers with workflow-change acceptance instead of requiring Ultrafuzz control seals, link journals, controller generations, current schema bindings, graph identity, or metadata projections. `--refresh-controller` renders the current controller beside the historical source and continues that same Smithers run; it no longer publishes an authenticated historical generation. Completed Smithers rows and artifact bytes are not reset, replayed, migrated, or rewritten, and missing optional final-report metadata renders as `unavailable`. Replay determinism after accepting changed workflow source is therefore Smithers' caller-visible responsibility. The controller re-finalization CLI option is removed; use native resume or the existing explicit reset/retry operations (#939). - **[config] [docs] [evals]** Removes the `full` audit profile and renames its packaged topology to `topologies/exhaustive.yml` (byte-identical, topology digest unchanged); the `exhaustive` profile now runs that complete specialist workflow at its existing maximum settings, and no longer follows an edited `.ultrafuzz/topology.yml` (a project `topology_path` or `--topology-path` still overrides it). Configurations that select `full` now fail with the standard unknown-profile diagnostic. The public `full` and `threat-model` benchmark lanes launch with the `exhaustive` profile; the only effective policy change on those lanes is `same_agent_attempts` 3 → 5, since lane and target pins outrank the remaining profile settings (#792). - **[config] [docs]** Renames the unmodified built-in audit profile from `balanced` to `default` and removes the separate default pointer. Configurations that select `balanced` now fail with the standard unknown-profile diagnostic and must select `default` or omit `audit_profile`. Historical run metadata retains its exact recorded id; eval-history and benchmark reporting do not key cohorts by audit-profile id, so otherwise matching runs remain in the same reporting cohort across the migration boundary (#534). @@ -12,6 +13,7 @@ ### Other changes +- **[runtime] [cli] [docs]** Agent adapters no longer hang or fail tasks on telemetry and environment quirks: DeepSeekAgent stops rejecting the Anthropic usage fields Claude Code always prints (every DeepSeek task waited for its timeout) and imports `path`, Kimi and Pi output, session, and usage parsing fails open instead of throwing inside the child's output listeners, and native continuations no longer blank adapter-owned paths under the target such as OpenCode's XDG roots. `ultrafuzz init` now rewrites the stock `.smithers/agents` closure on every run, because planning accepts only its byte-exact packaged copy, so upgrades no longer need `init --force` (which also resets `ultrafuzz.toml`, the topology, and prompts); a linked or special file at an adapter path now fails init instead of being preserved. ESLint checks runtime templates with `no-undef`, and the adapter-boundary test no longer pins template hashes or sizes. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From d29d86b4ec751650cbce76cd0c8ac207a8743661 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:13:46 +0000 Subject: [PATCH 070/206] docs: explain OpenRouter prompt-injection guardrail rejections OpenRouterAgent, PiAgent, and OpenCodeAgent with an openrouter/ model all use the key in OPENROUTER_API_KEY. When a guardrail that covers that key sets prompt-injection detection to Block, OpenRouter rejects each matching request with HTTP 403 "Request blocked: prompt injection patterns detected" before it reaches a model. #1149 blamed transcript-like examples in Ultrafuzz's prompts. OpenRouter's documented exact regexes match none of Ultrafuzz's prompt sources or the rendered prompts of a local run. They do match text Ultrafuzz does not write: OpenCode 1.18.18's default system prompt, used for models without a model-specific prompt (DeepSeek, Qwen, GLM), matches role_delimiter_injection, and so does ordinary Vyper or YAML source. Document the operator fix in docs/config.md. Set prompt-injection detection to Flag, or turn it off, on every guardrail that covers the key, because OpenRouter applies the most restrictive action across the workspace default and member or key guardrails. Do not use Redact, which forwards the request with each match replaced. Runtime behavior is unchanged: the rejection is an ordinary agent failure under the [retry] policy, and the doc points there instead of restating it. Closes #1149 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/config.md | 31 +++++++++++++++++++++++++++++++ 2 files changed, 32 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..195754c2f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[docs]** `docs/config.md` now explains the HTTP 403 `Request blocked: prompt injection patterns detected` that `OpenRouterAgent`, `OpenCodeAgent`, and `PiAgent` get when an OpenRouter guardrail on their key sets prompt-injection detection to Block: set it to Flag, or turn it off, on every guardrail that covers the key, and do not use Redact, which forwards altered content. Runtime behavior is unchanged; the rejection is an ordinary agent failure under the `[retry]` policy (#1149). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/config.md b/docs/config.md index 4ccc55e3e..1387a1fc2 100644 --- a/docs/config.md +++ b/docs/config.md @@ -376,6 +376,37 @@ execution. The Smithers type surface pinned by this release stops at `xhigh`, but pi's command surface also accepts `max`, so the adapter validates and forwards that final level without degrading it. +## OpenRouter guardrails + +`OpenRouterAgent`, `PiAgent`, and `OpenCodeAgent` with an `openrouter/` model +(as in the shipped profile) send their requests through OpenRouter with the key +in `OPENROUTER_API_KEY`. If any guardrail covering that key sets +[prompt-injection detection](https://openrouter.ai/docs/guides/features/guardrails/prompt-injection) +to **Block**, OpenRouter rejects each request its detector matches with HTTP +403 `Request blocked: prompt injection patterns detected` before it reaches a +model. + +The match need not be in Ultrafuzz's task prompt. OpenRouter scans every +message in a request, including base64- and hex-decoded text, and these +harnesses also send their own system prompts and the target source and test +output the agent reads. For example, the default system prompt of OpenCode +1.18.18, used for models without a model-specific prompt such as DeepSeek, +Qwen, or GLM, has an `assistant: [...]` line followed by a `user:` line, which +matches OpenRouter's documented `role_delimiter_injection` pattern. + +In the **Security** section of every guardrail that covers the key (the +workspace default and any member or API-key guardrail), set prompt-injection +detection to **Flag**, which records matches without enforcing them, or turn it +off. OpenRouter applies the most restrictive action when several guardrails +apply. Do not use **Redact** either: it replaces each match with +`[PROMPT_INJECTION]` and forwards the request, so the model can work from +altered source or tool output with no error for Ultrafuzz to report. + +Ultrafuzz has no special handling for this rejection: the attempt fails like any +other agent error and follows the `[retry]` policy above. A retry on the same +profile sends the same task prompt with the same key, so it is rejected again +when the match is in that prompt or in the harness's system prompt. + ## Forge process guard Worker environments put a run-scoped Forge wrapper ahead of the installed From 92726882780d139a68a3cc0aa7b60d2c85171f03 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:14:50 +0000 Subject: [PATCH 071/206] ci: run every release validation lane on pull requests and never cancel main runs Pull requests ran only the five runtime lanes. package-gates, cli and benchmark-history-typecheck (which holds the workspace typecheck) first ran after merge, so their regressions landed on main: the latest main run, 36495460966, fails package-gates. Every lane now runs on every event. The lane table keeps its test that it runs each gate scripts/validate-release.mjs defines exactly once; the pull-request filter, PULL_REQUEST_REQUIRED_GATES and the --event argument are gone. The lanes no longer wait for the build and super-linter jobs, and the PR-only runtime and CLI smoke runners are deleted because the lanes run those suites in full. release-gates still requires every job. Measured on existing runs: a pull request took about 60 minutes (8.0 min for build and smoke, then 51.5 min for the slowest lane in run 36279711018); on the latest main push the new PR lanes took 12.9 (package-gates), 41.5 (cli) and 9.8 minutes against 26.8-39.5 minutes for the runtime lanes. All pushes to main shared one concurrency group; 24 of the last 40 main push runs were cancelled (as of 2026-09-29). cancel-in-progress: false alone would not stop that, because GitHub also cancels a pending run in the group when a newer run queues, and main pushes arrive in bursts (three within 9 seconds on 2026-09-28). Runs other than pull requests now get their own group, keyed by github.run_id. Delete the ci.yml line-by-line assertions in packages/modal/test/ci-config.test.ts and the timeout_minutes === 120 change detector. The remaining guard parses ci.yml and checks that release validation is not gated off pull requests. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/ci.yml | 39 ++-- docs/reference/agent-adapter-boundaries.md | 4 +- docs/reference/development.md | 24 ++- packages/cli/package.json | 1 - packages/cli/scripts/run-pr-smoke-tests.mjs | 62 ------ packages/modal/test/ci-config.test.ts | 186 ------------------ packages/runtime/package.json | 1 - .../runtime/scripts/run-pr-smoke-tests.mjs | 159 --------------- scripts/ci/release-validation-lanes.mjs | 67 +------ scripts/ci/release-validation-lanes.test.ts | 123 ++++-------- 10 files changed, 67 insertions(+), 599 deletions(-) delete mode 100644 packages/cli/scripts/run-pr-smoke-tests.mjs delete mode 100644 packages/runtime/scripts/run-pr-smoke-tests.mjs diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index fb1b98a4f..0174be83f 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -16,18 +16,18 @@ on: permissions: contents: read +# A new push to a pull request cancels that pull request's older run. Every +# other run gets its own group: a shared group would still cancel a queued run +# on main whenever another push arrived. concurrency: - group: ci-${{ github.workflow }}-${{ github.event.pull_request.number || github.ref }} + group: ci-${{ github.workflow }}-${{ github.event.pull_request.number || github.run_id }} cancel-in-progress: true jobs: draft-and-build-gates: - name: "${{ github.event_name == 'pull_request' && 'PR build and Node.js 24 runtime smoke' || 'Build gates' }}" + name: Build gates runs-on: ubuntu-latest - # The PR path also runs the serial runtime and CLI contract smoke suites. - # Under runner contention they completed successfully just after the former - # 15-minute ceiling, so keep a bounded two-times completion budget. - timeout-minutes: 30 + timeout-minutes: 15 steps: - name: Check out repository @@ -84,14 +84,6 @@ jobs: - name: Enforce production dependency advisory policy run: pnpm -w security:dependency-advisories - - name: Run PR runtime smoke tests - if: github.event_name == 'pull_request' - run: pnpm --filter @ultrafuzz/runtime test:pr-smoke:prebuilt - - - name: Run PR CLI status contract smoke tests - if: github.event_name == 'pull_request' - run: pnpm --filter @ultrafuzz/cli test:pr-smoke:prebuilt - external-static-analysis: name: External static analysis runs-on: ubuntu-latest @@ -125,10 +117,10 @@ jobs: VALIDATE_SHELL_SHFMT: true VALIDATE_YAML: true - # The lane table lives in scripts/ci/release-validation-lanes.mjs so the - # pull-request policy is a tested artifact instead of an invisible `if:`. - # While this job was gated off pull requests, every runtime test reported - # `skipping` on a PR and resume-path regressions merged with all checks green. + # The lane table lives in scripts/ci/release-validation-lanes.mjs, whose test + # checks that it runs every gate scripts/validate-release.mjs defines. Every + # lane runs on pull requests too: while lanes were gated off pull requests, + # regressions in the skipped suites merged with all checks green. release-validation-lanes: name: Select release validation lanes runs-on: ubuntu-latest @@ -144,17 +136,14 @@ jobs: - name: Select release validation lanes id: select - env: - EVENT_NAME: ${{ github.event_name }} run: | - echo "lanes=$(node scripts/ci/release-validation-lanes.mjs --event "$EVENT_NAME")" >> "$GITHUB_OUTPUT" + echo "lanes=$(node scripts/ci/release-validation-lanes.mjs)" >> "$GITHUB_OUTPUT" + # The lanes start without waiting for the build gates; release-gates still + # requires every job. release-validation: name: Full release validation (${{ matrix.description }}) - needs: - - draft-and-build-gates - - external-static-analysis - - release-validation-lanes + needs: release-validation-lanes runs-on: ubuntu-latest timeout-minutes: ${{ matrix.timeout_minutes }} strategy: diff --git a/docs/reference/agent-adapter-boundaries.md b/docs/reference/agent-adapter-boundaries.md index 179ae3fcf..8c015098a 100644 --- a/docs/reference/agent-adapter-boundaries.md +++ b/docs/reference/agent-adapter-boundaries.md @@ -63,5 +63,5 @@ responsibility in the same reviewed pull request. The gate is `packages/runtime/test/agent-adapter-boundaries.test.ts`. It scans every TypeScript source file in the adapter directory, derives shipped adapters -from `agentFactories`, and runs in both the required pull-request runtime smoke -path and the full runtime supporting-test shard. +from `agentFactories`, and runs in the runtime supporting-test lane, which pull +requests require. diff --git a/docs/reference/development.md b/docs/reference/development.md index 0c441735a..6aabe8a24 100644 --- a/docs/reference/development.md +++ b/docs/reference/development.md @@ -40,19 +40,17 @@ The root CI script runs format check, lint, build, and release validation: pnpm -w run ci ``` -Pull requests run CI policy checks, dependency policy, formatting, lint, the -workspace build, and a curated runtime smoke suite. The smoke suite reuses the -built workspace and covers runtime sharding, workflow controls, source revision -binding, generated workflow input, and representative initialization, -validation, planning, and workflow-compilation behavior. Feature branches are -validated only by the pull-request event, avoiding a duplicate push run. - -Pull requests also require all five runtime validation lanes: one supporting -lane, including the Bun adapter contracts, and four deterministic integration -shards. Together these run the full runtime suite before merge. Pushes to `main` -and manual workflow dispatches run all eight release lanes, adding package, -CLI, and benchmark-history/typecheck checks, with at most eight jobs in parallel. -Their results are recorded in stable gate order in the JSON report. +Every CI run, pull requests included, runs the build gates (CI policy checks, +formatting, lint, dead-code checks, the workspace build, strict lint of changed +lines, bundle budgets, and dependency policy) and all eight release validation +lanes: package gates, one runtime supporting lane that includes the Bun adapter +contracts, four runtime integration shards, the CLI suite, and benchmark history +with the workspace typecheck. The lanes do not wait for the build gates and run +at most eight at a time. Feature branches are validated only by the +pull-request event, avoiding a duplicate push run. A newer push to a pull +request cancels that pull request's older run; a push to `main` never cancels +another run. On `main`, the lane results are merged into one JSON report in +stable gate order. ## Package Checks diff --git a/packages/cli/package.json b/packages/cli/package.json index 5b3a2fd9e..4b5f7fbd4 100644 --- a/packages/cli/package.json +++ b/packages/cli/package.json @@ -26,7 +26,6 @@ "scripts": { "build": "rm -rf dist && tsc -p tsconfig.json && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", "test": "pnpm --filter @ultrafuzz/cli... build && rm -rf dist-test && tsc -p tsconfig.test.json && node --test --test-concurrency=1 dist-test/test/*.test.js", - "test:pr-smoke:prebuilt": "rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-pr-smoke-tests.mjs", "typecheck": "pnpm --filter @ultrafuzz/cli^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/cli/scripts/run-pr-smoke-tests.mjs b/packages/cli/scripts/run-pr-smoke-tests.mjs deleted file mode 100644 index 71fbe9af1..000000000 --- a/packages/cli/scripts/run-pr-smoke-tests.mjs +++ /dev/null @@ -1,62 +0,0 @@ -import { spawnSync } from "node:child_process"; -import { readFileSync } from "node:fs"; - -const namedTests = new Map([ - ["test/cli.test.ts", ["status surfaces a terminal product and live workflow lifecycle divergence"]], - [ - "test/cli-contracts.test.ts", - [ - "known command failures can retain a valid typed data snapshot", - "status separates successful queries, partial reports, unavailable reports, and unknown completion" - ] - ], - [ - "test/status-report.test.ts", - [ - "watch stops at ended runs with unavailable reports and distinguishes attention stops", - "status shows complete execution with a partial unchecked report without changing the execution verdict", - "status shows report failure and unknown coverage without presenting a false report path" - ] - ] -]); -const selectedTestNames = [...namedTests.values()].flat(); - -for (const [sourcePath, names] of namedTests) { - const source = readFileSync(sourcePath, "utf8"); - for (const name of names) { - if (!source.includes(`test(${JSON.stringify(name)}`)) { - throw new Error(`PR CLI smoke test is not registered: ${name}`); - } - } -} - -const pattern = `^(?:${selectedTestNames.map(escapeRegExp).join("|")})$`; -const files = [...namedTests.keys()].map((sourcePath) => `dist-test/${sourcePath.replace(/\.ts$/u, ".js")}`); -const result = spawnSync( - process.execPath, - ["--test", "--test-concurrency=1", "--test-reporter=tap", `--test-name-pattern=${pattern}`, ...files], - { - encoding: "utf8", - stdio: ["inherit", "pipe", "pipe"] - } -); -process.stdout.write(result.stdout ?? ""); -process.stderr.write(result.stderr ?? ""); -if (result.error !== undefined) throw result.error; -if (result.status !== 0) process.exit(result.status ?? 1); -assertNamedTestsPassed(result.stdout ?? "", selectedTestNames); - -function escapeRegExp(value) { - return value.replace(/[.*+?^${}()|[\]\\]/gu, "\\$&"); -} - -function assertNamedTestsPassed(output, expectedTestNames) { - const passedTestNames = [...output.matchAll(/^ok [0-9]+ - (.+)$/gmu)] - .map((match) => match[1]) - .filter((name) => !name.includes(" # SKIP") && !name.includes(" # TODO")) - .sort(); - const expected = [...expectedTestNames].sort(); - if (passedTestNames.length !== expected.length || passedTestNames.some((name, index) => name !== expected[index])) { - throw new Error(`PR CLI smoke passed unexpected tests: ${JSON.stringify(passedTestNames)}`); - } -} diff --git a/packages/modal/test/ci-config.test.ts b/packages/modal/test/ci-config.test.ts index 6df0d84eb..c7edb8762 100644 --- a/packages/modal/test/ci-config.test.ts +++ b/packages/modal/test/ci-config.test.ts @@ -4,7 +4,6 @@ import os from "node:os"; import path from "node:path"; import { describe, expect, it } from "vitest"; -import { parse } from "yaml"; import { MODAL_PUBLIC_FULL_SANDBOX_TIMEOUT_MS } from "../src/defaults.js"; import { @@ -622,191 +621,6 @@ describe("public Modal benchmark configuration", () => { } }); - it("runs the release validation lanes the policy script selects, on pull requests too", () => { - const workspace = path.resolve("../.."); - const workflow = parse(fs.readFileSync(path.join(workspace, ".github/workflows/ci.yml"), "utf8")) as { - on: { - push: { branches: string[] }; - pull_request: { types: string[] }; - workflow_dispatch: null; - }; - concurrency: { group: string; "cancel-in-progress": boolean }; - jobs: Record< - string, - { - name?: string; - if?: string; - needs?: string[]; - "timeout-minutes"?: number | string; - outputs?: Record; - strategy?: { - "fail-fast": boolean; - "max-parallel": number; - // #994 replaced the hard-coded lane table with a matrix expression - // expanded from the `release-validation-lanes` job output. - matrix: { include: string }; - }; - steps: Array<{ - name?: string; - id?: string; - if?: string; - run?: string; - uses?: string; - env?: Record; - with?: Record; - }>; - } - >; - }; - - expect(workflow.on.push.branches).toEqual(["main"]); - expect(workflow.on).not.toHaveProperty("merge_group"); - expect(workflow.on.workflow_dispatch).toBeNull(); - expect(workflow.on.pull_request.types).toEqual([ - "opened", - "synchronize", - "reopened", - "ready_for_review", - "converted_to_draft" - ]); - expect(workflow.concurrency.group).toContain("github.event.pull_request.number"); - expect(workflow.concurrency.group).toContain("github.ref"); - expect(workflow.concurrency["cancel-in-progress"]).toBe(true); - - const draftAndBuild = workflow.jobs["draft-and-build-gates"]; - expect(draftAndBuild?.name).toBe( - "${{ github.event_name == 'pull_request' && 'PR build and Node.js 24 runtime smoke' || 'Build gates' }}" - ); - // The PR-only runtime and CLI smoke suites have reached the old 15-minute - // ceiling while still making progress, so pin the bounded completion budget. - expect(draftAndBuild?.["timeout-minutes"]).toBe(30); - const steps = draftAndBuild?.steps ?? []; - const bunSetup = steps.find((step) => step.name === "Set up Bun"); - expect(bunSetup?.uses).toBe("oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6"); - expect(bunSetup?.with?.["bun-version"]).toBe("1.3.14"); - for (const name of [ - "Test CI policy scripts", - "Enforce production dependency advisory policy", - "Check formatting", - "Lint", - "Build" - ]) { - expect(steps.find((step) => step.name === name)?.if, `${name} must run for drafts`).toBeUndefined(); - } - expect(steps.find((step) => step.name === "Test CI policy scripts")?.run).toBe("pnpm -w test:ci-scripts"); - expect(steps.find((step) => step.name === "Enforce production dependency advisory policy")?.run).toBe( - "pnpm -w security:dependency-advisories" - ); - const runtimeSmoke = steps.find((step) => step.name === "Run PR runtime smoke tests"); - expect(runtimeSmoke?.if).toBe("github.event_name == 'pull_request'"); - expect(runtimeSmoke?.run).toBe("pnpm --filter @ultrafuzz/runtime test:pr-smoke:prebuilt"); - const cliStatusSmoke = steps.find((step) => step.name === "Run PR CLI status contract smoke tests"); - expect(cliStatusSmoke?.if).toBe("github.event_name == 'pull_request'"); - expect(cliStatusSmoke?.run).toBe("pnpm --filter @ultrafuzz/cli test:pr-smoke:prebuilt"); - - expect(workflow.jobs).not.toHaveProperty("pull-request-validation"); - - // The selection job is the workflow's only source of lanes, so pin the wiring - // end to end: the job publishes what the policy script prints, and the - // matrix expands exactly that output. - const laneSelection = workflow.jobs["release-validation-lanes"]; - expect(laneSelection?.outputs?.lanes).toBe("${{ steps.select.outputs.lanes }}"); - const selectStep = laneSelection?.steps.find((step) => step.name === "Select release validation lanes"); - expect(selectStep?.id).toBe("select"); - expect(selectStep?.env?.EVENT_NAME).toBe("${{ github.event_name }}"); - expect(selectStep?.run).toContain('scripts/ci/release-validation-lanes.mjs --event "$EVENT_NAME"'); - - const releaseValidation = workflow.jobs["release-validation"]; - expect(releaseValidation?.name).toBe("Full release validation (${{ matrix.description }})"); - // Regression guard for the outage this design exists to prevent. While this - // job carried `if: github.event_name != 'pull_request'`, every runtime test - // reported `skipping` on pull requests, so resume-path regressions merged - // with all checks green. - expect(releaseValidation?.if, "release validation must not be gated off pull requests").toBeUndefined(); - expect(releaseValidation?.needs).toEqual([ - "draft-and-build-gates", - "external-static-analysis", - "release-validation-lanes" - ]); - expect(releaseValidation?.strategy).toEqual({ - "fail-fast": false, - "max-parallel": 8, - matrix: { include: "${{ fromJSON(needs.release-validation-lanes.outputs.lanes) }}" } - }); - expect(releaseValidation?.["timeout-minutes"]).toBe("${{ matrix.timeout_minutes }}"); - expect(releaseValidation?.steps.find((step) => step.name === "Validate release lane")?.run).toContain("--gates"); - const modalDependentLaneBuild = releaseValidation?.steps.find( - (step) => step.name === "Build Modal-dependent lane dependencies" - ); - expect(modalDependentLaneBuild?.if).toBe("matrix.build_modal_dependencies == true"); - expect(modalDependentLaneBuild?.run).toBe("pnpm --filter @ultrafuzz/modal... build"); - const releaseReporterBuild = releaseValidation?.steps.find( - (step) => step.name === "Build release reporter dependencies" - ); - expect(releaseReporterBuild?.if).toBe("matrix.build_release_reporter == true"); - expect(releaseReporterBuild?.run).toBe("pnpm --filter @ultrafuzz/artifacts... build"); - // The split benchmark-history lane no longer shares a job with the `cli` gate, - // so it has to build the CLI closure that `benchmark:check:prebuilt` executes. - const cliLaneBuild = releaseValidation?.steps.find((step) => step.name === "Build CLI lane dependencies"); - expect(cliLaneBuild?.if).toBe("matrix.build_cli == true"); - expect(cliLaneBuild?.run).toBe("pnpm --filter @ultrafuzz/cli... build"); - // The lane table moved out of this file, so read it back the way the - // workflow does — by running the policy script for an integration event — - // rather than dropping the coverage this test used to carry. - const selection = execFileSync( - process.execPath, - [path.join(workspace, "scripts/ci/release-validation-lanes.mjs"), "--event", "push"], - { cwd: workspace, encoding: "utf8" } - ); - const pushLanes = JSON.parse(selection) as Array<{ lane: string; gates: string; timeout_minutes: number }>; - const laneGateIds = pushLanes.flatMap((entry) => entry.gates.split(",")); - expect(new Set(laneGateIds).size, "release validation lanes must not repeat a gate").toBe(laneGateIds.length); - expect(laneGateIds).toContain("cli"); - expect(laneGateIds).toContain("benchmark-history"); - expect(laneGateIds).toContain("workspace-typecheck"); - expect(pushLanes.map((entry) => entry.lane)).toContain("package-gates"); - expect(releaseValidation?.steps.find((step) => step.name === "Validate benchmark history charts")).toBeUndefined(); - const releaseGates = workflow.jobs["release-gates"]; - expect(releaseGates?.needs).toEqual([ - "draft-and-build-gates", - "external-static-analysis", - "release-validation-lanes", - "release-validation" - ]); - expect(releaseGates?.steps.find((step) => step.name === "Require the release validation lane selection")?.if).toBe( - "needs.release-validation-lanes.result != 'success'" - ); - for (const name of [ - "Check out repository", - "Set up pnpm", - "Set up Node.js", - "Install dependencies", - "Build release reporter dependencies", - "Download release validation lanes", - "Merge release validation report" - ]) { - expect(releaseGates?.steps.find((step) => step.name === name)?.if).toBe("github.event_name != 'pull_request'"); - } - expect(releaseGates?.steps.find((step) => step.name === "Install dependencies")?.run).toBe( - "pnpm install --frozen-lockfile" - ); - expect(releaseGates?.steps.find((step) => step.name === "Build release reporter dependencies")?.run).toBe( - "pnpm --filter @ultrafuzz/artifacts... build" - ); - expect(releaseGates?.steps.find((step) => step.name === "Merge release validation report")?.run).toContain( - "--merge-report-dir" - ); - // Formerly gated on `github.event_name != 'pull_request'`; #994 made the - // lanes required on pull requests, so the requirement must apply to every - // event. - expect(releaseGates?.steps.find((step) => step.name === "Require release validation lanes")?.if).toBe( - "always() && needs.release-validation.result != 'success'" - ); - expect(releaseGates?.steps.find((step) => step.name === "Upload release validation report")?.if).toBe( - "always() && github.event_name != 'pull_request'" - ); - }); - it("keeps Actions limited to CI without provider secrets or paid benchmark launches", () => { const workspace = path.resolve("../.."); const workflowRoot = path.join(workspace, ".github/workflows"); diff --git a/packages/runtime/package.json b/packages/runtime/package.json index e80034ba5..a89203db4 100644 --- a/packages/runtime/package.json +++ b/packages/runtime/package.json @@ -18,7 +18,6 @@ "scripts": { "build": "tsc -p tsconfig.json && cp ../../.ultrafuzz/topology.yml dist/topology.yml && rm -rf dist/templates && cp -R src/templates dist/templates && rm -rf dist/schema && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", "test": "pnpm --filter @ultrafuzz/runtime... build && rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-tests.mjs && bun test dist-test/test/runtime.test.js --test-name-pattern '^Bun adapter contract:'", - "test:pr-smoke:prebuilt": "rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-pr-smoke-tests.mjs", "test:release:runtime-shard": "pnpm --filter @ultrafuzz/runtime... build && rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-runtime-test-shard.mjs", "test:release:supporting": "pnpm --filter @ultrafuzz/runtime... build && rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-tests.mjs supporting && bun test dist-test/test/runtime.test.js --test-name-pattern '^Bun adapter contract:'", "test:kimi-contract": "pnpm --filter @ultrafuzz/runtime... build && rm -rf dist-test && tsc -p tsconfig.test.json && bun test dist-test/test/runtime.test.js --test-name-pattern '^Bun adapter contract: generated Kimi'", diff --git a/packages/runtime/scripts/run-pr-smoke-tests.mjs b/packages/runtime/scripts/run-pr-smoke-tests.mjs deleted file mode 100644 index 5b6c249df..000000000 --- a/packages/runtime/scripts/run-pr-smoke-tests.mjs +++ /dev/null @@ -1,159 +0,0 @@ -import { spawnSync } from "node:child_process"; -import { readFileSync } from "node:fs"; - -const supportingTestFiles = [ - "dist-test/test/agent-adapter-boundaries.test.js", - "dist-test/test/runtime-test-shard.test.js", - "dist-test/test/report-publication-status.test.js", - "dist-test/test/source-revision.test.js", - "dist-test/test/workflow-control.test.js" -]; - -const namedTests = new Map([ - [ - "test/runtime.test.ts", - [ - "init resolves one-hour node and execution-resource timeout defaults", - "validate rejects unknown agent references before launch", - "plan creates run layout, graph fingerprint, and rendered prompt before Smithers submission", - "compileSmithersWorkflow gates native dependencies on deterministic artifact verification", - "compileSmithersWorkflow maps cloud attempts to portable provider sandboxes", - "resume, replay, and fork delegate linked runs to Smithers lifecycle verbs", - "ordinary resume checks active-run ownership before detached preflight", - "a refresh resume reuses its own ownership inspection instead of inspecting twice", - "native continuation does not use historical trusted CLI identity as an authorization gate", - // Reads the release that is actually pinned. This is the lane's tripwire - // for a runner bump that widens an enum or the inspect envelope, which - // otherwise only shows up as a failed production run. - "pinned runner state and envelope contracts match Ultrafuzz's mirrors", - "the pinned runner drops resume pointers only from dead attempts" - ] - ], - [ - "test/generated-workflow-verifier.test.ts", - ["generated workflow input is an exact current-only envelope with bounded JSON operator data"] - ] -]); -const bunTestFile = "test/runtime.test.ts"; -const bunTestNamePrefix = "Bun adapter contract: "; -const bunTestNames = [ - "generated DeepSeek adapter uses the official endpoint and preserves independent usage components", - "generated DeepSeek adapter cleans an upstream command when environment policy rejects it", - "generated DeepSeek adapter corrects Smithers result and failed-attempt telemetry", - "generated DeepSeek adapter rejects ambiguous or noncanonical result telemetry", - // Needs bun:sqlite, so it can only run in this lane. Gates the claim that the - // pinned runner's schema migrations are additive over a stopped 0.34.0 store. - "pinned store migrations are additive over a 0.34.0 database" -]; -const selectedBunTestNames = bunTestNames.map((name) => `${bunTestNamePrefix}${name}`); -const smokeEnvironment = { ...process.env }; -delete smokeEnvironment.ULTRAFUZZ_RUNTIME_TEST_SHARD; -const selectedTestNames = [...namedTests.values()].flat(); - -for (const [sourcePath, names] of namedTests) { - const source = readFileSync(sourcePath, "utf8"); - for (const name of names) { - if (!source.includes(`test(${JSON.stringify(name)}`)) { - throw new Error(`PR runtime smoke test is not registered: ${name}`); - } - } -} -const bunTestSource = readFileSync(bunTestFile, "utf8"); -for (const name of bunTestNames) { - if (!bunTestSource.includes(JSON.stringify(name))) { - throw new Error(`PR Bun runtime smoke test is not registered: ${name}`); - } -} - -runNodeTests(supportingTestFiles); -runNodeTests( - [...namedTests.keys()].map((sourcePath) => `dist-test/${sourcePath.replace(/\.ts$/u, ".js")}`), - selectedTestNames, - selectedTestNames -); -runBunTests(`dist-test/${bunTestFile.replace(/\.ts$/u, ".js")}`, selectedBunTestNames); - -function runNodeTests(files, testNames, expectedTestNames) { - const args = ["--test", "--test-reporter=tap"]; - if (testNames !== undefined) { - const pattern = `^(?:${testNames.map(escapeRegExp).join("|")})$`; - args.push(`--test-name-pattern=${pattern}`); - } - args.push(...files); - - const result = spawnSync(process.execPath, args, { - encoding: "utf8", - env: smokeEnvironment, - stdio: ["inherit", "pipe", "pipe"] - }); - process.stdout.write(result.stdout ?? ""); - process.stderr.write(result.stderr ?? ""); - if (result.error !== undefined) throw result.error; - if (result.status !== 0) process.exit(result.status ?? 1); - if (expectedTestNames !== undefined) assertNamedTestsPassed(result.stdout ?? "", expectedTestNames); -} - -function runBunTests(file, testNames) { - const pattern = `^(?:${testNames.map(escapeRegExp).join("|")})$`; - const result = spawnSync("bun", ["test", file, "--test-name-pattern", pattern], { - encoding: "utf8", - env: smokeEnvironment, - stdio: ["inherit", "pipe", "pipe"] - }); - process.stdout.write(result.stdout ?? ""); - process.stderr.write(result.stderr ?? ""); - if (result.error !== undefined) throw result.error; - if (result.status !== 0) process.exit(result.status ?? 1); - assertBunTestsPassed(`${result.stdout ?? ""}\n${result.stderr ?? ""}`, testNames); -} - -function escapeRegExp(value) { - return value.replace(/[.*+?^${}()|[\]\\]/gu, "\\$&"); -} - -function assertNamedTestsPassed(output, expectedTestNames) { - const passedTestNames = [...output.matchAll(/^ok [0-9]+ - (.+)$/gmu)] - .map((match) => match[1]) - .filter((name) => !name.includes(" # SKIP") && !name.includes(" # TODO")) - .sort(); - const expected = [...expectedTestNames].sort(); - if (passedTestNames.length !== expected.length || passedTestNames.some((name, index) => name !== expected[index])) { - throw new Error(`PR runtime smoke passed unexpected tests: ${JSON.stringify(passedTestNames)}`); - } -} - -// Bun's runner prints a `(fail) ` line per failure but no per-test line -// for a pass -- verified against the pinned `bun-version: 1.3.14` in ci.yml, and -// against 1.4.0. So the post-condition is read off the run summary instead: the -// name pattern selects exactly `expectedTestNames`, so a rename, a skip or a -// filtered-out test shows up as a pass count below the expected one, which is -// the property this check exists to enforce. `stripAnsi` because bun colourises -// the summary whenever it believes it has a terminal. -function assertBunTestsPassed(output, expectedTestNames) { - const plain = stripAnsi(output); - const failures = [...plain.matchAll(/^\(fail\) (.+?)(?: \[[^\]]*\])?$/gmu)].map((match) => match[1]); - if (failures.length > 0) { - throw new Error(`PR Bun runtime smoke failed tests: ${JSON.stringify(failures)}`); - } - const summary = (label) => { - const matches = [...plain.matchAll(new RegExp(`^\\s*(\\d+) ${label}$`, "gmu"))].map((match) => Number(match[1])); - if (matches.length !== 1) { - throw new Error(`PR Bun runtime smoke could not read the "${label}" count from the runner summary`); - } - return matches[0]; - }; - const passed = summary("pass"); - const failed = summary("fail"); - if (failed !== 0) throw new Error(`PR Bun runtime smoke reported ${failed} failing test(s)`); - if (passed !== expectedTestNames.length) { - throw new Error( - `PR Bun runtime smoke ran ${passed} of ${expectedTestNames.length} expected tests; one was renamed, skipped, or filtered out` - ); - } -} - -function stripAnsi(value) { - // Built with `new RegExp` rather than a literal: the CSI introducer is a - // control character, which a regex literal cannot carry past `no-control-regex`. - return value.replace(new RegExp(`${String.fromCharCode(27)}\\[[0-9;]*m`, "gu"), ""); -} diff --git a/scripts/ci/release-validation-lanes.mjs b/scripts/ci/release-validation-lanes.mjs index de1f5ffb5..afdfb6592 100644 --- a/scripts/ci/release-validation-lanes.mjs +++ b/scripts/ci/release-validation-lanes.mjs @@ -1,20 +1,13 @@ /** - * Release validation lane policy. + * Release validation lane table. * - * `.github/workflows/ci.yml` used to hard-code its release-validation matrix and - * gate the whole job on `github.event_name != 'pull_request'`. Every runtime test - * therefore reported `skipping` on pull requests, so a change could break resume - * and still merge with every check green. The lane table now lives here, in one - * tested place, and each lane declares whether it is required on pull requests. - * - * Lanes marked `pull_request: true` must together execute the whole - * `@ultrafuzz/runtime` test suite: `runtime-1`..`runtime-4` cover every test in - * `packages/runtime/test/runtime.test.ts` (the shard splitter assigns each test - * name to exactly one of the four shards) and `runtime-supporting` covers every - * other runtime test file plus the Bun adapter contracts. + * `.github/workflows/ci.yml` expands this table into its release-validation + * matrix on every event, pull requests included. While lanes were gated off + * pull requests, every skipped suite reported `skipping` and its regressions + * merged with all checks green, then failed on main. */ -/** @typedef {{ lane: string, description: string, gates: string, timeout_minutes: number, pull_request: boolean, build_modal_dependencies?: boolean, build_release_reporter?: boolean, build_cli?: boolean }} ReleaseValidationLane */ +/** @typedef {{ lane: string, description: string, gates: string, timeout_minutes: number, build_modal_dependencies?: boolean, build_release_reporter?: boolean, build_cli?: boolean }} ReleaseValidationLane */ /** @type {readonly ReleaseValidationLane[]} */ export const RELEASE_VALIDATION_LANES = Object.freeze([ @@ -24,7 +17,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ gates: "dependency-advisories,ci-scripts,docs,config,audit-profile-package,packed-install,security,references,topology,prompts,artifacts,dashboard,evals,evmbench,modal", timeout_minutes: 45, - pull_request: false, build_modal_dependencies: true }, { @@ -36,7 +28,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ // can push that complete pass beyond 75 minutes, so retain a bounded budget // that covers the full fail-closed validation instead of canceling it late. timeout_minutes: 120, - pull_request: true, build_modal_dependencies: true }, { @@ -46,7 +37,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ // A slow hosted runner passed 59 of shard 4's 67 tests before the old // 75-minute cutoff. Give every shard the same bounded completion budget. timeout_minutes: 120, - pull_request: true, build_modal_dependencies: true }, { @@ -54,7 +44,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ description: "Node.js 24 runtime integration tests, shard 2/4", gates: "runtime-2", timeout_minutes: 120, - pull_request: true, build_modal_dependencies: true }, { @@ -62,7 +51,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ description: "Node.js 24 runtime integration tests, shard 3/4", gates: "runtime-3", timeout_minutes: 120, - pull_request: true, build_modal_dependencies: true }, { @@ -70,7 +58,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ description: "Node.js 24 runtime integration tests, shard 4/4", gates: "runtime-4", timeout_minutes: 120, - pull_request: true, build_modal_dependencies: true }, { @@ -79,7 +66,6 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ gates: "cli", // The complete local CLI suite took 76 minutes before job setup overhead. timeout_minutes: 120, - pull_request: false, build_release_reporter: true }, { @@ -87,50 +73,11 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ description: "Benchmark history charts and workspace typecheck", gates: "benchmark-history,workspace-typecheck", timeout_minutes: 45, - pull_request: false, build_release_reporter: true, build_cli: true } ]); -/** - * Release validation gates that must run before a pull request can merge. - * Together these execute every `@ultrafuzz/runtime` test, which is where run - * resume, controller refresh, replay, and fork are covered. - * - * @type {readonly string[]} - */ -export const PULL_REQUEST_REQUIRED_GATES = Object.freeze([ - "runtime-supporting", - "runtime-1", - "runtime-2", - "runtime-3", - "runtime-4" -]); - -/** - * @param {string} eventName GitHub `github.event_name`. - * @returns {ReleaseValidationLane[]} lanes to expand into the workflow matrix. - */ -export function selectReleaseValidationLanes(eventName) { - if (typeof eventName !== "string" || eventName.length === 0) { - throw new Error("release validation lane selection requires a GitHub event name"); - } - const lanes = - eventName === "pull_request" - ? RELEASE_VALIDATION_LANES.filter((lane) => lane.pull_request) - : [...RELEASE_VALIDATION_LANES]; - const selectedGates = new Set(lanes.flatMap((lane) => lane.gates.split(","))); - const missing = PULL_REQUEST_REQUIRED_GATES.filter((gate) => !selectedGates.has(gate)); - if (missing.length > 0) { - throw new Error(`release validation lanes omit required gates: ${missing.join(", ")}`); - } - return lanes.map((lane) => Object.fromEntries(Object.entries(lane).filter(([key]) => key !== "pull_request"))); -} - if (import.meta.url === `file://${process.argv[1]}`) { - const index = process.argv.indexOf("--event"); - const eventName = index === -1 ? undefined : process.argv[index + 1]; - if (eventName === undefined) throw new Error("release-validation-lanes.mjs requires --event "); - process.stdout.write(`${JSON.stringify(selectReleaseValidationLanes(eventName))}\n`); + process.stdout.write(`${JSON.stringify(RELEASE_VALIDATION_LANES)}\n`); } diff --git a/scripts/ci/release-validation-lanes.test.ts b/scripts/ci/release-validation-lanes.test.ts index 0ebb6049b..1c157d0b4 100644 --- a/scripts/ci/release-validation-lanes.test.ts +++ b/scripts/ci/release-validation-lanes.test.ts @@ -4,15 +4,18 @@ import fs from "node:fs"; import path from "node:path"; import { fileURLToPath } from "node:url"; -import { - PULL_REQUEST_REQUIRED_GATES, - RELEASE_VALIDATION_LANES, - selectReleaseValidationLanes -} from "./release-validation-lanes.mjs"; +import { parse } from "yaml"; + +import { RELEASE_VALIDATION_LANES } from "./release-validation-lanes.mjs"; const repoRoot = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", ".."); -const workflow = fs.readFileSync(path.join(repoRoot, ".github", "workflows", "ci.yml"), "utf8"); -const runtimeTestDir = path.join(repoRoot, "packages", "runtime", "test"); +const laneScript = path.join(repoRoot, "scripts", "ci", "release-validation-lanes.mjs"); + +interface WorkflowJob { + if?: string; + strategy?: { matrix?: { include?: string } }; + steps: Array<{ name?: string; if?: string }>; +} function declaredGateIds(): string[] { const source = fs.readFileSync(path.join(repoRoot, "scripts", "validate-release.mjs"), "utf8"); @@ -29,103 +32,43 @@ function laneGates(lanes: ReadonlyArray<{ gates: string }>): string[] { return lanes.flatMap((lane) => lane.gates.split(",")); } -describe("release validation lane policy", () => { - it("names only gates that scripts/validate-release.mjs defines", () => { - const declared = new Set(declaredGateIds()); - expect(declared.size).toBeGreaterThan(0); - for (const gate of laneGates(RELEASE_VALIDATION_LANES)) expect(declared.has(gate)).toBe(true); - }); +function workflowJobs(): Record { + const workflow = parse(fs.readFileSync(path.join(repoRoot, ".github", "workflows", "ci.yml"), "utf8")) as { + jobs: Record; + }; + return workflow.jobs; +} - it("runs every declared gate exactly once on push", () => { - const gates = laneGates(selectReleaseValidationLanes("push")); +describe("release validation lanes", () => { + it("run every gate scripts/validate-release.mjs defines exactly once", () => { + const gates = laneGates(RELEASE_VALIDATION_LANES); + expect(gates.length).toBeGreaterThan(0); expect([...gates].sort()).toEqual([...declaredGateIds()].sort()); expect(new Set(gates).size).toBe(gates.length); }); - it("runs the whole runtime suite on pull requests", () => { - const gates = laneGates(selectReleaseValidationLanes("pull_request")); - expect([...gates].sort()).toEqual([...PULL_REQUEST_REQUIRED_GATES].sort()); - }); - - it("gives complete runtime and CLI suites a bounded budget beyond observed 75-minute runs", () => { - for (const name of [...PULL_REQUEST_REQUIRED_GATES, "cli"]) { - const lane = RELEASE_VALIDATION_LANES.find((candidate) => candidate.lane === name); - expect(lane?.timeout_minutes).toBe(120); - } - }); - - it("covers every runtime test file on pull requests", () => { - // runtime-1..runtime-4 shard packages/runtime/test/runtime.test.ts by test - // name; runtime-supporting runs every other runtime test file. Nothing in - // the runtime package may fall outside that set, because run resume, - // controller refresh, replay, and fork are only covered there. - const gates = new Set(laneGates(selectReleaseValidationLanes("pull_request"))); - const shardScript = fs.readFileSync( - path.join(repoRoot, "packages", "runtime", "scripts", "run-runtime-test-shard.mjs"), - "utf8" - ); - expect(shardScript).toContain("dist-test/test/runtime.test.js"); - const shardTotal = [...gates].filter((gate) => /^runtime-[0-9]+$/u.test(gate)).length; - for (let index = 1; index <= shardTotal; index += 1) expect(gates.has(`runtime-${index}`)).toBe(true); - const describedShards = RELEASE_VALIDATION_LANES.filter((lane) => /^runtime-[0-9]+$/u.test(lane.lane)).map((lane) => - /shard ([0-9]+)\/([0-9]+)$/u.exec(lane.description) - ); - expect(describedShards.length).toBe(shardTotal); - for (const match of describedShards) expect(Number(match?.[2])).toBe(shardTotal); - expect(gates.has("runtime-supporting")).toBe(true); - const runtimeTestFiles = fs.readdirSync(runtimeTestDir).filter((entry) => entry.endsWith(".test.ts")); - expect(runtimeTestFiles).toContain("runtime.test.ts"); - expect(runtimeTestFiles).toContain("dynamic-lifecycle.test.ts"); - }); - - it("does not let the workflow gate release validation off pull requests", () => { - // Regression guard for the outage this policy exists to prevent: while the - // release-validation job carried `if: github.event_name != 'pull_request'`, + it("are not gated off pull requests by the workflow", () => { + // While release validation carried `if: github.event_name != 'pull_request'`, // every runtime test reported `skipping` on pull requests and resume-path // regressions merged with all checks green. - const releaseValidation = workflow.slice(workflow.indexOf("\n release-validation:")); - const job = releaseValidation.slice(0, releaseValidation.indexOf("\n release-gates:")); - expect(job).not.toContain("if: github.event_name != 'pull_request'"); - expect(job).toContain("fromJSON(needs.release-validation-lanes.outputs.lanes)"); - expect(job).toContain("timeout-minutes: ${{ matrix.timeout_minutes }}"); - }); - - it("requires the release validation lanes on pull requests", () => { - const releaseGates = workflow.slice(workflow.indexOf("\n release-gates:")); - const requirement = releaseGates.slice(releaseGates.indexOf("Require release validation lanes")); - const step = requirement.slice(0, requirement.indexOf("\n - name:")); - expect(step).toContain("needs.release-validation.result != 'success'"); - expect(step).not.toContain("github.event_name != 'pull_request'"); + const jobs = workflowJobs(); + const releaseValidation = jobs["release-validation"]; + expect(releaseValidation?.if).toBeUndefined(); + expect(releaseValidation?.strategy?.matrix?.include).toBe( + "${{ fromJSON(needs.release-validation-lanes.outputs.lanes) }}" + ); + const requirement = jobs["release-gates"]?.steps.find((step) => step.name === "Require release validation lanes"); + expect(requirement?.if).toBe("always() && needs.release-validation.result != 'success'"); }); - it("emits the workflow matrix for an event", () => { - const result = spawnSync( - process.execPath, - [path.join(repoRoot, "scripts", "ci", "release-validation-lanes.mjs"), "--event", "pull_request"], - { cwd: repoRoot, encoding: "utf8" } - ); + it("are printed as the workflow matrix", () => { + const result = spawnSync(process.execPath, [laneScript], { cwd: repoRoot, encoding: "utf8" }); expect(result.status).toBe(0); const lanes = JSON.parse(result.stdout) as Array>; - expect(lanes.map((lane) => lane.lane)).toEqual([ - "runtime-supporting", - "runtime-1", - "runtime-2", - "runtime-3", - "runtime-4" - ]); + expect(lanes).toEqual(JSON.parse(JSON.stringify(RELEASE_VALIDATION_LANES))); for (const lane of lanes) { - expect(lane).not.toHaveProperty("pull_request"); expect(typeof lane.description).toBe("string"); expect(typeof lane.timeout_minutes).toBe("number"); } }); - - it("rejects a missing event name", () => { - expect(() => selectReleaseValidationLanes("")).toThrow(); - const result = spawnSync(process.execPath, [path.join(repoRoot, "scripts", "ci", "release-validation-lanes.mjs")], { - cwd: repoRoot, - encoding: "utf8" - }); - expect(result.status).not.toBe(0); - }); }); From 4f5e97732fd459256f72b99199697541c7eac9a0 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:13:24 +0000 Subject: [PATCH 072/206] ci: run the release validation lanes under eatmydata Every runtime and CLI test that launches a run copies a sealed execution snapshot and fsyncs each file, twice (at write and again when the snapshot is sealed). The deep analysis measured 55,998 fsyncs taking 222 s of a 163-261 s test locally; under eatmydata the same test took 23-33 s, and a CLI test went from 169 s to 35 s. A CI runner is discarded after the job, so durability across a crash buys nothing there. Install eatmydata in the release-validation job and run validate:release under it; its LD_PRELOAD reaches every child the gate commands spawn. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/ci.yml | 8 +++++++- docs/reference/development.md | 12 ++++++------ 2 files changed, 13 insertions(+), 7 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 0174be83f..6a8cc9e16 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -161,6 +161,12 @@ jobs: sudo rm -rf /usr/local/lib/android sudo rm -rf /usr/share/dotnet + # Every test that launches a run copies a sealed execution snapshot of + # thousands of files and fsyncs each one. eatmydata makes fsync a no-op + # for the lane; an ephemeral runner has nothing to protect across a crash. + - name: Install eatmydata + run: sudo apt-get install -y eatmydata + - name: Check out repository uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0 with: @@ -203,7 +209,7 @@ jobs: - name: Validate release lane run: >- - pnpm -w validate:release -- + eatmydata pnpm -w validate:release -- --gates "${{ matrix.gates }}" --report ".ultrafuzz/release-validation/${{ matrix.lane }}.json" diff --git a/docs/reference/development.md b/docs/reference/development.md index 6aabe8a24..906fc18ab 100644 --- a/docs/reference/development.md +++ b/docs/reference/development.md @@ -45,12 +45,12 @@ formatting, lint, dead-code checks, the workspace build, strict lint of changed lines, bundle budgets, and dependency policy) and all eight release validation lanes: package gates, one runtime supporting lane that includes the Bun adapter contracts, four runtime integration shards, the CLI suite, and benchmark history -with the workspace typecheck. The lanes do not wait for the build gates and run -at most eight at a time. Feature branches are validated only by the -pull-request event, avoiding a duplicate push run. A newer push to a pull -request cancels that pull request's older run; a push to `main` never cancels -another run. On `main`, the lane results are merged into one JSON report in -stable gate order. +with the workspace typecheck. The lanes do not wait for the build gates, run +their tests under `eatmydata` (which turns `fsync` into a no-op), and run at +most eight at a time. Feature branches are validated only by the pull-request +event, avoiding a duplicate push run. A newer push to a pull request cancels +that pull request's older run; a push to `main` never cancels another run. On +`main`, the lane results are merged into one JSON report in stable gate order. ## Package Checks From ca8b55fcf87e681f1adc59e5a5e5d40bde51954f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:13:41 +0000 Subject: [PATCH 073/206] ci: report unused exports without blocking the build `pnpm -w knip` leaves out knip's exports, types and duplicates categories, so unused exports accumulate unseen; on this tree knip reports 52 unused exports and 38 unused exported types. Report them in a continue-on-error step that reuses the pinned knip script. Several open pull requests delete dead code; once they land, these categories can move into the blocking knip step. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/ci.yml | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 6a8cc9e16..92d4f955c 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -67,6 +67,12 @@ jobs: - name: Find dead code and dependencies run: pnpm -w knip + # Informational until the pending dead-code removals land; then these + # categories move into the blocking `pnpm -w knip` step above. + - name: Report unused exports + continue-on-error: true + run: pnpm -w knip --include exports,types,duplicates + - name: Build run: pnpm -w build From 5c8b53e59d7e189f98747d2f1fbd9acf0a607c7a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:13:41 +0000 Subject: [PATCH 074/206] test: give vitest packages a 30 s default test timeout config, topology, prompts, evals, evmbench and modal run a bare `vitest run` with no config file, so every test gets vitest's 5 s default. Main went red on it in runs 34967618488 (modal ci-config and modal-documents: "Test timed out in 5000ms") and 34634052856 (modal worker-result), and #1126 raised individual cases instead. Pass --testTimeout=30000 in the six package test scripts. Cases that set their own timeout keep it. Co-Authored-By: Claude Opus 5.5 --- packages/config/package.json | 2 +- packages/evals/package.json | 2 +- packages/evmbench/package.json | 2 +- packages/modal/package.json | 2 +- packages/prompts/package.json | 2 +- packages/topology/package.json | 2 +- 6 files changed, 6 insertions(+), 6 deletions(-) diff --git a/packages/config/package.json b/packages/config/package.json index 4297e5bdf..ab5b6866d 100644 --- a/packages/config/package.json +++ b/packages/config/package.json @@ -18,7 +18,7 @@ ], "scripts": { "build": "tsc -p tsconfig.build.json && cp defaults.toml dist/ultrafuzz.toml && cp audit-profiles.yml dist/audit-profiles.yml && rm -rf dist/topologies && cp -R topologies dist/topologies && rm -rf dist/schema && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", - "test": "pnpm --filter @ultrafuzz/config^... build && vitest run", + "test": "pnpm --filter @ultrafuzz/config^... build && vitest run --testTimeout=30000", "typecheck": "pnpm --filter @ultrafuzz/config^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/evals/package.json b/packages/evals/package.json index 00074375f..c5f82eac2 100644 --- a/packages/evals/package.json +++ b/packages/evals/package.json @@ -19,7 +19,7 @@ "scripts": { "build": "rm -rf dist && tsc -p tsconfig.build.json && node scripts/copy-prompts.mjs && rm -rf dist/schema && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", "schema:check": "pnpm build", - "test": "pnpm --filter @ultrafuzz/evals... build && vitest run", + "test": "pnpm --filter @ultrafuzz/evals... build && vitest run --testTimeout=30000", "typecheck": "pnpm --filter @ultrafuzz/evals^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/evmbench/package.json b/packages/evmbench/package.json index 967cc11e3..9e80ee483 100644 --- a/packages/evmbench/package.json +++ b/packages/evmbench/package.json @@ -19,7 +19,7 @@ "benchmark": "pnpm run build && node dist/cli.js", "build": "tsc -p tsconfig.build.json && rm -rf dist/schema && cp -R src/schema dist/schema", "lock": "pnpm run build && node dist/update-lock.js", - "test": "vitest run", + "test": "vitest run --testTimeout=30000", "typecheck": "tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/modal/package.json b/packages/modal/package.json index 350d48ae2..735c26160 100644 --- a/packages/modal/package.json +++ b/packages/modal/package.json @@ -22,7 +22,7 @@ "scripts": { "build": "tsc -p tsconfig.build.json && rm -rf dist/schema && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", "smoke": "pnpm run build && node dist/cli.js smoke", - "test": "vitest run", + "test": "vitest run --testTimeout=30000", "typecheck": "pnpm --filter @ultrafuzz/modal^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/prompts/package.json b/packages/prompts/package.json index ee13cd892..0b3b04c9b 100644 --- a/packages/prompts/package.json +++ b/packages/prompts/package.json @@ -17,7 +17,7 @@ ], "scripts": { "build": "tsc -p tsconfig.json && rm -rf dist/assets/prompts && mkdir -p dist/assets && cp -R ../../.ultrafuzz/prompts dist/assets/prompts", - "test": "pnpm --filter @ultrafuzz/prompts^... build && vitest run", + "test": "pnpm --filter @ultrafuzz/prompts^... build && vitest run --testTimeout=30000", "typecheck": "pnpm --filter @ultrafuzz/prompts^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { diff --git a/packages/topology/package.json b/packages/topology/package.json index f07b01716..00c553b2d 100644 --- a/packages/topology/package.json +++ b/packages/topology/package.json @@ -18,7 +18,7 @@ ], "scripts": { "build": "tsc -p tsconfig.json", - "test": "pnpm --filter @ultrafuzz/topology^... build && vitest run", + "test": "pnpm --filter @ultrafuzz/topology^... build && vitest run --testTimeout=30000", "typecheck": "pnpm --filter @ultrafuzz/topology^... build && tsc -p tsconfig.json --noEmit --pretty false" }, "dependencies": { From 953f2226a9fa1d8bda72b0f6e475766effcca25a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:15:17 +0000 Subject: [PATCH 075/206] test(cli): remove each test's temporary project when the test ends cli.test.ts, lifecycle-commands.test.ts and audit-profile-commands.test.ts created their project with mkdtemp and never removed it. On main, one run- launching test ("status surfaces a terminal product and live workflow lifecycle divergence") left 446 MB and 34,091 entries in TMPDIR; the sealed snapshot directories are mode dr-x, so a plain `rm -rf` cannot delete them. A CI lane accumulated every test's run directories until the job ended. temporaryRoot(prefix, t) creates the canonical directory and registers a t.after hook that restores owner write permission on the way down and removes it, together with the `-fake-bin` directory the fake engine fixtures create beside it. The tests pass their context to tempProject. Co-Authored-By: Claude Opus 5.5 --- .../cli/test/audit-profile-commands.test.ts | 20 +- packages/cli/test/cli.test.ts | 207 +++++++++--------- packages/cli/test/lifecycle-commands.test.ts | 61 +++--- packages/cli/test/temporary-root.ts | 31 +++ 4 files changed, 176 insertions(+), 143 deletions(-) create mode 100644 packages/cli/test/temporary-root.ts diff --git a/packages/cli/test/audit-profile-commands.test.ts b/packages/cli/test/audit-profile-commands.test.ts index ab3979741..b1022d489 100644 --- a/packages/cli/test/audit-profile-commands.test.ts +++ b/packages/cli/test/audit-profile-commands.test.ts @@ -1,12 +1,12 @@ import assert from "node:assert/strict"; import fs from "node:fs"; -import os from "node:os"; import path from "node:path"; -import test from "node:test"; +import test, { type TestContext } from "node:test"; import { packagedTopology } from "@ultrafuzz/config"; import { runCli } from "../src/index.js"; +import { temporaryRoot } from "./temporary-root.js"; interface Capture { stdout: string; @@ -14,8 +14,8 @@ interface Capture { code: number; } -function tempProject(): string { - return fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-cli-profile-")); +function tempProject(t: TestContext): string { + return temporaryRoot("ufz-cli-profile-", t); } async function cli(project: string, argv: string[]): Promise { @@ -44,8 +44,8 @@ function data(capture: Capture): Record { return (JSON.parse(capture.stdout) as { data: Record }).data; } -test("profile list and detail expose the catalog and effective project policy", async () => { - const project = tempProject(); +test("profile list and detail expose the catalog and effective project policy", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); assert.deepEqual( fs.readFileSync(path.join(project, ".ultrafuzz", "topology.yml")), @@ -84,8 +84,8 @@ test("profile list and detail expose the catalog and effective project policy", assert.equal(detailData.effective_settings.strategy_loops, 1); }); -test("topology list, show, and copy use the packaged assets safely", async () => { - const project = tempProject(); +test("topology list, show, and copy use the packaged assets safely", async (t) => { + const project = tempProject(t); const listed = await cli(project, ["topology", "list", "--json"]); assert.equal(listed.code, 0, listed.stderr); const listData = data(listed) as { @@ -128,8 +128,8 @@ test("topology list, show, and copy use the packaged assets safely", async () => assert.equal(fs.existsSync(path.join(project, "..", "outside.yml")), false); }); -test("validate accepts CLI profile and topology overrides with the documented precedence", async () => { - const project = tempProject(); +test("validate accepts CLI profile and topology overrides with the documented precedence", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); fs.writeFileSync(path.join(project, ".ultrafuzz", "topology.yml"), "not: [valid\n", "utf8"); diff --git a/packages/cli/test/cli.test.ts b/packages/cli/test/cli.test.ts index 18676ca3a..8bcff2ec9 100644 --- a/packages/cli/test/cli.test.ts +++ b/packages/cli/test/cli.test.ts @@ -4,7 +4,7 @@ import crypto from "node:crypto"; import fs from "node:fs"; import os from "node:os"; import path from "node:path"; -import test from "node:test"; +import test, { type TestContext } from "node:test"; import { ARTIFACT_VERIFICATION_SCHEMA_VERSION, @@ -45,6 +45,7 @@ import AdmZip from "adm-zip"; import { validateReportBundleManifest } from "../src/cli-schema-registry.js"; import { runCli } from "../src/index.js"; import { formatStatusDuration } from "../src/status-rendering.js"; +import { temporaryRoot } from "./temporary-root.js"; interface Capture { stdout: string; @@ -52,8 +53,8 @@ interface Capture { code: number; } -function tempProject(): string { - return fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-cli-")); +function tempProject(t: TestContext): string { + return temporaryRoot("ufz-cli-", t); } function shellQuote(value: string): string { @@ -759,8 +760,8 @@ function writeCanonicalReportPair(reportDir: string, report: Record { - const project = tempProject(); +test("report render gives producers the exact canonical Markdown without host repair", async (t) => { + const project = tempProject(t); const reportPath = path.join(project, "report.json"); const markdownPath = path.join(project, "report.md"); const report = currentReport("producer-render", [currentReportIssue()]); @@ -778,8 +779,8 @@ test("report render gives producers the exact canonical Markdown without host re assert.deepEqual(fs.readFileSync(markdownPath), before); }); -test("report render renders goal coverage from the run-root census only through its flag", async () => { - const project = tempProject(); +test("report render renders goal coverage from the run-root census only through its flag", async (t) => { + const project = tempProject(t); const runId = "census-render"; const runRoot = path.join(project, "runs", runId); fs.mkdirSync(runRoot, { recursive: true }); @@ -1239,8 +1240,8 @@ function writeJsonRecord(filePath: string, value: Record): void fs.writeFileSync(filePath, `${JSON.stringify(value, null, 2)}\n`, "utf8"); } -test("init and validate emit schema-versioned launch JSON", async () => { - const project = tempProject(); +test("init and validate emit schema-versioned launch JSON", async (t) => { + const project = tempProject(t); fs.writeFileSync(path.join(project, "ultrafuzz.toml"), "# owned\n", "utf8"); const init = await cli(project, ["init", "--json"]); @@ -1280,8 +1281,8 @@ test("init and validate emit schema-versioned launch JSON", async () => { assert.match(tampered.stdout + tampered.stderr, /absent\.database/u); }); -test("plain init surfaces a customized stale agent adapter diagnostic", async () => { - const project = tempProject(); +test("plain init surfaces a customized stale agent adapter diagnostic", async (t) => { + const project = tempProject(t); const initial = await cli(project, ["init", "--force"]); assert.equal(initial.code, 0, initial.stderr); @@ -1297,8 +1298,8 @@ test("plain init surfaces a customized stale agent adapter diagnostic", async () assert.equal(fs.readFileSync(adapterPath, "utf8"), customAdapter); }); -test("run exposes the trusted reference expectation catalog option", async () => { - const project = tempProject(); +test("run exposes the trusted reference expectation catalog option", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeSmallTopology(project); const run = await cli( @@ -1316,8 +1317,8 @@ test("run exposes the trusted reference expectation catalog option", async () => ); }); -test("run input flags are explicit, strict, and never reinterpret malformed inline JSON as a path", async () => { - const project = tempProject(); +test("run input flags are explicit, strict, and never reinterpret malformed inline JSON as a path", async (t) => { + const project = tempProject(t); fs.writeFileSync(path.join(project, "looks-like-a-path.json"), '{"loaded":true}\n', "utf8"); const malformedInline = await cli(project, ["run", "--input-json", "looks-like-a-path.json", "--json"]); @@ -1352,8 +1353,8 @@ test("run input flags are explicit, strict, and never reinterpret malformed inli assert.match(JSON.stringify(parseJson(linkedInput).diagnostics), /cannot open regular file/iu); }); -test("run reads a bounded immutable workflow-input file", async () => { - const project = tempProject(); +test("run reads a bounded immutable workflow-input file", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeSmallTopology(project); fs.writeFileSync(path.join(project, "operator-input.json"), '{"ticket":3}\n', "utf8"); @@ -1376,8 +1377,8 @@ test("run reads a bounded immutable workflow-input file", async () => { assert.equal(savedConfig.retry.sameAgentAttempts, 3); }); -test("run, ps, status, inspect, report, materialize, clean, and lifecycle commands expose product workflow evidence", async () => { - const project = tempProject(); +test("run, ps, status, inspect, report, materialize, clean, and lifecycle commands expose product workflow evidence", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeReportTopology(project); @@ -1779,8 +1780,8 @@ test("run, ps, status, inspect, report, materialize, clean, and lifecycle comman assert.equal(pauseData.submitted, true); }); -test("status --watch --json keeps a failing poll on one NDJSON line", async () => { - const project = tempProject(); +test("status --watch --json keeps a failing poll on one NDJSON line", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); const init = await cli(project, ["init", "--json"], env); assert.equal(init.code, 0, init.stderr); @@ -1803,8 +1804,8 @@ test("status --watch --json keeps a failing poll on one NDJSON line", async () = assert.equal((body.diagnostics as Array<{ code: string }>)[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); }); -test("ps text prefers linked workflow terminal status over a stale local running projection", async () => { - const project = tempProject(); +test("ps text prefers linked workflow terminal status over a stale local running projection", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); assert.equal((await cli(project, ["init", "--json"], env)).code, 0); writeSmallTopology(project); @@ -1846,8 +1847,8 @@ test("ps text prefers linked workflow terminal status over a stale local running assert.equal(jsonData.runs[0]?.workflow_status, "failed"); }); -test("status observes an incomplete launch with a successful CLI envelope and unknown liveness", async () => { - const project = tempProject(); +test("status observes an incomplete launch with a successful CLI envelope and unknown liveness", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); assert.equal((await cli(project, ["init", "--json"], env)).code, 0); writeSmallTopology(project); @@ -1877,8 +1878,8 @@ test("status observes an incomplete launch with a successful CLI envelope and un assert.doesNotMatch(text.stdout, /Progress: 0%|ETA: 0|running-healthy/u); }); -test("status surfaces a terminal product and live workflow lifecycle divergence", async () => { - const project = tempProject(); +test("status surfaces a terminal product and live workflow lifecycle divergence", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); assert.equal((await cli(project, ["init", "--json"], env)).code, 0); writeSmallTopology(project); @@ -1918,8 +1919,8 @@ test("status surfaces a terminal product and live workflow lifecycle divergence" ); }); -test("status --watch stops immediately on a degraded verdict even while product state is nonterminal", async () => { - const project = tempProject(); +test("status --watch stops immediately on a degraded verdict even while product state is nonterminal", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); assert.equal((await cli(project, ["init", "--json"], env)).code, 0); writeSmallTopology(project); @@ -1940,8 +1941,8 @@ test("status --watch stops immediately on a degraded verdict even while product assert.equal(body.data?.report?.status, "unknown"); }); -test("status surfaces quota parking with preserved attempts and the resume remediation", async () => { - const project = tempProject(); +test("status surfaces quota parking with preserved attempts and the resume remediation", async (t) => { + const project = tempProject(t); const env = fakeSmithersEnv(project); assert.equal((await cli(project, ["init", "--json"], env)).code, 0); writeSmallTopology(project); @@ -1997,8 +1998,8 @@ test("status surfaces quota parking with preserved attempts and the resume remed assert.doesNotMatch(healthy.stdout, /^Quota:/mu); }); -test("old commands and backend flags are rejected instead of aliased or shimmed", async () => { - const project = tempProject(); +test("old commands and backend flags are rejected instead of aliased or shimmed", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); for (const argv of [ @@ -2018,8 +2019,8 @@ test("old commands and backend flags are rejected instead of aliased or shimmed" } }); -test("references status is restored and reports offline cache state", async () => { - const project = tempProject(); +test("references status is restored and reports offline cache state", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); const previousXdgCacheHome = process.env.XDG_CACHE_HOME; process.env.XDG_CACHE_HOME = path.join(project, "empty-cache"); @@ -2054,8 +2055,8 @@ test("references status is restored and reports offline cache state", async () = } }); -test("runtime command failures emit a failing exit code with JSON", async () => { - const project = tempProject(); +test("runtime command failures emit a failing exit code with JSON", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeSmallTopology(project); @@ -2068,8 +2069,8 @@ test("runtime command failures emit a failing exit code with JSON", async () => assert.match(JSON.stringify(body.diagnostics), /CONFIG_MODEL_AGENT_INVALID/); }); -test("run rejects an OpenRouter override whose effective model is invalid", async () => { - const project = tempProject(); +test("run rejects an OpenRouter override whose effective model is invalid", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeSmallTopology(project); @@ -2093,8 +2094,8 @@ test("run rejects an OpenRouter override whose effective model is invalid", asyn assert.equal(fs.existsSync(path.join(project, ".ultrafuzz", "runs", "invalid-openrouter-model")), false); }); -test("report accepts populated accounting snapshots and preserves partial-pricing marker", async () => { - const project = tempProject(); +test("report accepts populated accounting snapshots and preserves partial-pricing marker", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeReportTopology(project); @@ -2188,8 +2189,8 @@ test("report accepts populated accounting snapshots and preserves partial-pricin assert.equal(fs.readFileSync(metadataPath, "utf8"), "{"); }); -test("report validates current artifacts without rewriting agent-owned bytes", async () => { - const project = tempProject(); +test("report validates current artifacts without rewriting agent-owned bytes", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-current-artifacts"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2221,8 +2222,8 @@ test("report validates current artifacts without rewriting agent-owned bytes", a assertFinalReportUnchanged(reportDir, reportSnapshot); }); -test("report displays authenticated final-verifier warnings while preserving the report files", async () => { - const project = tempProject(); +test("report displays authenticated final-verifier warnings while preserving the report files", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-host-warnings"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); fs.mkdirSync(reportDir, { recursive: true }); @@ -2249,8 +2250,8 @@ test("report displays authenticated final-verifier warnings while preserving the assertFinalReportUnchanged(reportDir, before); }); -test("report exposes terminal partial coverage and bundles only the authenticated runtime publication", async () => { - const project = tempProject(); +test("report exposes terminal partial coverage and bundles only the authenticated runtime publication", async (t) => { + const project = tempProject(t); const { run, report, stateBytes } = await createTerminalPartialReport(project, "report-runtime-partial"); const staleDirectory = path.join(run.run_root, "review", "runtime-report", "unverified-generation"); fs.mkdirSync(staleDirectory, { recursive: true }); @@ -2288,8 +2289,8 @@ test("report exposes terminal partial coverage and bundles only the authenticate assert.deepEqual(fs.readFileSync(path.join(run.run_root, "state.json")), stateBytes); }); -test("report verification is optional and unchecked bundles contain only the labeled report pair", async () => { - const project = tempProject(); +test("report verification is optional and unchecked bundles contain only the labeled report pair", async (t) => { + const project = tempProject(t); const { run, report, stateBytes } = await createTerminalPartialReport(project, "report-optional-verification"); const receipt = report.publications?.find((publication) => path.basename(publication.path) === "terminal.json"); assert.ok(receipt); @@ -2346,8 +2347,8 @@ test("report verification is optional and unchecked bundles contain only the lab assert.match(JSON.stringify(parseJson(withoutAccounting).diagnostics), /REPORT_ACCOUNTING_UNAVAILABLE/u); }); -test("report bundle --require-verified rejects receipt bytes changed only during archive capture", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects receipt bytes changed only during archive capture", async (t) => { + const project = tempProject(t); const { run, report } = await createTerminalPartialReport(project, "report-runtime-receipt-change"); const receipt = report.publications?.find((publication) => path.basename(publication.path) === "terminal.json"); assert.ok(receipt); @@ -2362,8 +2363,8 @@ test("report bundle --require-verified rejects receipt bytes changed only during assert.equal(fs.existsSync(path.join(project, ".ultrafuzz", "bundles", `${run.run_id}-report-bundle.zip`)), false); }); -test("eval report validates the registered summary and never synthesizes missing Markdown", async () => { - const project = tempProject(); +test("eval report validates the registered summary and never synthesizes missing Markdown", async (t) => { + const project = tempProject(t); const evalRunId = "eval-report-strict"; const runRoot = path.join(project, ".ultrafuzz", "evals", "runs", evalRunId); const summaryPath = path.join(runRoot, "summary.json"); @@ -2451,8 +2452,8 @@ test("eval report validates the registered summary and never synthesizes missing assert.equal((parseJson(valid).data as { eval_run_id: string }).eval_run_id, evalRunId); }); -test("report --require-verified rejects final_severity compatibility aliases without rewriting artifacts", async () => { - const project = tempProject(); +test("report --require-verified rejects final_severity compatibility aliases without rewriting artifacts", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-rejects-severity-alias"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2474,8 +2475,8 @@ test("report --require-verified rejects final_severity compatibility aliases wit assertFinalReportUnchanged(reportDir, reportSnapshot); }); -test("report --require-verified does not synthesize missing Markdown", async () => { - const project = tempProject(); +test("report --require-verified does not synthesize missing Markdown", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-missing-markdown"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2497,8 +2498,8 @@ test("report --require-verified does not synthesize missing Markdown", async () assert.equal(digest(fs.readFileSync(reportPath)), jsonSha256Before); }); -test("report accepts canonical severity and complete proof without rewriting either artifact", async () => { - const project = tempProject(); +test("report accepts canonical severity and complete proof without rewriting either artifact", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-canonical-severity", writeBoundedDedupeReportTopology); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2521,8 +2522,8 @@ test("report accepts canonical severity and complete proof without rewriting eit assert.equal(Object.hasOwn(report.issues[0] ?? {}, "final_severity"), false); }); -test("report --require-verified rejects malformed issues without rewriting stale bytes", async () => { - const project = tempProject(); +test("report --require-verified rejects malformed issues without rewriting stale bytes", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-current-malformed"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2557,8 +2558,8 @@ test("report --require-verified rejects malformed issues without rewriting stale assertFinalReportUnchanged(reportDir, reportSnapshot); }); -test("report --require-verified rejects legacy report versions without a compatibility reader", async () => { - const project = tempProject(); +test("report --require-verified rejects legacy report versions without a compatibility reader", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-legacy-version"); const reportDir = path.join(runData.run_root, "artifacts", "final-report"); const reportPath = path.join(reportDir, "report.json"); @@ -2580,8 +2581,8 @@ test("report --require-verified rejects legacy report versions without a compati assertFinalReportUnchanged(reportDir, reportSnapshot); }); -test("agent-owned bytes stay identical across validation, sync, aggregation, report, dashboard, eval, and bundle reads", async () => { - const project = tempProject(); +test("agent-owned bytes stay identical across validation, sync, aggregation, report, dashboard, eval, and bundle reads", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeByteIdentityTopology(project); const env = fakeSmithersEnv(project, true, ["aggregate-test-files"]); @@ -2752,8 +2753,8 @@ test("agent-owned bytes stay identical across validation, sync, aggregation, rep assertAgentBytesUnchanged(); }); -test("report bundle creates a portable ZIP without workspaces", async () => { - const project = tempProject(); +test("report bundle creates a portable ZIP without workspaces", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeReportTopology(project); @@ -2956,8 +2957,8 @@ test("report bundle creates a portable ZIP without workspaces", async () => { assertFinalReportUnchanged(reportDir, reportSnapshot); }); -test("report bundle manifest records files it could not package", async () => { - const project = tempProject(); +test("report bundle manifest records files it could not package", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-omissions"); const engineLogDir = path.join(runData.run_root, "smithers", "logs"); @@ -3000,8 +3001,8 @@ test("report bundle manifest records files it could not package", async () => { ); }); -test("report bundle --require-verified rejects a changed authenticated publication from any finalized producer", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects a changed authenticated publication from any finalized producer", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-mutated-publication"); const artifactDir = path.join(runData.run_root, "artifacts", "project-discovery"); const supportPublicationPath = "evidence/discovery-trace.txt"; @@ -3024,8 +3025,8 @@ test("report bundle --require-verified rejects a changed authenticated publicati ); }); -test("report bundle --require-verified rejects manifest bytes injected only into the recursive archive read", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects manifest bytes injected only into the recursive archive read", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-manifest-read-injection"); const artifactDir = path.join(runData.run_root, "artifacts", "project-discovery"); fs.writeFileSync(path.join(artifactDir, "stdout.txt"), "generated stdout\n", "utf8"); @@ -3048,8 +3049,8 @@ test("report bundle --require-verified rejects manifest bytes injected only into ); }); -test("report bundle --require-verified rejects graph-fingerprint bytes injected only into the archive read", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects graph-fingerprint bytes injected only into the archive read", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-fingerprint-read-injection"); const artifactDir = path.join(runData.run_root, "artifacts", "project-discovery"); fs.writeFileSync(path.join(artifactDir, "stdout.txt"), "generated stdout\n", "utf8"); @@ -3072,8 +3073,8 @@ test("report bundle --require-verified rejects graph-fingerprint bytes injected ); }); -test("report bundle --require-verified rejects a persistently tampered sealed graph fingerprint before writing a ZIP", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects a persistently tampered sealed graph fingerprint before writing a ZIP", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-fingerprint-persistent-tamper"); const fingerprintPath = path.join(runData.run_root, "graph.fingerprint"); const tamperedBytes = Buffer.from(`${"a".repeat(64)}\n`, "utf8"); @@ -3090,8 +3091,8 @@ test("report bundle --require-verified rejects a persistently tampered sealed gr ); }); -test("report bundle --require-verified rejects a declared report producer that claims success without finalization authority", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects a declared report producer that claims success without finalization authority", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-invalid-success-claim"); const layout = layoutForRunRoot(runData.run_root, runData.run_id); updateNodeState(layout, "final-report", { @@ -3113,8 +3114,8 @@ test("report bundle --require-verified rejects a declared report producer that c ); }); -test("report bundle --require-verified applies canonical report-pair validation to custom declarations", async () => { - const project = tempProject(); +test("report bundle --require-verified applies canonical report-pair validation to custom declarations", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeCustomReportTopology(project); const run = await cli( @@ -3148,8 +3149,8 @@ test("report bundle --require-verified applies canonical report-pair validation ); }); -test("report bundle preserves the exact validated event-record-v2 journal snapshot", async () => { - const project = tempProject(); +test("report bundle preserves the exact validated event-record-v2 journal snapshot", async (t) => { + const project = tempProject(t); const runId = "report-bundle-event-snapshot"; const runRoot = (await createReportRun(project, runId)).run_root; const record = bundleFixtureEvent(runId); @@ -3165,8 +3166,8 @@ test("report bundle preserves the exact validated event-record-v2 journal snapsh assert.deepEqual(captured, journalBytes); }); -test("report bundle treats only an absent event journal as optional", async () => { - const project = tempProject(); +test("report bundle treats only an absent event journal as optional", async (t) => { + const project = tempProject(t); const runId = "report-bundle-no-event-journal"; const runRoot = (await createReportRun(project, runId)).run_root; fs.rmSync(path.join(runRoot, "events.jsonl"), { force: true }); @@ -3180,8 +3181,8 @@ test("report bundle treats only an absent event journal as optional", async () = assert.equal(zip.readAsText("graph.fingerprint"), fs.readFileSync(path.join(runRoot, "graph.fingerprint"), "utf8")); }); -test("report bundle --require-verified rejects historical runs without current sealed workflow authority", async () => { - const project = tempProject(); +test("report bundle --require-verified rejects historical runs without current sealed workflow authority", async (t) => { + const project = tempProject(t); assert.equal((await cli(project, ["init", "--force"])).code, 0); const runId = "report-bundle-historical-unsealed"; const runRoot = path.join(project, ".ultrafuzz", "runs", runId); @@ -3196,7 +3197,7 @@ test("report bundle --require-verified rejects historical runs without current s }); test("report bundle --require-verified fails closed on every present invalid event journal", async (context) => { - const project = tempProject(); + const project = tempProject(context); const cases: Array<{ name: string; prepare(eventsPath: string, runId: string): void; @@ -3274,8 +3275,8 @@ test("report bundle --require-verified fails closed on every present invalid eve } }); -test("report bundle packages incomplete runs without a final-report JSON", async () => { - const project = tempProject(); +test("report bundle packages incomplete runs without a final-report JSON", async (t) => { + const project = tempProject(t); const runData = await createReportRun(project, "report-bundle-incomplete"); const bundled = await cli(project, ["report", "bundle", runData.run_id, "--json"]); @@ -3295,8 +3296,8 @@ test("report bundle packages incomplete runs without a final-report JSON", async ); }); -test("stats reads strict local current evidence and reports genuinely absent ledgers", async () => { - const project = tempProject(); +test("stats reads strict local current evidence and reports genuinely absent ledgers", async (t) => { + const project = tempProject(t); const fixture = writeStatsFixture(project, "stats-local", { linked: false }); const captured = await cli(project, ["stats", fixture.runId, "--json"]); @@ -3347,7 +3348,7 @@ test("stats reads strict local current evidence and reports genuinely absent led }); test("stats falls back to unchanged local evidence when workflow event output is malformed", async (context) => { - const project = tempProject(); + const project = tempProject(context); assert.equal((await cli(project, ["init", "--force"])).code, 0); writeSmallTopology(project); const env = fakeSmithersEnv(project); @@ -3381,8 +3382,8 @@ test("stats falls back to unchanged local evidence when workflow event output is } }); -test("stats queries a current report bundle offline and reports genuinely missing ledgers", async () => { - const project = tempProject(); +test("stats queries a current report bundle offline and reports genuinely missing ledgers", async (t) => { + const project = tempProject(t); const fixture = writeStatsFixture(project, "stats-bundle"); const bundlePath = path.join(project, "stats-bundle.zip"); writeStatsBundle(bundlePath, fixture); @@ -3418,8 +3419,8 @@ test("stats queries a current report bundle offline and reports genuinely missin ); }); -test("stats rejects a historical report-bundle manifest", async () => { - const project = tempProject(); +test("stats rejects a historical report-bundle manifest", async (t) => { + const project = tempProject(t); const fixture = writeStatsFixture(project, "stats-v1-bundle"); const bundlePath = path.join(project, "stats-v1-bundle.zip"); writeStatsBundle(bundlePath, fixture, { @@ -3435,7 +3436,7 @@ test("stats rejects a historical report-bundle manifest", async () => { }); test("stats fails closed for malformed present bundle evidence", async (context) => { - const project = tempProject(); + const project = tempProject(context); const cases: Array<{ name: string; member?: string; @@ -3609,7 +3610,7 @@ test("stats fails closed for malformed present bundle evidence", async (context) }); test("stats binds the v3 manifest entry count and rejects ZIP path aliases", async (context) => { - const project = tempProject(); + const project = tempProject(context); await context.test("manifest entry count", async () => { const fixture = writeStatsFixture(project, "stats-count-mismatch"); const bundlePath = path.join(project, "stats-count-mismatch.zip"); @@ -3670,8 +3671,8 @@ test("stats binds the v3 manifest entry count and rejects ZIP path aliases", asy }); }); -test("stats rejects local evidence whose leaf is a symlink", async () => { - const project = tempProject(); +test("stats rejects local evidence whose leaf is a symlink", async (t) => { + const project = tempProject(t); const fixture = writeStatsFixture(project, "stats-symlink"); const usagePath = path.join(fixture.runRoot, "usage.jsonl"); const outsidePath = path.join(project, "outside-usage.jsonl"); @@ -3684,8 +3685,8 @@ test("stats rejects local evidence whose leaf is a symlink", async () => { assert.match(JSON.stringify(parseJson(captured).diagnostics), /symlink/u); }); -test("stats rejects linked local evidence without sealed workflow authority", async () => { - const project = tempProject(); +test("stats rejects linked local evidence without sealed workflow authority", async (t) => { + const project = tempProject(t); const fixture = writeStatsFixture(project, "stats-missing-control-authority"); const captured = await cli(project, ["stats", fixture.runId, "--json"]); diff --git a/packages/cli/test/lifecycle-commands.test.ts b/packages/cli/test/lifecycle-commands.test.ts index 66bc223b5..a1340b71f 100644 --- a/packages/cli/test/lifecycle-commands.test.ts +++ b/packages/cli/test/lifecycle-commands.test.ts @@ -1,10 +1,10 @@ import assert from "node:assert/strict"; import fs from "node:fs"; -import os from "node:os"; import path from "node:path"; -import test from "node:test"; +import test, { type TestContext } from "node:test"; import { runCli } from "../src/index.js"; +import { temporaryRoot } from "./temporary-root.js"; const WORKFLOW_RUN_ID = "ultrafuzz-lifecycle-cli-run"; const RUN_ID = "lifecycle-cli-run"; @@ -15,8 +15,8 @@ interface Capture { code: number; } -function tempProject(): string { - return fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-lifecycle-cli-")); +function tempProject(t: TestContext): string { + return temporaryRoot("ufz-lifecycle-cli-", t); } function shellQuote(value: string): string { @@ -372,9 +372,10 @@ function fakeEnv(project: string, options: { cancelStatus?: string } = {}): Reco } async function launchedProject( + t: TestContext, options: { cancelStatus?: string } = {} ): Promise<{ project: string; env: Record; runRoot: string }> { - const project = tempProject(); + const project = tempProject(t); const env = fakeEnv(project, options); const init = await cli(project, ["init", "--json"], env); assert.equal(init.code, 0, init.stderr); @@ -385,8 +386,8 @@ async function launchedProject( return { project, env, runRoot: runData.run_root }; } -test("why reports the diagnosis in human and JSON output", async () => { - const { project, env } = await launchedProject(); +test("why reports the diagnosis in human and JSON output", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["why", RUN_ID], env); assert.equal(human.code, 0, human.stderr); @@ -408,8 +409,8 @@ test("why reports the diagnosis in human and JSON output", async () => { assert.equal(data.blockers[1]?.node_id, "node:strategy"); }); -test("timeline surfaces frame numbers for fork --frame", async () => { - const { project, env } = await launchedProject(); +test("timeline surfaces frame numbers for fork --frame", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["timeline", RUN_ID], env); assert.equal(human.code, 0, human.stderr); @@ -429,8 +430,8 @@ test("timeline surfaces frame numbers for fork --frame", async () => { ); }); -test("snapshots lists checkpoints without engine-internal identifiers", async () => { - const { project, env } = await launchedProject(); +test("snapshots lists checkpoints without engine-internal identifiers", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["snapshots", RUN_ID], env); assert.equal(human.code, 0, human.stderr); @@ -444,8 +445,8 @@ test("snapshots lists checkpoints without engine-internal identifiers", async () assert.doesNotMatch(JSON.stringify(body), /commit-1|op-1|workspace/u); }); -test("events returns bounded lifecycle events in human and JSON output", async () => { - const { project, env } = await launchedProject(); +test("events returns bounded lifecycle events in human and JSON output", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["events", RUN_ID], env); assert.equal(human.code, 0, human.stderr); @@ -460,8 +461,8 @@ test("events returns bounded lifecycle events in human and JSON output", async ( assert.doesNotMatch(fs.readFileSync(path.join(project, "smithers-commands.log"), "utf8"), /--raw/u); }); -test("events --watch streams one line per event and terminates", async () => { - const { project, env } = await launchedProject(); +test("events --watch streams one line per event and terminates", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["events", RUN_ID, "--watch", "--interval", "1"], env); assert.equal(human.code, 0, human.stderr); @@ -482,8 +483,8 @@ test("events --watch streams one line per event and terminates", async () => { } }); -test("node reports focused status and only expands tool payloads with --tools", async () => { - const { project, env } = await launchedProject(); +test("node reports focused status and only expands tool payloads with --tools", async (t) => { + const { project, env } = await launchedProject(t); const human = await cli(project, ["node", RUN_ID, "node:project-discovery"], env); assert.equal(human.code, 0, human.stderr); @@ -510,8 +511,8 @@ test("node reports focused status and only expands tool payloads with --tools", assert.deepEqual(data.attempts[0]?.tool_calls[0]?.input, { command: "forge build" }); }); -test("node --watch emits NDJSON envelopes and terminates", async () => { - const { project, env } = await launchedProject(); +test("node --watch emits NDJSON envelopes and terminates", async (t) => { + const { project, env } = await launchedProject(t); const watched = await cli(project, ["node", RUN_ID, "node:project-discovery", "--watch", "--json"], env); @@ -527,8 +528,8 @@ test("node --watch emits NDJSON envelopes and terminates", async () => { ); }); -test("events rejects a raw event category instead of widening the view", async () => { - const { project, env } = await launchedProject(); +test("events rejects a raw event category instead of widening the view", async (t) => { + const { project, env } = await launchedProject(t); const rejected = await cli(project, ["events", RUN_ID, "--type", "agent", "--json"], env); @@ -544,8 +545,8 @@ test("events rejects a raw event category instead of widening the view", async ( assert.equal(accepted.code, 0, accepted.stderr); }); -test("events --watch --json keeps a stream failure on one NDJSON line", async () => { - const { project, env } = await launchedProject(); +test("events --watch --json keeps a stream failure on one NDJSON line", async (t) => { + const { project, env } = await launchedProject(t); fs.writeFileSync(path.join(project, "fake-events-failure"), "stream broke\n", "utf8"); const watched = await cli(project, ["events", RUN_ID, "--watch", "--json"], env); @@ -559,8 +560,8 @@ test("events --watch --json keeps a stream failure on one NDJSON line", async () assertNoEngineBranding(body); }); -test("cancel distinguishes a submitted request from a confirmed cancellation", async () => { - const requested = await launchedProject({ cancelStatus: "cancel-requested" }); +test("cancel distinguishes a submitted request from a confirmed cancellation", async (t) => { + const requested = await launchedProject(t, { cancelStatus: "cancel-requested" }); const human = await cli(requested.project, ["cancel", RUN_ID], requested.env); assert.equal(human.code, 0, human.stderr); @@ -570,7 +571,7 @@ test("cancel distinguishes a submitted request from a confirmed cancellation", a }; assert.equal(runningState.status, "running"); - const confirmed = await launchedProject({ cancelStatus: "cancelled" }); + const confirmed = await launchedProject(t, { cancelStatus: "cancelled" }); const json = await cli(confirmed.project, ["cancel", RUN_ID, "--json"], confirmed.env); assert.equal(json.code, 0, json.stderr); const body = parseJson(json); @@ -584,8 +585,8 @@ test("cancel distinguishes a submitted request from a confirmed cancellation", a assert.match(humanConfirmed.stdout, /^Cancellation confirmed: lifecycle-cli-run is canceled$/mu); }); -test("doctor reports install posture in human and JSON output", async () => { - const { project, env } = await launchedProject(); +test("doctor reports install posture in human and JSON output", async (t) => { + const { project, env } = await launchedProject(t); writeSmallTopology(project, "recon"); const doctorEnv = { ...env, PATH: path.dirname(env.SMITHERS_BIN!) }; @@ -642,8 +643,8 @@ test("doctor reports install posture in human and JSON output", async () => { assert.ok(overrideBody.diagnostics.some((diagnostic) => diagnostic.code === "DOCTOR_AGENT_CREDENTIAL_MISSING")); }); -test("status recommends ultrafuzz why instead of the engine command", async () => { - const project = tempProject(); +test("status recommends ultrafuzz why instead of the engine command", async (t) => { + const project = tempProject(t); const binDir = path.join(path.dirname(project), path.basename(project) + "-fake-bin"); fs.mkdirSync(binDir, { recursive: true }); const smithers = path.join(binDir, "smithers"); diff --git a/packages/cli/test/temporary-root.ts b/packages/cli/test/temporary-root.ts new file mode 100644 index 000000000..cca32c14e --- /dev/null +++ b/packages/cli/test/temporary-root.ts @@ -0,0 +1,31 @@ +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import type { TestContext } from "node:test"; + +/** + * Create a canonical temporary directory that is removed when test `t` ends, + * together with the `-fake-bin` directory the fake engine fixtures create + * beside it. A launched run leaves about 450 MB of sealed snapshot directories + * with mode `dr-x`, which `rmSync` cannot remove until owner write permission is + * restored on the way down. + */ +export function temporaryRoot(prefix: string, t: TestContext): string { + const root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), prefix)); + t.after(() => { + for (const directory of [root, `${root}-fake-bin`]) { + restoreOwnerWrite(directory); + fs.rmSync(directory, { recursive: true, force: true, maxRetries: 3 }); + } + }); + return root; +} + +function restoreOwnerWrite(directory: string): void { + const stat = fs.lstatSync(directory, { throwIfNoEntry: false }); + if (stat === undefined || !stat.isDirectory()) return; + fs.chmodSync(directory, stat.mode | 0o700); + for (const entry of fs.readdirSync(directory, { withFileTypes: true })) { + if (entry.isDirectory()) restoreOwnerWrite(path.join(directory, entry.name)); + } +} From 5bd2a00bbc1d091bdb930614f390509f034cc6ca Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:16:48 +0000 Subject: [PATCH 076/206] build(lint): enforce complexity and size budgets on all code against a recorded baseline The complexity and size rules ran only in the diff-limited strict lint, which reports a message only when its line changed. ESLint reports these rules on a function's first line (max-lines on the first line past the limit), so adding branches inside an over-budget function passed, while editing only its signature failed, forcing an unrelated refactor on a fix. Move complexity, max-depth, max-lines, max-lines-per-function, max-nested-callbacks, max-params and max-statements into the always-on config, with the same relaxed values for tests and templates, and record the 1,003 existing violations (197 files; 701 of them in 108 packages/*/src files) with ESLint's bulk suppressions in eslint-suppressions.json. `pnpm -w lint` now fails when a file gains a violation of a rule and when a recorded count is higher than the file's violations; `pnpm -w lint:prune` lowers the counts. The diff-limited strict lint keeps the type-aware strict rules, no-console and the TODO/FIXME check. It passes --pass-on-unpruned-suppressions because it sees only changed lines, so its counts are below the recorded ones by design. ESLint writes the suppressions file without a final newline, so prettier ignores it. Co-Authored-By: Claude Opus 5.5 --- .prettierignore | 2 + docs/contributing.md | 23 + eslint-suppressions.json | 1824 ++++++++++++++++++++++++++++++++++++++ eslint.config.js | 53 +- package.json | 5 +- 5 files changed, 1884 insertions(+), 23 deletions(-) create mode 100644 eslint-suppressions.json diff --git a/.prettierignore b/.prettierignore index e33b3438d..b43cc9dcf 100644 --- a/.prettierignore +++ b/.prettierignore @@ -18,3 +18,5 @@ packages/config/test/fixtures/resolved-config.valid.pre-same-agent-attempts.json prompts/**/*.md docs/reference/audit-profiles.md docs/reference/prompt-catalog.md + +eslint-suppressions.json diff --git a/docs/contributing.md b/docs/contributing.md index 7251c611a..efaef00a5 100644 --- a/docs/contributing.md +++ b/docs/contributing.md @@ -78,3 +78,26 @@ pnpm -r test Prefer narrow package checks while iterating, then run broader gates before PR handoff when a change touches shared behavior or release workflows. + +## Complexity and Size Budgets + +`pnpm -w lint` applies complexity and size budgets to all code under +`packages/` and `scripts/`: cyclomatic complexity 20, nesting depth 4, 500 lines +per file, 80 lines and 40 statements per function, 5 parameters, and 4 nested +callbacks. Tests and workflow templates allow complexity 25, 1000 lines per +file, 150 lines and 80 statements per function, and 6 parameters. + +Violations that predate the budgets are counted per file and rule in +`eslint-suppressions.json`, ESLint's bulk-suppressions file: + +- A change that adds a violation to a file fails lint, because that file's + count for the rule rises above the recorded count. Fix the new violation. +- A change that removes violations also fails lint, with "There are + suppressions left that do not occur anymore". Run `pnpm -w lint:prune` and + commit the smaller `eslint-suppressions.json`. +- Moving code that already exceeds a budget, including renaming its file, + needs its count moved to the new path in `eslint-suppressions.json`. + +Counts are per file and rule, so a function that already exceeds a budget can +grow without failing lint. `pnpm -w lint:strict:ci` adds the type-aware strict +rules, `no-console`, and the TODO/FIXME check for changed lines only. diff --git a/eslint-suppressions.json b/eslint-suppressions.json new file mode 100644 index 000000000..a59d150c2 --- /dev/null +++ b/eslint-suppressions.json @@ -0,0 +1,1824 @@ +{ + "packages/artifacts/src/analysis-bundle.ts": { + "max-lines": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/attempt-ledger.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/artifacts/src/cloud-selected-task.ts": { + "complexity": { + "count": 2 + }, + "max-lines-per-function": { + "count": 2 + } + }, + "packages/artifacts/src/coverage-evidence.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/events.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/artifacts/src/finding-provenance.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "packages/artifacts/src/findings-schema.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/goal-plan.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/artifacts/src/json-file-validator.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/planned-graph.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/property-provenance.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/run-layout.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/artifacts/src/safe-paths.ts": { + "max-lines": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/semantic-gates.ts": { + "complexity": { + "count": 19 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 20 + }, + "max-statements": { + "count": 14 + } + }, + "packages/artifacts/src/smithers-task-manifest.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/src/state-schema.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/artifacts/src/state.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/artifacts/src/threat-model.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/artifacts/src/workflow-contracts.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/test/artifacts.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/artifacts/test/contract-fixtures.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 5 + } + }, + "packages/artifacts/test/schema.test.ts": { + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 5 + } + }, + "packages/artifacts/test/semantic-gates.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 9 + }, + "max-statements": { + "count": 1 + } + }, + "packages/artifacts/test/threat-goal-artifacts.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/cli/src/benchmark-analysis/lib/analysis.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/cli/src/benchmark-analysis/lib/charts.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 5 + }, + "max-statements": { + "count": 1 + } + }, + "packages/cli/src/benchmark-analysis/lib/outputs.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/cli/src/cli-contracts.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/cli/src/commands/artifact/validate.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/cli/src/commands/eval/history.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/cli/src/commands/node.ts": { + "complexity": { + "count": 1 + } + }, + "packages/cli/src/commands/report.ts": { + "max-params": { + "count": 1 + } + }, + "packages/cli/src/commands/report/bundle.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/cli/src/commands/stats.ts": { + "max-lines": { + "count": 1 + }, + "max-params": { + "count": 1 + } + }, + "packages/cli/src/run-statistics.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/cli/test/benchmark-analysis.test.ts": { + "max-params": { + "count": 2 + } + }, + "packages/cli/test/cli.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 6 + }, + "max-statements": { + "count": 2 + } + }, + "packages/cli/test/json-validate.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/cli/test/lifecycle-commands.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/cli/test/run-statistics.test.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/config/src/loader.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/config/src/resolve.ts": { + "complexity": { + "count": 5 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/config/test/audit-profiles.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/config/test/config.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + } + }, + "packages/config/test/resolved-config-schema.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/dashboard/frontend/src/graphPhases.ts": { + "complexity": { + "count": 1 + } + }, + "packages/dashboard/frontend/src/main.tsx": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 9 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/dashboard/frontend/src/templateValidation.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/dashboard/frontend/src/useManagedPromptEditor.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/dashboard/src/index.ts": { + "complexity": { + "count": 4 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 5 + }, + "max-params": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/dashboard/test/dashboard.test.ts": { + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/src/benchmark-analysis-contracts.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/evals/src/benchmark-manifest.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/src/eval-durable.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/evals/src/eval-semantic-gates.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + } + }, + "packages/evals/src/history.ts": { + "complexity": { + "count": 4 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 8 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "packages/evals/src/node-telemetry.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 2 + } + }, + "packages/evals/src/recovery-equivalence.ts": { + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/evals/src/runner.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/evals/src/scoring.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "packages/evals/src/status.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/evals/src/suite.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/evals/src/types.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/evals/test/analysis-bundle.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/evals/test/benchmark-analysis-contracts.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/benchmark-manifest.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/evals/test/efficiency.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/eval-durable.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/expansion.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/ground-truth.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/helpers.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/history.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + } + }, + "packages/evals/test/node-telemetry.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/public-diagnostics.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/recovery-equivalence.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/runner-publish.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/scoring.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/status.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evals/test/suite.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/evals/test/threat-model-benchmark-lane.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evmbench/src/adapter.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evmbench/src/runner.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evmbench/src/semantic-gates.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/evmbench/test/adapter.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/evmbench/test/contracts.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/src/auth.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/modal/src/cli.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/deterministic-archive.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/launch-state.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/modal-contracts.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/modal/src/modal-documents.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/modal-semantic-gates.ts": { + "complexity": { + "count": 4 + }, + "max-lines": { + "count": 1 + } + }, + "packages/modal/src/node-provider.ts": { + "complexity": { + "count": 8 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 10 + }, + "max-params": { + "count": 4 + }, + "max-statements": { + "count": 6 + } + }, + "packages/modal/src/node-worker.ts": { + "complexity": { + "count": 4 + }, + "max-depth": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-params": { + "count": 2 + }, + "max-statements": { + "count": 3 + } + }, + "packages/modal/src/pinned-source.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/src/public-bundle.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "packages/modal/src/public-eval-diagnostics.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/src/public-worker.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + }, + "max-params": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/recovery.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/resume.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/runner.ts": { + "complexity": { + "count": 8 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 12 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 3 + } + }, + "packages/modal/src/safe-archive.ts": { + "max-lines-per-function": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/terminal-disposition.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/worker-lineage.ts": { + "max-statements": { + "count": 1 + } + }, + "packages/modal/src/worker-result.ts": { + "max-params": { + "count": 1 + } + }, + "packages/modal/src/worker.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + } + }, + "packages/modal/test/auth.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/ci-config.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/modal/test/config.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/current-artifact-fixtures.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/deterministic-archive.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/launch-state.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + } + }, + "packages/modal/test/modal-documents.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/modal/test/node-provider.test.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 2 + } + }, + "packages/modal/test/pinned-source.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/public-bundle.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/public-eval-diagnostics.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/public-worker.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/modal/test/recovery-lifecycle.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/recovery.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/resume.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/runner.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 5 + } + }, + "packages/modal/test/terminal-disposition.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/terminal-recovery.integration.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/worker-lineage.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/modal/test/worker-result.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/prompts/src/render.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/prompts/test/render.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/prompts/test/semantic-anchors.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/references/src/index.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/references/src/vulnerability-database.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/agent-registry.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + } + }, + "packages/runtime/src/aggregation-semantic-context.ts": { + "complexity": { + "count": 2 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/src/artifact-gates.ts": { + "complexity": { + "count": 23 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 29 + }, + "max-params": { + "count": 13 + }, + "max-statements": { + "count": 11 + } + }, + "packages/runtime/src/canonical-properties-markdown.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/clean.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/data-governance.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + } + }, + "packages/runtime/src/doctor.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/dynamic-expansion.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/src/dynamic-runtime.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/final-report-markdown.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + } + }, + "packages/runtime/src/init.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/lifecycle-inspection.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/materialize.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/model-pricing.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/pinned-submodules.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/plan-run.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 4 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/prompt-artifact-authority.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/runtime-semantic-gates.ts": { + "complexity": { + "count": 1 + } + }, + "packages/runtime/src/smithers-executable-capability.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + } + }, + "packages/runtime/src/smithers.ts": { + "complexity": { + "count": 7 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 14 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 8 + } + }, + "packages/runtime/src/start-run.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 6 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 3 + } + }, + "packages/runtime/src/state-export.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/templates/smithers/agents/kimi.tsx": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/runtime/src/templates/smithers/agents/openrouter.tsx": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/runtime/src/templates/smithers/workflows/workflow.tsx": { + "complexity": { + "count": 10 + }, + "max-depth": { + "count": 5 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 6 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/trusted-cli-closure.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/src/trusted-cli.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/types.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/runtime/src/validate.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/verified-output.ts": { + "complexity": { + "count": 2 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/workflow-control.ts": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/src/workflow-controller-generation.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-params": { + "count": 1 + } + }, + "packages/runtime/src/workflow-execution-snapshot-capability.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/src/workflow-integrity.ts": { + "complexity": { + "count": 9 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 8 + }, + "max-params": { + "count": 4 + }, + "max-statements": { + "count": 6 + } + }, + "packages/runtime/src/workflow-run-link.ts": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + } + }, + "packages/runtime/src/workflow-sync.ts": { + "complexity": { + "count": 16 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 18 + }, + "max-params": { + "count": 2 + }, + "max-statements": { + "count": 10 + } + }, + "packages/runtime/src/workspace-handoff.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/test/agent-adapter-boundaries.test.ts": { + "complexity": { + "count": 1 + }, + "max-depth": { + "count": 2 + }, + "max-lines": { + "count": 1 + } + }, + "packages/runtime/test/artifact-gates.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 15 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/test/cloud-worker-handoff.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/test/dynamic-lifecycle.test.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/runtime/test/final-report-markdown.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/runtime/test/generated-workflow-verifier.test.ts": { + "max-depth": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 11 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/test/invariant-suite-handoff-durability.test.ts": { + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/test/lifecycle-inspection.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/runtime/test/pinned-submodules.test.ts": { + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 2 + } + }, + "packages/runtime/test/runtime-document-contracts.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/runtime/test/runtime.test.ts": { + "complexity": { + "count": 8 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 22 + }, + "max-statements": { + "count": 10 + } + }, + "packages/runtime/test/verified-output.test.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 3 + }, + "max-params": { + "count": 1 + } + }, + "packages/runtime/test/vulnerability-database.test.ts": { + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "packages/runtime/test/workspace-handoff.test.ts": { + "max-lines": { + "count": 1 + } + }, + "packages/topology/src/expand.ts": { + "max-params": { + "count": 1 + } + }, + "packages/topology/src/validate.ts": { + "complexity": { + "count": 1 + }, + "max-lines": { + "count": 1 + } + }, + "packages/topology/test/artifact-handoffs.test.ts": { + "max-lines-per-function": { + "count": 2 + } + }, + "packages/topology/test/expand.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/topology/test/packaged-topologies.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/topology/test/schema.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "packages/topology/test/validate.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "scripts/check-production-dependency-advisories.mjs": { + "complexity": { + "count": 2 + }, + "max-depth": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-params": { + "count": 1 + }, + "max-statements": { + "count": 2 + } + }, + "scripts/ci/dependency-advisories.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "scripts/ci/prepare-eval-history-publication.mjs": { + "complexity": { + "count": 3 + }, + "max-lines": { + "count": 1 + }, + "max-lines-per-function": { + "count": 2 + }, + "max-statements": { + "count": 1 + } + }, + "scripts/ci/prepare-eval-history-publication.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "scripts/ci/prepare-modal-benchmark-cleanup.mjs": { + "max-lines-per-function": { + "count": 1 + } + }, + "scripts/ci/validate-modal-benchmark-launch.test.ts": { + "max-lines-per-function": { + "count": 1 + } + }, + "scripts/ci/validate-threat-model-benchmark-gate.mjs": { + "complexity": { + "count": 1 + }, + "max-lines-per-function": { + "count": 1 + }, + "max-statements": { + "count": 1 + } + }, + "scripts/docs-check.mjs": { + "max-depth": { + "count": 1 + } + } +} \ No newline at end of file diff --git a/eslint.config.js b/eslint.config.js index 615493e8d..3ccb57712 100644 --- a/eslint.config.js +++ b/eslint.config.js @@ -9,6 +9,37 @@ const browserFiles = ["packages/dashboard/frontend/{public,src}/**/*.{js,ts,tsx} const typescriptFiles = ["packages/**/*.{ts,tsx}", "scripts/**/*.{ts,tsx}"]; const typeCheckedFiles = ["packages/*/src/**/*.ts", "packages/dashboard/frontend/src/**/*.{ts,tsx}"]; const sourceFiles = ["packages/**/*.{js,mjs,cjs,ts,tsx}", "scripts/**/*.{js,mjs,cjs,ts,tsx}"]; +// Complexity and size budgets apply to all code. Violations that predate them are +// counted per file and rule in eslint-suppressions.json: a file that gains one +// fails lint, and `pnpm -w lint:prune` removes the counts a change pays down. +const sizeConfigs = [ + { + files: sourceFiles, + rules: { + complexity: ["error", 20], + "max-depth": ["error", 4], + "max-lines": ["error", { max: 500, skipBlankLines: true, skipComments: true }], + "max-lines-per-function": ["error", { max: 80, skipBlankLines: true, skipComments: true }], + "max-nested-callbacks": ["error", 4], + "max-params": ["error", 5], + "max-statements": ["error", 40] + } + }, + { + files: [ + "**/{test,tests}/**/*.{js,mjs,cjs,ts,tsx}", + "**/*.test.{js,mjs,cjs,ts,tsx}", + "packages/runtime/src/templates/**/*.tsx" + ], + rules: { + complexity: ["error", 25], + "max-lines": ["error", { max: 1000, skipBlankLines: true, skipComments: true }], + "max-lines-per-function": ["error", { max: 150, skipBlankLines: true, skipComments: true }], + "max-params": ["error", 6], + "max-statements": ["error", 80] + } + } +]; const strictConfigs = strictLint ? [ ...tseslint.configs.strict.map((config) => ({ @@ -31,30 +62,9 @@ const strictConfigs = strictLint { files: sourceFiles, rules: { - complexity: ["error", 20], - "max-depth": ["error", 4], - "max-lines": ["error", { max: 500, skipBlankLines: true, skipComments: true }], - "max-lines-per-function": ["error", { max: 80, skipBlankLines: true, skipComments: true }], - "max-nested-callbacks": ["error", 4], - "max-params": ["error", 5], - "max-statements": ["error", 40], "no-console": "error", "no-warning-comments": ["error", { terms: ["todo", "fixme"], location: "anywhere" }] } - }, - { - files: [ - "**/{test,tests}/**/*.{js,mjs,cjs,ts,tsx}", - "**/*.test.{js,mjs,cjs,ts,tsx}", - "packages/runtime/src/templates/**/*.tsx" - ], - rules: { - complexity: ["error", 25], - "max-lines": ["error", { max: 1000, skipBlankLines: true, skipComments: true }], - "max-lines-per-function": ["error", { max: 150, skipBlankLines: true, skipComments: true }], - "max-params": ["error", 6], - "max-statements": ["error", 80] - } } ] : []; @@ -77,6 +87,7 @@ export default tseslint.config( }, js.configs.recommended, ...tseslint.configs.recommended, + ...sizeConfigs, ...strictConfigs, { files: ["**/*.{js,mjs,cjs,ts,tsx}"], diff --git a/package.json b/package.json index fba00c0ea..546a5cfaf 100644 --- a/package.json +++ b/package.json @@ -23,8 +23,9 @@ "knip": "pnpm dlx knip@6.33.0 --include files,dependencies,unlisted,unresolved --treat-config-hints-as-errors", "lint": "eslint \"packages/**/*.{ts,tsx,js,mjs,cjs}\" \"scripts/**/*.{ts,tsx,js,mjs,cjs}\" \"*.{js,mjs,cjs}\"", "lint:fix": "eslint \"packages/**/*.{ts,tsx,js,mjs,cjs}\" \"scripts/**/*.{ts,tsx,js,mjs,cjs}\" \"*.{js,mjs,cjs}\" --fix", - "lint:strict": "ULTRAFUZZ_STRICT_LINT=1 eslint \"packages/**/*.{ts,tsx,js,mjs,cjs}\" \"scripts/**/*.{ts,tsx,js,mjs,cjs}\" \"*.{js,mjs,cjs}\" --max-warnings 0", - "lint:strict:ci": "bash -o pipefail -c ': \"${ESLINT_PLUGIN_DIFF_COMMIT:?required}\"; git diff --name-only --diff-filter=ACMR -z \"$ESLINT_PLUGIN_DIFF_COMMIT\" -- \"*.js\" \"*.mjs\" \"*.cjs\" \"*.ts\" \"*.tsx\" | xargs -0 --no-run-if-empty env ULTRAFUZZ_STRICT_LINT=1 eslint --max-warnings 0'", + "lint:prune": "pnpm -w lint --prune-suppressions", + "lint:strict": "ULTRAFUZZ_STRICT_LINT=1 eslint \"packages/**/*.{ts,tsx,js,mjs,cjs}\" \"scripts/**/*.{ts,tsx,js,mjs,cjs}\" \"*.{js,mjs,cjs}\" --max-warnings 0 --pass-on-unpruned-suppressions", + "lint:strict:ci": "bash -o pipefail -c ': \"${ESLINT_PLUGIN_DIFF_COMMIT:?required}\"; git diff --name-only --diff-filter=ACMR -z \"$ESLINT_PLUGIN_DIFF_COMMIT\" -- \"*.js\" \"*.mjs\" \"*.cjs\" \"*.ts\" \"*.tsx\" | xargs -0 --no-run-if-empty env ULTRAFUZZ_STRICT_LINT=1 eslint --max-warnings 0 --pass-on-unpruned-suppressions'", "security:dependency-advisories": "pnpm --filter @ultrafuzz/artifacts... build && node scripts/check-production-dependency-advisories.mjs", "size": "size-limit", "test": "pnpm run test:ci-scripts && pnpm -r test", From 87fa795a6dc274c04c02990718fdd3d77bec324f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:22:05 +0000 Subject: [PATCH 077/206] docs: disable only MD013 for the one-entry-per-line changelog Super-linter's markdownlint (MD013, 400 columns) validates every changed Markdown file in full. The gate landed (#1005) after CHANGELOG.md was last edited, and nearly every entry is one line longer than 400 columns, so any pull request that adds a changelog entry fails "External static analysis" on the file's existing lines. Line length is the only rule the file breaks under super-linter's configuration, so disable just MD013 for this file and keep every other Markdown rule. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7419964e2..07da09dd1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -127,3 +127,5 @@ ## v0.0.1 - First external private release. + + From 638ff30f70835b21dfbea9248d3666b47a3a95ae Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:22:31 +0000 Subject: [PATCH 078/206] docs: exempt CHANGELOG.md from markdownlint's line-length rule Super-Linter lints only the files a pull request changes, and every CHANGELOG.md entry is a single line far past markdownlint's 400-character MD013 limit, so any PR that adds an entry fails external static analysis. That failure also skips the release-validation lanes that depend on it. Other open PRs add the same directive. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index c0d40d502..0bad8fad3 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,7 @@ # Changelog + + ## Unreleased ### Breaking changes From 2aa9d696f7ce89c845823961380c8ce557d9d20b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:47:38 +0000 Subject: [PATCH 079/206] fix(runtime): a reused attempt number supersedes an unrecorded occurrence Smithers restarts attempt numbering after a timetravel/retry-task reset and upserts the (run, node, iteration, attempt) row, so a later NodeStarted can reopen an identity whose earlier terminal occurrence no sync recorded. terminalWorkflowAttempts threw in that case unless the occurrence was already ledgered or passed the #952 trace/published-replacement proof. The reducer runs in the authority pre-pass, so the throw returned ok:false before any node or run state was reconciled: a finished workflow stayed running/failed on every status poll, including after the supported `resume --retry-failed` of a pre-agent failure (#1099). Both allowed branches only deleted the superseded occurrence, and the proof guarded the optional `agent` field of a ledger row that nothing reads. The supersession is now unconditional, and the authorization machinery (recordedTerminalSequences/authorizeUnrecordedSuperseded, the activation tracking, the three trace/published-replacement authority helpers, TerminalAttemptSupersessionContext and appendTerminalTaskAttempts' authorityEvents parameter) is deleted. This reverses #951/#952's rule that missing historical authority stays fail-closed. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/workflow-sync.ts | 440 ++------------------------ packages/runtime/test/runtime.test.ts | 399 +++++++++-------------- 2 files changed, 176 insertions(+), 663 deletions(-) diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index b8810c84e..1ec3dc8dc 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -6,7 +6,6 @@ import { parseResolvedConfigJsonBytes } from "@ultrafuzz/config"; import { ARTIFACT_MANIFEST_FILE, - MAX_REFERENCE_ARTIFACT_MANIFEST_AUTHORITY_BYTES, appendUsageEvents, appendNodeAttempts, appendEvent, @@ -126,8 +125,7 @@ import { projectWorkflowControlState } from "./workflow-control.js"; import { runtimeSemanticGateDiagnostics } from "./semantic-gates.js"; import { inspectSmithersAttemptAgentSelection, - reconcileSmithersAttemptAgentSelection, - smithersTaskAgentId + reconcileSmithersAttemptAgentSelection } from "./smithers-attempt-authority.js"; import { isRecord } from "@ultrafuzz/artifacts"; @@ -176,11 +174,6 @@ interface TerminalWorkflowAttempt { failureMessage?: string; } -interface TerminalAttemptSupersessionContext { - crossedRunActivation: boolean; - supersedingStartedSequence: number; -} - type SmithersNodeAttemptAuthorities = ReadonlyMap; interface AccountingSummary { @@ -1073,13 +1066,7 @@ export async function preserveFailedWorkflowAttemptsBeforeReset( const tasksByNodeId = new Map( loaded.tasks.filter((task) => task.execution.mode === "local").map((task) => [task.smithersNodeId, task]) ); - const failures = resetTerminalFailureAttempts({ - layout: evidence.layout, - workflowRunId: evidence.smithersRunId, - tasksByNodeId, - events, - recordedSequences - }); + const failures = resetTerminalFailureAttempts({ tasksByNodeId, events }); if (failures.length === 0) return; const terminalSequences = new Set( failures @@ -1126,33 +1113,15 @@ export async function preserveFailedWorkflowAttemptsBeforeReset( } function resetTerminalFailureAttempts(input: { - layout: RunLayout; - workflowRunId: string; tasksByNodeId: ReadonlyMap; events: WorkflowEvent[]; - recordedSequences: ReadonlySet; }): TerminalWorkflowAttempt[] { const relevantEvents = input.events.filter( (event) => event.type === "RunStarted" || input.tasksByNodeId.has(stringField(event.payload, "nodeId") ?? "") ); - return terminalWorkflowAttempts(relevantEvents, { - recordedTerminalSequences: input.recordedSequences, - authorizeUnrecordedSuperseded: (attempt, context) => { - const task = input.tasksByNodeId.get(attempt.nodeId); - return ( - context.crossedRunActivation && - task !== undefined && - supersededSuccessfulAttemptHasTraceAuthority( - input.layout, - input.workflowRunId, - task, - attempt, - input.events, - context - ) - ); - } - }).filter((attempt) => attempt.outcome === "failed" || attempt.outcome === "timed-out"); + return terminalWorkflowAttempts(relevantEvents).filter( + (attempt) => attempt.outcome === "failed" || attempt.outcome === "timed-out" + ); } function failedResetAttemptInputs(input: { @@ -2045,26 +2014,7 @@ async function inspectTerminalAttemptAuthorities(input: { const nodeId = stringField(event.payload, "nodeId"); return nodeId !== undefined && localTaskNodeIds.has(nodeId); }); - const recordedTerminalSequences = recordedTerminalAttemptSequences(input.layout, input.workflowRunId); - const pending = terminalWorkflowAttempts(relevantEvents, { - tolerateMissingStarts: true, - recordedTerminalSequences, - authorizeUnrecordedSuperseded: (attempt, context) => { - const task = tasksByNodeId.get(attempt.nodeId); - return ( - context.crossedRunActivation && - task !== undefined && - supersededSuccessfulAttemptHasTraceAuthority( - input.layout, - input.workflowRunId, - task, - attempt, - input.events, - context - ) - ); - } - }).filter( + const pending = terminalWorkflowAttempts(relevantEvents, { tolerateMissingStarts: true }).filter( (attempt) => tasksByNodeId.get(attempt.nodeId)?.execution.mode === "local" && (input.terminalSequences === undefined || input.terminalSequences.has(attempt.finishedSequence)) @@ -4033,7 +3983,6 @@ async function synchronizeTasks(input: { events: [...runActivationEvents, ...(eventsByNode.get(task.smithersNodeId) ?? [])].sort( (left, right) => left.sourceEventSequence - right.sourceEventSequence ), - authorityEvents: input.events, currentAttempt: evidence.attempt, currentStatus: patchStatus, finalization, @@ -5047,7 +4996,6 @@ function appendTerminalTaskAttempts(input: { workflowRunId: string; controlGeneration: string; events: WorkflowEvent[]; - authorityEvents: WorkflowEvent[]; currentAttempt?: number; currentStatus: NodeStatus; finalization: NodeFinalization; @@ -5055,19 +5003,7 @@ function appendTerminalTaskAttempts(input: { forbiddenSecretValues: readonly string[]; }): { appended: boolean; executedAttempts: number; currentAttemptExecuted: boolean } { const allExisting = replayNodeAttempts(input.layout).entries; - const observedTerminalAttempts = terminalWorkflowAttempts(input.events, { - recordedTerminalSequences: recordedTerminalAttemptSequencesFromEntries(allExisting, input.workflowRunId), - authorizeUnrecordedSuperseded: (attempt, context) => - context.crossedRunActivation && - supersededSuccessfulAttemptHasTraceAuthority( - input.layout, - input.workflowRunId, - input.task, - attempt, - input.authorityEvents, - context - ) - }); + const observedTerminalAttempts = terminalWorkflowAttempts(input.events); const terminalAttempts = input.task.execution.mode === "local" ? observedTerminalAttempts.filter((attempt) => { @@ -5259,24 +5195,27 @@ function nodeAttemptAgentProvenance( }; } +/** + * Reduce the append-only Smithers event stream to the latest terminal + * occurrence of each (node, iteration, attempt) identity. + * + * A Smithers reset (timetravel / retry-task) restarts attempt numbering and + * upserts the mutable attempt row, so a later NodeStarted can reopen an + * identity that already has a terminal occurrence. Only the newest occurrence + * is still described by `smithers node`; an earlier one is either already in + * the immutable attempt ledger (keyed by its terminal event sequence) or is + * history nobody observed in time. Neither may block synchronizing the run's + * current state, so the superseded occurrence is simply dropped here (#1099). + */ function terminalWorkflowAttempts( events: WorkflowEvent[], - options: { - tolerateMissingStarts?: boolean; - recordedTerminalSequences?: ReadonlySet; - authorizeUnrecordedSuperseded?: ( - attempt: TerminalWorkflowAttempt, - context: TerminalAttemptSupersessionContext - ) => boolean; - } = {} + options: { tolerateMissingStarts?: boolean } = {} ): TerminalWorkflowAttempt[] { const active = new Map< string, Pick >(); const attempts = new Map(); - const terminalActivations = new Map(); - let activation = 0; for (const event of events) { if (event.type === "RunStarted") { // Smithers cancels stale in-progress rows before each resumed activation, @@ -5284,7 +5223,6 @@ function terminalWorkflowAttempts( // duplicate starts fail-closed within one activation while allowing the // next activation to reuse the same durable attempt number. active.clear(); - activation += 1; continue; } if (event.type !== "NodeStarted" && event.type !== "NodeFinished" && event.type !== "NodeFailed") continue; @@ -5296,26 +5234,7 @@ function terminalWorkflowAttempts( const timestamp = new Date(event.timestampMs).toISOString(); if (event.type === "NodeStarted") { if (active.has(identity)) throw new Error(`Smithers attempt ${identity} has multiple active NodeStarted events`); - const superseded = attempts.get(identity); - if (superseded !== undefined) { - const terminalActivation = terminalActivations.get(identity); - if (terminalActivation === undefined) { - throw new Error(`Smithers attempt ${identity} is missing its terminal activation authority`); - } - if ( - !options.recordedTerminalSequences?.has(superseded.finishedSequence) && - options.authorizeUnrecordedSuperseded?.(superseded, { - crossedRunActivation: activation > terminalActivation, - supersedingStartedSequence: event.sourceEventSequence - }) !== true - ) { - throw new Error( - `Smithers attempt ${identity} supersedes terminal event ${superseded.finishedSequence} before durable attempt recording` - ); - } - attempts.delete(identity); - terminalActivations.delete(identity); - } + attempts.delete(identity); active.set(identity, { retry, iteration, @@ -5341,15 +5260,10 @@ function terminalWorkflowAttempts( ...(terminal.failureCategory === undefined ? {} : { failureCategory: terminal.failureCategory }), ...(terminal.failureMessage === undefined ? {} : { failureMessage: terminal.failureMessage }) }); - terminalActivations.set(identity, activation); } return [...attempts.values()].sort((left, right) => left.finishedSequence - right.finishedSequence); } -function recordedTerminalAttemptSequences(layout: RunLayout, workflowRunId: string): ReadonlySet { - return recordedTerminalAttemptSequencesFromEntries(replayNodeAttempts(layout).entries, workflowRunId); -} - function recordedTerminalAttemptSequencesFromEntries( entries: readonly NodeAttemptLedgerEntry[], workflowRunId: string @@ -5359,318 +5273,6 @@ function recordedTerminalAttemptSequencesFromEntries( ); } -function supersededSuccessfulAttemptHasTraceAuthority( - layout: RunLayout, - workflowRunId: string, - task: StoredWorkflowTask, - attempt: TerminalWorkflowAttempt, - events: readonly WorkflowEvent[], - context: TerminalAttemptSupersessionContext -): boolean { - if (attempt.outcome !== "succeeded") return false; - if (!successfulAttemptHasExactTraceAuthority(workflowRunId, task, attempt, events)) return false; - - const current = readRunState(layout).nodes[task.attemptId]; - const manifestPath = path.join(getNodeArtifactDir(layout, task.attemptId), ARTIFACT_MANIFEST_FILE); - let manifestAuthority: - { status: "missing" } | { status: "invalid" } | { status: "valid"; manifest: ArtifactManifest; digest: string }; - try { - const stat = fs.lstatSync(manifestPath); - if (!stat.isFile()) { - manifestAuthority = { status: "invalid" }; - } else { - const bytes = readRegularFileSnapshot(manifestPath, MAX_REFERENCE_ARTIFACT_MANIFEST_AUTHORITY_BYTES); - try { - const value = parseStrictJsonBytes(bytes, { - maxBytes: MAX_REFERENCE_ARTIFACT_MANIFEST_AUTHORITY_BYTES - }); - const validation = validateArtifactManifest(value); - manifestAuthority = validation.ok - ? { status: "valid", manifest: value as ArtifactManifest, digest: sha256Bytes(bytes) } - : { status: "invalid" }; - } catch { - manifestAuthority = { status: "invalid" }; - } - } - } catch (error) { - if ( - typeof error !== "object" || - error === null || - !("code" in error) || - (error as { code?: unknown }).code !== "ENOENT" - ) { - throw error; - } - manifestAuthority = { status: "missing" }; - } - - if (manifestAuthority.status === "missing") { - // A failed task-output disposition is the controller's immutable rejection - // of this otherwise successful executor occurrence. It is exactly the - // state retry-failed may replace at a later run activation. - return !immutableTerminalFinalization(current) || current?.status === "failed"; - } - if (manifestAuthority.status === "invalid") return false; - return currentPublishedReplacementOccurrenceHasAuthority({ - layout, - workflowRunId, - task, - attempt, - events, - context, - current, - manifest: manifestAuthority.manifest, - manifestDigest: manifestAuthority.digest - }); -} - -function successfulAttemptHasExactTraceAuthority( - workflowRunId: string, - task: StoredWorkflowTask, - attempt: TerminalWorkflowAttempt, - events: readonly WorkflowEvent[] -): boolean { - const summaries = events.filter( - (event) => - event.type === "AgentTraceSummary" && - event.sourceEventSequence > attempt.startedSequence && - event.sourceEventSequence < attempt.finishedSequence && - event.workflowRunId === workflowRunId && - event.payload.nodeId === attempt.nodeId && - event.payload.iteration === attempt.iteration && - event.payload.attempt === attempt.retry - ); - if (summaries.length !== 1) return false; - const summaryEvent = summaries[0]!; - const summary = recordField(summaryEvent.payload, "summary"); - if ( - summary?.runId !== workflowRunId || - summary.nodeId !== attempt.nodeId || - summary.iteration !== attempt.iteration || - summary.attempt !== attempt.retry - ) { - return false; - } - const traceStartedAtMs = numberField(summary, "traceStartedAtMs"); - const traceFinishedAtMs = numberField(summary, "traceFinishedAtMs"); - if ( - traceStartedAtMs === undefined || - traceFinishedAtMs === undefined || - traceFinishedAtMs !== summaryEvent.timestampMs || - traceStartedAtMs < Date.parse(attempt.startedAt) || - traceFinishedAtMs > Date.parse(attempt.finishedAt) - ) { - return false; - } - const agentId = stringField(summary, "agentId"); - const model = stringField(summary, "model"); - if (agentId === undefined || model === undefined) return false; - const selections = task.agentChain - .map((profile, chainIndex) => ({ profile, chainIndex })) - .filter(({ chainIndex }) => smithersTaskAgentId(task, chainIndex) === agentId); - if (selections.length !== 1) return false; - const profile = selections[0]!.profile; - return profile.modelName === undefined || profile.modelName === model; -} - -function currentPublishedReplacementOccurrenceHasAuthority(input: { - layout: RunLayout; - workflowRunId: string; - task: StoredWorkflowTask; - attempt: TerminalWorkflowAttempt; - events: readonly WorkflowEvent[]; - context: TerminalAttemptSupersessionContext; - current: NodeState | undefined; - manifest: ArtifactManifest; - manifestDigest: string; -}): boolean { - if (!input.context.crossedRunActivation || input.current?.status !== "succeeded") return false; - const workflow = recordField(input.current.provenance, "workflow"); - const verifierAttempt = numberField(workflow, "attempt"); - if ( - workflow?.run_id !== input.workflowRunId || - workflow.task_id !== input.task.verifierSmithersNodeId || - workflow.agent_task_id !== input.task.smithersNodeId || - workflow.verifier_task_id !== input.task.verifierSmithersNodeId || - workflow.state !== "finished" || - verifierAttempt === undefined - ) { - return false; - } - - const supersedingStarts = input.events.filter( - (event) => - event.sourceEventSequence === input.context.supersedingStartedSequence && - event.type === "NodeStarted" && - event.payload.nodeId === input.attempt.nodeId && - event.payload.iteration === input.attempt.iteration && - event.payload.attempt === input.attempt.retry - ); - if (supersedingStarts.length !== 1) return false; - const supersedingStart = supersedingStarts[0]!; - const verifierIteration = input.task.metadata.loop.index; - const verifierStarts = input.events.filter( - (event) => - event.sourceEventSequence > supersedingStart.sourceEventSequence && - event.type === "NodeStarted" && - event.payload.nodeId === input.task.verifierSmithersNodeId && - event.payload.iteration === verifierIteration && - event.payload.attempt === verifierAttempt && - new Date(event.timestampMs).toISOString() === input.current?.started_at - ); - if (verifierStarts.length !== 1) return false; - const verifierStart = verifierStarts[0]!; - const verifierTerminals = input.events.filter( - (event) => - event.sourceEventSequence > verifierStart.sourceEventSequence && - event.type === "NodeFinished" && - event.payload.nodeId === input.task.verifierSmithersNodeId && - event.payload.iteration === verifierIteration && - event.payload.attempt === verifierAttempt && - new Date(event.timestampMs).toISOString() === input.current?.finished_at - ); - if (verifierTerminals.length !== 1) return false; - const verifierTerminal = verifierTerminals[0]!; - - // The occurrence that first reuses a historical attempt identity may itself - // be abandoned at the next activation. Bind authority to the final producer - // occurrence before the exact published verifier instead of assuming that - // the reused occurrence must be the publisher. Every intervening abandoned - // producer must cross a RunStarted boundary without a terminal event; this - // keeps overlapping starts and unrelated completed occurrences fail-closed. - const producerStarts = input.events - .filter( - (event) => - event.sourceEventSequence >= supersedingStart.sourceEventSequence && - event.sourceEventSequence < verifierStart.sourceEventSequence && - event.type === "NodeStarted" && - event.payload.nodeId === input.attempt.nodeId && - event.payload.iteration === input.attempt.iteration - ) - .sort((left, right) => left.sourceEventSequence - right.sourceEventSequence); - const replacementStart = producerStarts.at(-1); - if (replacementStart === undefined) return false; - let activeProducerAttempt: number | undefined = input.attempt.retry; - for (const event of input.events) { - if ( - event.sourceEventSequence <= supersedingStart.sourceEventSequence || - event.sourceEventSequence >= replacementStart.sourceEventSequence - ) { - continue; - } - if (event.type === "RunStarted") { - activeProducerAttempt = undefined; - continue; - } - if (event.payload.nodeId !== input.attempt.nodeId || event.payload.iteration !== input.attempt.iteration) { - continue; - } - if (event.type === "NodeStarted") { - if (activeProducerAttempt !== undefined) return false; - const retry = numberField(event.payload, "attempt"); - if (retry === undefined || !Number.isSafeInteger(retry) || retry < input.attempt.retry) return false; - activeProducerAttempt = retry; - continue; - } - if (event.type === "NodeFinished" || event.type === "NodeFailed") return false; - } - if (replacementStart !== supersedingStart && activeProducerAttempt !== undefined) return false; - const replacementRetry = numberField(replacementStart.payload, "attempt"); - if ( - replacementRetry === undefined || - !Number.isSafeInteger(replacementRetry) || - replacementRetry < input.attempt.retry - ) { - return false; - } - const nextBoundarySequence = input.events - .filter( - (event) => - event.sourceEventSequence > replacementStart.sourceEventSequence && - (event.type === "RunStarted" || - (event.type === "NodeStarted" && - event.payload.nodeId === input.attempt.nodeId && - event.payload.iteration === input.attempt.iteration)) - ) - .map((event) => event.sourceEventSequence) - .sort((left, right) => left - right)[0]; - const replacementTerminals = input.events.filter( - (event) => - event.sourceEventSequence > replacementStart.sourceEventSequence && - (nextBoundarySequence === undefined || event.sourceEventSequence < nextBoundarySequence) && - (event.type === "NodeFinished" || event.type === "NodeFailed") && - event.payload.nodeId === input.attempt.nodeId && - event.payload.iteration === input.attempt.iteration && - event.payload.attempt === replacementRetry - ); - if (replacementTerminals.length !== 1 || replacementTerminals[0]!.type !== "NodeFinished") return false; - const replacementTerminal = replacementTerminals[0]!; - const replacementAttempt: TerminalWorkflowAttempt = { - retry: replacementRetry, - iteration: input.attempt.iteration, - nodeId: input.attempt.nodeId, - startedSequence: replacementStart.sourceEventSequence, - finishedSequence: replacementTerminal.sourceEventSequence, - startedAt: new Date(replacementStart.timestampMs).toISOString(), - finishedAt: new Date(replacementTerminal.timestampMs).toISOString(), - outcome: "succeeded" - }; - if (!successfulAttemptHasExactTraceAuthority(input.workflowRunId, input.task, replacementAttempt, input.events)) { - return false; - } - - if (replacementTerminal.sourceEventSequence >= verifierStart.sourceEventSequence) return false; - // A finished producer may leave its verifier pending until a later run - // activation. Only a new producer occurrence before that verifier, or a - // boundary/restart inside the verifier occurrence itself, breaks the link. - if ( - input.events.some( - (event) => - event.sourceEventSequence > replacementTerminal.sourceEventSequence && - event.sourceEventSequence < verifierTerminal.sourceEventSequence && - event.type === "NodeStarted" && - event.payload.nodeId === input.attempt.nodeId && - event.payload.iteration === input.attempt.iteration - ) || - input.events.some( - (event) => - event.sourceEventSequence > verifierStart.sourceEventSequence && - event.sourceEventSequence < verifierTerminal.sourceEventSequence && - (event.type === "RunStarted" || - (event.type === "NodeStarted" && - event.payload.nodeId === input.task.verifierSmithersNodeId && - event.payload.iteration === verifierIteration)) - ) - ) { - return false; - } - - const outputContracts = recordField(input.current.provenance, "output_contracts"); - if ( - outputContracts?.ok !== true || - !Array.isArray(outputContracts.missing) || - outputContracts.missing.length !== 0 || - outputContracts.artifact_manifest_sha256 !== input.manifestDigest - ) { - return false; - } - const markerDigest = input.manifest.provenance.verification_marker_sha256; - if (markerDigest === undefined) return false; - const expectedProvenance = JSON.parse( - JSON.stringify({ - run_id: input.layout.runId, - ...artifactProvenance(input.task, input.workflowRunId, markerDigest) - }) - ) as ArtifactProvenance; - return ( - input.manifest.run_id === input.layout.runId && - input.manifest.node_id === input.task.attemptId && - input.manifest.producer_node_id === (input.task.metadata.node.producerNodeId ?? input.task.attemptId) && - Date.parse(input.manifest.created_at) >= Date.parse(input.current.finished_at ?? "") && - isDeepStrictEqual(input.manifest.provenance, expectedProvenance) - ); -} - function terminalOutcomeForEvent(event: WorkflowEvent): | { outcome: NodeAttemptOutcome; diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..429fd181e 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -22851,61 +22851,7 @@ test("syncRun preserves a recorded terminal occurrence when Smithers reuses its ); }); -test("syncRun accepts a superseded unadmitted success with exact sealed trace authority", async () => { - const project = tempProject(); - initProject({ projectRoot: project, force: true }); - writeSmallTopology(project); - const runId = "sync-traced-reused-attempt"; - const workflowRunId = `ultrafuzz-${runId}`; - const nodeId = "node:project-discovery"; - const base = Date.parse("2026-07-03T00:00:00.000Z"); - const env = fakeLifecycleSmithersEnv(project, { - inspect: workflowInspect({ - workflowRunId, - status: "running", - state: "running", - steps: [{ id: nodeId, state: "in-progress", attempt: 1 }] - }), - events: workflowEvents(workflowRunId, [ - { type: "NodeStarted", nodeId, attempt: 1 }, - { - type: "AgentTraceSummary", - nodeId, - extra: { - iteration: 0, - attempt: 1, - summary: { - runId: workflowRunId, - nodeId, - iteration: 0, - attempt: 1, - traceStartedAtMs: base + 50, - traceFinishedAtMs: base + 100, - agentId: "ultrafuzz-agent:project-discovery:0:default", - model: "gpt-5.5" - } - } - }, - { type: "NodeFinished", nodeId, attempt: 1 }, - { type: "RunStarted" }, - { type: "NodeStarted", nodeId, attempt: 1 } - ]) - }); - const run = await startRun({ projectRoot: project, runId, env }); - assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - - const sync = await syncRun({ projectRoot: project, runId, env }); - - assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); - assert.equal(sync.value?.status, "running"); - assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); - assert.equal( - fs.existsSync(path.join(run.value!.run_root, "artifacts", "project-discovery", "artifact-manifest.json")), - false - ); -}); - -test("syncRun accepts exact historical trace authority after an immutable output-validation failure", async () => { +test("syncRun keeps an immutable output-validation failure when its successful occurrence is superseded", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23150,172 +23096,13 @@ test("syncRun preserves a published replacement verified under a later activatio const rejected = await syncRun({ projectRoot: project, runId, env }); - assert.equal(rejected.ok, false); - assert.deepEqual( - rejected.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(rejected.diagnostics[0]?.message ?? "", /before durable attempt recording/u); - assert.deepEqual(fs.readFileSync(manifestPath), manifestBeforeReplay); - assert.deepEqual(readRunState(layout).nodes["project-discovery"], taskStateBeforeReplay); - assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), ledgerBeforeReplay); -}); - -test("syncRun binds an abandoned reused attempt to the final published higher retry", async () => { - const project = tempProject(); - initProject({ projectRoot: project, force: true }); - writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); - const runId = "sync-published-after-abandoned-reuse"; - const workflowRunId = `ultrafuzz-${runId}`; - const nodeId = "node:project-discovery"; - const verifierNodeId = "verify:project-discovery"; - const base = Date.parse("2026-07-03T00:00:00.000Z"); - const publishedEvents: Parameters[1] = [ - { type: "RunStarted", sequence: 6, timestampMs: base + 600 }, - { type: "NodeStarted", nodeId, attempt: 3, sequence: 7, timestampMs: base + 700 }, - { - type: "AgentTraceSummary", - nodeId, - sequence: 8, - timestampMs: base + 800, - extra: { - iteration: 0, - attempt: 3, - summary: { - runId: workflowRunId, - nodeId, - iteration: 0, - attempt: 3, - traceStartedAtMs: base + 750, - traceFinishedAtMs: base + 800, - agentId: "ultrafuzz-agent:project-discovery:0:default", - model: "gpt-5.5" - } - } - }, - { type: "NodeFinished", nodeId, attempt: 3, sequence: 9, timestampMs: base + 900 }, - { type: "NodeStarted", nodeId: verifierNodeId, attempt: 1, sequence: 10, timestampMs: base + 1_000 }, - { type: "NodeFinished", nodeId: verifierNodeId, attempt: 1, sequence: 11, timestampMs: base + 1_100 }, - { type: "RunFinished", sequence: 12, timestampMs: base + 1_200 } - ]; - const env = fakeLifecycleSmithersEnv(project, { - inspect: workflowInspect({ - workflowRunId, - steps: [ - { id: nodeId, state: "finished", attempt: 3 }, - { id: verifierNodeId, state: "finished", attempt: 1 } - ] - }), - events: workflowEvents(workflowRunId, publishedEvents) - }); - const run = await startRun({ projectRoot: project, runId, env }); - assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); - const published = await syncRun({ projectRoot: project, runId, env }); - assert.equal(published.ok, true, JSON.stringify(published.diagnostics)); - assert.equal(published.value?.status, "succeeded", JSON.stringify(published.diagnostics)); - const layout = layoutForRunRoot(run.value!.run_root, runId); - const manifestPath = path.join(layout.artifactsDir, "project-discovery", "artifact-manifest.json"); - const manifestBeforeReplay = fs.readFileSync(manifestPath); - const taskStateBeforeReplay = structuredClone(readRunState(layout).nodes["project-discovery"]); - const ledgerBeforeReplay = fs.readFileSync(layout.attemptLedgerPath); - - const historicalEvents: Parameters[1] = [ - { type: "RunStarted", sequence: 0, timestampMs: base }, - { type: "NodeStarted", nodeId, attempt: 2, sequence: 1, timestampMs: base + 100 }, - { - type: "AgentTraceSummary", - nodeId, - sequence: 2, - timestampMs: base + 200, - extra: { - iteration: 0, - attempt: 2, - summary: { - runId: workflowRunId, - nodeId, - iteration: 0, - attempt: 2, - traceStartedAtMs: base + 150, - traceFinishedAtMs: base + 200, - agentId: "ultrafuzz-agent:project-discovery:0:default", - model: "gpt-5.5" - } - } - }, - { type: "NodeFinished", nodeId, attempt: 2, sequence: 3, timestampMs: base + 300 }, - { type: "RunStarted", sequence: 4, timestampMs: base + 400 }, - { type: "NodeStarted", nodeId, attempt: 2, sequence: 5, timestampMs: base + 500 }, - ...publishedEvents - ]; - fs.writeFileSync(env.SMITHERS_FAKE_EVENTS!, workflowEvents(workflowRunId, historicalEvents), "utf8"); - - const replayed = await syncRun({ projectRoot: project, runId, env }); - - assert.equal(replayed.ok, true, JSON.stringify(replayed.diagnostics)); - assert.equal(replayed.value?.status, "succeeded"); + assert.equal(rejected.ok, true, JSON.stringify(rejected.diagnostics)); assert.deepEqual(fs.readFileSync(manifestPath), manifestBeforeReplay); assert.deepEqual(readRunState(layout).nodes["project-discovery"], taskStateBeforeReplay); assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), ledgerBeforeReplay); - - const expectRejected = async (events: Parameters[1]): Promise => { - fs.writeFileSync(env.SMITHERS_FAKE_EVENTS!, workflowEvents(workflowRunId, events), "utf8"); - const rejected = await syncRun({ projectRoot: project, runId, env }); - assert.equal(rejected.ok, false); - assert.deepEqual( - rejected.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(rejected.diagnostics[0]?.message ?? "", /before durable attempt recording/u); - assert.deepEqual(fs.readFileSync(manifestPath), manifestBeforeReplay); - assert.deepEqual(readRunState(layout).nodes["project-discovery"], taskStateBeforeReplay); - assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), ledgerBeforeReplay); - }; - - await expectRejected(historicalEvents.filter((event) => event.sequence !== 6)); - - const terminalBeforeBoundary = historicalEvents.map((event) => - event.sequence !== undefined && event.sequence >= 6 ? { ...event, sequence: event.sequence + 1 } : event - ); - terminalBeforeBoundary.push({ - type: "NodeFailed", - nodeId, - attempt: 2, - sequence: 6, - timestampMs: base + 550, - error: { message: "intervening occurrence terminated" } - }); - terminalBeforeBoundary.sort((left, right) => (left.sequence ?? 0) - (right.sequence ?? 0)); - await expectRejected(terminalBeforeBoundary); - - const ambiguousReplacementTrace = historicalEvents.map((event) => - event.sequence !== undefined && event.sequence >= 9 ? { ...event, sequence: event.sequence + 1 } : event - ); - ambiguousReplacementTrace.push({ - type: "AgentTraceSummary", - nodeId, - sequence: 9, - timestampMs: base + 850, - extra: { - iteration: 0, - attempt: 3, - summary: { - runId: workflowRunId, - nodeId, - iteration: 0, - attempt: 3, - traceStartedAtMs: base + 825, - traceFinishedAtMs: base + 850, - agentId: "ultrafuzz-agent:project-discovery:0:default", - model: "gpt-5.5" - } - } - }); - ambiguousReplacementTrace.sort((left, right) => (left.sequence ?? 0) - (right.sequence ?? 0)); - await expectRejected(ambiguousReplacementTrace); }); -test("syncRun rejects trace-only supersession of an immutable successful publication", async () => { +test("syncRun keeps an immutable successful publication when its attempt identity is reused", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23370,15 +23157,11 @@ test("syncRun rejects trace-only supersession of an immutable successful publica const sync = await syncRun({ projectRoot: project, runId, env }); - assert.equal(sync.ok, false); - assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /before durable attempt recording/u); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); }); -test("syncRun rejects exact trace authority when an attempt identity is reused within one activation", async () => { +test("syncRun drops a superseded occurrence when an attempt identity is reused within one activation", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23423,16 +23206,12 @@ test("syncRun rejects exact trace authority when an attempt identity is reused w const sync = await syncRun({ projectRoot: project, runId, env }); - assert.equal(sync.ok, false); - assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /before durable attempt recording/u); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); }); -test("syncRun replays bounded production lifecycle history with ledgered and trace-authorized supersessions", async () => { +test("syncRun replays bounded production lifecycle history with ledgered and unrecorded supersessions", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23549,7 +23328,7 @@ test("syncRun replays bounded production lifecycle history with ledgered and tra assert.doesNotMatch(commandLog, /--raw|--type agent/u); }); -test("syncRun rejects a superseded unadmitted success without exact trace authority", async () => { +test("syncRun drops a superseded unadmitted success without trace evidence", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23574,15 +23353,11 @@ test("syncRun rejects a superseded unadmitted success without exact trace author const sync = await syncRun({ projectRoot: project, runId, env }); - assert.equal(sync.ok, false); - assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /before durable attempt recording/u); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); }); -test("syncRun rejects an unrecorded terminal occurrence superseded by a reused attempt number", async () => { +test("syncRun drops an unrecorded terminal occurrence superseded by a reused attempt number", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23613,12 +23388,8 @@ test("syncRun rejects an unrecorded terminal occurrence superseded by a reused a const sync = await syncRun({ projectRoot: project, runId, env }); - assert.equal(sync.ok, false); - assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /supersedes terminal event 1 before durable attempt recording/u); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); }); @@ -27325,3 +27096,143 @@ test("stopped reset checkpoint runs once and skips active, ordinary and committe await runSmithersLifecycleCommand({ ...input, resetNode: nodeId }); assert.equal(checkpoints, 0, "a committed reset only needs its pending continuation"); }); + +for (const historical of ["failed", "succeeded"] as const) { + test(`syncRun reconciles a finished workflow after an unrecorded ${historical} occurrence is reused`, async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = `sync-1099-${historical}`; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const verifierNodeId = "verify:project-discovery"; + const base = Date.parse("2026-07-03T00:00:00.000Z"); + // Occurrence 1 finishes (or fails) and is never observed by Ultrafuzz. A + // supported reset (Smithers timetravel/retry-task) then restarts numbering + // at attempt 1 and upserts the mutable attempt row, so only the replacement + // occurrence remains described by `smithers node`. + const events: Parameters[1] = [ + { type: "RunStarted", sequence: 0, timestampMs: base }, + { type: "NodeStarted", nodeId, attempt: 1, sequence: 1, timestampMs: base + 100 }, + historical === "failed" + ? { + type: "NodeFailed", + nodeId, + attempt: 1, + sequence: 2, + timestampMs: base + 200, + error: { message: "historical occurrence failed" } + } + : { type: "NodeFinished", nodeId, attempt: 1, sequence: 2, timestampMs: base + 200 }, + { type: "RunStarted", sequence: 3, timestampMs: base + 300 }, + { type: "NodeStarted", nodeId, attempt: 1, sequence: 4, timestampMs: base + 400 }, + { type: "NodeFinished", nodeId, attempt: 1, sequence: 5, timestampMs: base + 500 }, + { type: "NodeStarted", nodeId: verifierNodeId, attempt: 1, sequence: 6, timestampMs: base + 600 }, + { type: "NodeFinished", nodeId: verifierNodeId, attempt: 1, sequence: 7, timestampMs: base + 700 }, + { type: "RunFinished", sequence: 8, timestampMs: base + 800 } + ]; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [ + { id: nodeId, state: "finished", attempt: 1 }, + { id: verifierNodeId, state: "finished", attempt: 1 } + ] + }), + events: workflowEvents(workflowRunId, events), + nodeDetails: { + [nodeId]: { + node: { nodeId, lastAttempt: 1 }, + attempts: [ + { + nodeId, + attempt: 1, + state: "finished", + meta: { + agentChainIndex: 0, + agentId: "ultrafuzz-agent:project-discovery:0:default", + agentModel: "gpt-5.5" + } + } + ] + } + } + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + const layout = layoutForRunRoot(run.value.run_root, runId); + + const first = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(first.ok, true, JSON.stringify(first.diagnostics)); + assert.ok( + !first.diagnostics.some((diagnostic) => diagnostic.severity === "error"), + JSON.stringify(first.diagnostics) + ); + assert.equal(first.value?.status, "succeeded", JSON.stringify(first.diagnostics)); + assert.equal(readRunState(layout).status, "succeeded"); + const ledger = fs.readFileSync(layout.attemptLedgerPath); + assert.deepEqual( + ledger + .toString("utf8") + .trim() + .split("\n") + .map((line) => (JSON.parse(line) as { source_event_sequence: number }).source_event_sequence), + [5] + ); + + const second = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(second.ok, true, JSON.stringify(second.diagnostics)); + assert.equal(second.value?.status, "succeeded"); + assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), ledger); + }); +} + +test("resume --retry-failed after a pre-agent failure keeps synchronizing the reused attempt", async () => { + const fixture = await unobservedFailedResetFixture("retry-pre-agent-failure-sync"); + // Smithers never selected an agent, so the failure is not a model attempt and + // the pre-reset checkpoint deliberately records nothing for it. + fs.writeFileSync( + path.join(fixture.detailRoot, `${fixture.nodeId}.json`), + JSON.stringify({ + node: { nodeId: fixture.nodeId, lastAttempt: 1 }, + attempts: [ + { + nodeId: fixture.nodeId, + attempt: 1, + state: "failed", + meta: { agentChainIndex: null, agentId: null, agentModel: null } + } + ] + }) + ); + const resumed = await resumeRun({ + projectRoot: fixture.project, + runId: fixture.runId, + retryFailed: true, + env: fixture.env + }); + assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); + assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), ""); + const activeEnv = fakeLifecycleSmithersEnv(fixture.project, { + inspect: workflowInspect({ + workflowRunId: fixture.workflowRunId, + status: "running", + state: "running", + steps: [{ id: fixture.nodeId, state: "in-progress", attempt: 1 }] + }), + events: workflowEvents(fixture.workflowRunId, [ + ...fixture.events, + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: fixture.nodeId, attempt: 1 } + ]) + }); + for (let observation = 0; observation < 2; observation += 1) { + const sync = await syncRun({ projectRoot: fixture.project, runId: fixture.runId, env: activeEnv }); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); + } +}); From 610011551e86c8bf8ced2537eb4bb7ba5cbf5fb7 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:25:50 +0000 Subject: [PATCH 080/206] docs: record the CI, lint baseline and smol-toml changes in the changelog Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..c91263a3d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,8 @@ ### Other changes +- **[ci] [cli] [docs]** Pull requests now run every release validation lane, including the package gates, the CLI suite, and the workspace typecheck, and the lanes start without waiting for the build job; the PR-only runtime and CLI smoke runners are removed. Pushes to `main` no longer cancel each other's runs, the lanes run under `eatmydata`, vitest packages default to a 30 s test timeout, and CLI tests delete their temporary projects when each test ends. `pnpm -w lint` now enforces the complexity and size budgets on all code, with existing violations counted per file and rule in `eslint-suppressions.json`: a new violation fails lint, and after one is fixed `pnpm -w lint:prune` lowers the recorded count (#960, #923). +- **[config]** `ultrafuzz init` and config loading no longer fail with "project must be a table" (and the same error for every other table) when `smol-toml` 1.9 is installed, as it is for a packed install without the workspace lockfile. That release builds parsed tables with `Object.create(null)`; the loader now reads them as ordinary objects. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From ff0d4ff650646b47a3f935e30af7486edb23ece5 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:32:11 +0000 Subject: [PATCH 081/206] ci: stop super-linter failing every edit to CHANGELOG.md on line length Super-Linter lints each changed Markdown file in full with its default markdownlint rules, including MD013 at 400 characters. Four of the 109 tracked Markdown files have longer lines: CHANGELOG.md keeps each entry on one line (44 such lines), and docs/reference/agent-adapter-boundaries.md, docs/explanation/provider-harness-research.md and docs/CODE_OF_CONDUCT.md have long table rows or paragraphs. Any pull request that touched one of them failed `External static analysis` on lines it did not change, and no merged pull request has edited CHANGELOG.md since Super-Linter was added. Use Super-Linter's default markdownlint rules from the pinned v8.7.0 template, with MD013 off, as .github/linters/.yamllint.yml already does for YAML line length. markdownlint-cli 0.49.0 (the version Super-Linter runs) reports 46 MD013 errors on this branch's changed Markdown with the default template and none with this file. Co-Authored-By: Claude Opus 5.5 --- .github/linters/.markdown-lint.yml | 14 ++++++++++++++ 1 file changed, 14 insertions(+) create mode 100644 .github/linters/.markdown-lint.yml diff --git a/.github/linters/.markdown-lint.yml b/.github/linters/.markdown-lint.yml new file mode 100644 index 000000000..08a64e071 --- /dev/null +++ b/.github/linters/.markdown-lint.yml @@ -0,0 +1,14 @@ +# Super-Linter's default markdownlint rules (v8.7.0 TEMPLATES/.markdown-lint.yml) +# with MD013 line length off. Super-Linter lints each changed file in full, and +# CHANGELOG.md keeps each entry on one line while docs tables have rows past +# 400 characters, so any edit to those files failed on lines nobody touched. +MD004: false +MD007: + indent: 2 +MD013: false +MD026: + punctuation: ".,;:!。,;:" +MD029: false +MD033: false +MD036: false +blank_lines: false From 08582375dd0e36a1ba769d6eba5041f8767e196c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:59:36 +0000 Subject: [PATCH 082/206] test(modal): stop pinning the package test script to an exact string The smoke-isolation test asserted `scripts.test === "vitest run"`, so adding --testTimeout=30000 failed package-gates in this pull request's first run. The property the test protects, that the unit-test script never runs the real Modal smoke, is still asserted by `not.toContain("smoke")` two lines later. Co-Authored-By: Claude Opus 5.5 --- packages/modal/test/smoke.test.ts | 1 - 1 file changed, 1 deletion(-) diff --git a/packages/modal/test/smoke.test.ts b/packages/modal/test/smoke.test.ts index c00634c73..5905cb3f9 100644 --- a/packages/modal/test/smoke.test.ts +++ b/packages/modal/test/smoke.test.ts @@ -128,7 +128,6 @@ describe("dedicated cloud command", () => { expect(cliSource.match(/import\("\.\/smoke-modal\.js"\)/gu)).toHaveLength(1); expect(cliSource).not.toMatch(/^import .*smoke-modal/mu); - expect(packageJson.scripts.test).toBe("vitest run"); expect(packageJson.scripts.smoke).toContain("dist/cli.js smoke"); expect(packageJson.scripts.typecheck).toBe( "pnpm --filter @ultrafuzz/modal^... build && tsc -p tsconfig.json --noEmit --pretty false" From f120561832f1f9f5e6750a79f6969142a80090d9 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:58:46 +0000 Subject: [PATCH 083/206] fix(runtime): record each Smithers attempt once and never let the ledger block sync attempts.jsonl was rebuilt from the whole event log on every pass and every anomaly failed closed (#1139): - terminalWorkflowAttempts threw on duplicate starts and on terminals without a start, and only knew NodeFinished/NodeFailed, so cancelled attempts were never recorded; - recorded rows were re-derived from mutable node status, finalization and the verifier's attempt number, then compared byte for byte, so any drift froze the task's ledger with "already recorded with different immutable data" on every later pass; - a finished attempt rejected host-side got category "unknown" unless a diagnostic code contained "ARTIFACT", which the in-sync join gate then rejected, so the attempt was never recorded; - the authority pre-pass ran `smithers node` for every terminal task on every pass and any failure returned ok:false before node or run state was reconciled. Pairing now never throws: a NodeStarted that reuses an attempt number marks the earlier occurrence superseded, a terminal without a live start in the same activation (including a NodeCancelled for an attempt that already ended, never started, or has `attempt: null`) is skipped, and NodeCancelled is a terminal with outcome/category canceled and the Smithers reason. Rows are keyed by (workflow_run_id, terminal sequence) and appended once; an existing identity is returned without re-deriving or comparing it, and reconcileNodeAttemptLedgerEntry plus the in-sync semantic-gate re-check are deleted. The current attempt is chosen by the agent task's own attempt number, and a host rejection of a finished attempt is failed with invalid-output or artifact-validation. `smithers node` runs only for unrecorded, unsuperseded local occurrences; a failed inspection, an unresolvable selection or a failed append is a warning, and a superseded occurrence is recorded without agent provenance because Smithers' row now describes its replacement. With superseded failures recorded from their own events, the #1117 pre-reset checkpoint is redundant: preserveFailedWorkflowAttemptsBeforeReset, its helpers and the beforeStoppedReset plumbing are deleted, so bookkeeping can no longer refuse `resume --retry-failed`/`--reset-node`. Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/src/attempt-ledger.ts | 62 +- packages/artifacts/test/artifacts.test.ts | 69 +- packages/runtime/src/smithers.ts | 16 - packages/runtime/src/start-run.ts | 31 - packages/runtime/src/workflow-sync.ts | 587 ++++----------- .../runtime/test/dynamic-lifecycle.test.ts | 2 +- packages/runtime/test/runtime.test.ts | 710 ++++++++++-------- 7 files changed, 558 insertions(+), 919 deletions(-) diff --git a/packages/artifacts/src/attempt-ledger.ts b/packages/artifacts/src/attempt-ledger.ts index 94c9ad677..692d140eb 100644 --- a/packages/artifacts/src/attempt-ledger.ts +++ b/packages/artifacts/src/attempt-ledger.ts @@ -1,7 +1,4 @@ -import { isDeepStrictEqual } from "node:util"; - import { - matchesRedactedText, redactSecretsInText, redactedTextSpanCodePointLengths, SENSITIVE_REDACTION_PLACEHOLDER @@ -629,46 +626,6 @@ export function appendNodeAttempt( return appendNodeAttempts(layout, [input])[0]!; } -/** Preserve an inserted redaction only when its exact spans and all surrounding evidence replay unchanged. */ -export function reconcileNodeAttemptLedgerEntry( - prior: NodeAttemptLedgerEntry, - candidate: NodeAttemptLedgerEntry, - replayEvidence: { failureMessage?: string } = {} -): NodeAttemptLedgerEntry | undefined { - if (isDeepStrictEqual(prior, candidate)) return prior; - - const { - failure_message: priorFailureMessage, - failure_message_redaction_span_code_points: priorRedactionSpanCodePoints, - failure_message_truncated: priorFailureMessageTruncated, - ...priorWithoutFailureMessage - } = prior; - const { - failure_message: candidateFailureMessage, - failure_message_redaction_span_code_points: _candidateRedactionSpanCodePoints, - failure_message_truncated: _candidateFailureMessageTruncated, - ...candidateWithoutFailureMessage - } = candidate; - const replayFailure = - replayEvidence.failureMessage === undefined || - replayEvidence.failureMessage.includes(SENSITIVE_REDACTION_PLACEHOLDER) - ? undefined - : normalizeNodeAttemptFailureText(replayEvidence.failureMessage); - if ( - typeof priorFailureMessage === "string" && - typeof candidateFailureMessage === "string" && - priorRedactionSpanCodePoints !== undefined && - replayFailure !== undefined && - isDeepStrictEqual(priorWithoutFailureMessage, candidateWithoutFailureMessage) && - matchesRedactedText(priorFailureMessage, replayFailure, priorRedactionSpanCodePoints, { - allowObservedSuffix: priorFailureMessageTruncated === true - }) - ) { - return prior; - } - return undefined; -} - export function appendNodeAttempts( layout: Pick, inputs: readonly AppendNodeAttemptInput[] @@ -679,18 +636,17 @@ export function appendNodeAttempts( const byIdentity = new Map(existing.map((entry) => [nodeAttemptLedgerIdentity(entry), entry])); const pending: NodeAttemptLedgerEntry[] = []; const results = inputs.map((input): AppendNodeAttemptResult => { + // A Smithers occurrence is recorded once: replaying its identity returns the + // stored entry without re-deriving or comparing it. + const prior = byIdentity.get( + nodeAttemptLedgerIdentity({ + workflow_run_id: input.workflowRunId, + source_event_sequence: input.sourceEventSequence + }) + ); + if (prior !== undefined) return { entry: prior, appended: false }; const candidate = createNodeAttemptLedgerEntry(layout, input); const identity = nodeAttemptLedgerIdentity(candidate); - const prior = byIdentity.get(identity); - if (prior !== undefined) { - const reconciled = reconcileNodeAttemptLedgerEntry(prior, candidate, { - failureMessage: input.failureMessage - }); - if (reconciled === undefined) { - throw new Error(`node attempt ${identity} was already recorded with different immutable data`); - } - return { entry: reconciled, appended: false }; - } byIdentity.set(identity, candidate); pending.push(candidate); return { entry: candidate, appended: true }; diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 390b021cc..043d28ab4 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -301,43 +301,33 @@ function failedAttemptInput(id: string, failureMessage: string) { }; } -test("node attempt ledger replay keeps a persisted failure redaction authoritative across credential rotation", () => { +test("node attempt ledger records an occurrence once and never re-derives it", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "attempt-ledger-redacted-replay" }); const oldCredential = "correct horse battery staple"; const input = failedAttemptInput("redacted-replay", `provider echoed ${oldCredential}`); const first = appendNodeAttempt(layout, { ...input, forbiddenSecretValues: [oldCredential] }); const persistedBytes = fs.readFileSync(layout.attemptLedgerPath); + assert.equal(first.appended, true); assert.equal(first.entry.failure_message, "provider echoed "); assert.deepEqual(first.entry.failure_message_redaction_span_code_points, [[...oldCredential].length]); assert.equal(first.entry.failure_message_truncated, undefined); - const replayed = appendNodeAttempt(layout, { - ...input, - forbiddenSecretValues: ["new credential value"] - }); - assert.equal(replayed.appended, false); - assert.deepEqual(replayed.entry, first.entry); - assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), persistedBytes); - - for (const failureMessage of [ - `different provider failure ${oldCredential}`, - `provider echoed ${oldCredential}; changed non-secret context` + // A replay of the same Smithers identity returns the stored entry whatever it + // would derive now: a rotated secret, a changed message, or a different outcome. + for (const replay of [ + { ...input, forbiddenSecretValues: ["new credential value"] }, + { ...input, failureMessage: `different provider failure ${oldCredential}` }, + { ...input, outcome: "canceled" as const, failureCategory: "canceled" as const, failureMessage: "run-cancelled" } ]) { - assert.throws( - () => - appendNodeAttempt(layout, { - ...input, - failureMessage, - forbiddenSecretValues: ["new credential value"] - }), - /already recorded with different immutable data/u - ); + const replayed = appendNodeAttempt(layout, replay); + assert.equal(replayed.appended, false); + assert.deepEqual(replayed.entry, first.entry); } assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), persistedBytes); }); -test("node attempt ledger uses explicit provenance for truncated redaction replay", () => { +test("node attempt ledger keeps truncation and redaction provenance for long failure messages", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "attempt-ledger-truncated-replay" }); const oldCredential = "old credential material ".repeat(30).trim(); const input = failedAttemptInput( @@ -346,42 +336,15 @@ test("node attempt ledger uses explicit provenance for truncated redaction repla ); const first = appendNodeAttempt(layout, { ...input, forbiddenSecretValues: [oldCredential] }); - const persistedBytes = fs.readFileSync(layout.attemptLedgerPath); assert.equal(Buffer.byteLength(first.entry.failure_message ?? "", "utf8"), MAX_NODE_ATTEMPT_FAILURE_MESSAGE_BYTES); assert.deepEqual(first.entry.failure_message_redaction_span_code_points, [[...oldCredential].length]); assert.equal(first.entry.failure_message_truncated, true); - const replayed = appendNodeAttempt(layout, { - ...input, - forbiddenSecretValues: ["rotated credential value"] + const literal = appendNodeAttempt(layout, { + ...failedAttemptInput("literal-placeholder", `provider echoed ; ${"x".repeat(1_500)}`) }); - assert.equal(replayed.appended, false); - assert.deepEqual(replayed.entry, first.entry); - assert.deepEqual(fs.readFileSync(layout.attemptLedgerPath), persistedBytes); - - assert.throws( - () => - appendNodeAttempt(layout, { - ...input, - failureMessage: `changed prefix ${oldCredential}; stable context ${"x".repeat(1_500)}`, - forbiddenSecretValues: ["rotated credential value"] - }), - /already recorded with different immutable data/u - ); -}); - -test("node attempt ledger does not infer replay-safe redaction from a literal placeholder", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "attempt-ledger-literal-placeholder" }); - const input = failedAttemptInput("literal-placeholder", `provider echoed ; ${"x".repeat(1_500)}`); - - const first = appendNodeAttempt(layout, input); - assert.equal(first.entry.failure_message_redaction_span_code_points, undefined); - assert.equal(first.entry.failure_message_truncated, true); - assert.throws( - () => - appendNodeAttempt(layout, { ...input, failureMessage: `provider echoed old credential; ${"x".repeat(1_500)}` }), - /already recorded with different immutable data/u - ); + assert.equal(literal.entry.failure_message_redaction_span_code_points, undefined); + assert.equal(literal.entry.failure_message_truncated, true); }); test("createRunLayout rejects symlinked run roots before creating outside writes", () => { diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..f653d2249 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -5209,8 +5209,6 @@ export async function runSmithersLifecycleCommand(input: { priorInspection?: SmithersResumeInspection; /** Prepare launch authority only after ruling out an idempotent active attach. */ prepareContinuationEnvironment?: () => Record; - /** Preserve stopped-run failure evidence before its mutable attempt rows are reset. */ - beforeStoppedReset?: (inspection: SmithersCommandSnapshot) => Promise; relaunchPaths?: { runRoot: string; inputJson?: string; @@ -5256,18 +5254,6 @@ export async function runSmithersLifecycleCommand(input: { let preResumeStderr = ""; let currentInspection: CurrentSmithersInspect | undefined; let inspection: SmithersCommandSnapshot | undefined; - let stoppedResetPreserved = false; - const preserveStoppedReset = async (): Promise => { - if ( - stoppedResetPreserved || - currentInspection === undefined || - inspection === undefined || - smithersRunStateIsActive(currentInspection) - ) - return; - await input.beforeStoppedReset?.(inspection); - stoppedResetPreserved = true; - }; // Detached admission renders the workflow before Smithers checks whether // this run already has an active owner. Inspect every resume first so an // idempotent attach cannot fail preflight or compete with that owner (#968). @@ -5336,7 +5322,6 @@ export async function runSmithersLifecycleCommand(input: { }) : undefined; if (failedTasks.length > 0) { - await preserveStoppedReset(); const resetStderr: string[] = []; for (const failedTask of failedTasks) { const producerTask = retryProducerForFailedVerifier(currentInspection, failedTask); @@ -5390,7 +5375,6 @@ export async function runSmithersLifecycleCommand(input: { : path.join(input.relaunchPaths.runRoot, "smithers", "reset-node-applied.json"); let resetStderr = ""; if (!resetNodeMarkerMatches(resetMarkerPath, input.smithersRunId, input.resetNode)) { - await preserveStoppedReset(); // A failure can be durable in the canonical node snapshot even when the // runner cannot resolve its implicit "latest attempt" lookup. Pinning the // iteration from that snapshot keeps --reset-node recoverable by node ID. diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index ba32e842b..bfee33e5f 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -623,22 +623,6 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { force: input.force, retryFailed: input.retryFailed, priorInspection: refreshInspection, - beforeStoppedReset: async (inspection) => { - if (!hasRetainedResetControlAuthority(projectRoot, layout, workflow)) return; - // Authenticated original v2 plans predate governance and the current - // linked-evidence contract; retain their native legacy reset behavior. - if (authenticatedContinuationGovernancePath(projectRoot, layout, workflow.control_generation) === undefined) - return; - // workflow-sync reads linked evidence through this module. Load the - // failure-only checkpoint after initialization, at an actual reset. - const { preserveFailedWorkflowAttemptsBeforeReset } = await import("./workflow-sync.js"); - await preserveFailedWorkflowAttemptsBeforeReset({ - projectRoot, - runId, - env: lifecycleEnvironment, - inspection - }); - }, relaunchPaths: { runRoot: layout.root, logsDir: path.join(smithersRoot, "logs") @@ -687,21 +671,6 @@ async function submitSmithersContinuation(input: WorkflowLifecycleInput) { } } -function hasRetainedResetControlAuthority( - projectRoot: string, - layout: RunLayout, - workflow: Record -): boolean { - if (Object.hasOwn(workflow, "control_generation") || Object.hasOwn(workflow, "control_integrity_path")) return true; - try { - fs.lstatSync(workflowControlPaths(projectRoot, layout).integrityPath); - return true; - } catch (error) { - if (error instanceof Error && "code" in error && error.code === "ENOENT") return false; - throw error; - } -} - function parseContinuationResolvedConfigBytes(bytes: Uint8Array): ResolvedConfig { try { return parseResolvedConfigJsonBytes(bytes); diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index 1ec3dc8dc..62b6277a6 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -18,7 +18,6 @@ import { assertPathInside, assertRegularFileInside, createNodeState, - createNodeAttemptLedgerEntry, createUsageLedgerEntry, getNodeArtifactDir, layoutForRunRoot, @@ -27,7 +26,6 @@ import { queryNodeAttempts, readRegularFileSnapshot, readRunMetadataDocument, - reconcileNodeAttemptLedgerEntry, replayEvents, replayNodeAttempts, readRunState, @@ -125,7 +123,7 @@ import { projectWorkflowControlState } from "./workflow-control.js"; import { runtimeSemanticGateDiagnostics } from "./semantic-gates.js"; import { inspectSmithersAttemptAgentSelection, - reconcileSmithersAttemptAgentSelection + type SmithersAttemptAgentSelection } from "./smithers-attempt-authority.js"; import { isRecord } from "@ultrafuzz/artifacts"; @@ -172,6 +170,11 @@ interface TerminalWorkflowAttempt { outcome: NodeAttemptOutcome; failureCategory?: NodeAttemptFailureCategory; failureMessage?: string; + /** + * A later NodeStarted reused this (node, iteration, attempt) numbering, as a + * reset does, so Smithers' attempt row now describes the replacement. + */ + superseded: boolean; } type SmithersNodeAttemptAuthorities = ReadonlyMap; @@ -324,6 +327,8 @@ interface AttemptWorkflowEvidence { evidence: NodeWorkflowEvidence; source: "agent" | "verifier" | "preparation"; taskId: string; + /** The agent task's own attempt number; a verifier numbers its attempts independently. */ + agentAttempt?: number; } interface ObservedTaskEvidence { @@ -587,7 +592,8 @@ export async function refinalizeControllerFailures( (attempt) => attempt.nodeId === task.verifierSmithersNodeId && attempt.retry === verifierAttempt && - attempt.outcome === "succeeded" + attempt.outcome === "succeeded" && + !attempt.superseded ); if (matchingAttempts.length !== 1) { throw new Error(`linked workflow does not contain one exact finished verifier attempt for ${task.attemptId}`); @@ -1031,229 +1037,6 @@ function parseFinishedVerifierOutput( }; } -/** - * Preserve failed local executor occurrences while their exact selected-agent - * detail still exists. Only a stopped-run reset calls this checkpoint; normal - * continuation, active force resets and success publication keep their existing - * paths. It never finalizes artifacts, projects run state or resolves pricing. - */ -export async function preserveFailedWorkflowAttemptsBeforeReset( - input: SyncRunInput & { inspection: SmithersCommandSnapshot } -): Promise { - const projectRoot = path.resolve(input.projectRoot); - const evidence = await readLinkedWorkflowEvidence(projectRoot, input.runId, { observeOnly: true }); - if (!evidence.ok) throw new Error(evidence.diagnostics.map((entry) => entry.message).join("; ")); - const loaded = loadSynchronizationInputs({ - graph: evidence.verifiedControl.contents.graph, - tasks: evidence.verifiedControl.contents.tasks - }); - if (!loaded.ok) throw new Error(loaded.diagnostics.map((entry) => entry.message).join("; ")); - const environment = linkedWorkflowExecutionEnvironment(evidence, input.env); - const inspect = parseCurrentSmithersInspect(input.inspection, evidence.smithersRunId); - const readEvents = async (): Promise => { - const snapshot = await runSmithersInspectionCommand({ - args: ["events", evidence.smithersRunId, "--limit", "100000", "--json"], - projectRoot, - env: environment - }); - if (!snapshot.ok) throw new Error(workflowSnapshotDiagnostic(snapshot, "WORKFLOW_EVENTS_FAILED").message); - return snapshot; - }; - const eventsSnapshot = await readEvents(); - const events = parseWorkflowEvents(eventsSnapshot.stdout, evidence.smithersRunId); - const allExisting = replayNodeAttempts(evidence.layout).entries; - const recordedSequences = recordedTerminalAttemptSequencesFromEntries(allExisting, evidence.smithersRunId); - const tasksByNodeId = new Map( - loaded.tasks.filter((task) => task.execution.mode === "local").map((task) => [task.smithersNodeId, task]) - ); - const failures = resetTerminalFailureAttempts({ tasksByNodeId, events }); - if (failures.length === 0) return; - const terminalSequences = new Set( - failures - .filter((attempt) => !recordedSequences.has(attempt.finishedSequence)) - .map((attempt) => attempt.finishedSequence) - ); - const authorities = await inspectTerminalAttemptAuthorities({ - projectRoot, - workflowRunId: evidence.smithersRunId, - tasks: loaded.tasks, - events, - layout: evidence.layout, - env: environment, - control: {}, - terminalSequences - }); - const forbiddenSecretValues = sensitiveEnvironmentValues( - input.env ?? process.env, - loaded.tasks.flatMap((task) => [ - ...task.execution.agentCredentialEnv, - ...(task.execution.modal?.credentialEnv ?? []) - ]) - ); - const pending = failedResetAttemptInputs({ - layout: evidence.layout, - workflowRunId: evidence.smithersRunId, - controlGeneration: evidence.controlGeneration, - tasksByNodeId, - failures, - authorities, - allExisting, - forbiddenSecretValues - }); - validateFailedResetAttemptInputs({ layout: evidence.layout, pending, allExisting, events }); - await assertStoppedResetAuthorityUnchanged({ - evidence, - projectRoot, - environment, - inspect, - events, - readEvents - }); - appendNodeAttempts(evidence.layout, pending); -} - -function resetTerminalFailureAttempts(input: { - tasksByNodeId: ReadonlyMap; - events: WorkflowEvent[]; -}): TerminalWorkflowAttempt[] { - const relevantEvents = input.events.filter( - (event) => event.type === "RunStarted" || input.tasksByNodeId.has(stringField(event.payload, "nodeId") ?? "") - ); - return terminalWorkflowAttempts(relevantEvents).filter( - (attempt) => attempt.outcome === "failed" || attempt.outcome === "timed-out" - ); -} - -function failedResetAttemptInputs(input: { - layout: RunLayout; - workflowRunId: string; - controlGeneration: string; - tasksByNodeId: ReadonlyMap; - failures: readonly TerminalWorkflowAttempt[]; - authorities: SmithersNodeAttemptAuthorities; - allExisting: readonly NodeAttemptLedgerEntry[]; - forbiddenSecretValues: readonly string[]; -}): AppendNodeAttemptInput[] { - const existing = new Map(input.allExisting.map((entry) => [nodeAttemptLedgerIdentity(entry), entry])); - return input.failures.flatMap((attempt): AppendNodeAttemptInput[] => { - const task = input.tasksByNodeId.get(attempt.nodeId); - if (task === undefined) throw new Error("failed reset attempt has no sealed task authority"); - const prior = existing.get(JSON.stringify([input.workflowRunId, attempt.finishedSequence])); - if ( - prior === undefined && - inspectSmithersAttemptAgentSelection( - task, - input.authorities.get(smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration)), - attempt.retry - ) === undefined - ) - return []; - return [ - { - workflowRunId: input.workflowRunId, - // Recompute immutable authority; only an already-recorded agent selection - // can replace mutable detail removed by an interrupted reset. - controlGeneration: input.controlGeneration, - nodeId: task.metadata.node.storageId ?? task.concreteNodeId, - strategyAttemptId: task.attemptId, - iteration: attempt.iteration, - attempt: attempt.retry, - startedEventSequence: attempt.startedSequence, - sourceEventSequence: attempt.finishedSequence, - startedAt: attempt.startedAt, - finishedAt: attempt.finishedAt, - outcome: attempt.outcome, - inputManifestDigest: taskAttemptInputManifestDigest(input.layout, task), - ...(prior === undefined - ? { agent: nodeAttemptAgentProvenance(task, attempt, input.authorities) } - : prior.agent === undefined - ? {} - : { agent: prior.agent }), - ...(attempt.failureCategory === undefined ? {} : { failureCategory: attempt.failureCategory }), - ...(attempt.failureMessage === undefined ? {} : { failureMessage: attempt.failureMessage }), - forbiddenSecretValues: input.forbiddenSecretValues - } - ]; - }); -} - -function validateFailedResetAttemptInputs(input: { - layout: RunLayout; - pending: readonly AppendNodeAttemptInput[]; - allExisting: readonly NodeAttemptLedgerEntry[]; - events: readonly WorkflowEvent[]; -}): void { - const existingByIdentity = new Map(input.allExisting.map((entry) => [nodeAttemptLedgerIdentity(entry), entry])); - const candidates = input.pending.map((entry) => { - const candidate = createNodeAttemptLedgerEntry(input.layout, entry); - const prior = existingByIdentity.get(nodeAttemptLedgerIdentity(candidate)); - if (prior === undefined) return candidate; - const reconciled = reconcileNodeAttemptLedgerEntry(prior, candidate, { failureMessage: entry.failureMessage }); - if (reconciled === undefined) - throw new Error("recorded failed attempt no longer matches its immutable event authority"); - return reconciled; - }); - const proposedEntries = [ - ...input.allExisting, - ...candidates.filter((entry) => !existingByIdentity.has(nodeAttemptLedgerIdentity(entry))) - ]; - for (const candidate of candidates) { - const diagnostics = runtimeSemanticGateDiagnostics({ - schemaFilename: "node-attempt-ledger.schema.json", - document: candidate, - artifactPath: input.layout.attemptLedgerPath, - context: { - attemptLedger: { entries: proposedEntries, sourceEntries: [] }, - eventLog: { - events: input.events.map((event) => ({ - workflow_run_id: event.workflowRunId, - source_event_sequence: event.sourceEventSequence, - timestamp_ms: event.timestampMs, - type: event.type, - payload: event.payload - })) - } - } - }); - const errors = diagnostics.filter((diagnostic) => diagnostic.severity === "error"); - if (errors.length > 0) throw new Error(errors.map((entry) => entry.message).join("; ")); - } -} - -async function assertStoppedResetAuthorityUnchanged(input: { - evidence: LinkedWorkflowEvidence; - projectRoot: string; - environment: Record; - inspect: CurrentSmithersInspect; - events: readonly WorkflowEvent[]; - readEvents: () => Promise; -}): Promise { - const currentEvents = await input.readEvents(); - const currentInspect = await runSmithersInspectionCommand({ - args: ["inspect", input.evidence.smithersRunId, "--format", "json", "--full-output"], - projectRoot: input.projectRoot, - env: input.environment - }); - if ( - !currentInspect.ok || - !isDeepStrictEqual(parseCurrentSmithersInspect(currentInspect, input.evidence.smithersRunId), input.inspect) || - !isDeepStrictEqual(parseWorkflowEvents(currentEvents.stdout, input.evidence.smithersRunId), input.events) - ) { - throw new Error("stopped workflow changed while preserving failed attempts before reset"); - } - const currentEvidence = await readLinkedWorkflowEvidence(input.projectRoot, input.evidence.layout.runId, { - observeOnly: true - }); - if ( - !currentEvidence.ok || - currentEvidence.smithersRunId !== input.evidence.smithersRunId || - currentEvidence.controlGeneration !== input.evidence.controlGeneration || - !currentEvidence.verifiedControl.contents.tasks.equals(input.evidence.verifiedControl.contents.tasks) - ) { - throw new Error("sealed workflow authority changed while preserving failed attempts before reset"); - } -} - export async function syncRun(input: SyncRunInput, control: WorkflowSynchronizationControl = {}) { const result = await synchronizeLinkedWorkflowRun(input, control); if (!result.ok) { @@ -1408,9 +1191,10 @@ export async function synchronizeLinkedWorkflowRun( diagnostics: [diagnosticFromError(error, "workflow", "WORKFLOW_TOKEN_EVENTS_INVALID")] }; } - let attemptAuthorities: SmithersNodeAttemptAuthorities; + // Attempt-ledger bookkeeping never blocks reconciling node and run state. + let attemptAuthorities: SmithersNodeAttemptAuthorities = new Map(); try { - attemptAuthorities = await inspectTerminalAttemptAuthorities({ + const inspected = await inspectTerminalAttemptAuthorities({ projectRoot, workflowRunId: evidence.smithersRunId, tasks: loaded.tasks, @@ -1419,13 +1203,15 @@ export async function synchronizeLinkedWorkflowRun( env: linkedWorkflowExecutionEnvironment(evidence, input.env), control }); + attemptAuthorities = inspected.authorities; + diagnostics.push(...inspected.diagnostics); } catch (error) { const interrupted = synchronizationInterruptionDiagnostic(error); if (interrupted !== undefined) return { ok: false, diagnostics: [interrupted] }; - return { - ok: false, - diagnostics: [diagnosticFromError(error, "workflow", "WORKFLOW_ATTEMPT_INSPECT_FAILED")] - }; + diagnostics.push({ + ...diagnosticFromError(error, "workflow", "WORKFLOW_ATTEMPT_INSPECT_FAILED"), + severity: "warning" + }); } let syncResult; try { @@ -1995,6 +1781,14 @@ function inspectionExecutionControl( }; } +/** + * Fetch `smithers node` detail for the local occurrences that still need agent + * provenance: those neither recorded nor superseded (a superseded occurrence's + * attempt row now describes its replacement). Synchronization therefore + * inspects only attempts it has not recorded; a pre-agent failure is never + * recorded, so it is inspected again on each pass. A failed inspection is a + * warning that defers the node's occurrences to a later pass. + */ async function inspectTerminalAttemptAuthorities(input: { projectRoot: string; workflowRunId: string; @@ -2003,44 +1797,29 @@ async function inspectTerminalAttemptAuthorities(input: { layout: RunLayout; env: Record; control: WorkflowSynchronizationControl; - terminalSequences?: ReadonlySet; -}): Promise { - const tasksByNodeId = new Map(input.tasks.map((task) => [task.smithersNodeId, task])); +}): Promise<{ authorities: SmithersNodeAttemptAuthorities; diagnostics: RuntimeDiagnostic[] }> { const localTaskNodeIds = new Set( input.tasks.filter((task) => task.execution.mode === "local").map((task) => task.smithersNodeId) ); - const relevantEvents = input.events.filter((event) => { - if (event.type === "RunStarted") return true; - const nodeId = stringField(event.payload, "nodeId"); - return nodeId !== undefined && localTaskNodeIds.has(nodeId); - }); - const pending = terminalWorkflowAttempts(relevantEvents, { tolerateMissingStarts: true }).filter( - (attempt) => - tasksByNodeId.get(attempt.nodeId)?.execution.mode === "local" && - (input.terminalSequences === undefined || input.terminalSequences.has(attempt.finishedSequence)) - ); - const grouped = new Map(); - for (const attempt of pending) { - const key = smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration); - const entries = grouped.get(key) ?? []; - entries.push(attempt); - grouped.set(key, entries); - } - if (grouped.size > input.tasks.length) { - throw new Error("Smithers attempt authority inspection exceeds the sealed task bound"); + const recorded = recordedTerminalAttemptSequences(replayNodeAttempts(input.layout).entries, input.workflowRunId); + const pending = new Map(); + for (const attempt of terminalWorkflowAttempts(input.events)) { + if (localTaskNodeIds.has(attempt.nodeId) && !attempt.superseded && !recorded.has(attempt.finishedSequence)) { + pending.set(smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration), attempt); + } } const authorities = new Map(); - for (const [key, attempts] of grouped) { + const diagnostics: RuntimeDiagnostic[] = []; + for (const [key, attempt] of pending) { assertSynchronizationBudget(input.control); - const first = attempts[0]!; const snapshot = await runSmithersInspectionCommand({ args: [ "node", - first.nodeId, + attempt.nodeId, "-r", input.workflowRunId, "-i", - String(first.iteration), + String(attempt.iteration), "--format", "json", "--full-output" @@ -2051,15 +1830,15 @@ async function inspectTerminalAttemptAuthorities(input: { }); assertSynchronizationBudget(input.control); if (!snapshot.ok || snapshot.json === undefined) { - throw new Error(workflowSnapshotDiagnostic(snapshot, "WORKFLOW_ATTEMPT_INSPECT_FAILED").message); - } - const task = tasksByNodeId.get(first.nodeId)!; - for (const attempt of attempts) { - inspectSmithersAttemptAgentSelection(task, snapshot.json, attempt.retry); + diagnostics.push({ + ...workflowSnapshotDiagnostic(snapshot, "WORKFLOW_ATTEMPT_INSPECT_FAILED"), + severity: "warning" + }); + continue; } authorities.set(key, snapshot.json); } - return authorities; + return { authorities, diagnostics }; } function smithersNodeAttemptAuthorityKey(nodeId: string, iteration: number): string { @@ -3983,16 +3762,21 @@ async function synchronizeTasks(input: { events: [...runActivationEvents, ...(eventsByNode.get(task.smithersNodeId) ?? [])].sort( (left, right) => left.sourceEventSequence - right.sourceEventSequence ), - currentAttempt: evidence.attempt, + currentAttempt: attemptEvidence.agentAttempt, currentStatus: patchStatus, finalization, attemptAuthorities: input.attemptAuthorities, forbiddenSecretValues: input.forbiddenSecretValues }); + diagnostics.push(...ledger.diagnostics); retryCount = Math.max(0, ledger.executedAttempts - (ledger.currentAttemptExecuted ? 1 : 0)); changed ||= ledger.appended; } catch (error) { - diagnostics.push(diagnosticFromError(error, "artifacts", "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + // The next pass retries the append; the node itself still synchronizes. + diagnostics.push({ + ...diagnosticFromError(error, "artifacts", "NODE_ATTEMPT_LEDGER_WRITE_FAILED"), + severity: "warning" + }); } if (previousIsImmutable) { // A successful publication and a terminal invalid-output disposition are @@ -4996,45 +4780,61 @@ function appendTerminalTaskAttempts(input: { workflowRunId: string; controlGeneration: string; events: WorkflowEvent[]; + /** The agent task's own Smithers attempt number, never its verifier's. */ currentAttempt?: number; currentStatus: NodeStatus; finalization: NodeFinalization; attemptAuthorities: SmithersNodeAttemptAuthorities; forbiddenSecretValues: readonly string[]; -}): { appended: boolean; executedAttempts: number; currentAttemptExecuted: boolean } { +}): { appended: boolean; executedAttempts: number; currentAttemptExecuted: boolean; diagnostics: RuntimeDiagnostic[] } { const allExisting = replayNodeAttempts(input.layout).entries; - const observedTerminalAttempts = terminalWorkflowAttempts(input.events); - const terminalAttempts = - input.task.execution.mode === "local" - ? observedTerminalAttempts.filter((attempt) => { - const detail = input.attemptAuthorities.get( - smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration) - ); - if (detail === undefined) { - throw new Error( - `Smithers attempt authority is unavailable for attempt ${String(attempt.retry)} of ${JSON.stringify(attempt.nodeId)}` - ); - } - return inspectSmithersAttemptAgentSelection(input.task, detail, attempt.retry) !== undefined; - }) - : observedTerminalAttempts; + const recordedByIdentity = new Map(allExisting.map((entry) => [nodeAttemptLedgerIdentity(entry), entry] as const)); + const terminalAttempts = terminalWorkflowAttempts(input.events); const state = readRunState(input.layout); const inputManifestDigest = taskAttemptInputManifestDigest(input.layout, input.task); const manifestPath = path.join(getNodeArtifactDir(input.layout, input.task.attemptId), "artifact-manifest.json"); const outputManifestDigest = fs.existsSync(manifestPath) ? sha256File(manifestPath) : undefined; const existing = allExisting.filter((entry) => entry.strategy_attempt_id === input.task.attemptId); - const existingByIdentity = new Map(allExisting.map((entry) => [nodeAttemptLedgerIdentity(entry), entry] as const)); const sourceEntries = sourceNodeAttempts(input.layout, state.source_run_id, input.task.attemptId); const currentTerminalAttempt = input.currentAttempt === undefined ? undefined - : terminalAttempts.filter((attempt) => attempt.retry === input.currentAttempt).at(-1); + : terminalAttempts.filter((attempt) => attempt.retry === input.currentAttempt && !attempt.superseded).at(-1); const pending: AppendNodeAttemptInput[] = []; - const candidates: NodeAttemptLedgerEntry[] = []; - const candidatesByIdentity = new Map(); + const diagnostics: RuntimeDiagnostic[] = []; let currentAttemptExecuted = false; for (const attempt of terminalAttempts) { const isCurrent = attempt === currentTerminalAttempt; + // An occurrence is recorded once and never re-derived, so later node + // status, finalization or manifest changes cannot turn it into a conflict. + const recorded = recordedByIdentity.get( + nodeAttemptLedgerIdentity({ + workflow_run_id: input.workflowRunId, + source_event_sequence: attempt.finishedSequence + }) + ); + if (recorded !== undefined) { + currentAttemptExecuted ||= isCurrent && recorded.reuse.status === "executed"; + continue; + } + let agent: NodeAttemptAgentProvenance | undefined; + if (input.task.execution.mode === "local" && !attempt.superseded) { + const detail = input.attemptAuthorities.get(smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration)); + // Inspection was unavailable this pass; a later pass retries the occurrence. + if (detail === undefined) continue; + try { + const selection = inspectSmithersAttemptAgentSelection(input.task, detail, attempt.retry); + // Smithers never selected an agent-chain rung, so no model ran. + if (selection === undefined) continue; + agent = nodeAttemptAgentProvenance(selection); + } catch (error) { + // Agent provenance is optional evidence: record the occurrence without it. + diagnostics.push({ + ...diagnosticFromError(error, "workflow", "WORKFLOW_ATTEMPT_INSPECT_FAILED"), + severity: "warning" + }); + } + } let outcome = attempt.outcome; let failureCategory = attempt.failureCategory; let failureMessage = attempt.failureMessage; @@ -5055,8 +4855,10 @@ function appendTerminalTaskAttempts(input: { // artifact-validation failure. The attempt ledger is immutable, so wait // for a later synchronization pass with the manifest instead of turning a // successful executor outcome into a permanent phantom failure (#352). - if (outcome === "succeeded" && outputDigest === undefined) continue; - if (isCurrent && ["failed", "timed-out", "canceled"].includes(outcome)) { + // A finished occurrence that a reset superseded before the host verified + // it has no output of its own left to record. + if (outcome === "succeeded" && (outputDigest === undefined || attempt.superseded)) continue; + if (isCurrent && (outcome === "failed" || outcome === "timed-out")) { failureMessage = input.finalization.lastError ?? failureMessage; } const reuseSource = @@ -5070,15 +4872,7 @@ function appendTerminalTaskAttempts(input: { if (reuseSource !== undefined && outputDigest === undefined) { outputDigest = reuseSource.outputManifestDigest; } - const reuse = - reuseSource === undefined - ? undefined - : { - status: "reused" as const, - sourceWorkflowRunId: reuseSource.workflowRunId, - sourceEventSequence: reuseSource.sourceEventSequence - }; - const appendInput: AppendNodeAttemptInput = { + pending.push({ workflowRunId: input.workflowRunId, controlGeneration: input.controlGeneration, nodeId: input.task.metadata.node.storageId ?? input.task.concreteNodeId, @@ -5091,96 +4885,35 @@ function appendTerminalTaskAttempts(input: { finishedAt: attempt.finishedAt, outcome, inputManifestDigest, - ...(input.task.execution.mode === "local" - ? { agent: nodeAttemptAgentProvenance(input.task, attempt, input.attemptAuthorities) } - : {}), + ...(agent === undefined ? {} : { agent }), ...(outputDigest === undefined ? {} : { outputManifestDigest: outputDigest }), - ...(reuse === undefined ? {} : { reuse }), + ...(reuseSource === undefined + ? {} + : { + reuse: { + status: "reused" as const, + sourceWorkflowRunId: reuseSource.workflowRunId, + sourceEventSequence: reuseSource.sourceEventSequence + } + }), ...(failureCategory === undefined ? {} : { failureCategory }), ...(failureMessage === undefined ? {} : { failureMessage }), forbiddenSecretValues: input.forbiddenSecretValues - }; - const candidate = createNodeAttemptLedgerEntry(input.layout, appendInput); - const identity = nodeAttemptLedgerIdentity(candidate); - const duplicateCandidate = candidatesByIdentity.get(identity); - if (duplicateCandidate !== undefined) { - if (!isDeepStrictEqual(duplicateCandidate, candidate)) { - throw new Error(`node attempt ${identity} appears with conflicting immutable data in the source snapshot`); - } - currentAttemptExecuted ||= isCurrent && duplicateCandidate.reuse.status === "executed"; - continue; - } - const recordedEntry = existingByIdentity.get(identity); - if (recordedEntry !== undefined) { - const reconciled = reconcileNodeAttemptLedgerEntry(recordedEntry, candidate, { - failureMessage - }); - if (reconciled === undefined) { - throw new Error(`node attempt ${identity} was already recorded with different immutable data`); - } - candidatesByIdentity.set(identity, reconciled); - candidates.push(reconciled); - currentAttemptExecuted ||= isCurrent && recordedEntry.reuse.status === "executed"; - continue; - } - candidatesByIdentity.set(identity, candidate); - candidates.push(candidate); - pending.push(appendInput); - currentAttemptExecuted ||= isCurrent && reuse?.status !== "reused"; - } - const proposedEntries = [ - ...allExisting, - ...candidates.filter((candidate) => !existingByIdentity.has(nodeAttemptLedgerIdentity(candidate))) - ]; - const gateContext = { - attemptLedger: { entries: proposedEntries, sourceEntries }, - eventLog: { - events: input.events.map((event) => ({ - workflow_run_id: event.workflowRunId, - source_event_sequence: event.sourceEventSequence, - timestamp_ms: event.timestampMs, - type: event.type, - payload: event.payload - })) - } - }; - for (const candidate of candidates) { - const diagnostics = runtimeSemanticGateDiagnostics({ - schemaFilename: "node-attempt-ledger.schema.json", - document: candidate, - artifactPath: input.layout.attemptLedgerPath, - context: gateContext }); - if (diagnostics.some((diagnostic) => diagnostic.severity === "error")) { - throw new Error( - diagnostics - .filter((diagnostic) => diagnostic.severity === "error") - .map((diagnostic) => diagnostic.message) - .join("; ") - ); - } + currentAttemptExecuted ||= isCurrent && reuseSource === undefined; } - const results = pending.length === 0 ? [] : appendNodeAttempts(input.layout, pending); - const allEntries = proposedEntries.filter((entry) => entry.strategy_attempt_id === input.task.attemptId); + const appended = appendNodeAttempts(input.layout, pending).filter((result) => result.appended); return { - appended: results.some((result) => result.appended), - executedAttempts: allEntries.filter((entry) => entry.reuse.status === "executed").length, - currentAttemptExecuted + appended: appended.length > 0, + executedAttempts: [...existing, ...appended.map((result) => result.entry)].filter( + (entry) => entry.reuse.status === "executed" + ).length, + currentAttemptExecuted, + diagnostics }; } -function nodeAttemptAgentProvenance( - task: StoredWorkflowTask, - attempt: Pick, - authorities: SmithersNodeAttemptAuthorities -): NodeAttemptAgentProvenance { - const detail = authorities.get(smithersNodeAttemptAuthorityKey(attempt.nodeId, attempt.iteration)); - if (detail === undefined) { - throw new Error( - `Smithers attempt authority is unavailable for attempt ${attempt.retry} of ${JSON.stringify(attempt.nodeId)}` - ); - } - const selection = reconcileSmithersAttemptAgentSelection(task, detail, attempt.retry); +function nodeAttemptAgentProvenance(selection: SmithersAttemptAgentSelection): NodeAttemptAgentProvenance { const agent = selection.profile; return { chain_index: selection.chainIndex, @@ -5196,75 +4929,68 @@ function nodeAttemptAgentProvenance( } /** - * Reduce the append-only Smithers event stream to the latest terminal - * occurrence of each (node, iteration, attempt) identity. + * Pair each Smithers attempt occurrence's NodeStarted with its terminal event. * - * A Smithers reset (timetravel / retry-task) restarts attempt numbering and - * upserts the mutable attempt row, so a later NodeStarted can reopen an - * identity that already has a terminal occurrence. Only the newest occurrence - * is still described by `smithers node`; an earlier one is either already in - * the immutable attempt ledger (keyed by its terminal event sequence) or is - * history nobody observed in time. Neither may block synchronizing the run's - * current state, so the superseded occurrence is simply dropped here (#1099). + * An occurrence is identified by its terminal event sequence, so it stays + * recordable after a reset (timetravel / retry-task) restarts the attempt + * numbering: the earlier occurrence is only marked `superseded`, because + * Smithers upserts the attempt row for the replacement (#1099). Unpairable + * events are skipped, never fatal. A start without a terminal was abandoned + * (Smithers cancels in-progress rows at the next RunStarted without an event), + * and a terminal without a live start in the same activation, such as a + * NodeCancelled for an attempt that already ended or never started, is not an + * occurrence (#1139). */ -function terminalWorkflowAttempts( - events: WorkflowEvent[], - options: { tolerateMissingStarts?: boolean } = {} -): TerminalWorkflowAttempt[] { +function terminalWorkflowAttempts(events: readonly WorkflowEvent[]): TerminalWorkflowAttempt[] { const active = new Map< string, Pick >(); - const attempts = new Map(); + const latestByIdentity = new Map(); + const attempts: TerminalWorkflowAttempt[] = []; for (const event of events) { if (event.type === "RunStarted") { - // Smithers cancels stale in-progress rows before each resumed activation, - // but does not emit NodeCancelled for those abandoned occurrences. Keep - // duplicate starts fail-closed within one activation while allowing the - // next activation to reuse the same durable attempt number. active.clear(); continue; } - if (event.type !== "NodeStarted" && event.type !== "NodeFinished" && event.type !== "NodeFailed") continue; - const payload = event.payload; - const nodeId = requiredWorkflowEventString(payload.nodeId, `${event.type} nodeId`); - const retry = requiredWorkflowEventCount(payload.attempt, `${event.type} attempt`); - const iteration = requiredWorkflowEventCount(payload.iteration, `${event.type} iteration`); + if (!["NodeStarted", "NodeFinished", "NodeFailed", "NodeCancelled"].includes(event.type)) continue; + const { nodeId, iteration, attempt: retry } = event.payload; + // A NodeCancelled for a node with no live attempt carries `attempt: null`. + if (typeof nodeId !== "string" || !Number.isSafeInteger(iteration) || !Number.isSafeInteger(retry)) continue; const identity = JSON.stringify([nodeId, iteration, retry]); const timestamp = new Date(event.timestampMs).toISOString(); if (event.type === "NodeStarted") { - if (active.has(identity)) throw new Error(`Smithers attempt ${identity} has multiple active NodeStarted events`); - attempts.delete(identity); + const previous = latestByIdentity.get(identity); + if (previous !== undefined) previous.superseded = true; active.set(identity, { - retry, - iteration, + retry: retry as number, + iteration: iteration as number, nodeId, startedSequence: event.sourceEventSequence, startedAt: timestamp }); continue; } - const terminal = terminalOutcomeForEvent(event); - if (terminal === undefined) continue; const started = active.get(identity); - if (started === undefined) { - if (options.tolerateMissingStarts === true) continue; - throw new Error(`Smithers terminal event has no preceding NodeStarted for ${identity}`); - } + const terminal = terminalOutcomeForEvent(event); + if (started === undefined || terminal === undefined) continue; active.delete(identity); - attempts.set(identity, { + const occurrence: TerminalWorkflowAttempt = { ...started, finishedSequence: event.sourceEventSequence, finishedAt: timestamp, outcome: terminal.outcome, ...(terminal.failureCategory === undefined ? {} : { failureCategory: terminal.failureCategory }), - ...(terminal.failureMessage === undefined ? {} : { failureMessage: terminal.failureMessage }) - }); + ...(terminal.failureMessage === undefined ? {} : { failureMessage: terminal.failureMessage }), + superseded: false + }; + latestByIdentity.set(identity, occurrence); + attempts.push(occurrence); } - return [...attempts.values()].sort((left, right) => left.finishedSequence - right.finishedSequence); + return attempts; } -function recordedTerminalAttemptSequencesFromEntries( +function recordedTerminalAttemptSequences( entries: readonly NodeAttemptLedgerEntry[], workflowRunId: string ): ReadonlySet { @@ -5289,6 +5015,11 @@ function terminalOutcomeForEvent(event: WorkflowEvent): ? { outcome: "timed-out", failureCategory: "timeout", ...(failureMessage ? { failureMessage } : {}) } : { outcome: "failed", failureCategory: "executor-error", ...(failureMessage ? { failureMessage } : {}) }; } + case "NodeCancelled": { + // Smithers names the cancellation: run-cancelled, unmounted or aborted. + const reason = stringField(event.payload, "reason"); + return { outcome: "canceled", failureCategory: "canceled", ...(reason ? { failureMessage: reason } : {}) }; + } default: return undefined; } @@ -5316,23 +5047,14 @@ function finalizationFailureCategory( if (outcome === "timed-out") { return "timeout"; } - if (outcome === "canceled") { - return "canceled"; - } if (outcome !== "failed") { return undefined; } - if (finalization.diagnostics.some((diagnostic) => diagnostic.code === "FINDINGS_VALIDATION_FAILED")) { - return "invalid-output"; - } - if ( - finalization.diagnostics.some( - (diagnostic) => diagnostic.code === "ARTIFACT_MANIFEST_WRITE_FAILED" || diagnostic.code.includes("ARTIFACT") - ) - ) { - return "artifact-validation"; - } - return "unknown"; + // Only a finished executor occurrence reaches this overlay, so its failed + // node was rejected host-side: by its verifier or by the artifact gates. + return finalization.diagnostics.some((diagnostic) => diagnostic.code === "FINDINGS_VALIDATION_FAILED") + ? "invalid-output" + : "artifact-validation"; } function reusedSourceAttempt(input: { @@ -5471,13 +5193,14 @@ function completionEvidenceForTask( if (agentEvidence === undefined) { return undefined; } + const agentAttempt = agentEvidence.attempt === undefined ? {} : { agentAttempt: agentEvidence.attempt }; if (agentEvidence.status !== "succeeded") { - return { evidence: agentEvidence, source: "agent", taskId: task.smithersNodeId }; + return { evidence: agentEvidence, source: "agent", taskId: task.smithersNodeId, ...agentAttempt }; } if (verifierEvidence === undefined) { return undefined; } - return { evidence: verifierEvidence, source: "verifier", taskId: task.verifierSmithersNodeId }; + return { evidence: verifierEvidence, source: "verifier", taskId: task.verifierSmithersNodeId, ...agentAttempt }; } function evidenceFromStep(step: WorkflowStep): NodeWorkflowEvidence { diff --git a/packages/runtime/test/dynamic-lifecycle.test.ts b/packages/runtime/test/dynamic-lifecycle.test.ts index 86e7679f8..8832b654d 100644 --- a/packages/runtime/test/dynamic-lifecycle.test.ts +++ b/packages/runtime/test/dynamic-lifecycle.test.ts @@ -1306,7 +1306,7 @@ test("dynamic child failure, skip, and timeout keep strict joins blocked with du const ledgerOutcome = readLedger(fixture).find( (entry) => entry.strategy_attempt_id === generated.attemptId )?.outcome; - // The current strict attempt ledger records only NodeFinished/NodeFailed terminal authorities. + // The attempt ledger records only NodeFinished/NodeFailed/NodeCancelled terminal events. // A skip or heartbeat timeout remains durable in state without inventing a ledger terminal // event that the workflow runner did not emit. assert.equal(ledgerOutcome, outcome === "failed" ? "failed" : undefined, outcome); diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 429fd181e..31930e6a7 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -2180,15 +2180,19 @@ function fakeLifecycleSmithersEnv( "utf8" ); fs.mkdirSync(nodeDetailsDirectory, { recursive: true }); - const terminalAttempts = new Map>(); + const terminalAttempts = new Map>(); + const terminalStates = { NodeFinished: "finished", NodeFailed: "failed", NodeCancelled: "cancelled" } as const; for (const line of (input.events ?? "").trim().split("\n").filter(Boolean)) { const event = JSON.parse(line) as { type?: unknown; payload?: Record }; - if (event.type !== "NodeFinished" && event.type !== "NodeFailed") continue; + if (event.type !== "NodeFinished" && event.type !== "NodeFailed" && event.type !== "NodeCancelled") continue; const nodeId = event.payload?.nodeId; const attempt = event.payload?.attempt; if (typeof nodeId !== "string" || !nodeId.startsWith("node:") || typeof attempt !== "number") continue; const rows = terminalAttempts.get(nodeId) ?? []; - rows.push({ attempt, state: event.type === "NodeFinished" ? "finished" : "failed" }); + // Smithers upserts one row per attempt number, so a reused number keeps its latest state. + const reused = rows.findIndex((row) => row.attempt === attempt); + if (reused !== -1) rows.splice(reused, 1); + rows.push({ attempt, state: terminalStates[event.type] }); terminalAttempts.set(nodeId, rows); } for (const [nodeId, attempts] of terminalAttempts) { @@ -22370,10 +22374,12 @@ test("syncRun keeps redacted failure state and attempt evidence stable across cr assert.doesNotMatch(fs.readFileSync(file, "utf8"), new RegExp(`${oldCredential}|${rotatedCredential}`, "u")); } + // A recorded occurrence is never re-derived, so even a different message for + // the same terminal event neither rewrites it nor reports a ledger conflict. fs.writeFileSync(lifecycleEnv.SMITHERS_FAKE_EVENTS!, failureEvents(`different provider failure ${oldCredential}`)); const changed = await synchronize(rotatedEnv); assert.equal(changed.ok, true, JSON.stringify(changed.diagnostics)); - assert.ok(changed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + assert.ok(!changed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); for (const [file, before] of snapshots) assert.deepEqual(fs.readFileSync(file), before); }); @@ -22523,10 +22529,10 @@ test("syncRun trusts non-ordinal Smithers selection when opaque profiles share a }); }); -test("syncRun fails closed when Smithers selection does not match the sealed chain", async () => { +test("syncRun records an attempt without agent provenance when Smithers selection does not match the sealed chain", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); - writeSmallTopology(project); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); const workflowRunId = "ultrafuzz-mismatched-selection"; const env = fakeLifecycleSmithersEnv(project, { inspect: workflowInspect({ @@ -22546,77 +22552,32 @@ test("syncRun fails closed when Smithers selection does not match the sealed cha }); const run = await startRun({ projectRoot: project, runId: "mismatched-selection", env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); const sync = await syncRun({ projectRoot: project, runId: "mismatched-selection", env }); - assert.equal(sync.ok, false); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "succeeded", JSON.stringify(sync.diagnostics)); + const warnings = sync.diagnostics.filter((diagnostic) => diagnostic.code === "WORKFLOW_ATTEMPT_INSPECT_FAILED"); + assert.equal(warnings.length, 1, JSON.stringify(sync.diagnostics)); + assert.equal(warnings[0]?.severity, "warning"); + assert.match(warnings[0]?.message ?? "", /agent ID does not match sealed chain rung 0/u); + const entries = fs + .readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8") + .trim() + .split("\n") + .map((line) => JSON.parse(line) as { outcome?: string; agent?: unknown }); assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /agent ID does not match sealed chain rung 0/u); - assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); -}); - -test("syncRun rejects forged provenance in an existing immutable attempt", async () => { - const project = tempProject(); - initProject({ projectRoot: project, force: true }); - writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); - const workflowRunId = "ultrafuzz-forged-recorded-provenance"; - const env = fakeLifecycleSmithersEnv(project, { - inspect: workflowInspect({ - workflowRunId, - steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] - }), - events: workflowEvents(workflowRunId, [ - { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, - { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, - { type: "RunFinished" } - ]) - }); - const run = await startRun({ projectRoot: project, runId: "forged-recorded-provenance", env }); - assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); - - const firstSync = await syncRun({ projectRoot: project, runId: "forged-recorded-provenance", env }); - assert.equal(firstSync.ok, true, JSON.stringify(firstSync.diagnostics)); - const ledgerPath = path.join(run.value!.run_root, "attempts.jsonl"); - const entry = JSON.parse(fs.readFileSync(ledgerPath, "utf8")) as Record; - entry.agent = { - chain_index: 0, - profile_id: "default", - agent_ref: "CodexAgent", - model_name: "forged-model", - reasoning_effort: "xhigh", - role: "primary", - selection: "observed" - }; - const forgedLedger = `${JSON.stringify(entry)}\n`; - fs.writeFileSync(ledgerPath, forgedLedger, "utf8"); - - const replayed = await syncRun({ projectRoot: project, runId: "forged-recorded-provenance", env }); - - assert.equal(replayed.ok, true, JSON.stringify(replayed.diagnostics)); - assert.ok(replayed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); - assert.match( - replayed.diagnostics.find((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")?.message ?? "", - /already recorded with different immutable data/u - ); - assert.equal(fs.readFileSync(ledgerPath, "utf8"), forgedLedger); - assert.equal( - fs - .readFileSync(env.SMITHERS_FAKE_LOG!, "utf8") - .split("\n") - .filter((line) => line.startsWith("node node:project-discovery ")).length, - 2 + entries.map((entry) => [entry.outcome, entry.agent]), + [["succeeded", undefined]] ); }); -test("syncRun rejects a legacy agentless immutable local attempt", async () => { +test("syncRun never re-derives or re-inspects a recorded attempt", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); - const workflowRunId = "ultrafuzz-agentless-recorded-provenance"; + const workflowRunId = "ultrafuzz-recorded-attempt-final"; const env = fakeLifecycleSmithersEnv(project, { inspect: workflowInspect({ workflowRunId, @@ -22628,33 +22589,33 @@ test("syncRun rejects a legacy agentless immutable local attempt", async () => { { type: "RunFinished" } ]) }); - const run = await startRun({ projectRoot: project, runId: "agentless-recorded-provenance", env }); + const run = await startRun({ projectRoot: project, runId: "recorded-attempt-final", env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); - const firstSync = await syncRun({ projectRoot: project, runId: "agentless-recorded-provenance", env }); + const firstSync = await syncRun({ projectRoot: project, runId: "recorded-attempt-final", env }); assert.equal(firstSync.ok, true, JSON.stringify(firstSync.diagnostics)); const ledgerPath = path.join(run.value!.run_root, "attempts.jsonl"); const entry = JSON.parse(fs.readFileSync(ledgerPath, "utf8")) as Record; + // A row that differs from what this build would derive now, such as one an + // older build wrote without agent provenance, stays exactly as recorded. delete entry.agent; - const legacyLedger = `${JSON.stringify(entry)}\n`; - fs.writeFileSync(ledgerPath, legacyLedger, "utf8"); + const recordedLedger = `${JSON.stringify(entry)}\n`; + fs.writeFileSync(ledgerPath, recordedLedger, "utf8"); - const replayed = await syncRun({ projectRoot: project, runId: "agentless-recorded-provenance", env }); + const replayed = await syncRun({ projectRoot: project, runId: "recorded-attempt-final", env }); assert.equal(replayed.ok, true, JSON.stringify(replayed.diagnostics)); - assert.ok(replayed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); - assert.match( - replayed.diagnostics.find((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")?.message ?? "", - /already recorded with different immutable data/u - ); - assert.equal(fs.readFileSync(ledgerPath, "utf8"), legacyLedger); + assert.ok(!replayed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + assert.equal(fs.readFileSync(ledgerPath, "utf8"), recordedLedger); + // Only the first pass needed Smithers attempt detail; a recorded attempt costs + // no runner subprocess on later passes. assert.equal( fs .readFileSync(env.SMITHERS_FAKE_LOG!, "utf8") .split("\n") .filter((line) => line.startsWith("node node:project-discovery ")).length, - 2 + 1 ); }); @@ -22972,8 +22933,20 @@ test("syncRun keeps an immutable output-validation failure when its successful o assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.ok(!sync.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_ATTEMPT_INSPECT_FAILED")); + assert.ok(!sync.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); assert.equal(fs.existsSync(path.join(layout.artifactsDir, "project-discovery", "artifact-manifest.json")), false); assert.match(fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"), /^node node:project-discovery /mu); + // The replacement executor finished but the host rejected its output, and the + // sealed disposition carries no finalization diagnostic to classify it by. + const entries = fs + .readFileSync(layout.attemptLedgerPath, "utf8") + .trim() + .split("\n") + .map((line) => JSON.parse(line) as Record); + assert.deepEqual( + entries.map((entry) => [entry.source_event_sequence, entry.outcome, entry.failure_category]), + [[7, "failed", "artifact-validation"]] + ); }); test("syncRun preserves a published replacement verified under a later activation", async () => { @@ -23357,7 +23330,7 @@ test("syncRun drops a superseded unadmitted success without trace evidence", asy assert.equal(sync.value?.status, "running"); }); -test("syncRun drops an unrecorded terminal occurrence superseded by a reused attempt number", async () => { +test("syncRun records an unrecorded failed occurrence superseded by a reused attempt number", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23390,10 +23363,27 @@ test("syncRun drops an unrecorded terminal occurrence superseded by a reused att assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.equal(sync.value?.status, "running"); - assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); + // Smithers' attempt row now describes the replacement, so the superseded + // failure is recorded from its own events, without agent provenance. + const entries = fs + .readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8") + .trim() + .split("\n") + .map((line) => JSON.parse(line) as Record); + assert.deepEqual( + entries.map((entry) => [ + entry.started_event_sequence, + entry.source_event_sequence, + entry.outcome, + entry.failure_message, + entry.agent + ]), + [[0, 1, "failed", "unrecorded occurrence", undefined]] + ); + assert.doesNotMatch(fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"), /^node /mu); }); -test("syncRun rejects duplicate active starts for one reused attempt identity", async () => { +test("syncRun tolerates duplicate active starts for one attempt identity", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project); @@ -23417,12 +23407,9 @@ test("syncRun rejects duplicate active starts for one reused attempt identity", const sync = await syncRun({ projectRoot: project, runId, env }); - assert.equal(sync.ok, false); - assert.deepEqual( - sync.diagnostics.map((diagnostic) => diagnostic.code), - ["WORKFLOW_ATTEMPT_INSPECT_FAILED"] - ); - assert.match(sync.diagnostics[0]?.message ?? "", /multiple active NodeStarted events/u); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "running"); + assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); }); test("syncRun abandons an unterminated occurrence at a later run activation boundary", async () => { @@ -26740,7 +26727,7 @@ function workspaceRoot(): string { } for (const recovery of ["reset", "retry", "refresh"] as const) { - test(`resume preserves unobserved failed executor history before stopped ${recovery}`, async () => { + test(`an unobserved failed occurrence stays in the attempt ledger across a stopped ${recovery}`, async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); @@ -26767,13 +26754,6 @@ for (const recovery of ["reset", "retry", "refresh"] as const) { const ledgerPath = path.join(run.value.run_root, "attempts.jsonl"); const sealPath = path.join(run.value.run_root, "smithers", "control-integrity.json"); const seal = fs.readFileSync(sealPath); - const metadata = JSON.parse(fs.readFileSync(path.join(run.value.run_root, "run.json"), "utf8")) as { - workflow: { control_generation: string }; - }; - assert.equal(fs.readFileSync(ledgerPath, "utf8"), ""); - const commandLog = env.SMITHERS_FAKE_LOG; - assert.ok(commandLog); - fs.writeFileSync(commandLog, "", "utf8"); const resumed = await resumeRun({ projectRoot: project, runId, @@ -26782,37 +26762,14 @@ for (const recovery of ["reset", "retry", "refresh"] as const) { ...(recovery === "refresh" ? { refreshController: true } : {}) }); assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - const retained = fs.readFileSync(ledgerPath, "utf8"); - const entries = retained - .trim() - .split("\n") - .filter(Boolean) - .map( - (line) => JSON.parse(line) as { source_event_sequence: number; control_generation: string; outcome: string } - ); - assert.equal( - entries.length, - 1, - "reset must preserve the unobserved failed occurrence before mutable attempt detail is replaced" - ); - const firstEntry = entries[0]; - assert.ok(firstEntry); - assert.equal(firstEntry.source_event_sequence, 2); - assert.equal(firstEntry.control_generation, metadata.workflow.control_generation); - assert.equal(firstEntry.outcome, "failed"); + assert.match(fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"), /^timetravel /mu); assert.deepEqual( fs.readFileSync(sealPath), seal, "rendering a continuation must retain original control authority" ); - const commands = fs.readFileSync(commandLog, "utf8"); - const nodeInspectionIndex = commands.indexOf(`node ${nodeId} `); - const resetIndex = commands.indexOf("timetravel "); - assert.ok(nodeInspectionIndex >= 0 && resetIndex >= 0 && nodeInspectionIndex < resetIndex); - assert.equal( - fs.existsSync(path.join(run.value.run_root, "artifacts", "project-discovery", "artifact-manifest.json")), - false - ); + // The reset restarts numbering at attempt 1 and Smithers upserts that row, so + // the failed occurrence survives only in the append-only event log. const activeEnv = fakeLifecycleSmithersEnv(project, { inspect: workflowInspect({ workflowRunId, @@ -26826,12 +26783,23 @@ for (const recovery of ["reset", "retry", "refresh"] as const) { { type: "NodeStarted", nodeId, attempt: 1 } ]) }); - for (let observation = 0; observation < 2; observation += 1) { - const sync = await syncRun({ projectRoot: project, runId, env: activeEnv }); - assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); - assert.equal(sync.value?.status, "running"); - assert.equal(fs.readFileSync(ledgerPath, "utf8"), retained); - } + const first = await syncRun({ projectRoot: project, runId, env: activeEnv }); + assert.equal(first.ok, true, JSON.stringify(first.diagnostics)); + assert.equal(first.value?.status, "running"); + const retained = fs.readFileSync(ledgerPath, "utf8"); + assert.deepEqual( + retained + .trim() + .split("\n") + .map((line) => { + const entry = JSON.parse(line) as { source_event_sequence: number; outcome: string; failure_message: string }; + return [entry.source_event_sequence, entry.outcome, entry.failure_message]; + }), + [[2, "failed", "failed before observation"]] + ); + const second = await syncRun({ projectRoot: project, runId, env: activeEnv }); + assert.equal(second.ok, true, JSON.stringify(second.diagnostics)); + assert.equal(fs.readFileSync(ledgerPath, "utf8"), retained); }); } @@ -26878,225 +26846,29 @@ async function unobservedFailedResetFixture(runId: string) { }; } -test("stopped reset refuses an incompatible recorded failure before any reset", async () => { - const fixture = await unobservedFailedResetFixture("reset-incompatible-ledger"); +test("resume resets a stopped run even when its attempt ledger holds an edited row", async () => { + const fixture = await unobservedFailedResetFixture("reset-edited-ledger"); const synced = await syncRun({ projectRoot: fixture.project, runId: fixture.runId, env: fixture.env }); assert.equal(synced.ok, true, JSON.stringify(synced.diagnostics)); const row = JSON.parse(fs.readFileSync(fixture.ledgerPath, "utf8")) as { manifests: { input_sha256: string } }; row.manifests.input_sha256 = "0".repeat(64); - const incompatible = `${JSON.stringify(row)}\n`; - fs.writeFileSync(fixture.ledgerPath, incompatible); + const edited = `${JSON.stringify(row)}\n`; + fs.writeFileSync(fixture.ledgerPath, edited); fs.writeFileSync(fixture.commandLog, ""); - const resumed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - resetNode: fixture.nodeId, - env: fixture.env - }); - assert.equal(resumed.ok, false); - assert.match(JSON.stringify(resumed.diagnostics), /immutable event authority/u); - assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), incompatible); - assert.doesNotMatch(fs.readFileSync(fixture.commandLog, "utf8"), /^timetravel |^up /mu); -}); -test("stopped reset cannot downgrade missing sealed tasks to native legacy", async () => { - const fixture = await unobservedFailedResetFixture("reset-missing-tasks"); - fs.rmSync(path.join(fixture.runRoot, "smithers", "tasks.json")); const resumed = await resumeRun({ projectRoot: fixture.project, runId: fixture.runId, resetNode: fixture.nodeId, env: fixture.env }); - assert.equal(resumed.ok, false); - assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), ""); - assert.doesNotMatch(fs.readFileSync(fixture.commandLog, "utf8"), /^timetravel |^up /mu); -}); -test("stopped reset uses authenticated config after its mutable presentation copy is lost", async () => { - const fixture = await unobservedFailedResetFixture("reset-missing-config-presentation"); - fs.rmSync(path.join(fixture.runRoot, "smithers", "resolved-config.json")); - const resumed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - resetNode: fixture.nodeId, - env: fixture.env - }); + // Attempt bookkeeping is not an admission check for the operator's recovery. assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - const entries = fs.readFileSync(fixture.ledgerPath, "utf8").trim().split("\n"); - assert.equal(entries.length, 1); - assert.equal((JSON.parse(entries[0] ?? "null") as { outcome: string }).outcome, "failed"); - assert.match(fs.readFileSync(fixture.commandLog, "utf8"), /^node node:project-discovery /mu); -}); - -test("stopped reset validates every selected attempt before appending a checkpoint", async () => { - const fixture = await unobservedFailedResetFixture("reset-invalid-later-selection"); - const env = fakeLifecycleSmithersEnv(fixture.project, { - inspect: workflowInspect({ - workflowRunId: fixture.workflowRunId, - status: "failed", - state: "failed", - steps: [{ id: fixture.nodeId, state: "failed", attempt: 2 }] - }), - events: workflowEvents(fixture.workflowRunId, [ - ...fixture.events, - { type: "NodeStarted", nodeId: fixture.nodeId, attempt: 2 }, - { type: "NodeFailed", nodeId: fixture.nodeId, attempt: 2, error: { message: "second failed occurrence" } } - ]), - nodeDetails: { - [fixture.nodeId]: { - node: { nodeId: fixture.nodeId, lastAttempt: 2 }, - attempts: [1, 2].map((attempt) => ({ - nodeId: fixture.nodeId, - attempt, - state: "failed", - meta: { - agentChainIndex: attempt === 1 ? 0 : 9, - agentId: "ultrafuzz-agent:project-discovery:0:default", - agentModel: "gpt-5.5" - } - })) - } - } - }); - const resumed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - resetNode: fixture.nodeId, - retryFailed: true, - env - }); - assert.equal(resumed.ok, false); - assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), ""); - assert.doesNotMatch(fs.readFileSync(fixture.commandLog, "utf8"), /^timetravel |^up /mu); -}); - -test("stopped reset retains pre-agent failure admission without inventing an executed attempt", async () => { - const fixture = await unobservedFailedResetFixture("reset-pre-agent-failure"); - fs.writeFileSync( - path.join(fixture.detailRoot, `${fixture.nodeId}.json`), - JSON.stringify({ - node: { nodeId: fixture.nodeId, lastAttempt: 1 }, - attempts: [ - { - nodeId: fixture.nodeId, - attempt: 1, - state: "failed", - meta: { agentChainIndex: null, agentId: null, agentModel: null } - } - ] - }) - ); - const resumed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - retryFailed: true, - env: fixture.env - }); - assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), ""); + assert.equal(fs.readFileSync(fixture.ledgerPath, "utf8"), edited); assert.match(fs.readFileSync(fixture.commandLog, "utf8"), /^timetravel /mu); }); -test("stopped reset checkpoint survives interruption with mutable attempt detail removed", async () => { - const fixture = await unobservedFailedResetFixture("reset-checkpoint-retry"); - const failed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - resetNode: fixture.nodeId, - env: { ...fixture.env, SMITHERS_FAKE_FAIL_UP: "1" } - }); - assert.equal(failed.ok, false); - const retained = fs.readFileSync(fixture.ledgerPath); - assert.ok(retained.length > 0); - // Emulate interruption after reset took effect but before its marker became durable. - fs.rmSync(path.join(fixture.runRoot, "smithers", "reset-node-applied.json")); - fs.writeFileSync( - path.join(fixture.detailRoot, `${fixture.nodeId}.json`), - JSON.stringify({ node: { nodeId: fixture.nodeId, lastAttempt: 1 }, attempts: [] }) - ); - fs.writeFileSync( - fixture.inspectPath, - JSON.stringify( - workflowInspect({ - workflowRunId: fixture.workflowRunId, - status: "failed", - state: "failed", - steps: [{ id: fixture.nodeId, state: "pending", attempt: 1 }] - }) - ) - ); - fs.writeFileSync(fixture.commandLog, ""); - const resumed = await resumeRun({ - projectRoot: fixture.project, - runId: fixture.runId, - resetNode: fixture.nodeId, - env: fixture.env - }); - assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - assert.deepEqual(fs.readFileSync(fixture.ledgerPath), retained); - const commands = fs.readFileSync(fixture.commandLog, "utf8"); - assert.doesNotMatch(commands, /^node /mu); - assert.match(commands, /^timetravel /mu); -}); - -test("stopped reset checkpoint runs once and skips active, ordinary and committed-marker continuations", async () => { - const project = tempProject(); - const workflowRunId = "ultrafuzz-reset-hook-selection"; - const nodeId = "node:fixture"; - const failedInspect = workflowInspect({ - workflowRunId, - status: "failed", - state: "failed", - steps: [{ id: nodeId, state: "failed", attempt: 1 }] - }); - const env = fakeLifecycleSmithersEnv(project, { inspect: failedInspect }); - const inspectPath = env.SMITHERS_FAKE_INSPECT; - assert.ok(inspectPath); - let checkpoints = 0; - const input = { - action: "resume" as const, - smithersRunId: workflowRunId, - workflowPath: path.join(project, ".smithers", "workflows", "workflow.tsx"), - projectRoot: project, - relaunchPaths: { runRoot: project, logsDir: path.join(project, "logs") }, - keepWorkspaces: false, - controllerLeaseSeconds: 60, - env, - environmentVariableNames: ["SMITHERS_FAKE_FAIL_UP"], - beforeStoppedReset: () => { - checkpoints += 1; - return Promise.resolve(); - } - }; - await runSmithersLifecycleCommand({ ...input, retryFailed: true, resetNode: nodeId }); - assert.equal(checkpoints, 1, "combined reset options preserve the stopped history once"); - checkpoints = 0; - await runSmithersLifecycleCommand(input); - assert.equal(checkpoints, 0, "ordinary continuation does not gain an attempt gate"); - fs.writeFileSync( - inspectPath, - JSON.stringify( - workflowInspect({ - workflowRunId, - status: "running", - state: "running", - steps: [{ id: nodeId, state: "in-progress", attempt: 1 }] - }) - ) - ); - await runSmithersLifecycleCommand({ ...input, resetNode: nodeId, force: true }); - assert.equal(checkpoints, 0, "active force reset retains its existing behavior"); - fs.writeFileSync(inspectPath, JSON.stringify(failedInspect)); - await assert.rejects( - runSmithersLifecycleCommand({ ...input, resetNode: nodeId, env: { ...env, SMITHERS_FAKE_FAIL_UP: "1" } }) - ); - assert.equal(checkpoints, 1); - checkpoints = 0; - await runSmithersLifecycleCommand({ ...input, resetNode: nodeId }); - assert.equal(checkpoints, 0, "a committed reset only needs its pending continuation"); -}); - for (const historical of ["failed", "succeeded"] as const) { test(`syncRun reconciles a finished workflow after an unrecorded ${historical} occurrence is reused`, async () => { const project = tempProject(); @@ -27174,13 +26946,15 @@ for (const historical of ["failed", "succeeded"] as const) { assert.equal(first.value?.status, "succeeded", JSON.stringify(first.diagnostics)); assert.equal(readRunState(layout).status, "succeeded"); const ledger = fs.readFileSync(layout.attemptLedgerPath); + // The superseded failure is recorded from its own events; the superseded + // success was never verified by the host, so it has no output to record. assert.deepEqual( ledger .toString("utf8") .trim() .split("\n") .map((line) => (JSON.parse(line) as { source_event_sequence: number }).source_event_sequence), - [5] + historical === "failed" ? [2, 5] : [5] ); const second = await syncRun({ projectRoot: project, runId, env }); @@ -27193,8 +26967,7 @@ for (const historical of ["failed", "succeeded"] as const) { test("resume --retry-failed after a pre-agent failure keeps synchronizing the reused attempt", async () => { const fixture = await unobservedFailedResetFixture("retry-pre-agent-failure-sync"); - // Smithers never selected an agent, so the failure is not a model attempt and - // the pre-reset checkpoint deliberately records nothing for it. + // Smithers never selected an agent, so resume has no model attempt to record. fs.writeFileSync( path.join(fixture.detailRoot, `${fixture.nodeId}.json`), JSON.stringify({ @@ -27236,3 +27009,274 @@ test("resume --retry-failed after a pre-agent failure keeps synchronizing the re assert.equal(sync.value?.status, "running"); } }); + +function attemptLedgerRows(runRoot: string): Array> { + const text = fs.readFileSync(path.join(runRoot, "attempts.jsonl"), "utf8").trim(); + return text === "" ? [] : text.split("\n").map((line) => JSON.parse(line) as Record); +} + +test("syncRun records a cancelled attempt with its Smithers reason", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = "sync-cancelled-attempt"; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "cancelled", + state: "cancelled", + includeVerifierSteps: false, + steps: [{ id: nodeId, state: "cancelled", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeCancelled", nodeId, attempt: 1, extra: { reason: "run-cancelled" } }, + { type: "RunCancelled" } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + + const sync = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.deepEqual( + attemptLedgerRows(run.value!.run_root).map((entry) => [ + entry.started_event_sequence, + entry.source_event_sequence, + entry.outcome, + entry.failure_category, + entry.failure_message, + (entry.agent as { profile_id?: string } | undefined)?.profile_id + ]), + [[1, 2, "canceled", "canceled", "run-cancelled", "default"]] + ); +}); + +test("syncRun ignores a cancellation that names no live attempt", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "sync-unmatched-cancellation"; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "failed", + state: "failed", + steps: [{ id: nodeId, state: "failed", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + // Smithers can persist a cancellation before the attempt's NodeStarted, + // after its terminal event, or with `attempt: null` for a waiting node. + { type: "NodeCancelled", nodeId, attempt: 1, extra: { reason: "run-cancelled" } }, + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeFailed", nodeId, attempt: 1, error: { message: "provider overloaded" } }, + { type: "NodeCancelled", nodeId, attempt: 1, extra: { reason: "run-cancelled" } }, + { type: "NodeCancelled", nodeId, extra: { attempt: null, reason: "run-cancelled" } } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + + const sync = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.ok(!sync.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + assert.deepEqual( + attemptLedgerRows(run.value!.run_root).map((entry) => [entry.source_event_sequence, entry.outcome]), + [[2, "failed"]] + ); +}); + +test("syncRun keeps a cancelled occurrence that a reset superseded before any sync", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = "sync-cancelled-then-reused"; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: nodeId, state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeCancelled", nodeId, attempt: 1, extra: { reason: "run-cancelled" } }, + { type: "RunCancelled" }, + { type: "RunStarted" }, + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeFinished", nodeId, attempt: 1 }, + { type: "NodeStarted", nodeId: "verify:project-discovery", attempt: 1 }, + { type: "NodeFinished", nodeId: "verify:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + + const sync = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "succeeded", JSON.stringify(sync.diagnostics)); + assert.deepEqual( + attemptLedgerRows(run.value!.run_root).map((entry) => [ + entry.started_event_sequence, + entry.source_event_sequence, + entry.outcome, + entry.agent === undefined ? "no agent" : "agent" + ]), + [ + [1, 2, "canceled", "no agent"], + [5, 6, "succeeded", "agent"] + ] + ); +}); + +test("syncRun attributes a verifier rejection to the agent's own attempt and never re-derives it", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); + const runId = "sync-verifier-rejects-retried-agent"; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const verifierNodeId = "verify:project-discovery"; + const failedPass: Parameters[1] = [ + { type: "RunStarted" }, + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeFailed", nodeId, attempt: 1, error: { message: "provider overloaded" } }, + { type: "NodeStarted", nodeId, attempt: 2 }, + { type: "NodeFinished", nodeId, attempt: 2 }, + { type: "NodeStarted", nodeId: verifierNodeId, attempt: 1 }, + { type: "NodeFailed", nodeId: verifierNodeId, attempt: 1, error: { message: "verifier process crashed" } }, + { type: "RunFailed" } + ]; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "failed", + state: "failed", + steps: [ + { id: nodeId, state: "finished", attempt: 2 }, + { id: verifierNodeId, state: "failed", attempt: 1 } + ] + }), + events: workflowEvents(workflowRunId, failedPass) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + + const failed = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(failed.ok, true, JSON.stringify(failed.diagnostics)); + assert.equal(failed.value?.status, "failed"); + assert.ok(!failed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + const recorded = fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"); + // The verifier numbers its attempts independently: its failure belongs to the + // agent's finished attempt 2, not to the agent's earlier failed attempt 1. + assert.deepEqual( + attemptLedgerRows(run.value!.run_root).map((entry) => [ + entry.attempt, + entry.outcome, + entry.failure_category, + entry.failure_message + ]), + [ + [1, "failed", "executor-error", "provider overloaded"], + [2, "failed", "artifact-validation", "verifier process crashed"] + ] + ); + + // Re-running only the verifier changes the node status after both attempts + // were recorded; the recorded rows stay final instead of becoming conflicts. + fs.writeFileSync( + env.SMITHERS_FAKE_EVENTS!, + workflowEvents(workflowRunId, [ + ...failedPass, + { type: "RunStarted" }, + { type: "NodeStarted", nodeId: verifierNodeId, attempt: 1 }, + { type: "NodeFinished", nodeId: verifierNodeId, attempt: 1 }, + { type: "RunFinished" } + ]) + ); + fs.writeFileSync( + env.SMITHERS_FAKE_INSPECT!, + `${JSON.stringify( + workflowInspect({ + workflowRunId, + steps: [ + { id: nodeId, state: "finished", attempt: 2 }, + { id: verifierNodeId, state: "finished", attempt: 1 } + ] + }) + )}\n` + ); + for (let observation = 0; observation < 2; observation += 1) { + const recovered = await syncRun({ projectRoot: project, runId, env }); + assert.equal(recovered.ok, true, JSON.stringify(recovered.diagnostics)); + assert.equal(recovered.value?.status, "succeeded", JSON.stringify(recovered.diagnostics)); + assert.ok(!recovered.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); + assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), recorded); + } +}); + +test("syncRun reconciles node state when Smithers attempt detail is unavailable and records it later", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "sync-attempt-detail-unavailable"; + const workflowRunId = `ultrafuzz-${runId}`; + const nodeId = "node:project-discovery"; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "failed", + state: "failed", + steps: [{ id: nodeId, state: "failed", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "NodeStarted", nodeId, attempt: 1 }, + { type: "NodeFailed", nodeId, attempt: 1, error: { message: "agent failed" } } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + const detailPath = path.join(env.SMITHERS_FAKE_NODE_DETAILS!, `${nodeId}.json`); + const detail = fs.readFileSync(detailPath); + fs.rmSync(detailPath); + + const unavailable = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(unavailable.ok, true, JSON.stringify(unavailable.diagnostics)); + assert.equal(unavailable.value?.status, "failed"); + assert.deepEqual( + unavailable.diagnostics + .filter((diagnostic) => diagnostic.code === "WORKFLOW_ATTEMPT_INSPECT_FAILED") + .map((diagnostic) => diagnostic.severity), + ["warning"] + ); + const layout = layoutForRunRoot(run.value!.run_root, runId); + assert.equal(readRunState(layout).nodes["project-discovery"]?.status, "failed"); + assert.deepEqual(attemptLedgerRows(run.value!.run_root), []); + + fs.writeFileSync(detailPath, detail); + const recovered = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(recovered.ok, true, JSON.stringify(recovered.diagnostics)); + assert.ok(!recovered.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_ATTEMPT_INSPECT_FAILED")); + assert.deepEqual( + attemptLedgerRows(run.value!.run_root).map((entry) => [ + entry.outcome, + (entry.agent as { profile_id?: string } | undefined)?.profile_id + ]), + [["failed", "default"]] + ); +}); From 4a44146e55db9c4103d6b1bab302d49b661fdeb0 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:21:19 +0000 Subject: [PATCH 084/206] fix(runtime): rebuild run.json accounting from usage.jsonl on every pass Every pass asserted the stored run.json accounting against usage.jsonl before recomputing it. usage.jsonl is appended before run.json is rewritten, so a pass stopped between the two writes (Ctrl-C on `status`, a timeout, an OOM) left the ledger one step ahead, and every later sync threw "run.json#$.accounting.segments[0] does not exactly correspond to the usage ledger". That throw was not caught, so `status`, `stats`, `why` and the eval poller failed for the rest of the run (#1138). run.json accounting is now a cache derived from usage.jsonl: the stored copy is parsed best-effort only for its cached prices and is always rebuilt from the ledger. A usage event is recorded once by its Smithers identity and never re-derived, so a row an earlier build wrote cannot become a conflict; the in-sync usage-gate re-check and the post-append exactness throw are deleted. A model that a successfully fetched pricing catalog does not list is settled instead of re-downloading the catalog on every sync. The TokenUsageReported exact-key check is dropped (per-field validation stays), so an extra upstream key is ignored instead of failing every sync. An accounting failure is now a WORKFLOW_ACCOUNTING_FAILED warning and node and run status still reconcile. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/workflow-sync.ts | 167 ++++++--------- packages/runtime/test/runtime.test.ts | 285 +++++++++++++++++++++----- 2 files changed, 300 insertions(+), 152 deletions(-) diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index 62b6277a6..1cd919a58 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -120,7 +120,6 @@ import { } from "./smithers.js"; import { runsRootForProject } from "./validate.js"; import { projectWorkflowControlState } from "./workflow-control.js"; -import { runtimeSemanticGateDiagnostics } from "./semantic-gates.js"; import { inspectSmithersAttemptAgentSelection, type SmithersAttemptAgentSelection @@ -1295,16 +1294,24 @@ export async function synchronizeLinkedWorkflowRun( if (preAccountingBudgetDiagnostic !== undefined) { return { ok: false, diagnostics: [preAccountingBudgetDiagnostic] }; } - const accountingResult = await synchronizeWorkflowAccounting({ - layout, - workflowRunId: evidence.smithersRunId, - controlGeneration: evidence.controlGeneration, - events: tokenEvents, - tasks: loaded.tasks, - attemptEvents: events, - control, - env: input.env ?? process.env - }); + // Usage accounting is a projection of the usage ledger: a failure is reported + // and retried on the next pass, never allowed to block run status. + let accountingResult: Awaited> = { + changed: false, + available: false + }; + try { + accountingResult = await synchronizeWorkflowAccounting({ + layout, + workflowRunId: evidence.smithersRunId, + controlGeneration: evidence.controlGeneration, + events: tokenEvents, + control, + env: input.env ?? process.env + }); + } catch (error) { + diagnostics.push({ ...diagnosticFromError(error, "workflow", "WORKFLOW_ACCOUNTING_FAILED"), severity: "warning" }); + } if (accountingResult.budgetDiagnostic !== undefined) { return { ok: false, diagnostics: [accountingResult.budgetDiagnostic] }; } @@ -1850,8 +1857,6 @@ async function synchronizeWorkflowAccounting(input: { workflowRunId: string; controlGeneration: string; events: WorkflowEvent[]; - tasks: readonly StoredWorkflowTask[]; - attemptEvents: WorkflowEvent[]; env?: Record; control: WorkflowSynchronizationControl; }): Promise<{ @@ -1866,18 +1871,17 @@ async function synchronizeWorkflowAccounting(input: { if (metadata.workflow.control_generation !== input.controlGeneration) { throw new Error("run.json workflow control generation does not match the linked Smithers run"); } - const storedAccounting = - metadata.accounting === undefined ? undefined : storedAccountingDocument(metadata.accounting, input.workflowRunId); - const existingUsageReplay = replayUsageEvents(input.layout); - if (storedAccounting !== undefined) { - assertAccountingMatchesUsageLedger(storedAccounting, existingUsageReplay.entries, "run.json#$.accounting"); - } + // run.json accounting is a cache derived from usage.jsonl: every pass rebuilds + // it from the ledger and reads the prior copy only for its cached prices. A + // pass stopped between the ledger append and the run.json write therefore + // leaves a stale cache that the next pass replaces (#1138). + const storedAccounting = cachedAccountingDocument(metadata.accounting, input.workflowRunId); const preparedUsage = prepareWorkflowUsageEvents( input.layout, input.workflowRunId, input.controlGeneration, input.events, - existingUsageReplay + replayUsageEvents(input.layout) ); const stateSourceRunId = readRunState(input.layout).source_run_id; @@ -1885,7 +1889,7 @@ async function synchronizeWorkflowAccounting(input: { throw new Error("run.json and state.json disagree on source_run_id"); } if (preparedUsage.entries.length === 0) { - if (storedAccounting !== undefined) { + if (metadata.accounting !== undefined) { throw new Error("run.json accounting cannot exist when the usage ledger is empty"); } return { changed: false, available: false }; @@ -1893,8 +1897,13 @@ async function synchronizeWorkflowAccounting(input: { const storedPricingCatalog = storedAccounting?.pricingCatalog; const storedPricing = storedAccounting?.pricing ?? new Map(); - const previouslyUnresolvedModels = - storedPricingCatalog?.status === "disabled" ? new Set(storedPricingCatalog.unresolved_models) : new Set(); + // A model that a disabled or successfully fetched catalog does not list stays + // unpriced; only an unavailable catalog is fetched again on a later pass. + const previouslyUnresolvedModels = new Set( + storedPricingCatalog === undefined || storedPricingCatalog.status === "unavailable" + ? [] + : storedPricingCatalog.unresolved_models + ); const latestLedgerEvents = workflowEventsFromUsageLedger(latestUsageLedgerEntriesByAttempt(preparedUsage.entries)); const requiredModels = modelsRequiringPricing(latestLedgerEvents); const missingModels = requiredModels.filter( @@ -1986,12 +1995,7 @@ async function synchronizeWorkflowAccounting(input: { return { changed: false, available: false, budgetDiagnostic: preAccountingMutationBudgetDiagnostic }; } - if (preparedUsage.inputs.length > 0) { - const appended = appendUsageEvents(input.layout, preparedUsage.inputs); - if (!isDeepStrictEqual(appended.replay.entries, preparedUsage.entries)) { - throw new Error("usage ledger changed after its immutable validation snapshot"); - } - } + if (preparedUsage.inputs.length > 0) appendUsageEvents(input.layout, preparedUsage.inputs); if (accountingChanged) { writeRunMetadataDocument(input.layout.runMetadataPath, nextMetadata); } @@ -2113,70 +2117,25 @@ function prepareWorkflowUsageEvents( events: WorkflowEvent[], existingReplay: UsageLedgerReplay ): PreparedUsageLedgerAppend { - const usageEvents = events.filter((event) => event.type === "TokenUsageReported"); - const existingByIdentity = new Map( - existingReplay.entries.map((entry) => [usageLedgerIdentity(entry), entry] as const) - ); - const candidateInputs = new Map(); - const candidates = new Map(); - for (const event of usageEvents) { - const candidateInput = normalizedUsageLedgerInput(workflowRunId, controlGeneration, event); - const candidate = createUsageLedgerEntry(layout, candidateInput); - const identity = usageLedgerIdentity(candidate); - const duplicateCandidate = candidates.get(identity); - if (duplicateCandidate !== undefined) { - if (!isDeepStrictEqual(duplicateCandidate, candidate)) { - throw new Error(`usage event ${identity} appears with conflicting immutable data in the source snapshot`); - } - continue; - } - const existing = existingByIdentity.get(identity); - if (existing !== undefined) { - if (!isDeepStrictEqual(existing, candidate)) { - throw new Error(`usage event ${identity} was already recorded with different immutable data`); - } - continue; - } - candidates.set(identity, candidate); - candidateInputs.set(identity, candidateInput); - } - - const entries = existingReplay.entries.map((entry) => candidates.get(usageLedgerIdentity(entry)) ?? entry); + // A usage event is recorded once by its Smithers identity and never + // re-derived, so a row written by an earlier build cannot become a conflict. + const recorded = new Set(existingReplay.entries.map((entry) => usageLedgerIdentity(entry))); + const entries = [...existingReplay.entries]; const pendingEntries: UsageLedgerEntry[] = []; const inputs: AppendUsageEventInput[] = []; - for (const [identity, candidate] of candidates) { - if (existingByIdentity.has(identity)) continue; - entries.push(candidate); - pendingEntries.push(candidate); - inputs.push(candidateInputs.get(identity)!); - } - const context = { - usageLedger: { entries }, - eventLog: { - events: events.map((event) => ({ - workflow_run_id: event.workflowRunId, - source_event_sequence: event.sourceEventSequence, - timestamp_ms: event.timestampMs, - type: event.type, - payload: event.payload - })) - } - }; - for (const candidate of candidates.values()) { - const diagnostics = runtimeSemanticGateDiagnostics({ - schemaFilename: "usage-ledger.schema.json", - document: candidate, - artifactPath: layout.usageLedgerPath, - context + for (const event of events) { + if (event.type !== "TokenUsageReported") continue; + const identity = usageLedgerIdentity({ + workflow_run_id: event.workflowRunId, + source_event_sequence: event.sourceEventSequence }); - if (diagnostics.some((diagnostic) => diagnostic.severity === "error")) { - throw new Error( - diagnostics - .filter((diagnostic) => diagnostic.severity === "error") - .map((diagnostic) => diagnostic.message) - .join("; ") - ); - } + if (recorded.has(identity)) continue; + recorded.add(identity); + const usageInput = normalizedUsageLedgerInput(workflowRunId, controlGeneration, event); + const entry = createUsageLedgerEntry(layout, usageInput); + entries.push(entry); + pendingEntries.push(entry); + inputs.push(usageInput); } return { inputs, entries, pendingEntries }; } @@ -2365,6 +2324,16 @@ function emptyAccountingSummary(): AccountingSummary { }; } +function cachedAccountingDocument(value: unknown, workflowRunId: string): StoredAccountingDocument | undefined { + if (value === undefined) return undefined; + try { + return storedAccountingDocument(value, workflowRunId); + } catch { + // An unreadable cache is recomputed from the usage ledger, not trusted. + return undefined; + } +} + function storedAccountingDocument(value: unknown, expectedWorkflowRunId?: string): StoredAccountingDocument { const label = "run.json#$.accounting"; const stored = exactStoredRecord( @@ -5651,16 +5620,15 @@ function parseInspectSnapshot(snapshot: SmithersCommandSnapshot, expectedWorkflo } /** - * The closed-world key contract Ultrafuzz enforces on the pinned runner's - * `TokenUsageReported` payload, which `synchronizeLinkedWorkflowRun` reads on - * every status, state, diagnose and evals row sync. Exported so a test can diff - * it against the engine's own emitters: the check is exact-key and re-throws - * everywhere except `stats`, so a key added upstream takes out synchronization - * for every run on its first agent task -- which is exactly what 0.35.0's - * `freshInputTokens` and `costUsd` did. + * The keys of the pinned runner's `TokenUsageReported` payload. Synchronization + * validates each accounting field it reads and ignores any other key, so a key + * added upstream no longer takes out synchronization the way 0.35.0's + * `freshInputTokens` and `costUsd` did. A test diffs this list against the + * engine's own emitters, so a new usage field is noticed at the pin bump + * rather than silently left out of accounting. */ export const CURRENT_SMITHERS_TOKEN_EVENT_KEY_CONTRACT = { - allowed: [ + known: [ "type", "runId", "nodeId", @@ -5764,9 +5732,6 @@ function validateSmithersEventPayload(type: string, payload: Record { const project = tempProject(); initProject({ projectRoot: project, force: true }); @@ -20079,12 +20076,13 @@ test("syncRun accepts the pinned 0.35.0 usage payload and still bounds its new f () => syncWithUsage("cost-not-a-number", { inputTokens: 5, outputTokens: 1, costUsd: "0.01" }), /costUsd is invalid/u ); - // A key outside the contract still fails closed; widening it for two fields - // must not have relaxed the assertion itself. - await assert.rejects( - () => syncWithUsage("unknown-usage-key", { inputTokens: 5, outputTokens: 1, totalTokens: 6 }), - /contains unsupported fields/u - ); + // A key synchronization does not read is ignored; it no longer fails the sync. + const unknownKey = await syncWithUsage("unknown-usage-key", { inputTokens: 5, outputTokens: 1, totalTokens: 60 }); + assert.equal(unknownKey.result.ok, true, JSON.stringify(unknownKey.result.diagnostics)); + const unknownKeyMetadata = JSON.parse(fs.readFileSync(path.join(unknownKey.runRoot, "run.json"), "utf8")) as { + accounting?: { current?: { total_tokens?: number } }; + }; + assert.equal(unknownKeyMetadata.accounting?.current?.total_tokens, 6); }); test("syncRun counts cache-only usage when aggregate input is explicitly zero", async () => { @@ -20416,10 +20414,10 @@ test("syncRun appends unseen usage events to the current control segment", async assert.equal(fs.readFileSync(path.join(runRoot, "usage.jsonl"), "utf8").trim().split("\n").length, 2); }); -test("syncRun rejects a malformed-present usage ledger without accounting fallback", async () => { +test("syncRun reports a malformed usage ledger without blocking run status", async () => { const project = tempProject(); initProject({ projectRoot: project, force: true }); - writeSmallTopology(project); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); const workflowRunId = "ultrafuzz-malformed-only-usage"; const env = fakeLifecycleSmithersEnv(project, { @@ -20435,19 +20433,26 @@ test("syncRun rejects a malformed-present usage ledger without accounting fallba }); const run = await startRun({ projectRoot: project, runId: "malformed-only-usage", env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", ["setup/project-discovery.md", "findings.json"]); - fs.writeFileSync(path.join(run.value!.run_root, "usage.jsonl"), "{malformed\n", "utf8"); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + fs.writeFileSync(path.join(run.value.run_root, "usage.jsonl"), "{malformed\n", "utf8"); - const usagePath = path.join(run.value!.run_root, "usage.jsonl"); + const usagePath = path.join(run.value.run_root, "usage.jsonl"); const usageBefore = fs.readFileSync(usagePath); - const metadataBefore = fs.readFileSync(path.join(run.value!.run_root, "run.json")); + const metadataBefore = fs.readFileSync(path.join(run.value.run_root, "run.json")); - await assert.rejects( - () => syncRun({ projectRoot: project, runId: "malformed-only-usage", env }), - /usage ledger record 1 is invalid strict JSON/u + const sync = await syncRun({ projectRoot: project, runId: "malformed-only-usage", env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + assert.equal(sync.value?.status, "succeeded", JSON.stringify(sync.diagnostics)); + const warnings = sync.diagnostics.filter((diagnostic) => diagnostic.code === "WORKFLOW_ACCOUNTING_FAILED"); + assert.deepEqual( + warnings.map((diagnostic) => diagnostic.severity), + ["warning"] ); + assert.match(warnings[0]?.message ?? "", /usage ledger record 1 is invalid strict JSON/u); assert.deepEqual(fs.readFileSync(usagePath), usageBefore); - assert.deepEqual(fs.readFileSync(path.join(run.value!.run_root, "run.json")), metadataBefore); + assert.deepEqual(fs.readFileSync(path.join(run.value.run_root, "run.json")), metadataBefore); }); test("syncRun records unavailable spend when workflow token events are unpriced", async () => { @@ -22552,7 +22557,8 @@ test("syncRun records an attempt without agent provenance when Smithers selectio }); const run = await startRun({ projectRoot: project, runId: "mismatched-selection", env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); const sync = await syncRun({ projectRoot: project, runId: "mismatched-selection", env }); @@ -22563,7 +22569,7 @@ test("syncRun records an attempt without agent provenance when Smithers selectio assert.equal(warnings[0]?.severity, "warning"); assert.match(warnings[0]?.message ?? "", /agent ID does not match sealed chain rung 0/u); const entries = fs - .readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8") + .readFileSync(path.join(run.value.run_root, "attempts.jsonl"), "utf8") .trim() .split("\n") .map((line) => JSON.parse(line) as { outcome?: string; agent?: unknown }); @@ -23358,6 +23364,7 @@ test("syncRun records an unrecorded failed occurrence superseded by a reused att }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); const sync = await syncRun({ projectRoot: project, runId, env }); @@ -23366,7 +23373,7 @@ test("syncRun records an unrecorded failed occurrence superseded by a reused att // Smithers' attempt row now describes the replacement, so the superseded // failure is recorded from its own events, without agent provenance. const entries = fs - .readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8") + .readFileSync(path.join(run.value.run_root, "attempts.jsonl"), "utf8") .trim() .split("\n") .map((line) => JSON.parse(line) as Record); @@ -23380,7 +23387,7 @@ test("syncRun records an unrecorded failed occurrence superseded by a reused att ]), [[0, 1, "failed", "unrecorded occurrence", undefined]] ); - assert.doesNotMatch(fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"), /^node /mu); + assert.doesNotMatch(fs.readFileSync(fakeRunnerPath(env, "SMITHERS_FAKE_LOG"), "utf8"), /^node /mu); }); test("syncRun tolerates duplicate active starts for one attempt identity", async () => { @@ -23404,12 +23411,13 @@ test("syncRun tolerates duplicate active starts for one attempt identity", async }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); const sync = await syncRun({ projectRoot: project, runId, env }); assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.equal(sync.value?.status, "running"); - assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), ""); + assert.equal(fs.readFileSync(path.join(run.value.run_root, "attempts.jsonl"), "utf8"), ""); }); test("syncRun abandons an unterminated occurrence at a later run activation boundary", async () => { @@ -26762,7 +26770,7 @@ for (const recovery of ["reset", "retry", "refresh"] as const) { ...(recovery === "refresh" ? { refreshController: true } : {}) }); assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - assert.match(fs.readFileSync(env.SMITHERS_FAKE_LOG!, "utf8"), /^timetravel /mu); + assert.match(fs.readFileSync(fakeRunnerPath(env, "SMITHERS_FAKE_LOG"), "utf8"), /^timetravel /mu); assert.deepEqual( fs.readFileSync(sealPath), seal, @@ -27010,6 +27018,12 @@ test("resume --retry-failed after a pre-agent failure keeps synchronizing the re } }); +function fakeRunnerPath(env: Record, name: string): string { + const value = env[name]; + assert.ok(value, `${name} is not set`); + return value; +} + function attemptLedgerRows(runRoot: string): Array> { const text = fs.readFileSync(path.join(runRoot, "attempts.jsonl"), "utf8").trim(); return text === "" ? [] : text.split("\n").map((line) => JSON.parse(line) as Record); @@ -27039,12 +27053,13 @@ test("syncRun records a cancelled attempt with its Smithers reason", async () => }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); const sync = await syncRun({ projectRoot: project, runId, env }); assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.deepEqual( - attemptLedgerRows(run.value!.run_root).map((entry) => [ + attemptLedgerRows(run.value.run_root).map((entry) => [ entry.started_event_sequence, entry.source_event_sequence, entry.outcome, @@ -27082,13 +27097,14 @@ test("syncRun ignores a cancellation that names no live attempt", async () => { }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); const sync = await syncRun({ projectRoot: project, runId, env }); assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.ok(!sync.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); assert.deepEqual( - attemptLedgerRows(run.value!.run_root).map((entry) => [entry.source_event_sequence, entry.outcome]), + attemptLedgerRows(run.value.run_root).map((entry) => [entry.source_event_sequence, entry.outcome]), [[2, "failed"]] ); }); @@ -27120,14 +27136,15 @@ test("syncRun keeps a cancelled occurrence that a reset superseded before any sy }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); const sync = await syncRun({ projectRoot: project, runId, env }); assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); assert.equal(sync.value?.status, "succeeded", JSON.stringify(sync.diagnostics)); assert.deepEqual( - attemptLedgerRows(run.value!.run_root).map((entry) => [ + attemptLedgerRows(run.value.run_root).map((entry) => [ entry.started_event_sequence, entry.source_event_sequence, entry.outcome, @@ -27172,18 +27189,19 @@ test("syncRun attributes a verifier rejection to the agent's own attempt and nev }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - writeRequiredArtifactSet(run.value!.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); const failed = await syncRun({ projectRoot: project, runId, env }); assert.equal(failed.ok, true, JSON.stringify(failed.diagnostics)); assert.equal(failed.value?.status, "failed"); assert.ok(!failed.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); - const recorded = fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"); + const recorded = fs.readFileSync(path.join(run.value.run_root, "attempts.jsonl"), "utf8"); // The verifier numbers its attempts independently: its failure belongs to the // agent's finished attempt 2, not to the agent's earlier failed attempt 1. assert.deepEqual( - attemptLedgerRows(run.value!.run_root).map((entry) => [ + attemptLedgerRows(run.value.run_root).map((entry) => [ entry.attempt, entry.outcome, entry.failure_category, @@ -27198,7 +27216,7 @@ test("syncRun attributes a verifier rejection to the agent's own attempt and nev // Re-running only the verifier changes the node status after both attempts // were recorded; the recorded rows stay final instead of becoming conflicts. fs.writeFileSync( - env.SMITHERS_FAKE_EVENTS!, + fakeRunnerPath(env, "SMITHERS_FAKE_EVENTS"), workflowEvents(workflowRunId, [ ...failedPass, { type: "RunStarted" }, @@ -27208,7 +27226,7 @@ test("syncRun attributes a verifier rejection to the agent's own attempt and nev ]) ); fs.writeFileSync( - env.SMITHERS_FAKE_INSPECT!, + fakeRunnerPath(env, "SMITHERS_FAKE_INSPECT"), `${JSON.stringify( workflowInspect({ workflowRunId, @@ -27224,7 +27242,7 @@ test("syncRun attributes a verifier rejection to the agent's own attempt and nev assert.equal(recovered.ok, true, JSON.stringify(recovered.diagnostics)); assert.equal(recovered.value?.status, "succeeded", JSON.stringify(recovered.diagnostics)); assert.ok(!recovered.diagnostics.some((diagnostic) => diagnostic.code === "NODE_ATTEMPT_LEDGER_WRITE_FAILED")); - assert.equal(fs.readFileSync(path.join(run.value!.run_root, "attempts.jsonl"), "utf8"), recorded); + assert.equal(fs.readFileSync(path.join(run.value.run_root, "attempts.jsonl"), "utf8"), recorded); } }); @@ -27249,7 +27267,8 @@ test("syncRun reconciles node state when Smithers attempt detail is unavailable }); const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); - const detailPath = path.join(env.SMITHERS_FAKE_NODE_DETAILS!, `${nodeId}.json`); + assert.ok(run.value); + const detailPath = path.join(fakeRunnerPath(env, "SMITHERS_FAKE_NODE_DETAILS"), `${nodeId}.json`); const detail = fs.readFileSync(detailPath); fs.rmSync(detailPath); @@ -27263,9 +27282,9 @@ test("syncRun reconciles node state when Smithers attempt detail is unavailable .map((diagnostic) => diagnostic.severity), ["warning"] ); - const layout = layoutForRunRoot(run.value!.run_root, runId); + const layout = layoutForRunRoot(run.value.run_root, runId); assert.equal(readRunState(layout).nodes["project-discovery"]?.status, "failed"); - assert.deepEqual(attemptLedgerRows(run.value!.run_root), []); + assert.deepEqual(attemptLedgerRows(run.value.run_root), []); fs.writeFileSync(detailPath, detail); const recovered = await syncRun({ projectRoot: project, runId, env }); @@ -27273,10 +27292,174 @@ test("syncRun reconciles node state when Smithers attempt detail is unavailable assert.equal(recovered.ok, true, JSON.stringify(recovered.diagnostics)); assert.ok(!recovered.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_ATTEMPT_INSPECT_FAILED")); assert.deepEqual( - attemptLedgerRows(run.value!.run_root).map((entry) => [ + attemptLedgerRows(run.value.run_root).map((entry) => [ entry.outcome, (entry.agent as { profile_id?: string } | undefined)?.profile_id ]), [["failed", "default"]] ); }); + +test("syncRun rebuilds run.json accounting after a pass stopped between the usage and run.json writes", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "usage-crash-window"; + const workflowRunId = `ultrafuzz-${runId}`; + const usage = (sequence: number, inputTokens: number) => ({ + type: "TokenUsageReported", + nodeId: "node:project-discovery", + attempt: 1, + sequence, + extra: { iteration: 0, inputTokens, outputTokens: 1, model: "test-model", agent: "test-agent" } + }); + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1, sequence: 0 }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1, sequence: 2 }, + { type: "RunFinished", sequence: 3 } + ]), + tokenEvents: workflowEvents(workflowRunId, [usage(1, 10)]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + writeRequiredArtifactSet(runRoot, "project-discovery", ["setup/project-discovery.md", "findings.json"]); + const first = await syncRun({ projectRoot: project, runId, env }); + assert.equal(first.ok, true, JSON.stringify(first.diagnostics)); + const metadataPath = path.join(runRoot, "run.json"); + const metadataBeforeLaterUsage = fs.readFileSync(metadataPath); + + // A later usage snapshot reaches usage.jsonl, but the pass stops before it + // rewrites run.json: restore run.json to the bytes that pass never replaced. + fs.writeFileSync( + fakeRunnerPath(env, "SMITHERS_FAKE_TOKEN_EVENTS"), + workflowEvents(workflowRunId, [usage(1, 10), usage(4, 30)]) + ); + const interrupted = await syncRun({ projectRoot: project, runId, env }); + assert.equal(interrupted.ok, true, JSON.stringify(interrupted.diagnostics)); + assert.equal(fs.readFileSync(path.join(runRoot, "usage.jsonl"), "utf8").trim().split("\n").length, 2); + fs.writeFileSync(metadataPath, metadataBeforeLaterUsage); + + const recovered = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(recovered.ok, true, JSON.stringify(recovered.diagnostics)); + assert.ok(!recovered.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_ACCOUNTING_FAILED")); + const metadata = JSON.parse(fs.readFileSync(metadataPath, "utf8")) as { + accounting?: { checkpoint?: { ledger_event_count?: number }; current?: { total_tokens?: number } }; + }; + assert.equal(metadata.accounting?.checkpoint?.ledger_event_count, 2); + assert.equal(metadata.accounting?.current?.total_tokens, 31); +}); + +test("syncRun keeps a recorded usage row that this build would derive differently", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "usage-row-recorded-once"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, + { + type: "TokenUsageReported", + nodeId: "node:project-discovery", + attempt: 1, + extra: { iteration: 0, inputTokens: 10, outputTokens: 5, model: "test-model", agent: "codex" } + }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ]) + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const runRoot = run.value.run_root; + writeRequiredArtifactSet(runRoot, "project-discovery", ["setup/project-discovery.md", "findings.json"]); + const first = await syncRun({ projectRoot: project, runId, env }); + assert.equal(first.ok, true, JSON.stringify(first.diagnostics)); + const usagePath = path.join(runRoot, "usage.jsonl"); + const row = JSON.parse(fs.readFileSync(usagePath, "utf8")) as { usage: { agent: string } }; + // Stands in for a row an earlier build normalized differently from this one. + row.usage.agent = "earlier-build-agent"; + const recordedUsage = `${JSON.stringify(row)}\n`; + fs.writeFileSync(usagePath, recordedUsage); + + const replayed = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(replayed.ok, true, JSON.stringify(replayed.diagnostics)); + assert.ok(!replayed.diagnostics.some((diagnostic) => diagnostic.code === "WORKFLOW_ACCOUNTING_FAILED")); + assert.equal(fs.readFileSync(usagePath, "utf8"), recordedUsage); + const metadata = JSON.parse(fs.readFileSync(path.join(runRoot, "run.json"), "utf8")) as { + accounting?: { current?: { agents?: string[] } }; + }; + assert.deepEqual(metadata.accounting?.current?.agents, ["earlier-build-agent"]); +}); + +test("syncRun settles a model the fetched pricing catalog does not list instead of refetching it", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "unlisted-model-pricing"; + const workflowRunId = `ultrafuzz-${runId}`; + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + steps: [{ id: "node:project-discovery", state: "finished", attempt: 1 }] + }), + events: workflowEvents(workflowRunId, [ + { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, + { + type: "TokenUsageReported", + nodeId: "node:project-discovery", + attempt: 1, + extra: { + iteration: 0, + inputTokens: 10, + outputTokens: 5, + costUsd: undefined, + model: "unlisted-model", + agent: "codex" + } + }, + { type: "NodeFinished", nodeId: "node:project-discovery", attempt: 1 }, + { type: "RunFinished" } + ]) + }); + env.ULTRAFUZZ_PRICING_CATALOG_URL = pricingCatalogDataUrl({ + test: { models: { "listed-model": { cost: { input: 1, output: 1 } } } } + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + writeRequiredArtifactSet(run.value.run_root, "project-discovery", ["setup/project-discovery.md", "findings.json"]); + let catalogFetches = 0; + const countingFetch: typeof fetch = (input, init) => { + catalogFetches += 1; + return testPricingFetch(input, init); + }; + + for (let pass = 0; pass < 2; pass += 1) { + const sync = await runtimeSyncRun( + { projectRoot: project, runId, env }, + { pricingFetch: countingFetch, pricingLookupHostname: async () => [{ address: "93.184.216.34", family: 4 }] } + ); + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics)); + } + + assert.equal(catalogFetches, 1); + const metadata = JSON.parse(fs.readFileSync(path.join(run.value.run_root, "run.json"), "utf8")) as { + accounting?: { pricing_catalog?: { status?: string; unresolved_models?: string[] } }; + }; + assert.equal(metadata.accounting?.pricing_catalog?.status, "available"); + assert.deepEqual(metadata.accounting?.pricing_catalog?.unresolved_models, ["unlisted-model"]); +}); From b4a192349bac6157c95c1315f6bc6e7c041ff4f3 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:23:08 +0000 Subject: [PATCH 085/206] fix(runtime): report a Smithers event stream that reaches the CLI event limit Synchronization reads both event streams with `smithers events --limit 100000`. The pinned 0.35.0 CLI returns the oldest 100,000 matching events and says nothing under --json when it stops there, so on a larger run every later lifecycle or usage event silently never reached node evidence, the attempt ledger or accounting. The `events` command has no sequence cursor to page past the limit (`--since` is a time window and `logs --from-seq` prints human text), so a stream that returns exactly the limit is now reported as a WORKFLOW_EVENTS_TRUNCATED warning naming the last sequence read. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/workflow-sync.ts | 25 ++++++++++++-- packages/runtime/test/runtime.test.ts | 50 +++++++++++++++++++++++++++ 2 files changed, 72 insertions(+), 3 deletions(-) diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index 1cd919a58..613aabab3 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -387,6 +387,12 @@ export interface WorkflowSynchronizationControl { } const MAX_OBSERVATION_SYNC_TIMEOUT_MS = 60_000; +/** + * `smithers events` returns at most this many events, oldest first, and says + * nothing when it stops there. The pinned CLI has no sequence cursor to page + * past it, so synchronization can only report that the rest is missing. + */ +const SMITHERS_EVENTS_LIMIT = 100_000; /** * Optionally bound the state refresh performed before read-only observer @@ -510,7 +516,7 @@ export async function refinalizeControllerFailures( } const eventsSnapshot = await runSmithersInspectionCommand({ - args: ["events", input.workflowRunId, "--limit", "100000", "--json"], + args: ["events", input.workflowRunId, "--limit", String(SMITHERS_EVENTS_LIMIT), "--json"], projectRoot: input.projectRoot, env: inspectionEnvironment }); @@ -1129,7 +1135,7 @@ export async function synchronizeLinkedWorkflowRun( }; } const eventsSnapshot = await runSmithersInspectionCommand({ - args: ["events", evidence.smithersRunId, "--limit", "100000", "--json"], + args: ["events", evidence.smithersRunId, "--limit", String(SMITHERS_EVENTS_LIMIT), "--json"], projectRoot, env: linkedWorkflowExecutionEnvironment(evidence, input.env), ...inspectionExecutionControl(control, synchronizationNowMs) @@ -1140,7 +1146,7 @@ export async function synchronizeLinkedWorkflowRun( return { ok: false, diagnostics: [postEventsBudgetDiagnostic] }; } const tokenEventsSnapshot = await runSmithersInspectionCommand({ - args: ["events", evidence.smithersRunId, "--type", "token", "--limit", "100000", "--json"], + args: ["events", evidence.smithersRunId, "--type", "token", "--limit", String(SMITHERS_EVENTS_LIMIT), "--json"], projectRoot, env: linkedWorkflowExecutionEnvironment(evidence, input.env), ...inspectionExecutionControl(control, synchronizationNowMs) @@ -1190,6 +1196,19 @@ export async function synchronizeLinkedWorkflowRun( diagnostics: [diagnosticFromError(error, "workflow", "WORKFLOW_TOKEN_EVENTS_INVALID")] }; } + for (const [stream, consequence, parsed] of [ + ["lifecycle", "node evidence and the attempt ledger can be incomplete", events], + ["token", "usage accounting can be incomplete", tokenEvents] + ] as const) { + const last = parsed.at(-1); + if (parsed.length < SMITHERS_EVENTS_LIMIT || last === undefined) continue; + diagnostics.push({ + code: "WORKFLOW_EVENTS_TRUNCATED", + message: `Smithers returned its ${String(SMITHERS_EVENTS_LIMIT)}-event limit for the ${stream} stream; events after sequence ${String(last.sourceEventSequence)} are not synchronized, so ${consequence}`, + severity: "warning", + source: "workflow" + }); + } // Attempt-ledger bookkeeping never blocks reconciling node and run state. let attemptAuthorities: SmithersNodeAttemptAuthorities = new Map(); try { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index a12399ed3..1c14cda8c 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -27463,3 +27463,53 @@ test("syncRun settles a model the fetched pricing catalog does not list instead assert.equal(metadata.accounting?.pricing_catalog?.status, "available"); assert.deepEqual(metadata.accounting?.pricing_catalog?.unresolved_models, ["unlisted-model"]); }); + +test("syncRun warns that a Smithers event stream at the CLI event limit may be truncated", async () => { + const project = tempProject(); + initProject({ projectRoot: project, force: true }); + writeSmallTopology(project); + const runId = "event-stream-limit"; + const workflowRunId = `ultrafuzz-${runId}`; + const base = Date.parse("2026-07-03T00:00:00.000Z"); + // `smithers events --limit 100000` stops at exactly this many events, oldest + // first, with no truncation marker in its JSON output. + const lifecycle = Array.from( + { length: 100_000 }, + (_, sequence) => + `${JSON.stringify({ + runId: workflowRunId, + seq: sequence, + timestampMs: base + sequence, + type: "RunStatusChanged", + payload: { type: "RunStatusChanged", runId: workflowRunId, timestampMs: base + sequence } + })}\n` + ).join(""); + const env = fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "in-progress", attempt: 1 }] + }), + events: lifecycle, + tokenEvents: "" + }); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + + const sync = await syncRun({ projectRoot: project, runId, env }); + + assert.equal(sync.ok, true, JSON.stringify(sync.diagnostics.slice(0, 3))); + assert.equal(sync.value?.status, "running"); + assert.deepEqual( + sync.diagnostics + .filter((diagnostic) => diagnostic.code === "WORKFLOW_EVENTS_TRUNCATED") + .map((diagnostic) => [diagnostic.severity, diagnostic.message]), + [ + [ + "warning", + "Smithers returned its 100000-event limit for the lifecycle stream; events after sequence 99999 are not synchronized, so node evidence and the attempt ledger can be incomplete" + ] + ] + ); +}); From 2b80ca737e3418936e08d276f7e7fb9ccfe619ee Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 00:28:02 +0000 Subject: [PATCH 086/206] docs: describe the attempt ledger and accounting cache semantics Documents that each attempt is recorded once by its terminal Smithers event, that cancelled and reset-superseded attempts are recorded, how host rejections are categorized, that bookkeeping failures are warnings, that run.json accounting is rebuilt from usage.jsonl, the pricing negative cache, and the WORKFLOW_EVENTS_TRUNCATED warning. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/artifacts-reports.md | 41 ++++++++++++++++++++++++----- 2 files changed, 35 insertions(+), 7 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..244c36e01 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [artifacts] [docs]** Attempt and usage bookkeeping no longer stops run synchronization. A reset that reuses an attempt number supersedes the earlier occurrence instead of failing every sync, cancelled attempts are recorded as `canceled`, each attempt and usage row is recorded once and never re-derived, and a failed `smithers node` inspection, ledger append, or accounting pass is a warning while node and run status still reconcile; `run.json` accounting is rebuilt from `usage.jsonl` on every pass, so a pass stopped between the two writes no longer strands the run. The pre-reset attempt checkpoint from #1117 is removed, a model missing from a fetched pricing catalog is not re-downloaded on every sync, unknown `TokenUsageReported` keys are ignored, and an event stream that reaches the 100,000-event `smithers events` limit is reported as `WORKFLOW_EVENTS_TRUNCATED` (#1099, #1139, #1138). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index 2683bb70e..975c0d2c0 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -204,17 +204,37 @@ manifests. Each new attempt also records its selected Smithers chain index, model-profile ID, agent reference, optional model/reasoning values, and -primary-or-fallback role. Selection is reconciled against Smithers' durable -attempt metadata and the sealed task chain; it is never inferred from the retry -number or token model. Failed primaries therefore remain visible even when a -later fallback produces the accepted output. +primary-or-fallback role when Smithers' durable attempt metadata identifies a +rung of the sealed task chain; selection is never inferred from the retry number +or token model. Failed primaries therefore remain visible even when a later +fallback produces the accepted output. An attempt that fails before Smithers +selects a rung ran no model and is not recorded. + +Each attempt is identified by its terminal Smithers event and is recorded once: +later synchronization never re-derives or rewrites it, even when the node's +status changes afterwards. Finished, failed, timed-out, and cancelled attempts +are recorded; a cancellation has outcome and category `canceled` and the +Smithers cancellation reason as its message. A finished attempt whose output the +verifier or artifact gates reject is recorded as failed with category +`invalid-output` for findings validation and `artifact-validation` otherwise. +`resume --retry-failed` and `--reset-node` restart Smithers' attempt numbering, +after which Smithers' attempt row describes only the replacement. An attempt +that such a reset superseded before any synchronization recorded it is therefore +recorded from its events without the agent block, and a superseded finished +attempt that the host never verified is not recorded. Attempt summaries and retry counts are derived from this ledger. Replaying a known transition does not append it again, so resume, replay, checkpoint continuation, and controller takeover preserve prior lifecycle history. Reused work points to its source attempt and is reported separately from executed work. The ledger stores typed failure categories but never raw diagnostics, inputs, -outputs, or configuration. +outputs, or configuration. Ledger bookkeeping does not block synchronization: +when Smithers attempt detail is unavailable or an append fails, synchronization +reports a warning, still reconciles node and run status, and retries on the next +pass. Synchronization reads each Smithers event stream with +`smithers events --limit 100000`, the CLI maximum, which returns the oldest +events first; a stream that returns exactly that many events is reported as +`WORKFLOW_EVENTS_TRUNCATED`, because any later attempts or usage cannot be read. For terminal report producers, `report.json#run_metadata.agent_execution` contains the full planned attempt chain, the attempts that failed before the @@ -769,7 +789,12 @@ source event sequence as audit evidence, but accounting uses only the latest cumulative usage snapshot for each workflow attempt. `accounting.cumulative` combines those canonical attempt snapshots with any source-run lineage. `accounting.checkpoint` records the raw ledger position used by the durable -metadata snapshot. Usage and pricing completeness are reported independently +metadata snapshot. The accounting block is a cache that every synchronization +rebuilds from `usage.jsonl`, and a usage row is recorded once and never +re-derived, so a synchronization interrupted between the usage append and the +`run.json` write is repaired by the next one. A failed accounting pass is +reported as a `WORKFLOW_ACCOUNTING_FAILED` warning and does not block run status. +Usage and pricing completeness are reported independently through `usage_complete`/`usage_incomplete_reasons` and `pricing_complete`/`pricing_incomplete_reasons`. @@ -796,7 +821,9 @@ a usable cost. Kimi-family models are priced from the pinned Moonshot provider entry, while DeepSeek-family models are priced from the pinned first-party DeepSeek entry. Either family stays listed in `pricing_catalog.unresolved_models` when its first-party entry is absent rather -than borrowing a same-named rate from another provider. +than borrowing a same-named rate from another provider. A model that a fetched +catalog does not list stays unresolved without another catalog download; only +an unavailable catalog is retried on a later synchronization. The final report is a review artifact. It is not an automatic vulnerability submission, repository mutation, or patch application. From ef54525db25e0f90190dddb11fe536eb9d1fa8d6 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:42:14 +0000 Subject: [PATCH 087/206] test(cli): run one campaign end to end with a controller kill and resume Every runtime and CLI test drove a fake `smithers` shell script or ran the engine on a hand-written workflow, so a break between the generated workflow, the pinned Smithers engine, the sealed execution snapshot, and the CLI only surfaced in a real campaign. The new test runs `init`, `run`, `status`, `resume`, `stats`, `report`, and `events` as separate CLI processes against the pinned engine that `run` and `resume` install and start under Bun. A stub `codex` on PATH writes the artifacts each prompt's output contract names, builds the final report from the host-injected authorities, and renders it with the prompt's `ultrafuzz report render` command. The stub holds the second node open while the test SIGKILLs the detached engine and supervisor; once status reports the run orphaned, the test resumes it. It asserts that the run succeeds with a verified report, that only the interrupted node's agent ran twice, that no engine task started again after it finished, and that status and stats describe the same complete run. A todo subtest records a gap it found: stats counts three agent attempts where status and the stub count four, because attempts.jsonl is built from NodeFinished/NodeFailed events and resume cancels the interrupted attempt without one. The file lives in test/e2e/ with its own `test:e2e` script, so the CLI suite glob does not run it twice. Co-Authored-By: Claude Opus 5.5 --- packages/cli/package.json | 1 + packages/cli/test/e2e/campaign-resume.test.ts | 485 ++++++++++++++++++ 2 files changed, 486 insertions(+) create mode 100644 packages/cli/test/e2e/campaign-resume.test.ts diff --git a/packages/cli/package.json b/packages/cli/package.json index 5b3a2fd9e..baadd9f19 100644 --- a/packages/cli/package.json +++ b/packages/cli/package.json @@ -26,6 +26,7 @@ "scripts": { "build": "rm -rf dist && tsc -p tsconfig.json && cp -R schema dist/schema && node scripts/verify-schema-registry.mjs", "test": "pnpm --filter @ultrafuzz/cli... build && rm -rf dist-test && tsc -p tsconfig.test.json && node --test --test-concurrency=1 dist-test/test/*.test.js", + "test:e2e": "pnpm --filter @ultrafuzz/cli... build && rm -rf dist-test && tsc -p tsconfig.test.json && node --test dist-test/test/e2e/*.test.js", "test:pr-smoke:prebuilt": "rm -rf dist-test && tsc -p tsconfig.test.json && node scripts/run-pr-smoke-tests.mjs", "typecheck": "pnpm --filter @ultrafuzz/cli^... build && tsc -p tsconfig.json --noEmit --pretty false" }, diff --git a/packages/cli/test/e2e/campaign-resume.test.ts b/packages/cli/test/e2e/campaign-resume.test.ts new file mode 100644 index 000000000..3e777615b --- /dev/null +++ b/packages/cli/test/e2e/campaign-resume.test.ts @@ -0,0 +1,485 @@ +import assert from "node:assert/strict"; +import { execFile, execFileSync } from "node:child_process"; +import fs from "node:fs"; +import os from "node:os"; +import path from "node:path"; +import test from "node:test"; +import { setTimeout as delay } from "node:timers/promises"; +import { fileURLToPath } from "node:url"; + +// One campaign through the shipped product path: `init`, `run`, `resume`, `status`, `stats`, +// `report`, and `events` each run as their own `ultrafuzz` process, and the generated workflow runs +// on the pinned Smithers engine that `run` and `resume` install from npm and start under Bun. Only +// the model is fake: a stub `codex` executable on PATH. Other runtime and CLI tests drive a fake +// `smithers` shell script, or run the engine on hand-written workflows. + +const CLI_ENTRYPOINT = fileURLToPath(new URL("../../../dist/index.js", import.meta.url)); +const MINUTE = 60_000; +const AGENT_NODES = ["project-discovery", "summarize", "final-report"] as const; +const INTERRUPTED_NODE = "summarize"; + +const TOPOLOGY = `version: 2 +defaults: + strategy_loops: 1 +nodes: + - id: __start__ + kind: meta + role: start + depends_on: [] + - id: project-discovery + kind: agentic + prompt: setup/project-discovery.md + depends_on: + - __start__ + outputs: + - path: discovery.txt + contract: ultrafuzz/text@1 + primary: true + - id: summarize + kind: agentic + prompt: setup/project-discovery.md + depends_on: + - project-discovery + outputs: + - path: summary.txt + contract: ultrafuzz/text@1 + primary: true + - id: final-report + kind: agentic + prompt: review/final-report.md + depends_on: + - summarize + outputs: + - path: report.md + contract: ultrafuzz/nonempty-markdown@1 + primary: true + - path: report.json + contract: ultrafuzz/report@3 + - id: __finish__ + kind: meta + role: finish + depends_on: + - final-report +`; + +// A public campaign needs no disclosure acknowledgement. CodexAgent routes to model:openai. +const PUBLIC_DATA_GOVERNANCE_POLICY = { + schema_version: "ultrafuzz.data-governance-policy.v1", + sensitivity: "public", + source_destinations: ["model:openai"], + artifact_destinations: [], + destination_policies: [ + { + destination: "model:openai", + processor: "stub codex", + region: "local", + retention_policy: "test fixture", + training_policy: "none", + dpa_status: "not-required", + minimization_policy: "synthetic fixture only", + data_handling_basis: "public test fixture" + } + ], + openrouter_model_allowlist: [] +}; + +interface StubConfig { + logPath: string; + holdPath: string; + holdNode: string; +} + +interface AgentCall { + node: string; + pid: number; + event: "started" | "held" | "completed"; +} + +/** + * The fake `codex` binary. The engine spawns it exactly as it spawns Codex (`codex exec ... --json + * -`, prompt on stdin). It writes every artifact the prompt's output contract names, renders the + * final report with the `ultrafuzz report render` command the prompt gives, and prints the Codex + * JSONL the engine parses. While `holdPath` exists, `holdNode` never finishes, so the test can kill + * the controller in the middle of it. The function is serialized into the binary, so it may use + * only globals. + */ +function stubCodex(config: StubConfig): void { + const fs = process.getBuiltinModule("node:fs"); + const path = process.getBuiltinModule("node:path"); + const { execFileSync } = process.getBuiltinModule("node:child_process"); + const args = process.argv.slice(2); + const prompt = fs.readFileSync(0, "utf8"); + const node = (process.env.SMITHERS_NODE_ID ?? "").replace(/^node:/u, ""); + const log = (event: AgentCall["event"]): void => + fs.appendFileSync(config.logPath, `${JSON.stringify({ node, pid: process.pid, event })}\n`); + log("started"); + if (node === config.holdNode && fs.existsSync(config.holdPath)) { + log("held"); + setTimeout(() => process.exit(1), 10 * 60_000); + return; + } + const outputs = new Map(); + for (const [, file, contract] of prompt.matchAll(/^- Path: `([^`]+)`.*\n\s+Contract: `([^`]+)`/gmu)) { + if (file === undefined || contract === undefined) continue; + if (!["ultrafuzz/text@1", "ultrafuzz/report@3", "ultrafuzz/nonempty-markdown@1"].includes(contract)) { + throw new Error(`stub codex cannot write ${contract}`); + } + fs.mkdirSync(path.dirname(file), { recursive: true }); + outputs.set(contract, file); + } + const text = outputs.get("ultrafuzz/text@1"); + if (text !== undefined) fs.writeFileSync(text, `stub output for ${node}\n`); + const report = outputs.get("ultrafuzz/report@3"); + if (report !== undefined) { + // Copy the host-injected authorities the prompt names, as the final-report prompt instructs. + const authority = (suffix: string): Record => { + const file = [...prompt.matchAll(/workspace-relative file "([^"]+)"/gu)] + .map((match) => match[1]) + .find((candidate) => candidate?.endsWith(suffix) === true); + if (file === undefined) throw new Error(`stub codex found no ${suffix} authority in the prompt`); + return JSON.parse(fs.readFileSync(path.resolve(file), "utf8")) as Record; + }; + const data = authority(".final-report-prompt.json"); + const document = { + schema_version: "ultrafuzz.report.v3", + run_metadata: { ...authority(".final-report-run-metadata.json"), agent_execution: data.agent_execution }, + issues: [], + non_production_outcomes: [], + property_provenance: [], + property_implementation_coverage: data.property_implementation_coverage + }; + fs.writeFileSync(report, `${JSON.stringify(document, null, 2)}\n`); + const render = /^ultrafuzz report render (.+)$/mu.exec(prompt)?.[1]; + if (render === undefined) throw new Error("stub codex found no report render command in the prompt"); + const renderArgs = [...render.matchAll(/(--[a-z-]+) '([^']+)'/gu)].flatMap(([, flag, value]) => [ + flag ?? "", + value ?? "" + ]); + // The renderer writes report.md; its stdout must not interleave with the JSONL below. + execFileSync("ultrafuzz", ["report", "render", ...renderArgs], { stdio: ["ignore", "ignore", "inherit"] }); + } + const emit = (event: object): boolean => process.stdout.write(`${JSON.stringify(event)}\n`); + emit({ type: "thread.started", thread_id: `stub-${node}` }); + emit({ type: "turn.started" }); + emit({ type: "item.completed", item: { id: "answer", type: "agent_message", text: `stub completed ${node}` } }); + const lastMessage = args.indexOf("--output-last-message"); + if (lastMessage >= 0) fs.writeFileSync(args[lastMessage + 1] ?? "", `stub completed ${node}`); + emit({ type: "turn.completed", usage: { input_tokens: 100, cached_input_tokens: 0, output_tokens: 10 } }); + log("completed"); +} + +interface Campaign { + root: string; + project: string; + runId: string; + env: NodeJS.ProcessEnv; + logPath: string; + holdPath: string; +} + +interface HealthValue { + status: string; + verdict: string; + ended: boolean; + progress: { finished: number; in_progress: number; pending: number; failed: number; skipped: number; total: number }; + model_mix: Array<{ attempts: number }>; + report: { status: string; completion: string; verification: string }; +} + +interface StatsValue { + status: string; + nodes: Array<{ + node_id: string; + status: string; + attempt_count: number | null; + executed_attempt_count: number | null; + }>; + totals: { node_count: number; status_counts: Record; attempts_complete: boolean }; +} + +interface WorkflowEvent { + sequence: number; + category: string; + node_id: string | null; +} + +function prepareCampaign(): Campaign { + const root = fs.mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ufz-e2e-")); + const project = path.join(root, "target"); + const bin = path.join(root, "bin"); + const codexHome = path.join(root, "codex-home"); + const tmp = path.join(root, "tmp"); + for (const directory of [project, bin, codexHome, tmp]) fs.mkdirSync(directory, { mode: 0o700 }); + const campaign: Campaign = { + root, + project, + runId: `e2e-${process.pid}-${Date.now().toString(36)}`, + logPath: path.join(root, "agent-calls.jsonl"), + holdPath: path.join(root, `hold-${INTERRUPTED_NODE}`), + env: { + ...Object.fromEntries( + Object.entries(process.env).filter( + ([name]) => name !== "NODE_TEST_CONTEXT" && !/^(?:ULTRAFUZZ|SMITHERS|CODEX|OPENAI)_/u.test(name) + ) + ), + PATH: `${bin}${path.delimiter}${process.env.PATH ?? ""}`, + // Subscription auth reads auth.json here, so the engine's agent preflight stays offline. + CODEX_HOME: codexHome, + // Keeps the operator controller that each engine command installs inside the fixture root. + TMPDIR: tmp, + ULTRAFUZZ_DATA_GOVERNANCE_POLICY: JSON.stringify(PUBLIC_DATA_GOVERNANCE_POLICY), + ULTRAFUZZ_PRICING_CATALOG_URL: "off" + } + }; + fs.writeFileSync( + path.join(codexHome, "auth.json"), + `${JSON.stringify({ auth_mode: "chatgpt", tokens: { access_token: "stub" } })}\n` + ); + const stubConfig: StubConfig = { logPath: campaign.logPath, holdPath: campaign.holdPath, holdNode: INTERRUPTED_NODE }; + const stub = `#!/usr/bin/env node\n(${String(stubCodex)})(${JSON.stringify(stubConfig)});\n`; + fs.writeFileSync(path.join(bin, "codex"), stub, { mode: 0o755 }); + fs.writeFileSync(path.join(project, "README.md"), "# Target\n"); + for (const args of [ + ["init", "-q"], + ["add", "README.md"], + ["commit", "-q", "-m", "fixture"] + ]) { + execFileSync("git", ["-c", "user.name=Ultrafuzz", "-c", "user.email=e2e@ultrafuzz.invalid", ...args], { + cwd: project, + stdio: "ignore" + }); + } + return campaign; +} + +async function ultrafuzz(campaign: Campaign, args: string[], timeoutMs = 5 * MINUTE): Promise { + const command = `ultrafuzz ${args.join(" ")}`; + const { error, stdout, stderr } = await new Promise<{ error: Error | null; stdout: string; stderr: string }>( + (resolve) => { + execFile( + process.execPath, + [CLI_ENTRYPOINT, ...args, "--project", campaign.project, "--json"], + { cwd: campaign.project, env: campaign.env, timeout: timeoutMs, killSignal: "SIGKILL", maxBuffer: 64 << 20 }, + // A failed envelope exits non-zero; the envelope itself is the evidence. + (failure, out, err) => resolve({ error: failure, stdout: out, stderr: err }) + ); + } + ); + let envelope: { ok: boolean; data: unknown; diagnostics: unknown[] }; + try { + envelope = JSON.parse(stdout) as typeof envelope; + } catch { + assert.fail(`${command} printed no JSON envelope (${error?.message ?? "exit 0"})\nstderr: ${stderr}`); + } + assert.equal(envelope.ok, true, `${command}: ${JSON.stringify(envelope.diagnostics)}`); + return envelope.data as T; +} + +async function waitFor(label: string, timeoutMs: number, probe: () => Promise | T | undefined) { + const deadline = Date.now() + timeoutMs; + for (;;) { + const value = await probe(); + if (value !== undefined) return value; + if (Date.now() > deadline) assert.fail(`timed out after ${timeoutMs / MINUTE} minutes waiting for ${label}`); + await delay(1_000); + } +} + +function agentCalls(campaign: Campaign): AgentCall[] { + if (!fs.existsSync(campaign.logPath)) return []; + // The last element is empty or a line the stub is still appending. + const lines = fs.readFileSync(campaign.logPath, "utf8").split("\n").slice(0, -1); + return lines.map((line) => JSON.parse(line) as AgentCall); +} + +/** Linux PIDs whose command line mentions `needle`. */ +function processesMentioning(needle: string): number[] { + return fs.readdirSync("/proc").flatMap((entry) => { + if (!/^\d+$/u.test(entry) || Number(entry) === process.pid) return []; + try { + return fs.readFileSync(`/proc/${entry}/cmdline`, "utf8").includes(needle) ? [Number(entry)] : []; + } catch { + return []; + } + }); +} + +function signal(pid: number, name: NodeJS.Signals | 0): boolean { + try { + process.kill(pid, name); + return true; + } catch { + return false; + } +} + +function removeTree(root: string): void { + // Sealed execution snapshots are read-only directories; a leaked fixture must not fail the test. + const restore = (directory: string): void => { + fs.chmodSync(directory, 0o700); + for (const entry of fs.readdirSync(directory, { withFileTypes: true })) { + if (entry.isDirectory()) restore(path.join(directory, entry.name)); + } + }; + try { + restore(root); + fs.rmSync(root, { recursive: true, force: true }); + } catch { + // Best effort. + } +} + +/** A task the engine finished must not start again: resume reuses it. */ +function assertNoFinishedTaskRestarted(events: WorkflowEvent[]): void { + const finished = new Set(); + for (const event of events) { + if (event.node_id === null) continue; + if (event.category === "NodeStarted") { + assert.equal(finished.has(event.node_id), false, `${event.node_id} started again at event ${event.sequence}`); + } + if (event.category === "NodeFinished") finished.add(event.node_id); + } + assert.ok(finished.size > 0, "the workflow event stream has no finished tasks"); +} + +/** Launches the campaign and SIGKILLs its detached controller while `INTERRUPTED_NODE` runs. */ +async function interruptMidRun(campaign: Campaign, mark: (phase: string) => void): Promise { + await ultrafuzz(campaign, ["init"], 2 * MINUTE); + const configPath = path.join(campaign.project, "ultrafuzz.toml"); + const generated = fs.readFileSync(configPath, "utf8"); + // init configures Codex API-key auth, whose preflight calls api.openai.com. + const config = generated.replace( + /^\[agents\.CodexAgent\]\n(?:[^[\n].*\n)*/mu, + '[agents.CodexAgent]\nauth = "subscription"\n' + ); + assert.notEqual(config, generated, "init no longer writes an [agents.CodexAgent] table"); + fs.writeFileSync(configPath, config); + fs.writeFileSync(path.join(campaign.project, ".ultrafuzz", "topology.yml"), TOPOLOGY); + fs.writeFileSync(campaign.holdPath, ""); + + const launched = await ultrafuzz<{ status: string; workflow_ids: string[] }>( + campaign, + ["run", "--run-id", campaign.runId], + 20 * MINUTE + ); + mark("run submitted"); + assert.equal(launched.status, "running"); + const [workflowRunId] = launched.workflow_ids; + assert.ok(workflowRunId !== undefined, "run did not report its workflow run ID"); + + const held = await waitFor(`${INTERRUPTED_NODE} to start`, 15 * MINUTE, () => + agentCalls(campaign).find((call) => call.node === INTERRUPTED_NODE && call.event === "held") + ); + // A host crash takes down the detached engine and the supervisor that would otherwise restart it. + const controller = processesMentioning(workflowRunId); + assert.ok(controller.length > 0, "no detached controller process is running the workflow"); + for (const pid of controller) signal(pid, "SIGKILL"); + await waitFor("the controller and its agent to exit", MINUTE, () => + processesMentioning(workflowRunId).length === 0 && !signal(held.pid, 0) ? true : undefined + ); + mark("controller killed"); +} + +test( + "a campaign whose controller is SIGKILLed mid-node resumes on the pinned engine without re-running finished nodes", + { timeout: 45 * MINUTE, skip: process.platform === "linux" ? false : "finds the detached controller through /proc" }, + async (t) => { + const campaign = prepareCampaign(); + const { runId } = campaign; + // The supervisor's command line names only the run ID; everything else names the fixture root. + const killLeftovers = (): void => { + for (const pid of [...processesMentioning(campaign.root), ...processesMentioning(runId)]) signal(pid, "SIGKILL"); + }; + process.once("exit", killLeftovers); + // Phase timings in the test output show where CI time goes. + const started = Date.now(); + const mark = (phase: string): void => t.diagnostic(`${phase} after ${Math.round((Date.now() - started) / 1000)} s`); + try { + await interruptMidRun(campaign, mark); + + // Until the dead engine's heartbeat lease lapses, the engine still reports the run as + // running, and resume leaves a running run alone (`submitted: false`). + await waitFor("status to report the run orphaned", 5 * MINUTE, async () => { + const health = await ultrafuzz(campaign, ["status", runId]); + assert.equal(health.ended, false); + return health.verdict === "orphaned" ? health : undefined; + }); + fs.rmSync(campaign.holdPath); + const resumed = await ultrafuzz<{ submitted: boolean }>(campaign, ["resume", runId], 15 * MINUTE); + mark("resume submitted"); + assert.equal(resumed.submitted, true); + + const health = await waitFor("the resumed run to end", 15 * MINUTE, async () => { + const current = await ultrafuzz(campaign, ["status", runId]); + return current.ended ? current : undefined; + }); + mark("run ended"); + assert.equal(health.status, "succeeded"); + assert.deepEqual( + { + status: health.report.status, + completion: health.report.completion, + verification: health.report.verification + }, + { status: "available", completion: "complete", verification: "verified" } + ); + + const report = await ultrafuzz<{ source: string; json_path: string; markdown_path: string }>(campaign, [ + "report", + runId + ]); + assert.equal(report.source, "verified-runtime-report"); + const reportJson = JSON.parse(fs.readFileSync(report.json_path, "utf8")) as { run_metadata: { run_id: string } }; + assert.equal(reportJson.run_metadata.run_id, runId); + assert.match(fs.readFileSync(report.markdown_path, "utf8"), /\S/u); + + // Every agent ran once, except the node the kill interrupted, which ran again after resume. + const starts = agentCalls(campaign).filter((call) => call.event === "started"); + assert.deepEqual( + AGENT_NODES.map((node) => [node, starts.filter((call) => call.node === node).length]), + AGENT_NODES.map((node) => [node, node === INTERRUPTED_NODE ? 2 : 1]) + ); + const events = await ultrafuzz<{ events: WorkflowEvent[]; truncated: boolean }>(campaign, ["events", runId]); + assert.equal(events.truncated, false); + assert.equal(events.events.filter((event) => event.category === "RunStarted").length, 2); + assertNoFinishedTaskRestarted(events.events); + + // `status` counts engine tasks and `stats` counts topology nodes; both must describe the same + // complete run. + const { finished, in_progress, pending, failed, skipped } = health.progress; + assert.deepEqual( + { finished, in_progress, pending, failed, skipped }, + { finished: health.progress.total, in_progress: 0, pending: 0, failed: 0, skipped: 0 } + ); + const stats = await ultrafuzz(campaign, ["stats", runId]); + assert.equal(stats.status, "succeeded"); + assert.equal(stats.totals.attempts_complete, true); + assert.deepEqual( + stats.nodes.map((node) => [node.node_id, node.status]), + AGENT_NODES.map((node) => [node, "succeeded"]) + ); + assert.equal(stats.totals.node_count, AGENT_NODES.length); + assert.equal(stats.totals.status_counts.succeeded, AGENT_NODES.length); + assert.equal(stats.nodes.find((node) => node.node_id === "project-discovery")?.executed_attempt_count, 1); + // status counts every agent invocation, including the one the kill interrupted. + const statusAttempts = health.model_mix.reduce((total, entry) => total + entry.attempts, 0); + assert.equal(statusAttempts, starts.length); + mark("checks done"); + + await t.test( + "stats counts the agent attempt the controller crash interrupted", + { + todo: "attempts.jsonl is built from NodeFinished/NodeFailed events, and resume cancels this attempt without one" + }, + () => { + const statsAttempts = stats.nodes.reduce((total, node) => total + (node.attempt_count ?? 0), 0); + assert.equal(statsAttempts, statusAttempts); + } + ); + } finally { + killLeftovers(); + process.removeListener("exit", killLeftovers); + removeTree(campaign.root); + } + } +); From f0b53bfc41ac8ffe2aff7177d40d33b71a0ebf81 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:42:14 +0000 Subject: [PATCH 088/206] ci: require the end-to-end campaign lane on pull requests Adds a `cli-e2e` release-validation gate that runs `pnpm --filter @ultrafuzz/cli test:e2e`, and a lane for it that is required on pull requests. The lane has a 60-minute budget and the test itself a 45-minute timeout. It runs beside the runtime lanes rather than inside the push-only CLI lane. The budget test in release-validation-lanes.test.ts keeps its 120-minute expectation for the complete runtime and CLI suites only. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/reference/development.md | 19 +++++++++++++++---- scripts/ci/release-validation-lanes.mjs | 17 ++++++++++++++--- scripts/ci/release-validation-lanes.test.ts | 5 +++-- scripts/validate-release.mjs | 7 +++++++ 5 files changed, 40 insertions(+), 9 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..aea90a889 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[cli] [ci]** Adds a pull-request-required `cli-e2e` lane that runs one campaign end to end. `ultrafuzz init`, `run`, `resume`, `status`, `stats`, `report`, and `events` run as separate processes on the pinned Smithers engine under Bun, and a stub `codex` executable writes the declared artifacts, including a final report the host verifies. The test SIGKILLs the detached engine and supervisor while a node runs, resumes the run, and asserts that it succeeds with a verified report and that no finished task starts again; until now, the runtime and CLI tests used a fake `smithers` script or hand-written workflows. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/development.md b/docs/reference/development.md index 0c441735a..265462bc5 100644 --- a/docs/reference/development.md +++ b/docs/reference/development.md @@ -49,10 +49,12 @@ validated only by the pull-request event, avoiding a duplicate push run. Pull requests also require all five runtime validation lanes: one supporting lane, including the Bun adapter contracts, and four deterministic integration -shards. Together these run the full runtime suite before merge. Pushes to `main` -and manual workflow dispatches run all eight release lanes, adding package, -CLI, and benchmark-history/typecheck checks, with at most eight jobs in parallel. -Their results are recorded in stable gate order in the JSON report. +shards. Together these run the full runtime suite before merge. They also +require the `cli-e2e` lane, which runs one campaign end to end (see below). +Pushes to `main` and manual workflow dispatches run all nine release lanes, +adding package, CLI, and benchmark-history/typecheck checks, with at most eight +jobs in parallel. Their results are recorded in stable gate order in the JSON +report. ## Package Checks @@ -69,6 +71,15 @@ pnpm --filter @ultrafuzz/cli test pnpm --filter @ultrafuzz/modal test ``` +`pnpm --filter @ultrafuzz/cli test:e2e` runs the end-to-end campaign test in +`packages/cli/test/e2e/`. It drives `init`, `run`, `resume`, `status`, `stats`, +`report`, and `events` as separate CLI processes against the pinned Smithers +engine under Bun, with a stub `codex` executable in place of the model. It +SIGKILLs the detached controller while one node is running, resumes the run, and +checks that it succeeds with a verified report and that no finished task started +again. It needs Linux, Bun, Git, and access to the npm registry, because `run` +and `resume` install the pinned engine from npm as they do for any campaign. + Package-local `typecheck` and `test` scripts may build direct workspace dependencies first because package exports point at `dist/**`. diff --git a/scripts/ci/release-validation-lanes.mjs b/scripts/ci/release-validation-lanes.mjs index de1f5ffb5..c1c8ce769 100644 --- a/scripts/ci/release-validation-lanes.mjs +++ b/scripts/ci/release-validation-lanes.mjs @@ -82,6 +82,14 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ pull_request: false, build_release_reporter: true }, + { + lane: "cli-e2e", + description: "End-to-end campaign with controller kill and resume on the pinned engine", + gates: "cli-e2e", + timeout_minutes: 60, + pull_request: true, + build_release_reporter: true + }, { lane: "benchmark-history-typecheck", description: "Benchmark history charts and workspace typecheck", @@ -95,8 +103,10 @@ export const RELEASE_VALIDATION_LANES = Object.freeze([ /** * Release validation gates that must run before a pull request can merge. - * Together these execute every `@ultrafuzz/runtime` test, which is where run - * resume, controller refresh, replay, and fork are covered. + * The runtime gates execute every `@ultrafuzz/runtime` test, which is where run + * resume, controller refresh, replay, and fork are covered. `cli-e2e` runs one + * campaign through the CLI on the pinned workflow engine, killing and resuming + * its controller. * * @type {readonly string[]} */ @@ -105,7 +115,8 @@ export const PULL_REQUEST_REQUIRED_GATES = Object.freeze([ "runtime-1", "runtime-2", "runtime-3", - "runtime-4" + "runtime-4", + "cli-e2e" ]); /** diff --git a/scripts/ci/release-validation-lanes.test.ts b/scripts/ci/release-validation-lanes.test.ts index 0ebb6049b..9b3a509a1 100644 --- a/scripts/ci/release-validation-lanes.test.ts +++ b/scripts/ci/release-validation-lanes.test.ts @@ -48,7 +48,7 @@ describe("release validation lane policy", () => { }); it("gives complete runtime and CLI suites a bounded budget beyond observed 75-minute runs", () => { - for (const name of [...PULL_REQUEST_REQUIRED_GATES, "cli"]) { + for (const name of [...PULL_REQUEST_REQUIRED_GATES.filter((gate) => gate.startsWith("runtime-")), "cli"]) { const lane = RELEASE_VALIDATION_LANES.find((candidate) => candidate.lane === name); expect(lane?.timeout_minutes).toBe(120); } @@ -111,7 +111,8 @@ describe("release validation lane policy", () => { "runtime-1", "runtime-2", "runtime-3", - "runtime-4" + "runtime-4", + "cli-e2e" ]); for (const lane of lanes) { expect(lane).not.toHaveProperty("pull_request"); diff --git a/scripts/validate-release.mjs b/scripts/validate-release.mjs index 341d855b6..095da8782 100644 --- a/scripts/validate-release.mjs +++ b/scripts/validate-release.mjs @@ -64,6 +64,13 @@ const gates = [ gate("evmbench", "EVMBench package tests", "pnpm", ["--filter", "@ultrafuzz/evmbench", "test"], ["G-EVMBENCH"]), gate("modal", "Modal package tests", "pnpm", ["--filter", "@ultrafuzz/modal", "test"], ["G-MODAL"]), gate("cli", "CLI package tests", "pnpm", ["--filter", "@ultrafuzz/cli", "test"], ["G-CLI"]), + gate( + "cli-e2e", + "End-to-end campaign on the pinned engine", + "pnpm", + ["--filter", "@ultrafuzz/cli", "test:e2e"], + ["G-CLI", "G-RUNTIME"] + ), gate("benchmark-history", "Benchmark history charts", "pnpm", ["-w", "benchmark:check:prebuilt"], ["G-CLI"]), gate("workspace-typecheck", "Workspace typecheck", "pnpm", ["-w", "typecheck"], ["G-WORKSPACE-TYPECHECK"]) ]; From 131ff3eddea3089c1dfe0829e818fd928de796c5 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:46:42 +0000 Subject: [PATCH 089/206] fix: record validator_build as provenance instead of comparing it VALIDATOR_BUILD_IDENTITY hashes five compiled validator modules and the ajv/ajv-formats versions. Planning records it, with the contract digest and schema binding, in every planned output, and readers required those values to equal what the reading build computes. A rebuild that changes any hashed byte therefore stranded in-flight runs (#921): strict run evidence failed with "planned graph output schema binding changed", so pause, cancel, replay and fork refused the run; status reported a divergence and skipped synchronization; host artifact gates, verified reads, the terminal report and the validator preflight refused it after resume. The planned schema content already pins what an artifact must satisfy, so the recorded validator build and contract digest are kept as provenance and no longer compared with the reading build: - planned-graph: sealed and fresh graphs are checked for shape and internal consistency only. The function is split into per-node helpers so the touched code passes the diff-limited strict lint. - artifact gates and verified reads validate each artifact against the schema content its planned output names: the installed schemas when the bundle digest matches, otherwise the bundle sealed in the run's execution snapshot, at the planned schema file. - the sealed schema loader loads a bundle as it was sealed, so a later build that adds, removes or re-versions a schema file no longer makes it throw; the gate's bundle digest check decides. - the terminal report no longer refuses on contract or binding drift. - the validator preflight identity gate compares schema id, digest, bundle digest and fixture digest, not the validator build, which also lets controller refresh adopt a rebuilt trusted CLI of the same schemas. - the workflow template stops comparing validatorBuild during task preparation, and dependency admission keeps the marker's sha256 and schema content binding but stops comparing contract digest, bundle digest and validator build. Tests that pinned VALIDATOR_BUILD_IDENTITY or asserted the stranding are rewritten; the new rotation test runs the operator's commands in a child process whose rebuilt strict-json.js yields a different identity. Refs #921 Co-Authored-By: Claude Opus 5.5 --- .../artifacts/src/json-validator-preflight.ts | 3 - packages/artifacts/src/planned-graph.ts | 170 ++++++------ .../artifacts/src/sealed-schema-registry.ts | 35 ++- packages/artifacts/src/semantic-gates.ts | 5 +- .../test/json-validator-preflight.test.ts | 20 +- packages/artifacts/test/schema.test.ts | 62 +++-- .../artifacts/test/semantic-gates.test.ts | 1 - packages/runtime/src/artifact-gates.ts | 67 ++--- packages/runtime/src/start-run.ts | 27 +- .../templates/smithers/workflows/workflow.tsx | 30 +-- packages/runtime/src/terminal-report.ts | 20 -- packages/runtime/src/trusted-cli.ts | 9 +- packages/runtime/src/verified-output.ts | 59 ++-- packages/runtime/src/workflow-integrity.ts | 10 +- packages/runtime/test/artifact-gates.test.ts | 28 +- .../test/generated-workflow-verifier.test.ts | 92 ++++++- packages/runtime/test/runtime.test.ts | 255 +++++++++++++----- packages/runtime/test/trusted-cli.test.ts | 42 ++- packages/runtime/test/verified-output.test.ts | 23 ++ 19 files changed, 576 insertions(+), 382 deletions(-) diff --git a/packages/artifacts/src/json-validator-preflight.ts b/packages/artifacts/src/json-validator-preflight.ts index 2464ad75a..64036fbcf 100644 --- a/packages/artifacts/src/json-validator-preflight.ts +++ b/packages/artifacts/src/json-validator-preflight.ts @@ -108,7 +108,6 @@ export interface JsonValidatorPreflightExpectedIdentity { schemaId: string; schemaSha256: string; schemaBundleSha256: string; - validatorBuild: string; artifactSha256: string; } @@ -134,7 +133,6 @@ export function parseJsonValidatorPreflightSuccessEnvelope( schemaId: binding!.schema_id, schemaSha256: binding!.schema_sha256, schemaBundleSha256: binding!.schema_bundle_sha256, - validatorBuild: binding!.validator_build, artifactSha256: ARTIFACT_VALIDATOR_SMOKE_FIXTURE_SHA256 }; const gates = executeSchemaSemanticGates(JSON_VALIDATOR_PREFLIGHT_SUCCESS_SCHEMA_FILENAME, { @@ -144,7 +142,6 @@ export function parseJsonValidatorPreflightSuccessEnvelope( schemaId: expected.schemaId, schemaSha256: expected.schemaSha256, schemaBundleSha256: expected.schemaBundleSha256, - validatorBuild: expected.validatorBuild, artifactSha256: expected.artifactSha256 } } diff --git a/packages/artifacts/src/planned-graph.ts b/packages/artifacts/src/planned-graph.ts index ebbaf72dd..3506866d3 100644 --- a/packages/artifacts/src/planned-graph.ts +++ b/packages/artifacts/src/planned-graph.ts @@ -5,7 +5,6 @@ import { } from "./artifact-contract-ids.js"; import { CANONICAL_ARTIFACT_RELATIVE_PATH_PATTERN } from "./artifact-path-primitives.js"; import { MAX_RETRY_CHAIN_ATTEMPTS } from "./artifact-limits.js"; -import { artifactContractDefinition, artifactContractSchemaBinding } from "./artifact-contracts.js"; import { validateRegisteredJsonSchema, type JsonSchemaValidationResult } from "./json-schema-validator.js"; import { readRegularFileSnapshot } from "./schema-registry.js"; import { parseStrictJsonBytes } from "./strict-json.js"; @@ -386,22 +385,11 @@ export function assertPlannedGraph(value: unknown): PlannedGraphDocument { /** * Parse a graph whose exact bytes are already authenticated by workflow-control - * evidence. A controller-only upgrade may change the digest of the complete - * schema bundle without changing this graph's contract schema. Keep every - * contract-specific binding strict and admit only that historical bundle ID. + * evidence. Sealed and freshly planned graphs are checked identically: shape and + * internal consistency, never the reading build's contract registry. */ export function assertSealedPlannedGraph(value: unknown): PlannedGraphDocument { - const shape = validatePlannedGraph(value); - if (!shape.ok) { - throw new Error( - `planned graph is schema-invalid: ${shape.issues - .map((issue) => `${issue.instancePath || "/"} ${issue.message}`) - .join("; ")}` - ); - } - const graph = value as PlannedGraphDocument; - assertPlannedGraphSemantics(graph, { allowHistoricalSchemaBundle: true }); - return graph; + return assertPlannedGraph(value); } export function readPlannedGraphDocument(filePath: string): PlannedGraphDocument { @@ -415,90 +403,13 @@ export function readPlannedGraphDocument(filePath: string): PlannedGraphDocument ); } -export function assertPlannedGraphSemantics( - graph: PlannedGraphDocument, - options: { allowHistoricalSchemaBundle?: boolean } = {} -): void { +export function assertPlannedGraphSemantics(graph: PlannedGraphDocument): void { const nodes = new Map(); const workflowTaskIds = new Set(); for (const node of graph.nodes) { if (nodes.has(node.id)) throw new Error(`planned graph repeats node ID ${JSON.stringify(node.id)}`); nodes.set(node.id, node); - const artifactIdentity = node.dynamic_generated?.storage_id ?? node.id; - const expectedArtifactDirs = node.model_fanout.map((model) => `artifacts/${model.attempt_id ?? artifactIdentity}`); - const expectedPrimaryArtifactDir = - node.dynamic_generated === undefined - ? `artifacts/${node.id}` - : (expectedArtifactDirs[0] ?? `artifacts/${artifactIdentity}`); - if (node.artifact_dir !== expectedPrimaryArtifactDir) { - throw new Error(`planned graph artifact_dir does not match node ID ${JSON.stringify(node.id)}`); - } - if ( - node.dynamic_generated !== undefined && - JSON.stringify(node.artifact_dirs ?? []) !== JSON.stringify(expectedArtifactDirs) - ) { - throw new Error(`planned graph dynamic artifact directories do not match its generated attempts`); - } - if (node.loop.index >= node.loop.count || node.loop.attempt_index !== node.loop.index) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} has inconsistent loop coordinates`); - } - - const outputPaths = new Set(); - let primaryCount = 0; - for (const output of node.outputs) { - if (outputPaths.has(output.path)) { - throw new Error( - `planned graph node ${JSON.stringify(node.id)} repeats output path ${JSON.stringify(output.path)}` - ); - } - outputPaths.add(output.path); - if (output.primary) primaryCount += 1; - const definition = artifactContractDefinition(output.contract); - if (output.contract_digest !== definition.digest) { - throw new Error(`planned graph output contract digest changed for ${JSON.stringify(output.path)}`); - } - const binding = artifactContractSchemaBinding(output.contract); - if ( - (binding === undefined && output.schema_file !== undefined) || - (binding !== undefined && - (output.schema_file !== binding.schema_file || - output.schema_id !== binding.schema_id || - output.schema_sha256 !== binding.schema_sha256 || - (options.allowHistoricalSchemaBundle !== true && - output.schema_bundle_sha256 !== binding.schema_bundle_sha256) || - output.validator_build !== binding.validator_build)) - ) { - throw new Error(`planned graph output schema binding changed for ${JSON.stringify(output.path)}`); - } - } - if (primaryCount !== 1) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} must identify exactly one primary output`); - } - - const modelKeys = new Set(); - const modelAttemptIds = new Set(); - for (const model of node.model_fanout) { - const key = `${model.model_profile_id}\u0000${model.model_index}\u0000${model.loop_index}\u0000${model.attempt_index}`; - if (modelKeys.has(key)) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} repeats a model-fanout identity`); - } - modelKeys.add(key); - if (model.loop_index !== node.loop.index) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} has a model bound to another loop`); - } - const expectedAttemptId = - node.model_fanout.length <= 1 - ? artifactIdentity - : `${artifactIdentity}__model_${model.model_index}__attempt_${model.attempt_index}`; - if (model.attempt_id !== undefined && model.attempt_id !== expectedAttemptId) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} has an inconsistent model attempt ID`); - } - const attemptId = model.attempt_id ?? expectedAttemptId; - if (modelAttemptIds.has(attemptId)) { - throw new Error(`planned graph node ${JSON.stringify(node.id)} repeats a model attempt ID`); - } - modelAttemptIds.add(attemptId); - } + assertPlannedNodeSemantics(node); for (const taskId of node.workflow?.task_node_ids ?? []) { if (workflowTaskIds.has(taskId)) throw new Error(`planned graph repeats workflow task ID ${JSON.stringify(taskId)}`); @@ -532,3 +443,74 @@ export function assertPlannedGraphSemantics( }; for (const nodeId of nodes.keys()) visit(nodeId); } + +function assertPlannedNodeSemantics(node: PlannedGraphNodeDocument): void { + const artifactIdentity = node.dynamic_generated?.storage_id ?? node.id; + const expectedArtifactDirs = node.model_fanout.map((model) => `artifacts/${model.attempt_id ?? artifactIdentity}`); + const expectedPrimaryArtifactDir = + node.dynamic_generated === undefined + ? `artifacts/${node.id}` + : (expectedArtifactDirs[0] ?? `artifacts/${artifactIdentity}`); + if (node.artifact_dir !== expectedPrimaryArtifactDir) { + throw new Error(`planned graph artifact_dir does not match node ID ${JSON.stringify(node.id)}`); + } + if ( + node.dynamic_generated !== undefined && + JSON.stringify(node.artifact_dirs ?? []) !== JSON.stringify(expectedArtifactDirs) + ) { + throw new Error(`planned graph dynamic artifact directories do not match its generated attempts`); + } + if (node.loop.index >= node.loop.count || node.loop.attempt_index !== node.loop.index) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} has inconsistent loop coordinates`); + } + assertPlannedNodeOutputs(node); + assertPlannedNodeModelFanout(node, artifactIdentity); +} + +function assertPlannedNodeOutputs(node: PlannedGraphNodeDocument): void { + const outputPaths = new Set(); + let primaryCount = 0; + for (const output of node.outputs) { + if (outputPaths.has(output.path)) { + throw new Error( + `planned graph node ${JSON.stringify(node.id)} repeats output path ${JSON.stringify(output.path)}` + ); + } + outputPaths.add(output.path); + if (output.primary) primaryCount += 1; + // An output's contract digest and schema binding record the build that planned it, and the + // planned-graph schema already requires a binding exactly for schema-backed contracts. They + // are not re-derived against the reading build: that stranded in-flight runs whenever a + // rebuild changed any of them (#921). Artifact gates validate against the recorded schema. + } + if (primaryCount !== 1) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} must identify exactly one primary output`); + } +} + +function assertPlannedNodeModelFanout(node: PlannedGraphNodeDocument, artifactIdentity: string): void { + const modelKeys = new Set(); + const modelAttemptIds = new Set(); + for (const model of node.model_fanout) { + const key = [model.model_profile_id, model.model_index, model.loop_index, model.attempt_index].join("\u0000"); + if (modelKeys.has(key)) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} repeats a model-fanout identity`); + } + modelKeys.add(key); + if (model.loop_index !== node.loop.index) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} has a model bound to another loop`); + } + const expectedAttemptId = + node.model_fanout.length <= 1 + ? artifactIdentity + : `${artifactIdentity}__model_${String(model.model_index)}__attempt_${String(model.attempt_index)}`; + if (model.attempt_id !== undefined && model.attempt_id !== expectedAttemptId) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} has an inconsistent model attempt ID`); + } + const attemptId = model.attempt_id ?? expectedAttemptId; + if (modelAttemptIds.has(attemptId)) { + throw new Error(`planned graph node ${JSON.stringify(node.id)} repeats a model attempt ID`); + } + modelAttemptIds.add(attemptId); + } +} diff --git a/packages/artifacts/src/sealed-schema-registry.ts b/packages/artifacts/src/sealed-schema-registry.ts index 7b7925ea4..82c5a84da 100644 --- a/packages/artifacts/src/sealed-schema-registry.ts +++ b/packages/artifacts/src/sealed-schema-registry.ts @@ -4,6 +4,7 @@ import path from "node:path"; import { artifactSchemaRegistry, + DEFAULT_MAX_JSON_INSTANCE_BYTES, readRegularFileSnapshot, type ArtifactSchemaRegistryEntry } from "./schema-registry.js"; @@ -15,7 +16,13 @@ const MAX_REGISTERED_PATTERNS = 256; const MAX_REGISTERED_BUNDLE_PATTERNS = 1_024; const MAX_REGISTERED_PATTERN_LENGTH = 1_024; -/** Load a complete, physical schema bundle from an authenticated execution snapshot. */ +/** + * Load the complete physical schema bundle sealed in an execution snapshot, exactly as it was sealed. + * A later build may add, remove or re-version schema files, so the bundle is not compared with the + * installed registry; callers compare its bundle digest with the one an artifact was planned against. + * A sealed bundle is only used to validate, so a file carries the installed build's contract, gate + * and export metadata only when that build still registers it under the same `$id`. + */ export function artifactSchemaRegistryFromDirectory(directory: string): readonly ArtifactSchemaRegistryEntry[] { const resolved = path.resolve(directory); const lexical = fs.lstatSync(resolved); @@ -28,19 +35,11 @@ export function artifactSchemaRegistryFromDirectory(directory: string): readonly .readdirSync(resolved) .filter((filename) => filename.endsWith(".schema.json")) .sort(); - const unknown = filenames.filter((filename) => !currentByFilename.has(filename)); - const missing = [...currentByFilename.keys()].filter((filename) => !filenames.includes(filename)); - if (unknown.length > 0 || missing.length > 0) { - throw new Error( - `schema registry mismatch${unknown.length > 0 ? `; unregistered: ${unknown.join(", ")}` : ""}${missing.length > 0 ? `; missing: ${missing.join(", ")}` : ""}` - ); - } let bundleBytes = 0; let bundlePatterns = 0; return Object.freeze( filenames.map((filename): ArtifactSchemaRegistryEntry => { - const current = currentByFilename.get(filename)!; const snapshot = readRegularFileSnapshot(path.join(resolved, filename), MAX_REGISTERED_SCHEMA_BYTES); bundleBytes += snapshot.byteLength; if (bundleBytes > MAX_REGISTERED_BUNDLE_BYTES) { @@ -56,7 +55,10 @@ export function artifactSchemaRegistryFromDirectory(directory: string): readonly if (parsed.$schema !== "https://json-schema.org/draft/2020-12/schema") { throw new Error(`schema must declare Draft 2020-12: ${filename}`); } - if (parsed.$id !== current.id) throw new Error(`sealed schema $id changed for ${filename}`); + const id = parsed.$id; + if (typeof id !== "string" || id.length === 0 || id.includes("#")) { + throw new Error(`schema must have a fragment-free non-empty $id: ${filename}`); + } bundlePatterns += assertRegisteredPatternLimits(parsed, filename); if (bundlePatterns > MAX_REGISTERED_BUNDLE_PATTERNS) { throw new Error(`registered schema bundle exceeds the ${MAX_REGISTERED_BUNDLE_PATTERNS}-pattern limit`); @@ -65,8 +67,19 @@ export function artifactSchemaRegistryFromDirectory(directory: string): readonly for (const reference of localReferences) { if (/^https?:/iu.test(reference)) throw new Error(`remote schema reference is forbidden: ${reference}`); } + const current = currentByFilename.get(filename); return Object.freeze({ - ...current, + ...(current?.id === id + ? current + : { + filename, + id, + role: "subschema" as const, + contractIds: Object.freeze([]), + maxInstanceBytes: DEFAULT_MAX_JSON_INSTANCE_BYTES, + semanticGates: Object.freeze([]), + typescriptExport: "" + }), sha256: sha256(snapshot), schema: deepFreezeJson(parsed), localReferences: Object.freeze(localReferences) diff --git a/packages/artifacts/src/semantic-gates.ts b/packages/artifacts/src/semantic-gates.ts index c0bcc42a1..c12f0886c 100644 --- a/packages/artifacts/src/semantic-gates.ts +++ b/packages/artifacts/src/semantic-gates.ts @@ -228,7 +228,6 @@ export interface SemanticValidatorPreflightContext { schemaId: string; schemaSha256: string; schemaBundleSha256: string; - validatorBuild: string; artifactSha256: string; } @@ -429,7 +428,8 @@ function jsonValidatorPreflightIdentityIssues(document: unknown, context: Semant ["$.data.schema.id", at(document, ["data", "schema", "id"]), expected.schemaId], ["$.data.schema.sha256", at(document, ["data", "schema", "sha256"]), expected.schemaSha256], ["$.data.schema.bundle_sha256", at(document, ["data", "schema", "bundle_sha256"]), expected.schemaBundleSha256], - ["$.data.schema.validator_build", at(document, ["data", "schema", "validator_build"]), expected.validatorBuild], + // `validator_build` is provenance: the validator that reported the identity may come from another + // build of the same schemas, which must not stop an in-flight run's preflight (#921). ["$.data.artifact_sha256", at(document, ["data", "artifact_sha256"]), expected.artifactSha256] ] as const; return checks.flatMap(([pathValue, actual, wanted]) => @@ -7808,7 +7808,6 @@ const gateSpecifications = { "validatorPreflight.schemaId", "validatorPreflight.schemaSha256", "validatorPreflight.schemaBundleSha256", - "validatorPreflight.validatorBuild", "validatorPreflight.artifactSha256" ], jsonValidatorPreflightIdentityIssues diff --git a/packages/artifacts/test/json-validator-preflight.test.ts b/packages/artifacts/test/json-validator-preflight.test.ts index 9401182a7..7555f0351 100644 --- a/packages/artifacts/test/json-validator-preflight.test.ts +++ b/packages/artifacts/test/json-validator-preflight.test.ts @@ -63,7 +63,6 @@ function identityGate(document: unknown) { schemaId: binding.schema_id, schemaSha256: binding.schema_sha256, schemaBundleSha256: binding.schema_bundle_sha256, - validatorBuild: binding.validator_build, artifactSha256: ARTIFACT_VALIDATOR_SMOKE_FIXTURE_SHA256 } } @@ -86,6 +85,19 @@ test("validator preflight parser accepts only the exact non-transforming success ); }); +test("validator preflight parser accepts the same schemas reported by another validator build", () => { + // #921: a rebuild that only changes the validator modules must not fail an in-flight run's + // preflight, whose trusted CLI was sealed by the earlier build. The schema identity still binds. + const value = successEnvelope(); + objectField(objectField(value, "data"), "schema").validator_build = `ultrafuzz-json-validator.v1:${"9".repeat(64)}`; + + assert.deepEqual(parseJsonValidatorPreflightSuccessEnvelope(encode(value)), value); + assert.equal(identityGate(value)?.status, "passed"); + objectField(objectField(value, "data"), "schema").sha256 = "0".repeat(64); + assert.throws(() => parseJsonValidatorPreflightSuccessEnvelope(encode(value)), /mismatched identity/u); + assert.equal(identityGate(value)?.status, "failed"); +}); + test("validator preflight parser can authenticate a sealed historical identity explicitly", () => { const value = successEnvelope(); const schema = objectField(objectField(value, "data"), "schema"); @@ -97,7 +109,6 @@ test("validator preflight parser can authenticate a sealed historical identity e schemaId: binding.schema_id, schemaSha256: binding.schema_sha256, schemaBundleSha256: schema.bundle_sha256 as string, - validatorBuild: schema.validator_build as string, artifactSha256: ARTIFACT_VALIDATOR_SMOKE_FIXTURE_SHA256 }; @@ -172,11 +183,6 @@ const contractMutations: ReadonlyArray<{ name: "missing validator build", mutate: (value) => void delete objectField(objectField(value, "data"), "schema").validator_build }, - { - name: "wrong validator build", - structurallyValid: true, - mutate: (value) => void (objectField(objectField(value, "data"), "schema").validator_build = "legacy") - }, { name: "missing registration status", mutate: (value) => void delete objectField(objectField(value, "data"), "schema").registered diff --git a/packages/artifacts/test/schema.test.ts b/packages/artifacts/test/schema.test.ts index a6694a9b3..e67e59a92 100644 --- a/packages/artifacts/test/schema.test.ts +++ b/packages/artifacts/test/schema.test.ts @@ -178,7 +178,7 @@ test("current-controller schema materialization replaces only an older physical } }); -test("loads only a complete physical sealed schema bundle", () => { +test("loads a sealed schema bundle exactly as it was sealed, even when the installed build has moved on", () => { const root = mkdtempSync(path.join(fs.realpathSync(os.tmpdir()), "ultrafuzz-sealed-schema-bundle-")); const destination = path.join(root, "schemas"); try { @@ -192,9 +192,32 @@ test("loads only a complete physical sealed schema bundle", () => { fs.symlinkSync(destination, linked, "dir"); assert.throws(() => artifactSchemaRegistryFromDirectory(linked), /snapshot directory is unsafe/u); + // #921: a run sealed by another build may hold a schema this build dropped, lack one it added, + // and carry an older `$id` for a file this build re-versioned. The bundle still loads as sealed; + // the artifact gate tells bundles apart by digest, not by agreement with the installed registry. fs.chmodSync(destination, 0o700); + const retired = { $schema: "https://json-schema.org/draft/2020-12/schema", $id: "urn:ultrafuzz:schema:retired:1" }; + fs.writeFileSync(path.join(destination, "retired.schema.json"), JSON.stringify(retired), "utf8"); + fs.chmodSync(path.join(destination, "report.schema.json"), 0o600); + fs.rmSync(path.join(destination, "report.schema.json")); + const reversioned = path.join(destination, "usage-ledger.schema.json"); + fs.chmodSync(reversioned, 0o600); + const usageLedger = JSON.parse(fs.readFileSync(reversioned, "utf8")) as Record; + fs.writeFileSync(reversioned, JSON.stringify({ ...usageLedger, $id: `${String(usageLedger.$id)}-sealed` }), "utf8"); + const sealed = artifactSchemaRegistryFromDirectory(destination); + assert.equal(sealed.find((entry) => entry.filename === "retired.schema.json")?.id, retired.$id); + assert.equal( + sealed.find((entry) => entry.filename === "usage-ledger.schema.json")?.id, + `${String(usageLedger.$id)}-sealed` + ); + assert.equal( + sealed.some((entry) => entry.filename === "report.schema.json"), + false + ); + assert.notEqual(schemaRegistryBundleDigest(sealed), artifactSchemaBundleDigest()); + fs.writeFileSync(path.join(destination, "foreign.schema.json"), "{}\n", "utf8"); - assert.throws(() => artifactSchemaRegistryFromDirectory(destination), /registry mismatch/u); + assert.throws(() => artifactSchemaRegistryFromDirectory(destination), /must declare Draft 2020-12/u); } finally { fs.chmodSync(destination, 0o700); for (const file of readdirSync(destination)) fs.chmodSync(path.join(destination, file), 0o600); @@ -2171,27 +2194,32 @@ test("planned graph v4 validates whole documents and executes every registered d ...artifactContractSchemaBinding("ultrafuzz/findings@2"), primary: true }; - const preUpgradeValidatorBuild = - "ultrafuzz-json-validator.v1:028be3251e9ac213ad6e1c037d8da47c9d903e8be563c8f4fa2149839565bcab"; - assert.equal(VALIDATOR_BUILD_IDENTITY, preUpgradeValidatorBuild); - const historicalBundle = structuredClone(graph); - historicalBundle.nodes = [ + // #921: a graph planned by another build records that build's contract digest, schema content and + // validator build. None of it is re-derived against the reading build, so a rebuild or upgrade + // after launch cannot make the run's own graph unreadable. + const otherBuild = structuredClone(graph); + otherBuild.nodes = [ { ...node, outputs: [ { ...findingsOutput, + contract_digest: "d".repeat(64), + schema_sha256: "e".repeat(64), schema_bundle_sha256: "f".repeat(64), - validator_build: preUpgradeValidatorBuild + validator_build: `ultrafuzz-json-validator.v1:${"9".repeat(64)}` } ] } ]; - assert.throws(() => assertPlannedGraph(historicalBundle), /schema binding changed/u); - assert.deepEqual(assertSealedPlannedGraph(historicalBundle), historicalBundle); - const historicalSchemaDrift = structuredClone(historicalBundle); - historicalSchemaDrift.nodes[0]!.outputs[0]!.schema_sha256 = "e".repeat(64); - assert.throws(() => assertSealedPlannedGraph(historicalSchemaDrift), /schema binding changed/u); + assert.notEqual(otherBuild.nodes[0]?.outputs[0]?.validator_build, VALIDATOR_BUILD_IDENTITY); + assert.deepEqual(assertPlannedGraph(otherBuild), otherBuild); + assert.deepEqual(assertSealedPlannedGraph(otherBuild), otherBuild); + // The shape still requires a schema binding exactly for schema-backed contracts. + const { schema_sha256: _schemaSha256, ...unboundFindings } = findingsOutput; + assert.equal(validatePlannedGraph({ ...graph, nodes: [{ ...node, outputs: [unboundFindings] }] }).ok, false); + const boundMarkdown = { ...node.outputs[0], ...artifactContractSchemaBinding("ultrafuzz/findings@2") }; + assert.equal(validatePlannedGraph({ ...graph, nodes: [{ ...node, outputs: [boundMarkdown] }] }).ok, false); const documentGateFailures: Array<{ name: string; graph: PlannedGraphDocument; message: RegExp }> = [ { @@ -2276,14 +2304,6 @@ test("planned graph v4 validates whole documents and executes every registered d graph: { ...graph, nodes: [{ ...node, loop: { ...node.loop, index: 1, count: 1, attempt_index: 1 } }] }, message: /inconsistent loop coordinates/u }, - { - name: "planned-graph-contract-identity", - graph: { - ...graph, - nodes: [{ ...node, outputs: [{ ...findingsOutput, contract_digest: "f".repeat(64) }] }] - }, - message: /contract digest changed/u - }, { name: "planned-graph-model-loop-coupling", graph: { diff --git a/packages/artifacts/test/semantic-gates.test.ts b/packages/artifacts/test/semantic-gates.test.ts index dbdded71b..5dc653a62 100644 --- a/packages/artifacts/test/semantic-gates.test.ts +++ b/packages/artifacts/test/semantic-gates.test.ts @@ -3317,7 +3317,6 @@ test("every contextual registration executes real positive and negative checks", schemaId: "schema-id", schemaSha256: "a".repeat(64), schemaBundleSha256: "b".repeat(64), - validatorBuild: "validator-build", artifactSha256: "c".repeat(64) } } diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 8bebe8a23..66e4d54e9 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -3691,49 +3691,40 @@ function readStrictRegisteredDocument(artifactPath: string, schemaFilename: Arti return document; } -function verifyRequiredArtifactSchemaBinding( +/** + * Validate one artifact against the schema content its planned output names. Verified reads use the + * same check, so a verified output stays valid for as long as its own schema snapshot exists. + */ +export function verifyRequiredArtifactSchemaBinding( layout: RunLayout, absolutePath: string, output: PlannedGraphNode["outputs"][number], artifactBytes: Uint8Array ): RuntimeDiagnostic[] { const current = artifactContractSchemaBinding(output.contract); - const actual = { + const planned = { schema_file: output.schema_file, schema_id: output.schema_id, schema_sha256: output.schema_sha256, schema_bundle_sha256: output.schema_bundle_sha256, validator_build: output.validator_build }; - if (current === undefined) { - if (Object.values(actual).every((value) => value === undefined)) return []; - return [schemaBindingMismatchDiagnostic(absolutePath, output, null, actual)]; - } - if ( - current.schema_file !== actual.schema_file || - current.schema_id !== actual.schema_id || - current.schema_sha256 !== actual.schema_sha256 || - current.validator_build !== actual.validator_build || - typeof actual.schema_bundle_sha256 !== "string" || - !/^[0-9a-f]{64}$/u.test(actual.schema_bundle_sha256) - ) { - return [schemaBindingMismatchDiagnostic(absolutePath, output, current, actual)]; + if (current === undefined && Object.values(planned).every((value) => value === undefined)) return []; + if (current === undefined || planned.schema_file === undefined || planned.schema_bundle_sha256 === undefined) { + return [schemaBindingMismatchDiagnostic(absolutePath, output, planned)]; } - const expected = { - schema_file: actual.schema_file, - schema_id: actual.schema_id, - schema_sha256: actual.schema_sha256, - schema_bundle_sha256: actual.schema_bundle_sha256, - validator_build: actual.validator_build - } as const; + // The planned binding names schema content. Its validator build only records which build planned + // the output and is never compared: validate with the installed schemas when they are the planned + // bundle, and otherwise with the bundle sealed into this run's execution snapshot, so a rebuild or + // upgrade does not change which schema an in-flight run's artifacts must satisfy (#921). let schemaPath: string; let schemaRegistry: ReturnType | undefined; try { - if (current.schema_bundle_sha256 === expected.schema_bundle_sha256) { - schemaPath = path.join(artifactSchemaDirectory(), current.schema_file); + if (current.schema_bundle_sha256 === planned.schema_bundle_sha256) { + schemaPath = safeResolveInside(artifactSchemaDirectory(), planned.schema_file, "planned artifact schema"); } else { - schemaPath = sealedArtifactSchemaPath(layout, current.schema_file); + schemaPath = sealedArtifactSchemaPath(layout, planned.schema_file); schemaRegistry = artifactSchemaRegistryFromDirectory(path.dirname(schemaPath)); } } catch (error) { @@ -3755,19 +3746,18 @@ function verifyRequiredArtifactSchemaBinding( ...(schemaRegistry === undefined ? {} : { schemaRegistry }) }); if ( - validation.schema?.id !== expected.schema_id || - validation.schema?.sha256 !== expected.schema_sha256 || - validation.schema?.bundle_sha256 !== expected.schema_bundle_sha256 || - validation.schema?.validator_build !== expected.validator_build + validation.schema?.id !== planned.schema_id || + validation.schema?.sha256 !== planned.schema_sha256 || + validation.schema?.bundle_sha256 !== planned.schema_bundle_sha256 ) { return [ { code: "ARTIFACT_VALIDATOR_IDENTITY_MISMATCH", - message: `Host validator identity for ${output.path} does not match the planned schema binding`, + message: `The schema available for ${output.path} does not match its planned schema binding`, severity: "error", source: "artifact-schema", path: absolutePath, - details: { contract: output.contract, expected, actual: validation.schema } + details: { contract: output.contract, expected: planned, actual: validation.schema } } ]; } @@ -3780,10 +3770,10 @@ function verifyRequiredArtifactSchemaBinding( path: `${absolutePath}${diagnostic.instancePath === undefined ? "" : `#${diagnostic.instancePath || "/"}`}`, details: { contract: output.contract, - schema_id: expected.schema_id, - schema_sha256: expected.schema_sha256, - schema_bundle_sha256: expected.schema_bundle_sha256, - validator_build: expected.validator_build, + schema_id: planned.schema_id, + schema_sha256: planned.schema_sha256, + schema_bundle_sha256: planned.schema_bundle_sha256, + validator_build: planned.validator_build, ...(diagnostic.schemaPath === undefined ? {} : { schema_path: diagnostic.schemaPath }), ...(diagnostic.keyword === undefined ? {} : { keyword: diagnostic.keyword }) } @@ -3793,16 +3783,15 @@ function verifyRequiredArtifactSchemaBinding( function schemaBindingMismatchDiagnostic( absolutePath: string, output: PlannedGraphNode["outputs"][number], - expected: ReturnType | null, - actual: Readonly> + planned: Readonly> ): RuntimeDiagnostic { return { code: "ARTIFACT_SCHEMA_BINDING_MISMATCH", - message: `Planned schema identity for ${output.path} does not match validator build ${expected?.validator_build ?? "unbound"}`, + message: `Planned schema binding for ${output.path} does not fit contract ${output.contract}`, severity: "error", source: "artifact-schema", path: absolutePath, - details: { contract: output.contract, expected, actual } + details: { contract: output.contract, planned } }; } diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index ba32e842b..6a0f35e0c 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -1094,23 +1094,14 @@ function parseSealedTaskManifestForObserver(contents: Readonly<{ graph: Buffer; /** * `assertSealedPlannedGraph` does two different jobs behind one name. The first is structural: the * bytes must validate against the planned-graph schema, and nothing can report on a document that is - * not a planned graph at all. The second, `assertPlannedGraphSemantics`, re-derives the graph against - * *this build* — it looks every output contract up in the running process's artifact-contract registry - * and insists the digests, schema IDs and validator build recorded at compile time still match what - * this checkout produces. + * not a planned graph at all. The second, `assertPlannedGraphSemantics`, checks the graph's internal + * consistency (unique IDs, dependency joins, artifact directories, loop and model coordinates) with + * the rules of the build doing the reading. * - * That second job is not a property of the run; it is a property of the tree observing the run. An - * operator whose checkout has moved on since the run was submitted -- a rebased branch, a newer - * release, a contract whose schema was revised -- gets `planned graph output schema binding changed` - * and loses `status` for a run that is otherwise intact and possibly still executing. That is the same - * failure as issue #866, one throw further along the same read-only path: the run is fine, the - * observer's registry disagrees, and the operator is the one punished. `packages/artifacts`'s own - * semantic-gate collects exactly this condition as an issue rather than raising it, so the softer - * reading already exists in the codebase. - * - * Execution must still refuse: running a node whose output contract no longer matches the registry - * that will validate its artifacts would produce evidence nothing can check. So the downgrade is - * observer-only, and schema invalidity stays fatal for everyone. + * Those rules belong to the tree observing the run, not to the run: a sealed graph that a newer + * checkout's rules reject is still the graph the run is executing, and losing `status` for it is the + * same failure as issue #866, one throw further along the same read-only path. So the downgrade is + * observer-only: execution callers still refuse, and schema invalidity stays fatal for everyone. */ function parseSealedPlannedGraphForObserver(graphBytes: Buffer): { graph: PlannedGraphDocument; @@ -1121,12 +1112,12 @@ function parseSealedPlannedGraphForObserver(graphBytes: Buffer): { if (!validatePlannedGraph(value).ok) return { graph: assertSealedPlannedGraph(value), divergences: [] }; const graph = value as PlannedGraphDocument; try { - assertPlannedGraphSemantics(graph, { allowHistoricalSchemaBundle: true }); + assertPlannedGraphSemantics(graph); } catch (error) { return { graph, divergences: [ - `sealed planned graph no longer re-derives against this build's artifact contracts: ${error instanceof Error ? error.message : String(error)}` + `sealed planned graph fails this build's planned-graph semantics: ${error instanceof Error ? error.message : String(error)}` ] }; } diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index ceab70669..f2d1a39c7 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -38,7 +38,6 @@ const runtimeModule = process.env.ULTRAFUZZ_RUNTIME_MODULE ?? new URL("../../modules/@ultrafuzz/runtime/dist/index.js", import.meta.url).href; const { - artifactContractDefinition, artifactContractSchemaBinding, artifactSchemaRegistry, artifactValidatorSmokeFixturePath, @@ -4067,6 +4066,9 @@ function prepareArtifactMirror( } function assertTaskOutputSchemaBindings(task: (typeof taskSpecs)[number]): void { + // The verifier validates with this process's schemas and records the planned binding in its marker, + // so the schema content must be the planned one. `validatorBuild` is provenance and is not compared: + // a rebuild of the validator modules must not stop an in-flight run (#921). for (const output of task.outputs) { const binding = artifactContractSchemaBinding( output.contract as Parameters[0] @@ -4075,8 +4077,7 @@ function assertTaskOutputSchemaBindings(task: (typeof taskSpecs)[number]): void binding?.schema_file !== output.schemaFile || binding?.schema_id !== output.schemaId || binding?.schema_sha256 !== output.schemaSha256 || - binding?.schema_bundle_sha256 !== output.schemaBundleSha256 || - binding?.validator_build !== output.validatorBuild + binding?.schema_bundle_sha256 !== output.schemaBundleSha256 ) { throw new Error(`artifact-contract failure: planned schema binding changed for ${output.path}`); } @@ -5855,15 +5856,16 @@ function assertVerifiedDependency( } seenPaths.add(entry.path); const expected = expectedArtifacts.get(entry.path); + // The marker must name the declared output and its schema content. Its contract digest, bundle + // digest and validator build only record the build that planned the output, so none of them is + // compared, with the declaration or with this build (#921). The bytes stay pinned by the + // marker's sha256 and are revalidated below. if ( expected === undefined || expected.contract !== entry.contract || - expected.contractDigest !== entry.contract_digest || expected.schemaFile !== entry.schema_file || expected.schemaId !== entry.schema_id || expected.schemaSha256 !== entry.schema_sha256 || - expected.schemaBundleSha256 !== entry.schema_bundle_sha256 || - expected.validatorBuild !== entry.validator_build || expected.primary !== entry.primary ) { throw new Error(`verification marker artifact is not a declared output ${entry.path}`); @@ -5880,22 +5882,6 @@ function assertVerifiedDependency( `artifact-contract failure: verified dependency artifact is missing ${entry.path}`, MAX_VERIFIED_ARTIFACT_BYTES ); - const definition = artifactContractDefinition(entry.contract as Parameters[0]); - if (definition.digest !== entry.contract_digest) { - throw new Error(`verified dependency contract changed ${entry.path}`); - } - const currentBinding = artifactContractSchemaBinding( - entry.contract as Parameters[0] - ); - if ( - currentBinding?.schema_file !== entry.schema_file || - currentBinding?.schema_id !== entry.schema_id || - currentBinding?.schema_sha256 !== entry.schema_sha256 || - currentBinding?.schema_bundle_sha256 !== entry.schema_bundle_sha256 || - currentBinding?.validator_build !== entry.validator_build - ) { - throw new Error(`verified dependency schema binding changed ${entry.path}`); - } const artifactSha = createHash("sha256").update(artifactSnapshot.bytes).digest("hex"); if (artifactSha !== entry.sha256) { throw new Error(`verified dependency artifact changed ${entry.path}`); diff --git a/packages/runtime/src/terminal-report.ts b/packages/runtime/src/terminal-report.ts index 46eb8862a..862b59d1b 100644 --- a/packages/runtime/src/terminal-report.ts +++ b/packages/runtime/src/terminal-report.ts @@ -8,8 +8,6 @@ import { assertRunMetadataDocument, assertRunStateDocument, assertSealedPlannedGraph, - artifactContractDefinition, - artifactContractSchemaBinding, executeSemanticGate, layoutForRunRoot, parseSmithersTaskManifestBytes, @@ -236,7 +234,6 @@ function loadTerminalReportInputs(runRoot: string): TerminalReportInputs { if (state.provenance?.workflow.controlGeneration !== sha256Bytes(authority.workflow_control_seal.bytes)) { throw invalidAuthority("terminal report control generation differs from the current seal"); } - assertCurrentContractBindings(authority); const events = replayEvents(layout, Number.MAX_SAFE_INTEGER).records; const stopped = requireStoppedWorkflowEvent(layout, state, events, authority); const completion = deriveTerminalReportCompletion(authority); @@ -262,23 +259,6 @@ function loadTerminalReportInputs(runRoot: string): TerminalReportInputs { }; } -function assertCurrentContractBindings(authority: VerifiedRunOutputAuthoritySnapshot): void { - const graph = assertSealedPlannedGraph(parseStrictJsonBytes(authority.graph.bytes)); - for (const node of graph.nodes) { - for (const output of node.outputs) { - const binding = artifactContractSchemaBinding(output.contract); - if ( - output.contract_digest !== artifactContractDefinition(output.contract).digest || - (binding !== undefined && - ["schema_file", "schema_id", "schema_sha256", "validator_build"].some( - (field) => Reflect.get(output, field) !== Reflect.get(binding, field) - )) - ) - throw invalidAuthority(`terminal report found schema/control drift for ${node.id}`); - } - } -} - function createTerminalReceipt( inputs: TerminalReportInputs, jsonBytes: Buffer, diff --git a/packages/runtime/src/trusted-cli.ts b/packages/runtime/src/trusted-cli.ts index 68053919d..37905c800 100644 --- a/packages/runtime/src/trusted-cli.ts +++ b/packages/runtime/src/trusted-cli.ts @@ -268,7 +268,6 @@ export function runTrustedJsonValidatorPreflight(input: { layout: RunLayout; tru schemaId: paths.schemaId, schemaSha256: paths.schemaSha256, schemaBundleSha256: metadata.schema_bundle_sha256, - validatorBuild: metadata.validator_build, artifactSha256: paths.fixtureSha256 }); assertTrustedCliLauncher({ layout: input.layout, launcherPath: input.trusted.launcherPath }); @@ -310,11 +309,8 @@ export function assertTrustedCliLauncher(input: { layout: RunLayout; launcherPat // closure on every dispatch. const authenticatedClosurePreflights = new Set(); -function preflightTrustedCliClosure( - closure: TrustedCliClosure, - expected: { validator_build: string; schema_bundle_sha256: string } -): void { - const authenticated = [closure.digest, expected.validator_build, expected.schema_bundle_sha256].join("\0"); +function preflightTrustedCliClosure(closure: TrustedCliClosure, expected: { schema_bundle_sha256: string }): void { + const authenticated = [closure.digest, expected.schema_bundle_sha256].join("\0"); if (authenticatedClosurePreflights.has(authenticated)) return; const paths = closurePreflightPaths(closure); const stdout = execFileSync( @@ -344,7 +340,6 @@ function preflightTrustedCliClosure( schemaId: paths.schemaId, schemaSha256: paths.schemaSha256, schemaBundleSha256: expected.schema_bundle_sha256, - validatorBuild: expected.validator_build, artifactSha256: paths.fixtureSha256 }); authenticatedClosurePreflights.add(authenticated); diff --git a/packages/runtime/src/verified-output.ts b/packages/runtime/src/verified-output.ts index beabc6301..e893a540a 100644 --- a/packages/runtime/src/verified-output.ts +++ b/packages/runtime/src/verified-output.ts @@ -9,8 +9,6 @@ import { assertNoSymlinkComponents, assertPathInside, assertRegularFileInside, - artifactContractDefinition, - artifactContractSchemaBinding, boundArtifactValidationWarnings, layoutForRunRoot, parseStrictJsonBytes, @@ -39,7 +37,11 @@ import { type SmithersTaskManifestTask } from "@ultrafuzz/artifacts"; -import { verifyRequiredArtifactsForAttempt, type ArtifactGateAttemptAuthority } from "./artifact-gates.js"; +import { + verifyRequiredArtifactSchemaBinding, + verifyRequiredArtifactsForAttempt, + type ArtifactGateAttemptAuthority +} from "./artifact-gates.js"; import { authenticatedDependencyAdmissionAttemptIds } from "./dependency-admission.js"; import { loadGoalSearchCoverageSnapshot, projectCanonicalFinalReport } from "./final-report-markdown.js"; import { verifySealedTaskManifestSnapshot, type VerifiedSealedTaskManifestSnapshot } from "./workflow-integrity.js"; @@ -369,7 +371,7 @@ function loadFinalizedNodeOutputAuthority(input: LoadVerifiedNodeOutputInput): F const artifactDir = safeResolveInside(layout.artifactsDir, candidate.attemptId, "verified artifact directory"); const publicationSnapshots = readAndBindPublications(artifactDir, plannedNode, documents); - const outputSnapshots = validatePlannedOutputSnapshots(artifactDir, plannedNode, publicationSnapshots); + const outputSnapshots = validatePlannedOutputSnapshots(layout, artifactDir, plannedNode, publicationSnapshots); assertPropertyCampaignEvidencePublications(outputSnapshots, publicationSnapshots); const manifestSeal = finalizedManifestSeal(candidate.state); @@ -1170,15 +1172,28 @@ function readAndBindGateContextFiles( } function validatePlannedOutputSnapshots( + layout: RunLayout, artifactDir: string, plannedNode: PlannedGraphNodeDocument, publications: ReadonlyMap ): VerifiedOutputArtifactSnapshot[] { return plannedNode.outputs.map((output) => { - assertCurrentContractBinding(output); const publication = publications.get(output.path); if (publication === undefined) throw invalidAuthority(`planned output is not published: ${output.path}`); - const validation = validateArtifactContractBytes(output.contract, publication.bytes, publication.absolutePath); + // A schema-backed output is validated against the schema content it was planned with, not this + // build's, so a later rebuild or upgrade cannot turn an already verified output invalid (#921). + const schemaIssues = verifyRequiredArtifactSchemaBinding( + layout, + publication.absolutePath, + output, + publication.bytes + ); + const validation = + schemaIssues.length > 0 + ? { ok: false, issues: schemaIssues, value: undefined } + : output.schema_file === undefined + ? validateArtifactContractBytes(output.contract, publication.bytes, publication.absolutePath) + : { ok: true, issues: [], value: parseStrictJsonBytes(publication.bytes) }; if (!validation.ok) { throw invalidOutput( `verified output failed ${output.contract} validation for ${output.path}: ${validation.issues @@ -1343,27 +1358,6 @@ function assertRunAuthorityIdentity(layout: RunLayout, state: RunState): void { } } -function assertCurrentContractBinding(output: PlannedGraphOutput): void { - const definition = artifactContractDefinition(output.contract); - if (definition.digest !== output.contract_digest) { - throw invalidAuthority(`current contract digest changed for ${output.path}`); - } - const binding = artifactContractSchemaBinding(output.contract); - const actualBinding = schemaBinding(output); - if ( - binding === undefined - ? actualBinding !== undefined - : actualBinding === undefined || - binding.schema_file !== actualBinding.schema_file || - binding.schema_id !== actualBinding.schema_id || - binding.schema_sha256 !== actualBinding.schema_sha256 || - binding.validator_build !== actualBinding.validator_build || - !/^[0-9a-f]{64}$/u.test(actualBinding.schema_bundle_sha256) - ) { - throw invalidAuthority(`current JSON Schema binding changed for ${output.path}`); - } -} - function sameOutputContracts( actual: readonly ArtifactManifestOutputContract[], expected: readonly PlannedGraphOutput[] @@ -1388,17 +1382,6 @@ function withoutSha256(entry: ArtifactVerificationEntry): Omit { - if (output.schema_file === undefined) return undefined; - return Object.freeze({ - schema_file: output.schema_file, - schema_id: output.schema_id!, - schema_sha256: output.schema_sha256!, - schema_bundle_sha256: output.schema_bundle_sha256!, - validator_build: output.validator_build! - }); -} - function concreteNodeIdFromManifest(manifest: ArtifactManifest): string | undefined { const metadata: unknown = manifest.provenance.metadata; return isRecord(metadata) && typeof metadata.concrete_node_id === "string" ? metadata.concrete_node_id : undefined; diff --git a/packages/runtime/src/workflow-integrity.ts b/packages/runtime/src/workflow-integrity.ts index 4dcae3d23..c5048d05c 100644 --- a/packages/runtime/src/workflow-integrity.ts +++ b/packages/runtime/src/workflow-integrity.ts @@ -672,8 +672,7 @@ export function verifyWorkflowControlSnapshot( contents, readBoundedRegularFile(layout.root, layout.statePath, "run state"), executionFiles.find((file) => file.snapshotPath === "controls/plan.json")?.contents, - runtimeStateNodeIds, - true + runtimeStateNodeIds ); } catch (error) { reportDivergence( @@ -2085,12 +2084,9 @@ function deriveWorkflowControlBindings( >, stateContents: Buffer, planContents: Buffer | undefined, - runtimeStateNodeIds?: readonly string[], - allowHistoricalSchemaBundle = false + runtimeStateNodeIds?: readonly string[] ): WorkflowControlBindings { - const graph = (allowHistoricalSchemaBundle ? assertSealedPlannedGraph : assertPlannedGraph)( - parseStrictJsonBytes(contents.graph) - ); + const graph = assertPlannedGraph(parseStrictJsonBytes(contents.graph)); const expandedGraph = assertExpandedGraphSchema(parseStrictJsonBytes(contents.expanded_graph)); const tasksDocument = parseSmithersTaskManifestBytes(contents.tasks); assertSmithersTaskManifestMatchesPlannedGraph(tasksDocument, graph); diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 703a32d38..508f2dc84 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -4674,16 +4674,32 @@ test("project discovery gate requires ledger evidence to survive in the markdown ); }); -test("artifact validation rejects a persisted schema binding that differs from the current registry", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-schema-binding-mismatch" }); +test("artifact validation binds schema content and treats the validator build as provenance", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-schema-binding-provenance" }); const artifactDir = getNodeArtifactDir(layout, "strategy-a", { create: true }); const node = plannedNode(["findings.json"]); - node.outputs[0]!.schema_sha256 = "0".repeat(64); + const [findings] = node.outputs; + assert.ok(findings); + // #921: an output planned by another validator build of the same schemas is still gated on its + // content. A validator rebuild after launch must neither pass an invalid artifact nor fail a valid one. + findings.validator_build = `ultrafuzz-json-validator.v1:${"9".repeat(64)}`; fs.writeFileSync(path.join(artifactDir, "findings.json"), "[]\n", "utf8"); + const rebuilt = verifyRequiredArtifactsForAttempt(layout, node, node.id); + assert.equal(rebuilt.ok, true, JSON.stringify(rebuilt.diagnostics)); + fs.writeFileSync(path.join(artifactDir, "findings.json"), "{}\n", "utf8"); + const invalid = verifyRequiredArtifactsForAttempt(layout, node, node.id); + assert.equal(invalid.ok, false); + assert.ok(invalid.diagnostics.some((diagnostic) => diagnostic.code === "JSON_SCHEMA_VIOLATION")); - const result = verifyRequiredArtifactsForAttempt(layout, node, node.id); - assert.equal(result.ok, false); - assert.ok(result.diagnostics.some((diagnostic) => diagnostic.code === "ARTIFACT_SCHEMA_BINDING_MISMATCH")); + // A planned schema digest that the planned bundle does not contain cannot be validated at all. + fs.writeFileSync(path.join(artifactDir, "findings.json"), "[]\n", "utf8"); + findings.schema_sha256 = "0".repeat(64); + const unknownSchema = verifyRequiredArtifactsForAttempt(layout, node, node.id); + assert.equal(unknownSchema.ok, false); + assert.ok( + unknownSchema.diagnostics.some((diagnostic) => diagnostic.code === "ARTIFACT_VALIDATOR_IDENTITY_MISMATCH"), + JSON.stringify(unknownSchema.diagnostics) + ); }); test("project discovery gate accepts a repository-root scan probe", () => { diff --git a/packages/runtime/test/generated-workflow-verifier.test.ts b/packages/runtime/test/generated-workflow-verifier.test.ts index 839f7f256..56927575f 100644 --- a/packages/runtime/test/generated-workflow-verifier.test.ts +++ b/packages/runtime/test/generated-workflow-verifier.test.ts @@ -1632,7 +1632,7 @@ test("generated Smithers verifier rejects zero-byte generated-test companions", const helper = source.slice(helperStart, verifierStart); assert.match( source, - /artifactContractDefinition,[\s\S]*assertRegularFileInside,[\s\S]*validateArtifactContractBytes/u + /artifactContractSchemaBinding,[\s\S]*assertRegularFileInside,[\s\S]*validateArtifactContractBytes/u ); assert.match(source, /validateArtifactContractBytes,[\s\S]*writeFileDurable[\s\S]*= await import/u); assert.doesNotMatch(source, /validateArtifactContract\(/u); @@ -6702,7 +6702,7 @@ test("generated Smithers workflow binds every planned output to the preflighted ); assert.match(binding, /for \(const output of task\.outputs\)/u); assert.match(binding, /artifactContractSchemaBinding\(/u); - for (const field of ["schemaFile", "schemaId", "schemaSha256", "schemaBundleSha256", "validatorBuild"]) { + for (const field of ["schemaFile", "schemaId", "schemaSha256", "schemaBundleSha256"]) { assert.match(binding, new RegExp(`output\\.${field}`, "u")); } assert.match(source.slice(artifactsImportStart, artifactsImportEnd), /parseJsonValidatorPreflightSuccessEnvelope/u); @@ -6713,6 +6713,43 @@ test("generated Smithers workflow binds every planned output to the preflighted ); }); +test("generated task preparation binds output schema content but not the validator build", () => { + const source = fs.readFileSync(workflowTemplatePath, "utf8"); + const start = source.indexOf("function assertTaskOutputSchemaBindings"); + const end = source.indexOf("\n}\n", start) + 2; + assert.ok(start >= 0 && end > start, source); + const helper = ts.transpileModule(source.slice(start, end), { + compilerOptions: { module: ts.ModuleKind.None, target: ts.ScriptTarget.ES2022 } + }).outputText; + const planned = artifactContractSchemaBinding("ultrafuzz/findings@2"); + assert.ok(planned); + let current = planned; + const assertTaskOutputSchemaBindings = new Function( + "artifactContractSchemaBinding", + `${helper}; return assertTaskOutputSchemaBindings;` + )(() => current) as (task: unknown) => void; + const task = { + outputs: [ + { + path: "findings.json", + contract: "ultrafuzz/findings@2", + schemaFile: planned.schema_file, + schemaId: planned.schema_id, + schemaSha256: planned.schema_sha256, + schemaBundleSha256: planned.schema_bundle_sha256, + validatorBuild: planned.validator_build + } + ] + }; + assert.doesNotThrow(() => assertTaskOutputSchemaBindings(task)); + // #921: a rebuild after launch changed only the validator modules. The schemas the verifier will + // use are still the planned ones, so preparation proceeds. + current = { ...planned, validator_build: `ultrafuzz-json-validator.v1:${"9".repeat(64)}` }; + assert.doesNotThrow(() => assertTaskOutputSchemaBindings(task)); + current = { ...planned, schema_sha256: "0".repeat(64) }; + assert.throws(() => assertTaskOutputSchemaBindings(task), /planned schema binding changed for findings\.json/u); +}); + test("generated validator preflight budgets a contended CLI start and reports the wall time it spent", () => { const timedOut = Object.assign(new Error("spawnSync ultrafuzz ETIMEDOUT"), { code: "ETIMEDOUT" }); const harness = loadJsonValidatorPreflight({ failure: timedOut }); @@ -9870,6 +9907,9 @@ test("generated Smithers dependency verification fails closed before descendant schema_bundle_sha256: schemaBinding.schemaBundleSha256, validator_build: schemaBinding.validatorBuild }; + // What the running build reports for these contracts; a rebuild or upgrade changes it mid-run. + let currentContractDigest = "a".repeat(64); + let currentSchemaBinding: Record = markerSchemaBinding; const taskSpecs = [ { attemptId: "property-specification-fanin", @@ -9957,8 +9997,11 @@ test("generated Smithers dependency verification fails closed before descendant taskSpecs, () => ({ ok: true, issues: [] }), () => undefined, - (contract: string) => ({ digest: "a".repeat(64), format: contract === "ultrafuzz/text@1" ? "text" : "json" }), - (contract: string) => (contract === "ultrafuzz/text@1" ? undefined : markerSchemaBinding), + (contract: string) => ({ + digest: currentContractDigest, + format: contract === "ultrafuzz/text@1" ? "text" : "json" + }), + (contract: string) => (contract === "ultrafuzz/text@1" ? undefined : currentSchemaBinding), (contract: string, contents: Uint8Array) => ({ ok: true, issues: [], @@ -10068,12 +10111,6 @@ test("generated Smithers dependency verification fails closed before descendant ...validArtifact, contract: "ultrafuzz/findings@2" } - ], - [ - { - ...validArtifact, - contract_digest: "b".repeat(64) - } ] ]) { writeMarker(artifacts); @@ -10089,6 +10126,12 @@ test("generated Smithers dependency verification fails closed before descendant /artifact dependency has not passed verification property-specification-fanin/u ); fs.writeFileSync(path.join(dependency, "properties.json"), verifiedBytes); + // #921: the marker's contract digest records the build that planned the output. A different one, + // from the declaration or from the running build, does not make the verified producer inadmissible. + writeMarker([{ ...validArtifact, contract_digest: "b".repeat(64) }]); + currentContractDigest = "e".repeat(64); + assert.doesNotThrow(() => assertVerifiedDependency(task, dependency)); + currentContractDigest = "a".repeat(64); writeMarker([validArtifact]); const verifiedAuthority = assertVerifiedDependency(task, dependency); assert.equal(verifiedAuthority.attemptId, "property-specification-fanin"); @@ -10155,6 +10198,35 @@ test("generated Smithers dependency verification fails closed before descendant [{ path: generatedArtifact.path, sha256: generatedArtifact.sha256 }, companionPublication, supportPublication] ); assert.doesNotThrow(() => assertVerifiedDependency(task, generatedDependency)); + // #921: a rebuilt validator or an unrelated schema edit changes the running build's bundle digest + // and validator build, and a refreshed controller may declare them too. Both are provenance; the + // marker must still name the declared schema content, and its bytes stay pinned by sha256. + currentSchemaBinding = { + ...markerSchemaBinding, + schema_bundle_sha256: "f".repeat(64), + validator_build: `ultrafuzz-json-validator.v1:${"9".repeat(64)}` + }; + assert.doesNotThrow(() => assertVerifiedDependency(task, generatedDependency)); + currentSchemaBinding = markerSchemaBinding; + const generatedPublications = [ + { path: generatedArtifact.path, sha256: generatedArtifact.sha256 }, + companionPublication, + supportPublication + ]; + const rebuiltMarkerArtifact = { + ...generatedArtifact, + schema_bundle_sha256: "f".repeat(64), + validator_build: "ultrafuzz-json-validator.v1:rebuilt" + }; + writeAttemptMarker("generated-tests-fanin", [rebuiltMarkerArtifact], generatedPublications); + assert.doesNotThrow(() => assertVerifiedDependency(task, generatedDependency)); + const otherSchemaMarkerArtifact = { ...generatedArtifact, schema_sha256: "0".repeat(64) }; + writeAttemptMarker("generated-tests-fanin", [otherSchemaMarkerArtifact], generatedPublications); + assert.throws( + () => assertVerifiedDependency(task, generatedDependency), + /artifact dependency has not passed verification generated-tests-fanin/u + ); + writeAttemptMarker("generated-tests-fanin", [generatedArtifact], generatedPublications); fs.writeFileSync(companionPath, "contract Tampered {}\n", "utf8"); assert.throws( () => assertVerifiedDependency(task, generatedDependency), diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 357f3466d..cfb8856c7 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -118,7 +118,7 @@ import { sealedBunStartupControlDrift } from "../src/workflow-integrity.js"; import { linkedWorkflowExecutionEnvironment } from "../src/start-run.js"; -import { verifyRequiredArtifactsForAttempt } from "../src/artifact-gates.js"; +import { verifyRequiredArtifactSchemaBinding, verifyRequiredArtifactsForAttempt } from "../src/artifact-gates.js"; import { projectCanonicalFinalReport } from "../src/final-report-markdown.js"; import { loadReportSnapshot } from "../src/unverified-report.js"; import { projectArtifactSchemaDir, projectArtifactSchemaJson } from "../src/init.js"; @@ -13423,72 +13423,181 @@ test("a sealed manifest that stops re-deriving leaves status readable while nati ); }); -test("a planned graph that stops matching this build's contracts leaves status readable", async () => { +// Appends a comment to one of the five compiled modules VALIDATOR_BUILD_IDENTITY hashes, exactly as a +// comment-only rebuild does, but only for the child process that loads it. +const REBUILT_VALIDATOR_PRELOAD = String.raw` +import fs from "node:fs"; +const rebuilt = process.env.ULTRAFUZZ_TEST_REBUILT_VALIDATOR_MODULE; +const readFileSync = fs.readFileSync; +fs.readFileSync = function (file, ...rest) { + const contents = readFileSync.call(this, file, ...rest); + return file === rebuilt && Buffer.isBuffer(contents) ? Buffer.concat([contents, Buffer.from("\n// rebuilt\n")]) : contents; +}; +`; + +const REBUILT_VALIDATOR_OPERATOR = String.raw` +import { execFileSync } from "node:child_process"; +import fs from "node:fs"; +import path from "node:path"; +const input = JSON.parse(fs.readFileSync(process.argv[2], "utf8")); +const artifacts = await import(input.artifactsModule); +const runtime = await import(input.runtimeModule); +const lifecycle = { projectRoot: input.project, runId: input.runId, env: input.env }; +const describe = (diagnostics) => diagnostics.map((diagnostic) => diagnostic.code + ": " + diagnostic.message); +const messages = (outcome) => (outcome.ok ? [] : describe(outcome.diagnostics)); +const result = { validatorBuild: artifacts.VALIDATOR_BUILD_IDENTITY }; +// Each step records its own outcome, so one failure cannot hide what the others do. +const step = async (name, run) => { + try { + result[name] = await run(); + } catch (error) { + result[name] = { thrown: error instanceof Error ? error.message : String(error) }; + } +}; +const layout = artifacts.layoutForRunRoot(input.runRoot); +const findingsPath = path.join(layout.artifactsDir, "project-discovery", "findings.json"); +const gate = () => { + const graph = artifacts.readPlannedGraphDocument(layout.graphPath); + const node = graph.nodes.find((entry) => entry.id === "project-discovery"); + const tasks = JSON.parse(fs.readFileSync(path.join(layout.root, "smithers", "tasks.json"), "utf8")).tasks; + const task = tasks.find((entry) => entry.attemptId === node.id); + const outcome = runtime.verifyRequiredArtifactsForAttempt(layout, node, node.id, { task, tasks }); + return { ok: outcome.ok, diagnostics: describe(outcome.diagnostics) }; +}; +await step("strict", async () => messages(await runtime.readLinkedWorkflowEvidence(input.project, input.runId))); +await step("health", async () => { + const health = await runtime.getRunHealth(lifecycle); + return { ok: health.ok, diagnostics: describe(health.diagnostics) }; +}); +await step("validGate", gate); +const validFindings = fs.readFileSync(findingsPath); +fs.writeFileSync(findingsPath, "{}\n"); +await step("invalidGate", gate); +fs.writeFileSync(findingsPath, validFindings); +await step("preflight", () => { + const envelope = execFileSync(path.join(layout.root, "trusted-bin", "ultrafuzz"), ["json", "validate", "--json"], { + encoding: "utf8" + }); + artifacts.parseJsonValidatorPreflightSuccessEnvelope(Buffer.from(envelope, "utf8")); + return "ok"; +}); +await step("resume", async () => + messages(await runtime.resumeRun({ ...lifecycle, ultrafuzzCliEntrypoint: input.cliEntrypoint })) +); +await step("pause", async () => { + const paused = await runtime.pauseRun(lifecycle); + return paused.ok ? paused.value.status : messages(paused); +}); +await step("cancel", async () => { + const cancelled = await runtime.cancelRun({ ...lifecycle, env: input.cancelEnv }); + return cancelled.ok ? cancelled.value.status : messages(cancelled); +}); +process.stdout.write(JSON.stringify(result)); +`; + +test("a validator rebuild after launch leaves lifecycle commands, status and schema gates working", async () => { + // #921: VALIDATOR_BUILD_IDENTITY hashes five compiled validator modules and is recorded in every + // planned output. A comment-only rebuild of any of them used to make strict evidence, and so + // pause/cancel/replay/fork, fail with "planned graph output schema binding changed", and made every + // schema-backed gate and validator preflight of a resumed run fail. The operator's commands run in + // a child process whose rebuilt module yields a different identity; the run was sealed by this one. const project = tempProject(); initProject({ projectRoot: project, force: true }); - writeSmallTopology(project); + writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); const env = fakeSmithersEnv(project); - const runId = "diverged-contract-binding-status"; + const runId = "rebuilt-validator-lifecycle"; const run = await startRun({ projectRoot: project, runId, env }); assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + const runRoot = run.value?.run_root; + assert.ok(runRoot); + writeRequiredArtifactSet(runRoot, "project-discovery", [GENERIC_RUNTIME_MARKDOWN_PATH, "findings.json"]); + const plannedBuilds = readPlannedGraphDocument(path.join(runRoot, "graph.json")).nodes.flatMap((node) => + node.outputs.flatMap((output) => (output.validator_build === undefined ? [] : [output.validator_build])) + ); + assert.ok(plannedBuilds.length > 0); + assert.ok(plannedBuilds.every((build) => build === VALIDATOR_BUILD_IDENTITY)); - // `assertSealedPlannedGraph` looks every output contract up in the *running process's* artifact - // registry and insists the schema identity recorded at compile time still matches. That is a - // property of the tree doing the observing, not of the run: an operator whose checkout has moved on - // since the run was submitted loses `status` for a run that is intact and possibly still executing - // (issue #866). Rewriting the recorded schema digest reproduces exactly that disagreement without - // needing two builds. - const graphPath = path.join(run.value!.run_root, "graph.json"); - const graph = JSON.parse(fs.readFileSync(graphPath, "utf8")) as { - nodes: { outputs: { path: string; schema_sha256?: string }[] }[]; + const artifactsModule = import.meta.resolve("@ultrafuzz/artifacts"); + const scripts = temporaryRoot("ufz-rebuilt-validator-"); + fs.writeFileSync(path.join(scripts, "preload.mjs"), REBUILT_VALIDATOR_PRELOAD, "utf8"); + fs.writeFileSync(path.join(scripts, "operator.mjs"), REBUILT_VALIDATOR_OPERATOR, "utf8"); + fs.writeFileSync( + path.join(scripts, "input.json"), + JSON.stringify({ + project, + runId, + runRoot, + env, + artifactsModule, + runtimeModule: new URL("../src/index.js", import.meta.url).href, + cliEntrypoint: fakeUltrafuzzCliEntrypoint(project), + cancelEnv: fakeLifecycleSmithersEnv(project, { + inspect: workflowInspect({ + workflowRunId: `ultrafuzz-${runId}`, + status: "running", + state: "running", + steps: [{ id: "node:project-discovery", state: "in-progress" }] + }) + }) + }), + "utf8" + ); + const operator = JSON.parse( + execFileSync( + process.execPath, + [ + "--import", + pathToFileURL(path.join(scripts, "preload.mjs")).href, + path.join(scripts, "operator.mjs"), + path.join(scripts, "input.json") + ], + { + cwd: project, + encoding: "utf8", + env: { + ...process.env, + ULTRAFUZZ_TEST_REBUILT_VALIDATOR_MODULE: path.join( + path.dirname(fileURLToPath(artifactsModule)), + "strict-json.js" + ) + }, + maxBuffer: 16 * 1024 * 1024, + timeout: 600_000 + } + ) + ) as { + validatorBuild: string; + strict: unknown; + health: { ok: boolean; diagnostics: string[] }; + validGate: { ok: boolean; diagnostics: string[] }; + invalidGate: { ok: boolean; diagnostics: string[] }; + preflight: unknown; + resume: unknown; + pause: unknown; + cancel: unknown; }; - const output = graph.nodes.flatMap((node) => node.outputs).find((candidate) => candidate.path === "findings.json"); - assert.ok(output, "fixture has no schema-bound output to diverge"); - output!.schema_sha256 = "a".repeat(64); - fs.writeFileSync(graphPath, `${JSON.stringify(graph, null, 2)}\n`, "utf8"); + const report = JSON.stringify(operator, null, 2); - // Strict linked-evidence callers keep reporting the schema drift. Ordinary resume intentionally - // does not use current artifact bindings as authorization for same-ID Smithers continuation. - const strict = await readLinkedWorkflowEvidence(project, runId); - assert.equal(strict.ok, false); - const resumed = await resumeRun({ projectRoot: project, runId, env }); - assert.equal(resumed.ok, true, JSON.stringify(resumed.diagnostics)); - const cancelled = await cancelRun({ projectRoot: project, runId, env }); - assert.equal(cancelled.ok, false); - assert.equal(cancelled.diagnostics[0]?.code, "WORKFLOW_CONTROL_EVIDENCE_INVALID"); - - const observed = await readLinkedWorkflowEvidence(project, runId, { tolerateControlDivergence: true }); - assert.equal(observed.ok, true, JSON.stringify(observed.ok ? [] : observed.diagnostics)); - if (observed.ok) { - assert.ok( - observed.verifiedControl.divergences.some((divergence) => - /sealed planned graph no longer re-derives against this build's artifact contracts: planned graph output schema binding changed for "findings\.json"/u.test( - divergence - ) - ), - JSON.stringify(observed.verifiedControl.divergences) - ); - } - - const health = await getRunHealth({ projectRoot: project, runId, env }); - assert.equal(health.ok, true, JSON.stringify(health.diagnostics)); - assert.equal(health.value?.run_id, runId); - const diverged = health.diagnostics.filter((diagnostic) => diagnostic.code === "WORKFLOW_CONTROL_EVIDENCE_DIVERGED"); - assert.ok( - diverged.some((diagnostic) => /output schema binding changed/u.test(diagnostic.message)), - JSON.stringify(health.diagnostics) - ); - assert.ok(diverged.every((diagnostic) => diagnostic.severity === "warning")); - assert.equal( - health.diagnostics.filter((diagnostic) => diagnostic.code === "WORKFLOW_CONTROL_EVIDENCE_INVALID").length, - 0, - JSON.stringify(health.diagnostics) + assert.match(operator.validatorBuild, /^ultrafuzz-json-validator\.v1:[0-9a-f]{64}$/u); + assert.notEqual(operator.validatorBuild, VALIDATOR_BUILD_IDENTITY, "the child must run a rebuilt validator"); + assert.deepEqual(operator.strict, [], report); + assert.equal(operator.health.ok, true, report); + assert.deepEqual( + operator.health.diagnostics.filter((diagnostic) => /^WORKFLOW_CONTROL_EVIDENCE_/u.test(diagnostic)), + [], + report ); - assert.equal( - health.diagnostics.filter((diagnostic) => diagnostic.code === "WORKFLOW_STATE_SYNC_SKIPPED").length, - 1, - JSON.stringify(health.diagnostics) + assert.equal(operator.validGate.ok, true, report); + // Provenance is not a weaker gate: an artifact its planned schema rejects is still rejected. + assert.equal(operator.invalidGate.ok, false, report); + assert.ok( + operator.invalidGate.diagnostics.some((diagnostic) => diagnostic.startsWith("JSON_SCHEMA_VIOLATION")), + report ); + assert.equal(operator.preflight, "ok", report); + assert.deepEqual(operator.resume, [], report); + assert.equal(operator.pause, "pause-requested", report); + assert.equal(operator.cancel, "cancel-requested", report); }); test("sealed Bun startup controls from another build are reported, not equated", () => { @@ -24497,11 +24606,6 @@ test("controller refresh admits a new stock bootstrap module but rejects semanti }); test("artifact gates validate a historical bundle through its active sealed schema snapshot", async () => { - assert.equal( - VALIDATOR_BUILD_IDENTITY, - "ultrafuzz-json-validator.v1:028be3251e9ac213ad6e1c037d8da47c9d903e8be563c8f4fa2149839565bcab", - "a compatibility-only bundle loader must retain the pre-upgrade validator identity" - ); const project = tempProject(); initProject({ projectRoot: project, force: true }); writeSmallTopology(project, GENERIC_RUNTIME_MARKDOWN_PATH); @@ -24538,6 +24642,22 @@ test("artifact gates validate a historical bundle through its active sealed sche const artifactPath = path.join(layout.artifactsDir, "project-discovery", "findings.json"); const schemaPath = path.join(historicalRoot, "modules", "@ultrafuzz", "artifacts", "schema", "findings.schema.json"); + // #921: the planning build also differed from this one in the output's own schema, in a schema file + // this build no longer ships, and in its validator build. None of that may strand the run: the gate + // validates against the planned schema content that the run sealed. + const historicalSchemaRoot = path.dirname(schemaPath); + fs.chmodSync(historicalSchemaRoot, 0o700); + fs.chmodSync(schemaPath, 0o600); + const findingsSchema = JSON.parse(fs.readFileSync(schemaPath, "utf8")) as Record; + findingsSchema.$comment = "historical findings schema fixture"; + fs.writeFileSync(schemaPath, `${JSON.stringify(findingsSchema, null, 2)}\n`, "utf8"); + fs.chmodSync(schemaPath, 0o400); + fs.writeFileSync( + path.join(historicalSchemaRoot, "retired.schema.json"), + `${JSON.stringify({ $schema: "https://json-schema.org/draft/2020-12/schema", $id: "urn:ultrafuzz:schema:retired:1" })}\n`, + { encoding: "utf8", mode: 0o400 } + ); + fs.chmodSync(historicalSchemaRoot, 0o500); const historicalValidation = validateRegisteredJsonBytesSync({ schemaPath, instanceBytes: fs.readFileSync(artifactPath), @@ -24554,10 +24674,21 @@ test("artifact gates validate a historical bundle through its active sealed sche for (const output of node.outputs) { if (output.schema_bundle_sha256 !== undefined) { output.schema_bundle_sha256 = historicalValidation.schema.bundle_sha256; + output.validator_build = `ultrafuzz-json-validator.v1:${"9".repeat(64)}`; } + if (output.path === "findings.json") output.schema_sha256 = historicalValidation.schema.sha256; } const verified = verifyRequiredArtifactsForAttempt(layout, node, node.id); assert.equal(verified.ok, true, JSON.stringify(verified.diagnostics)); + const findingsOutput = node.outputs.find((output) => output.path === "findings.json"); + assert.ok(findingsOutput); + assert.notEqual(findingsOutput.schema_sha256, artifactContractSchemaBinding("ultrafuzz/findings@2")?.schema_sha256); + // Validating against the sealed schema is not a weaker gate: its violations still reject. + const invalid = verifyRequiredArtifactSchemaBinding(layout, artifactPath, findingsOutput, Buffer.from("{}\n")); + assert.ok( + invalid.some((diagnostic) => diagnostic.code === "JSON_SCHEMA_VIOLATION"), + JSON.stringify(invalid) + ); state.provenance!.workflow.controllerExecutionSnapshot = `smithers/execution-snapshots/${"d".repeat(64)}`; writeRunState(layout, state); diff --git a/packages/runtime/test/trusted-cli.test.ts b/packages/runtime/test/trusted-cli.test.ts index 4c0b642f4..f174099f3 100644 --- a/packages/runtime/test/trusted-cli.test.ts +++ b/packages/runtime/test/trusted-cli.test.ts @@ -482,7 +482,11 @@ test("resume rejects missing, tampered, and stale trusted CLI identity", () => { const stale = prepareTrustedCliEnvironment({ layout: staleLayout, cliEntrypoint: entrypoint }); const metadataPath = path.join(staleLayout.root, "trusted-cli.json"); const metadata = JSON.parse(fs.readFileSync(metadataPath, "utf8")) as Record; - metadata.validator_build = "ultrafuzz-json-validator.v1:stale"; + // The recorded validator build is provenance (#921); the schema bundle it validates with still binds. + metadata.validator_build = "ultrafuzz-json-validator.v1:another-build"; + fs.writeFileSync(metadataPath, `${JSON.stringify(metadata, null, 2)}\n`, "utf8"); + runTrustedJsonValidatorPreflight({ layout: staleLayout, trusted: stale }); + metadata.schema_bundle_sha256 = "0".repeat(64); fs.writeFileSync(metadataPath, `${JSON.stringify(metadata, null, 2)}\n`, "utf8"); assert.throws( () => runTrustedJsonValidatorPreflight({ layout: staleLayout, trusted: stale }), @@ -660,7 +664,7 @@ test("legacy migration rejects stale evidence and recovers only the matching sta runTrustedJsonValidatorPreflight({ layout, trusted: recovered }); }); -test("active trusted CLI stays on its sealed transitive build and rejects an incompatible refresh", () => { +test("active trusted CLI stays on its sealed transitive build; refresh adopts a rebuilt validator of the same schemas", () => { const root = temporaryRoot("ultrafuzz-trusted-cli-"); const layout = createRunLayout({ projectRoot: root, runId: "sealed-transitive-build" }); const originalEnvelope = preflightEnvelope(); @@ -670,19 +674,20 @@ test("active trusted CLI stays on its sealed transitive build and rejects an inc const metadataBefore = fs.readFileSync(metadataPath); const closuresRoot = path.join(layout.root, "trusted-cli-closures"); const closuresBefore = fs.readdirSync(closuresRoot).sort(); - - const rebuiltEnvelope = structuredClone(originalEnvelope) as { - data: { schema: { validator_build: string } }; + const rebuild = (mutate: (schema: Record) => void): void => { + const envelope = structuredClone(originalEnvelope) as { data: { schema: Record } }; + mutate(envelope.data.schema); + fs.chmodSync(source.dependencyEntrypoint, 0o600); + fs.writeFileSync( + source.dependencyEntrypoint, + `export const output = ${JSON.stringify(JSON.stringify(envelope))};\n`, + "utf8" + ); + fs.chmodSync(source.dependencyEntrypoint, 0o400); }; - rebuiltEnvelope.data.schema.validator_build = "ultrafuzz-json-validator.v1:rebuilt-transitive"; - fs.chmodSync(source.dependencyEntrypoint, 0o600); - fs.writeFileSync( - source.dependencyEntrypoint, - `export const output = ${JSON.stringify(JSON.stringify(rebuiltEnvelope))};\n`, - "utf8" - ); - fs.chmodSync(source.dependencyEntrypoint, 0o400); + // A rebuild that reports different schemas is refused and leaves the sealed CLI in place. + rebuild((schema) => void (schema.bundle_sha256 = "0".repeat(64))); runTrustedJsonValidatorPreflight({ layout, trusted }); assert.deepEqual(fs.readFileSync(metadataPath), metadataBefore); assert.throws( @@ -696,6 +701,17 @@ test("active trusted CLI stays on its sealed transitive build and rejects an inc ); assert.deepEqual(fs.readFileSync(metadataPath), metadataBefore); assert.deepEqual(fs.readdirSync(closuresRoot).sort(), closuresBefore); + + // #921: a rebuild that only changed the validator build validates the same schemas, so controller + // refresh adopts it instead of leaving the continued run without a trusted validator. + rebuild((schema) => void (schema.validator_build = "ultrafuzz-json-validator.v1:rebuilt-transitive")); + const refreshed = prepareTrustedCliEnvironment({ + layout, + cliEntrypoint: source.entrypoint, + allowIdentityRotation: true + }); + assert.notDeepEqual(fs.readdirSync(closuresRoot).sort(), closuresBefore); + runTrustedJsonValidatorPreflight({ layout, trusted: refreshed }); }); test("trusted CLI closure prefers the authenticated execution generation for workspace dependencies", () => { diff --git a/packages/runtime/test/verified-output.test.ts b/packages/runtime/test/verified-output.test.ts index 89a23a7cc..125cee5f5 100644 --- a/packages/runtime/test/verified-output.test.ts +++ b/packages/runtime/test/verified-output.test.ts @@ -699,6 +699,29 @@ test("verified final-report reader binds immutable current bytes to verifier and ); }); +test("verified reads and the terminal report accept outputs planned by another validator build", () => { + // #921: `validator_build` and `contract_digest` record the build that planned the run. After a + // rebuild changes them, verified reads validate the same schema content and the terminal report is + // still published; both used to refuse with "current JSON Schema binding changed". + const rebuilt = `ultrafuzz-json-validator.v1:${"9".repeat(64)}`; + const outputs = finalReportOutputs().map((output) => ({ + ...output, + contract_digest: "d".repeat(64), + ...(output.validator_build === undefined ? {} : { validator_build: rebuilt }) + })); + assert.ok(outputs.some((output) => output.validator_build === rebuilt)); + const fixture = createVerifiedReportFixture("verified-report-rebuilt-validator", { outputs }); + + assert.deepEqual(loadVerifiedFinalReportSnapshot(fixture.layout.root).json_bytes, fixture.reportBytes); + recordStoppedReportFixture(fixture, "succeeded"); + const published = publishTerminalReport(fixture.layout.root, { + workflowRunId: WORKFLOW_RUN_ID, + workflowState: "succeeded" + }); + assert.equal(published?.terminal, true); + assert.deepEqual(fs.readFileSync(fixture.reportPath), fixture.reportBytes); +}); + test("final-report omissions remain visible as authenticated host diagnostics without rewriting the report", () => { const runId = "verified-report-own-warnings"; const issue = currentIssue(); From 76837712e8961f5e333de423d534b8dc73be9765 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:46:42 +0000 Subject: [PATCH 090/206] docs: describe validator_build as provenance The specification, schema reference, topology reference, CLI reference and artifacts explanation said the host and preflight verify the validator build. They now say the host validates against the planned schema content (the run's sealed snapshot when the installed bundle differs) and records the validator build without comparing it. Adds the CHANGELOG entry for the change. Refs #921 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + docs/SPECS.md | 14 ++++++++----- .../explanation/topology-prompts-artifacts.md | 8 ++++---- docs/reference/cli.md | 5 +++-- docs/reference/topology-yaml.md | 6 ++++-- docs/schemas.md | 20 ++++++++++++------- 6 files changed, 34 insertions(+), 20 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..aa1ace44c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[artifacts] [runtime] [docs]** A rebuild that changes `VALIDATOR_BUILD_IDENTITY` (a comment in one of the five modules it hashes, an ajv bump) no longer strands in-flight runs: the strict run evidence that `pause`, `cancel`, `replay` and `fork` read stops failing with `planned graph output schema binding changed`, `status` stops reporting it as a divergence, and host artifact gates, verified reads, the terminal report and the validator preflight stop failing on the identity difference. The recorded `validator_build` is now provenance only; host gates and verified reads validate each artifact against the schema content it was planned with, using the run's sealed schema snapshot when the installed bundle differs, and that snapshot now loads even when a later build added, removed or re-versioned schema files. Workflows rendered by this version also stop comparing contract digests, bundle digests and validator builds at task preparation and dependency admission; workflows rendered earlier keep their own checks (#921). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/SPECS.md b/docs/SPECS.md index 0fe0c4cea..6bf2d0be2 100644 --- a/docs/SPECS.md +++ b/docs/SPECS.md @@ -243,8 +243,11 @@ Every planned JSON output MUST resolve through the checked-in schema registry. The planned and expanded graph representations MUST persist the schema filename, fragment-free schema ID, schema SHA-256, package schema-bundle SHA-256, and validator build identity. Missing or partial bindings MUST fail planning or host -verification. Operators declare the versioned contract in topology; they MUST -NOT supply these trust identities manually in YAML. +verification. Hosts MUST validate against the schema content a binding names, +using the run's sealed schema snapshot when the installed bundle differs, and +MUST NOT require the recorded validator build or contract digest to equal their +own. Operators declare the versioned contract in topology; they MUST NOT supply +these trust identities manually in YAML. ## Prompts @@ -447,9 +450,10 @@ every non-builtin module whose lexical or physical resolution escapes the closure. Ordinary artifact and schema data reads remain outside this module boundary. Modal MUST provide the equivalent root-owned, read-only entrypoint. Both environments MUST run a real known-valid fixture and -verify the returned schema ID, schema digest, bundle digest, and validator build; -`command -v` alone is insufficient. A missing, tampered, or stale launcher or -closure is a setup failure for new model work. It MUST NOT turn historical +verify the returned schema ID, schema digest, and bundle digest; the returned +validator build is provenance and MUST NOT be compared. `command -v` alone is +insufficient. A missing, tampered, or stale launcher or closure is a setup +failure for new model work. It MUST NOT turn historical seals or schema identities into resume authorization. A current-controller continuation MAY select current validator packages while retaining historical source and artifacts as provenance. diff --git a/docs/explanation/topology-prompts-artifacts.md b/docs/explanation/topology-prompts-artifacts.md index 5c7ad740f..0f3e3b24a 100644 --- a/docs/explanation/topology-prompts-artifacts.md +++ b/docs/explanation/topology-prompts-artifacts.md @@ -67,8 +67,8 @@ This makes handoffs explicit: - The prompt tells the agent what to write. - The topology declares a named, versioned contract for the file. -- Planning binds that contract to an exact schema ID, schema digest, bundle - digest, and validator build. +- Planning binds that contract to an exact schema ID, schema digest, and bundle + digest, and records the validator build that planned it. - Every JSON producer runs the displayed schema-validation command after its final write and before returning. - A deterministic workflow task validates the artifact before dependents start. @@ -77,8 +77,8 @@ This makes handoffs explicit: - Downstream prompts reference it through typed template helpers. The displayed command resolves through a host-managed launcher placed before -target-controlled `PATH` entries. A real fixture preflight checks that launcher, -the schema registry, and the validator build before model work. The command is +target-controlled `PATH` entries. A real fixture preflight checks that launcher +and the schema identity it validates with before model work. The command is producer feedback, not a new inter-node message or validation-receipt schema. The runtime repeats shape validation and then applies contextual gates. diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..d0030be90 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -161,8 +161,9 @@ package-local registries. A registered schema whose filename or bytes differ from its pinned entry is a setup failure. Successful JSON output reports whether the schema was registered plus its fragment-free ID, schema SHA-256, owning package's bundle SHA-256, validator build identity, and the artifact -SHA-256. These identities bind the producer command to the later host check; -they are not a mutable validation receipt. +SHA-256. The schema identities bind the producer command to the later host +check, and the validator build is reported as provenance only; none of them is +a mutable validation receipt. For schema-backed producer tasks, Ultrafuzz places a run-owned trusted launcher before target-controlled `PATH` entries and validates a real known-valid fixture diff --git a/docs/reference/topology-yaml.md b/docs/reference/topology-yaml.md index 65708ce28..1ee29c44a 100644 --- a/docs/reference/topology-yaml.md +++ b/docs/reference/topology-yaml.md @@ -420,8 +420,10 @@ the planned output: - validator build identity. The expanded graph, run state, verification marker, and -`artifact-manifest.json` carry the same identity. The host rejects a missing, -partial, stale, or mismatched binding before publication. +`artifact-manifest.json` carry the same identity. The host rejects a missing or +partial binding, and an artifact its planned schema content rejects, before +publication. The validator build identity is recorded as provenance and is not +compared with the build doing the checking. Topology YAML remains version `2`; the persisted expanded graph uses `graphVersion: "4"` and schema ID diff --git a/docs/schemas.md b/docs/schemas.md index c8e8caedc..ab5840dd6 100644 --- a/docs/schemas.md +++ b/docs/schemas.md @@ -66,20 +66,26 @@ ultrafuzz json validate \ Use repeatable `--ref` flags only for explicitly supplied local dependencies. Bundled sibling schemas resolve offline without flags. The CLI and host use the same non-mutating parser, schema registry, Ajv configuration, resource limits, -schema digest, bundle digest, and validator-build identity. +schema digest, and bundle digest, and each reports its validator-build identity. Every planned JSON output persists the registered schema filename, `$id`, schema SHA-256, owning package's schema-bundle SHA-256, and validator build. The expanded graph, run state, `ultrafuzz.artifact-verification.v2` marker, and -`ultrafuzz.artifact-manifest.v3` repeat that binding. A missing, partial, stale, -or mismatched identity is a host setup/verification failure even when the JSON -would match a different schema with the same general shape. +`ultrafuzz.artifact-manifest.v3` repeat that binding. The host validates each +artifact against the schema content that binding names: the installed schemas +when their bundle digest is the planned one, otherwise the bundle sealed in the +run's execution snapshot. A missing or partial binding, or schema content that +neither holds, is a host verification failure even when the JSON would match a +different schema with the same general shape. The validator build is +provenance: graph reads, host artifact gates, and the validator preflight do not +compare it with the build doing the checking. Schema-backed producers receive a run-owned trusted launcher ahead of target-controlled `PATH` entries. Local and Modal environments use that launcher -to validate a real known-valid fixture and compare the returned schema, bundle, -and validator-build identity before model work. Local launchers resolve only a -verified content-addressed snapshot of the CLI and every transitive package, so +to validate a real known-valid fixture and compare the returned schema and +bundle identity before model work; the returned validator build is provenance. +Local launchers resolve only a verified content-addressed snapshot of the CLI +and every transitive package, so a working-tree rebuild cannot change an active run. Ambient Node loader/search variables are removed and both ESM and CommonJS module resolution must stay inside that snapshot; document reads are unaffected. A path lookup alone is not From 8239556f9cdb542303b4b2aab07bb6f935a21999 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:52:25 +0000 Subject: [PATCH 091/206] docs: note the Super-Linter line-length change in the changelog Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c91263a3d..f22391a4c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[ci] [cli] [docs]** Pull requests now run every release validation lane, including the package gates, the CLI suite, and the workspace typecheck, and the lanes start without waiting for the build job; the PR-only runtime and CLI smoke runners are removed. Pushes to `main` no longer cancel each other's runs, the lanes run under `eatmydata`, vitest packages default to a 30 s test timeout, and CLI tests delete their temporary projects when each test ends. `pnpm -w lint` now enforces the complexity and size budgets on all code, with existing violations counted per file and rule in `eslint-suppressions.json`: a new violation fails lint, and after one is fixed `pnpm -w lint:prune` lowers the recorded count (#960, #923). +- **[ci] [cli] [docs]** Pull requests now run every release validation lane, including the package gates, the CLI suite, and the workspace typecheck, and the lanes start without waiting for the build job; the PR-only runtime and CLI smoke runners are removed. Pushes to `main` no longer cancel each other's runs, the lanes run under `eatmydata`, vitest packages default to a 30 s test timeout, CLI tests delete their temporary projects when each test ends, and Super-Linter's Markdown check no longer enforces line length, which had failed every pull request that edited `CHANGELOG.md`. `pnpm -w lint` now enforces the complexity and size budgets on all code, with existing violations counted per file and rule in `eslint-suppressions.json`: a new violation fails lint, and after one is fixed `pnpm -w lint:prune` lowers the recorded count (#960, #923). - **[config]** `ultrafuzz init` and config loading no longer fail with "project must be a table" (and the same error for every other table) when `smol-toml` 1.9 is installed, as it is for a packed install without the workspace lockfile. That release builds parsed tables with `Object.create(null)`; the loader now reads them as ordinary objects. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. From 93b3419b7accc5e5b189488cd3273cdec14d16ed Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 01:55:44 +0000 Subject: [PATCH 092/206] docs: disable markdownlint line length for CHANGELOG.md Super-linter lints every changed file in full, so any pull request that adds a CHANGELOG entry fails MD013 on the file's existing entries, which are single lines of up to 1,800 characters. This is the same file-level directive #1179 adds, byte for byte, so the two merge without conflict. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index aea90a889..02e556846 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -127,3 +127,5 @@ ## v0.0.1 - First external private release. + + From bbe614a4c92b9a62422a774aa0c8e06adf1645dd Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:17:07 +0000 Subject: [PATCH 093/206] docs(runtime): say what lock-free dynamic expansion relies on and what a race leaves behind The comment that replaced withDynamicExpansionLock said the re-read after publication "refuses" a set a racing publisher made inconsistent. It does, but only after both manifests are durable: two renderers that create different groups at once can take the same sequence, and every later render and strict admission then fails with DYNAMIC_MANIFEST_SET_INVALID until the set is repaired by hand. The comment now says so, and says that creation assumes one renderer per run (Smithers refuses to resume a run whose driver is live before it renders), that renderers which load the same outputs publish identical bytes, and that on a filesystem without hard links a concurrent reader can catch a manifest mid-publication. The retry-archive docstring still justified the archive by the render needing the source artifact, which this branch made untrue. Its remaining job is to let the group expand again from the retried source's new output. Refs #1142 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/dynamic-expansion-retry.ts | 12 ++++++------ packages/runtime/src/dynamic-expansion.ts | 13 ++++++++----- 2 files changed, 14 insertions(+), 11 deletions(-) diff --git a/packages/runtime/src/dynamic-expansion-retry.ts b/packages/runtime/src/dynamic-expansion-retry.ts index a5b570b95..7bb3ce17f 100644 --- a/packages/runtime/src/dynamic-expansion-retry.ts +++ b/packages/runtime/src/dynamic-expansion-retry.ts @@ -61,12 +61,12 @@ export interface DynamicExpansionRetryArchive { * generation, and validate everything that could refuse it. * * Smithers resets the producer and all of its dependents, but the expansion - * manifests live outside Smithers state. Leaving them active makes the next - * workflow render require the canonical source artifact during the gap between - * producer completion and verifier publication. The decision is taken here, - * before the first `timetravel`, so ambiguous or unrecognized manifest state - * fails closed while Smithers state is still untouched. Returns `undefined` - * when no published manifest belongs to a retried source. + * manifests live outside Smithers state. Left in place, they keep the group at + * its published items; withdrawing them lets the group expand again from the + * retried source's new output. The decision is taken here, before the first + * `timetravel`, so ambiguous or unrecognized manifest state fails closed while + * Smithers state is still untouched. Returns `undefined` when no published + * manifest belongs to a retried source. */ export function planDynamicExpansionRetryArchive(input: { projectRoot: string; diff --git a/packages/runtime/src/dynamic-expansion.ts b/packages/runtime/src/dynamic-expansion.ts index 0ea3da7d8..5597b266a 100644 --- a/packages/runtime/src/dynamic-expansion.ts +++ b/packages/runtime/src/dynamic-expansion.ts @@ -244,11 +244,14 @@ export function loadOrCreateDynamicExpansion(input: { assertNoSymlinkComponents(runRoot, manifestDir, "dynamic expansion manifest directory"); const manifestPath = path.join(manifestDir, `${validateSafeId(input.groupNodeId, "dynamic group node ID")}.json`); const sourceArtifactRelativePath = path.relative(runRoot, input.sourceArtifactPath).split(path.sep).join("/"); - // No lock. Reading a published manifest needs none, and creation happens in the workflow render, - // which Smithers runs for one live driver per run. publishFileDurableExclusive never replaces a - // published file, and the re-read after publication refuses a set that a racing second publisher - // made inconsistent. The lock this replaced was never reclaimed, so a process killed while - // holding it failed every later render and lifecycle admission (#1142). + // No lock (#1142): the one this replaced was never reclaimed, so a process killed while holding it + // failed every later render and lifecycle admission. Published manifests are never rewritten, so + // reads need none; on a filesystem without hard links a concurrent reader can still catch one + // mid-publication and fail that read. Creation assumes one renderer per run (Smithers refuses to + // resume a run whose driver is live before it renders). Renderers that load the same outputs + // publish identical bytes, which publishFileDurableExclusive accepts. Two that create different + // groups at once can take the same `sequence`; the re-read after publication detects that but + // cannot undo it, and the run then needs manual repair. const priorManifests = readExpansionManifests(manifestDir); assertManifestSetMatchesInput(priorManifests, { runId: input.runId, From b822474833bae1f6901ba6a06de00f25f4ca2a9a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:17:07 +0000 Subject: [PATCH 094/206] test(runtime): cover the identity check a published expansion now relies on With the source no longer re-hashed, assertCompatibleManifest is the only check between a changed group definition and silently reusing the published items, and the source-change block this branch deleted was the only test that reached DYNAMIC_EXPANSION_CHANGED. The prompt-template and limit test now also changes the source JSON path, key path, node-ID template and template fingerprint of a published group and expects DYNAMIC_EXPANSION_CHANGED for each. It passes on main and on this branch, and fails when that call is made a no-op. The lifecycle test's rewritten source now carries a goal with a different key instead of the published plan plus a newline, so it exercises a different plan rather than only different bytes. The source re-run test and the base-controls retry test built the same sealed run by hand; they now share sealedDynamicRun. Refs #1142 Co-Authored-By: Claude Opus 5.5 --- .../runtime/test/dynamic-expansion.test.ts | 199 ++++++++---------- .../runtime/test/dynamic-lifecycle.test.ts | 7 +- 2 files changed, 90 insertions(+), 116 deletions(-) diff --git a/packages/runtime/test/dynamic-expansion.test.ts b/packages/runtime/test/dynamic-expansion.test.ts index edf19c8b2..5bb5279ea 100644 --- a/packages/runtime/test/dynamic-expansion.test.ts +++ b/packages/runtime/test/dynamic-expansion.test.ts @@ -81,6 +81,69 @@ function item(index: number): Record { }; } +/** A run whose sealed runtime controls fan `goals` from `planner` out into `fanout`, joined by `join`. */ +function sealedDynamicRun(runId: string, goals: Array>) { + const projectRoot = tempDirectory(); + const runRoot = path.join(projectRoot, "runs", runId); + const sourceArtifactPath = path.join(runRoot, "artifacts", "planner", "plan.json"); + const templatePath = path.join(runRoot, "templates", "worker.md"); + const graphPath = path.join(runRoot, "graph.json"); + const tasksPath = path.join(runRoot, "smithers", "tasks.json"); + const baseGraphPath = path.join(runRoot, "smithers", "runtime-base-graph.json"); + const baseTasksPath = path.join(runRoot, "smithers", "runtime-base-tasks.json"); + fs.mkdirSync(path.dirname(sourceArtifactPath), { recursive: true }); + fs.mkdirSync(path.dirname(templatePath), { recursive: true }); + fs.mkdirSync(path.dirname(tasksPath), { recursive: true }); + fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals })}\n`, "utf8"); + fs.writeFileSync(templatePath, "Investigate {{item.goal_prompt}}.\n", "utf8"); + const templateTask = compiledTask(projectRoot, runRoot, "fanout", "fanout", templatePath); + const joinTask = compiledTask(projectRoot, runRoot, "join", "join", undefined, ["fanout"]); + const group: CompiledSmithersDynamicGroup = { + groupNodeId: "fanout", + logicalNodeId: "fanout", + source: { + concreteNodeId: "planner", + attemptId: "planner", + verifierSmithersNodeId: "verify:planner", + artifactPath: sourceArtifactPath + }, + sourcePath: "$.goals", + keyPath: "id", + nodeIdTemplate: "dynamic:item:{{ item.id }}", + templatePath, + templateDigest: digest(fs.readFileSync(templatePath)), + templateFingerprint: digest("template-fingerprint"), + continueOnFail: true, + maxDynamicNodes: 100, + reservedNodeIds: ["planner", "fanout", "join"], + taskTemplates: [templateTask], + promptContext: promptContext(projectRoot, runRoot) + }; + // The seal keeps byte copies of the pre-expansion controls beside the mutable ones. + fs.writeFileSync(graphPath, `${JSON.stringify(plannedGraph(runId))}\n`, "utf8"); + fs.writeFileSync( + tasksPath, + `${JSON.stringify({ schema_version: "1.0", run_id: runId, tasks: [joinTask], dynamic_groups: [group] })}\n`, + "utf8" + ); + fs.copyFileSync(graphPath, baseGraphPath); + fs.copyFileSync(tasksPath, baseTasksPath); + return { + sourceArtifactPath, + controls: { + runId, + projectRoot, + runRoot, + graphPath, + tasksPath, + baseGraphPath, + baseTasksPath, + baseTasks: [joinTask], + groups: [group] + } + }; +} + test("dynamic expansion deterministically persists empty, one-item, and 100-item manifests", () => { for (const count of [0, 1, 100]) { const fixture = expansionFixture({ @@ -245,7 +308,7 @@ test("persisted expansion rejects tampering, transplantation, and symlink manife ); }); -test("persisted expansion rejects prompt-template and dynamic-limit changes", () => { +test("persisted expansion rejects prompt-template, topology-contract, and dynamic-limit changes", () => { const changedTemplate = expansionFixture({ runId: "changed-template", items: [item(0)] }); changedTemplate.invoke(); fs.writeFileSync(changedTemplate.templatePath, "Changed {{item.goal_prompt}}.\n", "utf8"); @@ -254,6 +317,23 @@ test("persisted expansion rejects prompt-template and dynamic-limit changes", () (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_TEMPLATE_CHANGED" ); + // A published manifest is reused without reading the source again, so this comparison is what + // keeps a changed group definition from silently adopting the old items. + const changedContract = expansionFixture({ runId: "changed-contract", items: [item(0)] }); + changedContract.invoke(); + for (const overrides of [ + { sourcePath: "$.other_goals" }, + { keyPath: "goal_prompt" }, + { nodeIdTemplate: "dynamic:other:{{ item.id }}" }, + { templateFingerprint: digest("fingerprint:other") } + ]) { + assert.throws( + () => changedContract.invoke(overrides), + (error: unknown) => error instanceof DynamicExpansionError && error.code === "DYNAMIC_EXPANSION_CHANGED", + JSON.stringify(overrides) + ); + } + const changedLimit = expansionFixture({ runId: "changed-limit", items: [item(0)], maxDynamicNodes: 100 }); changedLimit.invoke(); assert.throws( @@ -382,63 +462,8 @@ test("explicit source retry archives a complete expansion generation and rejects }); test("explicit source retry re-derives the base runtime controls after archiving an expansion", () => { - const runId = "retry-rematerialize"; - const projectRoot = tempDirectory(); - const runRoot = path.join(projectRoot, "runs", runId); - const sourceArtifactPath = path.join(runRoot, "artifacts", "planner", "plan.json"); - const templatePath = path.join(runRoot, "templates", "worker.md"); - const graphPath = path.join(runRoot, "graph.json"); - const tasksPath = path.join(runRoot, "smithers", "tasks.json"); - const baseGraphPath = path.join(runRoot, "smithers", "runtime-base-graph.json"); - const baseTasksPath = path.join(runRoot, "smithers", "runtime-base-tasks.json"); - fs.mkdirSync(path.dirname(sourceArtifactPath), { recursive: true }); - fs.mkdirSync(path.dirname(templatePath), { recursive: true }); - fs.mkdirSync(path.dirname(tasksPath), { recursive: true }); - fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals: [item(0)] })}\n`, "utf8"); - fs.writeFileSync(templatePath, "Investigate {{item.goal_prompt}}.\n", "utf8"); - const templateTask = compiledTask(projectRoot, runRoot, "fanout", "fanout", templatePath); - const joinTask = compiledTask(projectRoot, runRoot, "join", "join", undefined, ["fanout"]); - const group: CompiledSmithersDynamicGroup = { - groupNodeId: "fanout", - logicalNodeId: "fanout", - source: { - concreteNodeId: "planner", - attemptId: "planner", - verifierSmithersNodeId: "verify:planner", - artifactPath: sourceArtifactPath - }, - sourcePath: "$.goals", - keyPath: "id", - nodeIdTemplate: "dynamic:item:{{ item.id }}", - templatePath, - templateDigest: digest(fs.readFileSync(templatePath)), - templateFingerprint: digest("template-fingerprint"), - continueOnFail: true, - maxDynamicNodes: 100, - reservedNodeIds: ["planner", "fanout", "join"], - taskTemplates: [templateTask], - promptContext: promptContext(projectRoot, runRoot) - }; - // The seal keeps byte copies of the pre-expansion controls beside the mutable ones. - fs.writeFileSync(graphPath, `${JSON.stringify(plannedGraph(runId))}\n`, "utf8"); - fs.writeFileSync( - tasksPath, - `${JSON.stringify({ schema_version: "1.0", run_id: runId, tasks: [joinTask], dynamic_groups: [group] })}\n`, - "utf8" - ); - fs.copyFileSync(graphPath, baseGraphPath); - fs.copyFileSync(tasksPath, baseTasksPath); - const controls = { - runId, - projectRoot, - runRoot, - graphPath, - tasksPath, - baseGraphPath, - baseTasksPath, - baseTasks: [joinTask], - groups: [group] - }; + const { controls } = sealedDynamicRun("retry-rematerialize", [item(0)]); + const { runId, projectRoot, runRoot } = controls; const expanded = materializeDynamicRuntime({ ...controls, readyGroupIds: ["fanout"] }); assert.deepEqual(expanded.expandedGroupIds, ["fanout"]); const generated = expanded.tasks.find((task) => task.metadata.node.dynamic !== undefined); @@ -614,61 +639,7 @@ test("runtime materialization preserves required inputs and allows partial revie }); test("a re-run dynamic source keeps the published fan-out for renders and admission", () => { - const runId = "source-rerun"; - const projectRoot = tempDirectory(); - const runRoot = path.join(projectRoot, "runs", runId); - const sourceArtifactPath = path.join(runRoot, "artifacts", "planner", "plan.json"); - const templatePath = path.join(runRoot, "templates", "worker.md"); - const graphPath = path.join(runRoot, "graph.json"); - const tasksPath = path.join(runRoot, "smithers", "tasks.json"); - const baseGraphPath = path.join(runRoot, "smithers", "runtime-base-graph.json"); - const baseTasksPath = path.join(runRoot, "smithers", "runtime-base-tasks.json"); - fs.mkdirSync(path.dirname(sourceArtifactPath), { recursive: true }); - fs.mkdirSync(path.dirname(templatePath), { recursive: true }); - fs.mkdirSync(path.dirname(tasksPath), { recursive: true }); - fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals: [item(0), item(1)] })}\n`, "utf8"); - fs.writeFileSync(templatePath, "Investigate {{item.goal_prompt}}.\n", "utf8"); - const joinTask = compiledTask(projectRoot, runRoot, "join", "join", undefined, ["fanout"]); - const group: CompiledSmithersDynamicGroup = { - groupNodeId: "fanout", - logicalNodeId: "fanout", - source: { - concreteNodeId: "planner", - attemptId: "planner", - verifierSmithersNodeId: "verify:planner", - artifactPath: sourceArtifactPath - }, - sourcePath: "$.goals", - keyPath: "id", - nodeIdTemplate: "dynamic:item:{{ item.id }}", - templatePath, - templateDigest: digest(fs.readFileSync(templatePath)), - templateFingerprint: digest("template-fingerprint"), - continueOnFail: true, - maxDynamicNodes: 100, - reservedNodeIds: ["planner", "fanout", "join"], - taskTemplates: [compiledTask(projectRoot, runRoot, "fanout", "fanout", templatePath)], - promptContext: promptContext(projectRoot, runRoot) - }; - fs.writeFileSync(graphPath, `${JSON.stringify(plannedGraph(runId))}\n`, "utf8"); - fs.writeFileSync( - tasksPath, - `${JSON.stringify({ schema_version: "1.0", run_id: runId, tasks: [joinTask], dynamic_groups: [group] })}\n`, - "utf8" - ); - fs.copyFileSync(graphPath, baseGraphPath); - fs.copyFileSync(tasksPath, baseTasksPath); - const controls = { - runId, - projectRoot, - runRoot, - graphPath, - tasksPath, - baseGraphPath, - baseTasksPath, - baseTasks: [joinTask], - groups: [group] - }; + const { sourceArtifactPath, controls } = sealedDynamicRun("source-rerun", [item(0), item(1)]); const render = (readyGroupIds: string[]) => dynamicRuntimeFingerprint(materializeDynamicRuntime({ ...controls, readyGroupIds })); @@ -683,7 +654,7 @@ test("a re-run dynamic source keeps the published fan-out for renders and admiss assert.equal(render([]), published); assert.equal(render(["fanout"]), published); assert.equal(dynamicRuntimeFingerprint(verifyDynamicRuntimeMaterialization(controls)), published); - const tasks = JSON.parse(fs.readFileSync(tasksPath, "utf8")) as { tasks: CompiledSmithersTask[] }; + const tasks = JSON.parse(fs.readFileSync(controls.tasksPath, "utf8")) as { tasks: CompiledSmithersTask[] }; assert.deepEqual( tasks.tasks.flatMap((task) => task.metadata.node.dynamic?.expansionKey ?? []), ["goal-0", "goal-1"] diff --git a/packages/runtime/test/dynamic-lifecycle.test.ts b/packages/runtime/test/dynamic-lifecycle.test.ts index 8878bcca7..d66d887d0 100644 --- a/packages/runtime/test/dynamic-lifecycle.test.ts +++ b/packages/runtime/test/dynamic-lifecycle.test.ts @@ -781,8 +781,11 @@ test("a re-running dynamic source keeps published controls admissible and observ // Observation must not recreate or otherwise repair the dynamic source on the controller's behalf. assert.equal(fs.existsSync(sourcePath), false); - // The re-run then writes a different plan; admission still follows the published manifest. - fs.writeFileSync(sourcePath, `${publishedPlan}\n`, "utf8"); + // The re-run then writes a plan whose goal has a different key; admission still follows the + // published manifest. + const plan = JSON.parse(publishedPlan) as { threat_goals: Array> }; + plan.threat_goals = plan.threat_goals.map((goal) => ({ ...goal, id: "liquidation:early" })); + fs.writeFileSync(sourcePath, `${JSON.stringify(plan, null, 2)}\n`, "utf8"); const rewritten = await readLinkedWorkflowEvidence(fixture.project, fixture.runId); assert.equal(rewritten.ok, true, "diagnostics" in rewritten ? JSON.stringify(rewritten.diagnostics) : ""); }); From de764c0ee03feff207d3dd13085d8259b8f49d49 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:22:21 +0000 Subject: [PATCH 095/206] fix(runtime): stop Smithers persisting a TaskHeartbeat event per heartbeat write The pinned @smthrs/engine flushHeartbeat() writes the fenced attempt-row heartbeat and, when that write succeeds, appends a TaskHeartbeat event: an _smithers_events row and a stream.ndjson line. Ultrafuzz never attaches heartbeat data, so the event carries nothing the row lacks. A quiet agent writes one per throttled liveness pulse (the watchdog pulses every 250ms and writes are throttled to 500ms). An agent that streams output writes one per ownership check its stdout, stderr and tool callbacks force, and forced writes bypass the throttle: in a scratch run with a stdout write every 10ms, one 3.5s task appended 384-588 of them. Ultrafuzz never acts on these events: it handles only TaskHeartbeatTimeout, the engine's heartbeat timeout advances only when the attempt-row write succeeds, and `smithers why` reads the attempt row. Add an engine_task_heartbeat_event compatibility patch that deletes the emit and keeps the fenced write, registered in SMITHERS_COMPATIBILITY_PATCHES and in the fresh operator-controller patch list. The new integration test runs a one-task workflow under Bun with every registered compatibility patch applied as each Smithers module loads. Its agent owns a quiet 3.5s child under a 3s heartbeat timeout. On main the run records TaskHeartbeat rows (6 to 8 in the runs observed); with this patch it records none, the task still finishes, and the attempt row's heartbeat_at_ms keeps advancing. Refs #1147 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + packages/runtime/src/smithers.ts | 39 +++++ ...hers-controller-engine.integration.test.ts | 135 ++++++++++++++++++ 3 files changed, 175 insertions(+) create mode 100644 packages/runtime/test/smithers-controller-engine.integration.test.ts diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..83bac40cc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime]** The operator controller's Smithers engine no longer appends a `TaskHeartbeat` event (an `_smithers_events` row and a `stream.ndjson` line) after each successful attempt-row heartbeat write. A quiet agent task wrote up to two a second, and a streaming agent one per ownership check that its output forced past the throttle. Liveness, heartbeat timeouts and `smithers why` rest on the fenced attempt-row write, which is kept. Runs launched after upgrading get this (#1147). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..2b6d4ecff 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -1011,6 +1011,35 @@ const SMITHERS_ENGINE_AGENT_EVENT_OWNERSHIP_PATCH = ` const pendingOwnershipChe pendingOwnershipChecks.add(check); };`; +// After each successful fenced attempt-row heartbeat write, the engine appends a +// TaskHeartbeat event (an `_smithers_events` row plus a stream.ndjson line) that +// carries no heartbeat data, because Ultrafuzz never passes any. A quiet agent +// task writes one per throttled liveness pulse, up to two a second; an agent that +// streams output writes one per ownership check its stdout, stderr and tool +// callbacks force past the throttle. In a baseline campaign they were 83% of the +// event rows (#1147), and Ultrafuzz never acts on them (it handles only +// TaskHeartbeatTimeout). Liveness is the attempt-row write: the heartbeat-timeout +// watchdog advances only when it succeeds, and `smithers why` reads the row. Keep +// the write and drop the event. +const SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_SOURCE = ` "heartbeat:record", + ); + await eventBus.emitEventQueued({ + type: "TaskHeartbeat", + runId, + nodeId: desc.nodeId, + iteration: desc.iteration, + attempt: attemptNo, + hasData: heartbeatDataJson !== null, + dataSizeBytes, + intervalMs: intervalMs ?? undefined, + timestampMs: heartbeatAtMs, + }); + } catch (error) {`; +const SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH = ` "heartbeat:record", + ); + // ultrafuzz: the fenced attempt row above is the liveness record (#1147). + } catch (error) {`; + // Every event the engine persists first runs an idempotency probe that // filters `_smithers_events` on (run_id, timestamp_ms, type, payload_json). // The table's only index is its (run_id, seq) primary key, and the probe's @@ -2438,6 +2467,7 @@ export type SmithersCompatibilityPatchId = | "terminal_state_restore" | "resume_hydration" | "engine_agent_event_ownership" + | "engine_task_heartbeat_event" | "engine_agent_usage_progress" | "engine_main_usage_invocation" | "engine_json_correction_usage_invocation" @@ -2585,6 +2615,14 @@ export const SMITHERS_COMPATIBILITY_PATCHES: readonly SmithersCompatibilityPatch // Upstream coalescing its own in-flight proof retires this patch. upstreamAbsent: ["heartbeatOwnershipCheckInFlight"] }, + { + id: "engine_task_heartbeat_event", + packageName: "@smthrs/engine", + sourceRelativePath: "src/engine.js", + patchable: SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_SOURCE, + patched: SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH, + upstreamAbsent: [] + }, { id: "engine_agent_usage_progress", packageName: "@smthrs/engine", @@ -7067,6 +7105,7 @@ export function applySmithersCompatibilityPatches(projectRoot: string): void { SMITHERS_ENGINE_AGENT_EVENT_OWNERSHIP_PATCH, "agent event ownership coalescing" ], + [SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_SOURCE, SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH, "task heartbeat event"], [ SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_SOURCE, SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_PATCH, diff --git a/packages/runtime/test/smithers-controller-engine.integration.test.ts b/packages/runtime/test/smithers-controller-engine.integration.test.ts new file mode 100644 index 000000000..f87261b32 --- /dev/null +++ b/packages/runtime/test/smithers-controller-engine.integration.test.ts @@ -0,0 +1,135 @@ +import assert from "node:assert/strict"; +import { spawnSync } from "node:child_process"; +import fs from "node:fs"; +import { createRequire } from "node:module"; +import path from "node:path"; +import test from "node:test"; + +import { SMITHERS_COMPATIBILITY_PATCHES } from "../src/smithers.js"; +import { temporaryRoot } from "./temporary-root.js"; + +interface ControllerRun { + status: string; + patchedModules: string[]; + eventTypes: Record; + attempts: Array<{ state: string; started_at_ms: number; heartbeat_at_ms: number | null }>; +} + +// Rewrites every Smithers module the operator controller patches as Bun loads it, +// with the same exactly-once replacement, so a workflow runs on the engine +// Ultrafuzz ships without copying or mutating the shared package store. +const CONTROLLER_PATCH_PLUGIN = `import fs from "node:fs"; +import { plugin } from "bun"; + +const modules = JSON.parse(fs.readFileSync(process.env.CONTROLLER_PATCHES, "utf8")); +globalThis.ultrafuzzPatchedModules = []; +plugin({ + name: "ultrafuzz-controller-patches", + setup(build) { + for (const { target, filter, patches } of modules) { + build.onLoad({ filter: new RegExp(filter) }, (args) => { + let source = fs.readFileSync(args.path, "utf8"); + for (const [id, patchable, patched] of patches) { + if (source.split(patchable).length !== 2) throw new Error(id + " does not anchor exactly once"); + source = source.replace(patchable, patched); + } + globalThis.ultrafuzzPatchedModules.push(target); + return { contents: source, loader: "js" }; + }); + } + } +}); +`; + +function runOnControllerEngine(root: string, workflow: string, env: Record = {}): ControllerRun { + const modules = new Map>(); + for (const patch of SMITHERS_COMPATIBILITY_PATCHES) { + const target = `${patch.packageName}/${patch.sourceRelativePath}`; + modules.set(target, [...(modules.get(target) ?? []), [patch.id, patch.patchable, patch.patched]]); + } + fs.writeFileSync( + path.join(root, "controller-patches.json"), + JSON.stringify( + [...modules].map(([target, patches]) => ({ + target, + filter: `/${target.replace(/[.*+?^${}()|[\]\\]/gu, "\\$&")}$`, + patches + })) + ) + ); + fs.writeFileSync(path.join(root, "controller-patches.mjs"), CONTROLLER_PATCH_PLUGIN); + const runner = createRequire(import.meta.url).resolve("smthrs"); + fs.symlinkSync(path.dirname(path.dirname(path.dirname(runner))), path.join(root, "node_modules"), "dir"); + fs.writeFileSync( + path.join(root, "workflow.mjs"), + `import { spawn } from "node:child_process"; +import fs from "node:fs"; +import { Database } from "bun:sqlite"; +import { Effect } from "effect"; +import React from "react"; +import { createSmithers, runWorkflow } from "smthrs"; +import { z } from "zod/v4"; + +const h = React.createElement; +const { Workflow, Worktree, Task, smithers, outputs } = createSmithers( + { result: z.object({ value: z.string() }) }, + { dbPath: "smithers.db" } +); +${workflow} +const run = await Effect.runPromise(runWorkflow(workflow, { input: {}, runId: "controller-engine", rootDir: process.cwd() })); +const db = new Database("smithers.db", { readonly: true }); +const eventTypes = Object.fromEntries( + db.query("SELECT type, count(*) AS n FROM _smithers_events GROUP BY type").all().map((row) => [row.type, row.n]) +); +const attempts = db.query("SELECT state, started_at_ms, heartbeat_at_ms FROM _smithers_attempts").all(); +fs.writeFileSync("result.json", JSON.stringify({ status: run.status, patchedModules: globalThis.ultrafuzzPatchedModules, eventTypes, attempts })); +process.exit(0); +` + ); + const result = spawnSync("bun", ["--preload", "./controller-patches.mjs", "./workflow.mjs"], { + cwd: root, + env: { ...process.env, ...env, CONTROLLER_PATCHES: path.join(root, "controller-patches.json") }, + encoding: "utf8", + maxBuffer: 16 * 1024 * 1024, + timeout: 120_000 + }); + const output = [result.error, result.stdout, result.stderr].map((part) => String(part ?? "").slice(-4_000)); + assert.equal(result.status, 0, output.join("\n")); + const run = JSON.parse(fs.readFileSync(path.join(root, "result.json"), "utf8")) as ControllerRun; + assert.ok(run.patchedModules.includes("@smthrs/engine/src/engine.js"), "the engine was not patched"); + return run; +} + +// #1147: each successful attempt-row heartbeat write also appended a TaskHeartbeat +// event row and stream.ndjson line. This agent is quiet, so only the throttled +// liveness pulse keeps it live, through the fenced attempt-row write that the +// heartbeat timeout and `smithers why` depend on. +test("controller agent tasks stay live on their attempt row without a TaskHeartbeat event per pulse", () => { + const root = temporaryRoot("ufz-controller-heartbeat-"); + const run = runOnControllerEngine( + root, + `const agent = { + id: "sleeping-agent", + async generate(options) { + const child = spawn("sleep", ["3.5"], { stdio: "ignore" }); + options.onProcess?.({ phase: "started", pid: child.pid }); + const exitCode = await new Promise((resolve) => child.on("exit", resolve)); + options.onProcess?.({ phase: "exited", pid: child.pid, exitCode }); + return { text: JSON.stringify({ value: "done" }) }; + } +}; +const workflow = smithers(() => + h(Workflow, { name: "heartbeat" }, + h(Task, { id: "agent", output: outputs.result, agent, heartbeatTimeoutMs: 3_000, retries: 0 }, "work")));` + ); + + // The agent outlived its 3s heartbeat timeout, so the fenced write kept it live. + assert.equal(run.status, "finished"); + assert.equal(run.eventTypes.TaskHeartbeat ?? 0, 0, JSON.stringify(run.eventTypes)); + const [attempt, ...others] = run.attempts; + assert.ok(attempt !== undefined && others.length === 0, JSON.stringify(run.attempts)); + assert.ok( + attempt.heartbeat_at_ms !== null && attempt.heartbeat_at_ms - attempt.started_at_ms >= 2_000, + `the attempt row stopped recording liveness: ${JSON.stringify(attempt)}` + ); +}); From 82b008d913e0eda278343d7dd89f8db08e241838 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:22:41 +0000 Subject: [PATCH 096/206] fix(runtime): stop Smithers fetching and rebasing task worktrees Smithers' treats its value as a branch to track. Creating a worktree runs `git fetch origin` first. Each time a task re-enters an existing worktree, the pinned engine retries `git rebase origin/`, after another `git fetch origin` unless one succeeded for that repository in the last 60s (the worktree sync cache's default TTL). Ultrafuzz passes the recorded launch commit (or the pinned source branch), and origin/ never resolves, so every re-entry by an agent or verifier task logs "worktree sync rebase failed"; a failed rebase is never recorded, so it repeats. The fetches have no timeout, update the user's remote-tracking refs, and are retried on every re-entry while they fail; on this host `git fetch origin` against an unreachable HTTPS remote took 136s to fail. Task worktrees must stay on the launch commit, which assertWorkspaceSourceRevision enforces. Add two engine compatibility patches: engine_worktree_sync makes getWorktreeSyncCache() return an inert cache, so the re-entry path never fetches or rebases (git and jj), and engine_worktree_create_fetch removes the fetch before `git worktree add`, which already tried the local base first. Both are registered in SMITHERS_COMPATIBILITY_PATCHES and in the fresh operator-controller patch list. The new integration test runs a seeded from a local-only launch commit with a preparation task that creates it and a verifier task that re-enters it, on the controller-patched engine, with every git invocation recorded. On main it records `fetch origin`, `fetch origin` and `rebase origin/`; with this patch it records no fetch or rebase, the run finishes, and the worktree is on the launch commit and its task branch. #1148 also asks for worktree retention and missing-object diagnostics, which this does not change. Refs #1148 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + packages/runtime/src/smithers.ts | 52 ++++++++++++++++++ ...hers-controller-engine.integration.test.ts | 54 ++++++++++++++++++- 3 files changed, 106 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 83bac40cc..dcd12ee5a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime]** The operator controller's Smithers engine no longer fetches `origin` or rebases task worktrees. Creating a worktree ran an untimed `git fetch origin` first, and each later task entering it retried a `git rebase` onto an `origin/` ref that does not exist, logging a failed rebase every time, after another fetch unless one had succeeded in the last minute. Worktrees are still created from the local launch commit and are now re-entered unchanged. Runs launched after upgrading get this (#1148). - **[runtime]** The operator controller's Smithers engine no longer appends a `TaskHeartbeat` event (an `_smithers_events` row and a `stream.ndjson` line) after each successful attempt-row heartbeat write. A quiet agent task wrote up to two a second, and a streaming agent one per ownership check that its output forced past the throttle. Liveness, heartbeat timeouts and `smithers why` rest on the fenced attempt-row write, which is kept. Runs launched after upgrading get this (#1147). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 2b6d4ecff..dd011cb67 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -1040,6 +1040,34 @@ const SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH = ` "heartbeat:record", // ultrafuzz: the fenced attempt row above is the liveness record (#1147). } catch (error) {`; +// Smithers treats as a branch to track. Creating a +// worktree runs `git fetch origin` first, and each re-entry retries +// `git rebase origin/` until one succeeds, after another fetch unless one +// succeeded for that repository in the last 60 s. Ultrafuzz passes the recorded +// launch commit (or a pinned source branch), and `origin/` never resolves, +// so every re-entry logs a failed rebase; each fetch is an untimed call to the +// user's remote (#1148). Task worktrees must stay on the launch commit +// (assertWorkspaceSourceRevision), so never synchronize them. +const SMITHERS_ENGINE_WORKTREE_SYNC_SOURCE = `function getWorktreeSyncCache() { + if (!worktreeSyncCacheSingleton) { + worktreeSyncCacheSingleton = createWorktreeSyncCache({ ttlMs: resolveWorktreeFetchTtlMs() }); + } + return worktreeSyncCacheSingleton; +}`; +const SMITHERS_ENGINE_WORKTREE_SYNC_PATCH = `function getWorktreeSyncCache() { + // ultrafuzz: task worktrees stay on their recorded launch commit (#1148). + return { shouldFetch: () => false, recordFetch() {}, shouldRebase: () => false, recordRebase() {} }; +}`; +const SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_SOURCE = ` // Best effort: refresh remote refs for git so origin/main can be used as a + // base when local main is absent. + if (vcs.type === "git") { + await runGitCommand(vcs.root, ["fetch", "origin"]); + } +`; +const SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_PATCH = ` // ultrafuzz: task worktrees start from a local recorded commit, so creating + // one never fetches origin (#1148). +`; + // Every event the engine persists first runs an idempotency probe that // filters `_smithers_events` on (run_id, timestamp_ms, type, payload_json). // The table's only index is its (run_id, seq) primary key, and the probe's @@ -2468,6 +2496,8 @@ export type SmithersCompatibilityPatchId = | "resume_hydration" | "engine_agent_event_ownership" | "engine_task_heartbeat_event" + | "engine_worktree_sync" + | "engine_worktree_create_fetch" | "engine_agent_usage_progress" | "engine_main_usage_invocation" | "engine_json_correction_usage_invocation" @@ -2623,6 +2653,22 @@ export const SMITHERS_COMPATIBILITY_PATCHES: readonly SmithersCompatibilityPatch patched: SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH, upstreamAbsent: [] }, + { + id: "engine_worktree_sync", + packageName: "@smthrs/engine", + sourceRelativePath: "src/engine.js", + patchable: SMITHERS_ENGINE_WORKTREE_SYNC_SOURCE, + patched: SMITHERS_ENGINE_WORKTREE_SYNC_PATCH, + upstreamAbsent: [] + }, + { + id: "engine_worktree_create_fetch", + packageName: "@smthrs/engine", + sourceRelativePath: "src/engine.js", + patchable: SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_SOURCE, + patched: SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_PATCH, + upstreamAbsent: [] + }, { id: "engine_agent_usage_progress", packageName: "@smthrs/engine", @@ -7106,6 +7152,12 @@ export function applySmithersCompatibilityPatches(projectRoot: string): void { "agent event ownership coalescing" ], [SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_SOURCE, SMITHERS_ENGINE_TASK_HEARTBEAT_EVENT_PATCH, "task heartbeat event"], + [SMITHERS_ENGINE_WORKTREE_SYNC_SOURCE, SMITHERS_ENGINE_WORKTREE_SYNC_PATCH, "task worktree sync"], + [ + SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_SOURCE, + SMITHERS_ENGINE_WORKTREE_CREATE_FETCH_PATCH, + "task worktree creation fetch" + ], [ SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_SOURCE, SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_PATCH, diff --git a/packages/runtime/test/smithers-controller-engine.integration.test.ts b/packages/runtime/test/smithers-controller-engine.integration.test.ts index f87261b32..3dd605c4f 100644 --- a/packages/runtime/test/smithers-controller-engine.integration.test.ts +++ b/packages/runtime/test/smithers-controller-engine.integration.test.ts @@ -1,5 +1,5 @@ import assert from "node:assert/strict"; -import { spawnSync } from "node:child_process"; +import { execFileSync, spawnSync } from "node:child_process"; import fs from "node:fs"; import { createRequire } from "node:module"; import path from "node:path"; @@ -100,6 +100,10 @@ process.exit(0); return run; } +function git(cwd: string, ...args: string[]): string { + return execFileSync("git", args, { cwd, encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] }).trim(); +} + // #1147: each successful attempt-row heartbeat write also appended a TaskHeartbeat // event row and stream.ndjson line. This agent is quiet, so only the throttled // liveness pulse keeps it live, through the fenced attempt-row write that the @@ -133,3 +137,51 @@ const workflow = smithers(() => `the attempt row stopped recording liveness: ${JSON.stringify(attempt)}` ); }); + +// #1148: Ultrafuzz seeds each from the recorded launch commit, which can +// exist only locally. Smithers fetched origin before creating a worktree, and when a +// task re-entered it fetched again (unless a fetch had succeeded in the last minute) +// and rebased onto `origin/`, which is never a ref. +test("controller task worktrees stay on a local-only launch commit without fetching or rebasing", () => { + const root = temporaryRoot("ufz-controller-worktree-"); + git(root, "init", "--quiet", "--initial-branch=main"); + git(root, "config", "user.name", "Ultrafuzz Synthetic Test"); + git(root, "config", "user.email", "synthetic@example.invalid"); + fs.writeFileSync(path.join(root, "source.txt"), "published\n"); + git(root, "add", "source.txt"); + git(root, "commit", "--quiet", "-m", "published base"); + git(root, "init", "--quiet", "--bare", path.join(root, "origin.git")); + git(root, "remote", "add", "origin", path.join(root, "origin.git")); + git(root, "push", "--quiet", "origin", "main"); + fs.writeFileSync(path.join(root, "source.txt"), "launch\n"); + git(root, "commit", "--quiet", "-am", "local-only launch commit"); + const launch = git(root, "rev-parse", "HEAD"); + const gitLog = path.join(root, "git.log"); + const recorder = path.join(root, "git-recorder.sh"); + const realGit = execFileSync("sh", ["-c", "command -v git"], { encoding: "utf8" }).trim(); + fs.writeFileSync(recorder, `#!/bin/sh\nprintf '%s\\n' "$*" >> '${gitLog}'\nexec '${realGit}' "$@"\n`, { + mode: 0o755 + }); + const workspace = path.join(root, "workspaces", "task-a"); + + // Preparation creates the worktree; the verifier re-enters it, as in Ultrafuzz. + const run = runOnControllerEngine( + root, + `const lane = { path: ${JSON.stringify(workspace)}, branch: "ultrafuzz/r1/task-a", baseBranch: ${JSON.stringify(launch)} }; +const workflow = smithers(() => + h(Workflow, { name: "worktree" }, + h(Worktree, lane, + h(Task, { id: "prepare", output: outputs.result, retries: 0 }, () => ({ value: "prepared" })), + h(Task, { id: "verify", output: outputs.result, dependsOn: ["prepare"], retries: 0 }, () => ({ value: "verified" })))));`, + { SMITHERS_GIT_PATH: recorder, SMITHERS_KEEP_WORKTREES: "1" } + ); + + assert.equal(run.status, "finished"); + const synchronization = fs + .readFileSync(gitLog, "utf8") + .split("\n") + .filter((command) => /(?:^| )(?:fetch|rebase)(?: |$)/u.test(command)); + assert.deepEqual(synchronization, [], "a task worktree fetched origin or rebased onto origin/"); + assert.equal(git(workspace, "rev-parse", "HEAD"), launch); + assert.equal(git(workspace, "branch", "--show-current"), "ultrafuzz/r1/task-a"); +}); From 7d1600a86b289e8b47b9c154b116bb8e120d2ca7 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:26:26 +0000 Subject: [PATCH 097/206] test(cli): check that temporaryRoot removes a sealed tree and its fake-bin sibling The per-test cleanup was only measured, so a regression would bring the temporary-directory leak back silently. The new test builds the shape a launched run leaves (read-only files in dr-x directories) plus the `-fake-bin` sibling inside a subtest and asserts both are gone once the subtest ends. It fails if the helper skips restoring owner write permission (the hook throws EACCES) or skips the sibling. Co-Authored-By: Claude Opus 5.5 --- packages/cli/test/temporary-root.test.ts | 26 ++++++++++++++++++++++++ 1 file changed, 26 insertions(+) create mode 100644 packages/cli/test/temporary-root.test.ts diff --git a/packages/cli/test/temporary-root.test.ts b/packages/cli/test/temporary-root.test.ts new file mode 100644 index 000000000..9cd8ba9f1 --- /dev/null +++ b/packages/cli/test/temporary-root.test.ts @@ -0,0 +1,26 @@ +import assert from "node:assert/strict"; +import fs from "node:fs"; +import path from "node:path"; +import test from "node:test"; + +import { temporaryRoot } from "./temporary-root.js"; + +test("temporaryRoot removes a sealed run tree and the fake-bin sibling when the test ends", async (t) => { + let root = ""; + await t.test("fixture", (inner) => { + root = temporaryRoot("ufz-cli-cleanup-", inner); + // The shape a launched run leaves: read-only files in dr-x directories. + const sealed = path.join(root, ".ultrafuzz", "runs", "run-1", "snapshot"); + fs.mkdirSync(sealed, { recursive: true }); + fs.writeFileSync(path.join(sealed, "module.js"), "export {};\n", { mode: 0o400 }); + for (let directory = sealed; directory !== root; directory = path.dirname(directory)) { + fs.chmodSync(directory, 0o500); + } + fs.mkdirSync(`${root}-fake-bin`); + fs.writeFileSync(path.join(`${root}-fake-bin`, "smithers"), "#!/bin/sh\n", { mode: 0o500 }); + }); + + assert.notEqual(root, ""); + assert.equal(fs.existsSync(root), false); + assert.equal(fs.existsSync(`${root}-fake-bin`), false); +}); From 8c832d84b941dad9d77863301a7ee0c14cca88ec Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:37:54 +0000 Subject: [PATCH 098/206] fix(runtime): read the runner's clean summary and bound run-error redaction Review follow-ups on the run-level error reader added for #272. - Prefer the runner's structured `summary` over `message` when the run error has no `cause`. The runner appends a docs link to every SmithersError message and, for a failed run, raw `smithers up ... --resume true` / `smithers replay ...` commands. Those commands bypass the attested operator controller and contradict the documented "fix the cause, then resume" recovery. `summary` carries neither, so no text matching is needed. A plain Error without a cause (no `summary`) still falls back to `message`. - Redact only a prefix eight times the kept length. The runner stores error text untruncated and the redactor's cost grows with the square of one long token, so a 200,000-character hex token took about 7 s per parse, paid by every sync, status and resume of a failed run. It now takes about 13 ms. Eight times (not four) keeps a 4096-bit RSA PEM block that starts in the kept 1,000 characters whole: with a 4,000-character prefix a probe leaked two of its body lines. - Apply the same redaction, runner-name scrub and a 100-character cap to `code`, which was copied verbatim. - Reword the reader comment (it only guarantees the parse never fails) and the CLI reference, which now names both shapes the diagnostic covers and scopes the resume advice to run-level errors. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- docs/reference/cli.md | 14 ++++++++------ packages/runtime/src/smithers.ts | 21 ++++++++++++++++----- packages/runtime/test/runtime.test.ts | 24 ++++++++++++++++++++++-- 4 files changed, 47 insertions(+), 14 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 88a607f50..0cdf295d9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime]** A run that ends `failed` with no failed durable node now carries the runner's own run-level error in its `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE` diagnostic message (for example `workflow run error WORKFLOW_RENDER_FAILED: `), redacted and capped at 1,000 characters; the durable `workflow-failure-unattributed` event stays ids-only. A new pinned-runner integration test covers recovering such a run through `resume`: while its render-time cause persists the resume is refused and the run is left unchanged, and once the cause is removed the same run ID resumes and runs only its pending task. Removes the unused `smithersSnapshotUnverifiedDependencies` helper (#272). +- **[runtime]** A run that ends `failed` with no failed durable node now carries the runner's run-level error in its `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE` diagnostic message (for example `workflow run error WORKFLOW_RENDER_FAILED: `), redacted, with the message capped at 1,000 characters; the durable `workflow-failure-unattributed` event stays ids-only. A new pinned-runner integration test covers recovering such a run through `resume`: while its render-time cause persists the resume is refused and the run's inspected state is unchanged, and once the cause is removed the same run ID resumes and runs only its pending task. Removes the unused `smithersSnapshotUnverifiedDependencies` helper (#272). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 43eb6323b..a74b9b013 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -326,12 +326,14 @@ rewrite historical artifacts or automatically reset, replay, timetravel, or fork completed work. A run that ends `failed` with no failed durable node was stopped by something -no task owns, for example an exception thrown while rendering the workflow. -`status` reports it as `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE` and appends the -workflow runner's own error to the message. Fix that cause, then `resume` the -same run; it continues from the tasks that already finished. While the cause -persists, the resume either fails with `WORKFLOW_LIFECYCLE_FAILED` or the run -fails again. +no durable node owns: a run-level workflow runner error, such as an exception +thrown while rendering the workflow, or a failed workflow task outside the +durable graph. `status` reports it as `WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE`, +names any failed workflow tasks, and appends the runner's run-level error to +the message. For a run-level error, fix that cause, then `resume` the same run; +it continues from the tasks that already finished. While the cause persists, +the resume either fails with `WORKFLOW_LIFECYCLE_FAILED` or the run fails +again. `resume --refresh-controller` first renders the currently installed Ultrafuzz controller and stock adapters beside the historical source, then delegates to diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index f037ba5d2..270313c71 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -5980,19 +5980,30 @@ export function parseCurrentSmithersInspect( // A run-level failure (for example a render exception) names no task, so the // run row's error is the only record of why the run stopped. Read it loosely: -// it explains a stop and never gates one, so an unexpected shape yields nothing. +// an unexpected shape yields nothing and never fails the parse. function currentSmithersRunError(value: unknown): CurrentSmithersInspect["runError"] { if (!isObjectRecord(value)) return undefined; - // The runner records the thrown error as `cause` under its own summary. + const nonBlank = (field: unknown): field is string => typeof field === "string" && field.trim() !== ""; + // The runner records what was thrown as `cause`. Its `summary` is its own + // message before it appends a docs link and raw runner resume commands. const cause = isObjectRecord(value.cause) ? value.cause.message : undefined; - const message = [cause, value.message].find((text): text is string => typeof text === "string" && text.trim() !== ""); + const message = [cause, value.summary, value.message].find(nonBlank); if (message === undefined) return undefined; return { - ...(typeof value.code === "string" ? { code: value.code } : {}), - message: scrubWorkflowRunnerText(redactSecretsInText(message)).slice(0, 1_000) + ...(nonBlank(value.code) ? { code: runErrorText(value.code, 100) } : {}), + message: runErrorText(message, 1_000) }; } +// Redacted and scrubbed like other runner text, since it is printed. The runner +// stores error text untruncated and redaction cost grows with the square of one +// long token, so only a prefix is redacted. For the message, eight times the +// kept length still holds a whole PEM private key (3,300 characters at RSA-4096) +// that starts in the kept text, so it is redacted as one block. +function runErrorText(value: string, limit: number): string { + return scrubWorkflowRunnerText(redactSecretsInText(value.slice(0, limit * 8))).slice(0, limit); +} + function parseCurrentSmithersExhaustedLoops(value: unknown): CurrentSmithersExhaustedLoop[] { if (value === undefined) return []; if (!Array.isArray(value) || value.length === 0) { diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 578c4b987..d58b2d515 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -18358,17 +18358,37 @@ test("parseCurrentSmithersInspect reads the run-level error loosely and never fa assert.deepEqual( runError({ code: "WORKFLOW_RENDER_FAILED", + summary: 'Rendering workflow "w.tsx" threw: boom.', message: 'Rendering workflow "w.tsx" threw: boom. Resume with: smithers up w.tsx --resume true', cause: { name: "Error", message: "boom" } }), { code: "WORKFLOW_RENDER_FAILED", message: "boom" } ); + // Without a cause, `summary` is the message minus the docs link and raw runner commands appended to it. + assert.deepEqual( + runError({ + code: "DUPLICATE_ID", + summary: "Duplicate Task id detected: dependent", + message: + "Duplicate Task id detected: dependent See https://smithers.sh/reference/errors Resume with: smithers up w.tsx --run-id r --resume true" + }), + { code: "DUPLICATE_ID", message: "Duplicate Task id detected: dependent" } + ); assert.deepEqual(runError({ message: "Task failed: prepare:x", cause: "opaque" }), { message: "Task failed: prepare:x" }); - // The text reaches a printed diagnostic, so it is redacted and bounded like other runner output. - assert.deepEqual(runError({ message: "token=sk-private-secret" }), { message: "token=" }); + // Both fields reach a printed diagnostic, so they are redacted and bounded like other runner output. + assert.deepEqual(runError({ code: "sk-private-code", message: "token=sk-private-secret" }), { + code: "", + message: "token=" + }); + assert.equal(runError({ code: "C".repeat(500), message: "m" })?.code?.length, 100); assert.equal(runError({ message: "x".repeat(5_000) })?.message.length, 1_000); + // The runner stores error text untruncated and every sync, status and resume parses it, so one + // long unbroken token (calldata, bytecode) must not make each parse take seconds. + const started = performance.now(); + assert.equal(runError({ message: `0x${"0123456789abcdef".repeat(12_500)}` })?.message.length, 1_000); + assert.ok(performance.now() - started < 2_000, `parse took ${Math.round(performance.now() - started)} ms`); // Advisory only: an unexpected shape yields nothing rather than a new parse failure. for (const opaque of ["opaque", 7, null, [], { code: 7 }, { code: "X", message: " " }]) { assert.equal(runError(opaque), undefined, JSON.stringify(opaque)); From e6ededbadb679d10b9a143e3352eed16afd9a028 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:49:27 +0000 Subject: [PATCH 099/206] fix(topology): restore the review group time budget The triage and severity-classification prompts tell the agent that "the topology gives this review node an extended timeout", but no shipped topology did. The review group lost its pin in v0.0.2, so dedupe, triage, severity classification, test aggregation and the final report ran on run.default_timeout_seconds (3600) while every goal, strategy and specialist lane had 7200. The review chain is the widest fan-in, so a review stage that needs more than the default window times out deterministically and is retried with the same window; the review group halts on failure, so every later review stage and the final-report node are then skipped (#1150). Pin the review group in the four non-smoke topologies to the same 7200 seconds as the goal, strategy and specialist groups. smoke.yml has no triage or severity node and is left alone. Like the existing group pins, this one takes precedence over run.default_timeout_seconds and model-profile timeouts (#675). The new packaged-topology test expands every shipped topology and holds each node whose prompt promises an extended timeout to a window above the largest shipped default. It fails on the previous topologies with "default.yml node `triage` promises an extended timeout but resolves to ultrafuzz.toml `run.default_timeout_seconds`=3600". Co-Authored-By: Claude Opus 5.5 --- .ultrafuzz/topology.yml | 2 ++ packages/config/topologies/default.yml | 2 ++ packages/config/topologies/exhaustive.yml | 2 ++ packages/config/topologies/invariant-only.yml | 2 ++ .../topology/test/packaged-topologies.test.ts | 35 +++++++++++++++++++ 5 files changed, 43 insertions(+) diff --git a/.ultrafuzz/topology.yml b/.ultrafuzz/topology.yml index 3e627b41d..c04090e05 100644 --- a/.ultrafuzz/topology.yml +++ b/.ultrafuzz/topology.yml @@ -41,6 +41,8 @@ groups: review: label: Review color: "#0f766e" + defaults: + timeout_seconds: 7200 nodes: - id: __start__ kind: meta diff --git a/packages/config/topologies/default.yml b/packages/config/topologies/default.yml index 3e627b41d..c04090e05 100644 --- a/packages/config/topologies/default.yml +++ b/packages/config/topologies/default.yml @@ -41,6 +41,8 @@ groups: review: label: Review color: "#0f766e" + defaults: + timeout_seconds: 7200 nodes: - id: __start__ kind: meta diff --git a/packages/config/topologies/exhaustive.yml b/packages/config/topologies/exhaustive.yml index ac088eaff..18b69c454 100644 --- a/packages/config/topologies/exhaustive.yml +++ b/packages/config/topologies/exhaustive.yml @@ -41,6 +41,8 @@ groups: review: label: Review color: "#0f766e" + defaults: + timeout_seconds: 7200 nodes: - id: __start__ kind: meta diff --git a/packages/config/topologies/invariant-only.yml b/packages/config/topologies/invariant-only.yml index 90c069448..c69fd4974 100644 --- a/packages/config/topologies/invariant-only.yml +++ b/packages/config/topologies/invariant-only.yml @@ -23,6 +23,8 @@ groups: review: label: Review color: "#0f766e" + defaults: + timeout_seconds: 7200 nodes: - id: __start__ kind: meta diff --git a/packages/topology/test/packaged-topologies.test.ts b/packages/topology/test/packaged-topologies.test.ts index 453599f4e..08c2ccdeb 100644 --- a/packages/topology/test/packaged-topologies.test.ts +++ b/packages/topology/test/packaged-topologies.test.ts @@ -10,6 +10,7 @@ import { NON_JSON_ARTIFACT_CONTRACT_IDS, artifactContractSchemaBinding } from "@ultrafuzz/artifacts"; +import { loadBuiltInPromptAssets } from "@ultrafuzz/prompts"; import { expandTopology, loadTopology } from "../src/index.js"; @@ -489,6 +490,40 @@ describe("packaged topology collection", () => { expect(pinsChecked, "the shipped topologies must still pin an agentic timeout somewhere").toBeGreaterThan(0); }); + // #1150. The triage and severity-classification prompts tell the agent that "the topology gives + // this review node an extended timeout". No shipped topology did: the review group lost its pin in + // v0.0.2, so the widest fan-in and the per-finding panel stages ran on the run default while every + // strategy lane had the reviewed window. Each promise is checked against the expanded graph, so a + // prompt and its topology cannot drift apart again. + it("gives every node whose prompt promises an extended timeout more than the default", { timeout: 30_000 }, () => { + const shadowed = largestShadowedDefault(); + const promisingPrompts = new Set( + loadBuiltInPromptAssets() + .filter((asset) => /extended\s+timeout/u.test(asset.markdown)) + .map((asset) => asset.relativePath) + ); + let promisesChecked = 0; + for (const name of PACKAGED_TOPOLOGY_IDS) { + const topologyPath = path.join(TOPOLOGY_ROOT, `${name}.yml`); + const graph = expandTopology(loadTopology(REPOSITORY_ROOT, { topologyPath, requirePromptFiles: true }), { + projectRoot: REPOSITORY_ROOT + }); + for (const node of graph.nodes) { + if (node.promptPath === undefined || !promisingPrompts.has(node.promptPath)) { + continue; + } + promisesChecked += 1; + expect( + node.timeoutSeconds ?? shadowed.seconds, + `${name}.yml node \`${node.id}\` promises an extended timeout but resolves to ${ + node.timeoutSeconds ?? `${shadowed.source}=${shadowed.seconds}` + }` + ).toBeGreaterThan(shadowed.seconds); + } + } + expect(promisesChecked, "no shipped node promises an extended timeout; retire this guard").toBeGreaterThan(0); + }); + // The other half of the #672/#677 window decision, updated for #708's config-owned retry budget. // A node window may be spent once per topology attempt and once per configured same-agent attempt. // That complete product is spent out of `workflow_deadline_seconds`, while the timeout, retry From 7dd5ee6ee8cf313e5e768e4aa26f6f24ac2ed566 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:49:34 +0000 Subject: [PATCH 100/206] chore(topology): drop the unused timeout on pinned reference nodes Every property reference node in the four non-smoke topologies set timeout_seconds: 300 (36 lines). Reference nodes are materialized by plan-run before launch and compileSmithersWorkflow only compiles agentic nodes into Smithers tasks, so the value never bounded anything. It was only copied into the expanded graph fingerprint and the planned graph, and displayed by the dashboard. The vulnerability-database reference node already had no timeout. Existing project topologies that still carry the lines keep validating; the reference docs now say the field has no effect on reference nodes. Co-Authored-By: Claude Opus 5.5 --- .ultrafuzz/topology.yml | 9 --------- docs/reference/topology-yaml.md | 3 +++ packages/config/topologies/default.yml | 9 --------- packages/config/topologies/exhaustive.yml | 9 --------- packages/config/topologies/invariant-only.yml | 9 --------- 5 files changed, 3 insertions(+), 36 deletions(-) diff --git a/.ultrafuzz/topology.yml b/.ultrafuzz/topology.yml index c04090e05..c763c203a 100644 --- a/.ultrafuzz/topology.yml +++ b/.ultrafuzz/topology.yml @@ -200,7 +200,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/0kn0t.md contract: ultrafuzz/nonempty-markdown@1 @@ -213,7 +212,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-thinking.md contract: ultrafuzz/nonempty-markdown@1 @@ -226,7 +224,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-sanity.md contract: ultrafuzz/nonempty-markdown@1 @@ -239,7 +236,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/aviggiano.md contract: ultrafuzz/nonempty-markdown@1 @@ -252,7 +248,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/rounding.md contract: ultrafuzz/nonempty-markdown@1 @@ -265,7 +260,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/crytic.md contract: ultrafuzz/nonempty-markdown@1 @@ -278,7 +272,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/runtime-verification.md contract: ultrafuzz/nonempty-markdown@1 @@ -291,7 +284,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/a16z-erc4626.md contract: ultrafuzz/nonempty-markdown@1 @@ -304,7 +296,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/recon.md contract: ultrafuzz/nonempty-markdown@1 diff --git a/docs/reference/topology-yaml.md b/docs/reference/topology-yaml.md index 65708ce28..2f246be55 100644 --- a/docs/reference/topology-yaml.md +++ b/docs/reference/topology-yaml.md @@ -225,6 +225,9 @@ Reference nodes must: - Mark exactly one non-manifest output as primary. - Avoid `prompt`, `role`, and `model_profiles`. +Reference nodes are materialized when the run is planned and never run as +agent tasks, so `timeout_seconds` has no effect on their execution. + Downstream prompts can consume the normalized Markdown primary artifact with: ```md diff --git a/packages/config/topologies/default.yml b/packages/config/topologies/default.yml index c04090e05..c763c203a 100644 --- a/packages/config/topologies/default.yml +++ b/packages/config/topologies/default.yml @@ -200,7 +200,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/0kn0t.md contract: ultrafuzz/nonempty-markdown@1 @@ -213,7 +212,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-thinking.md contract: ultrafuzz/nonempty-markdown@1 @@ -226,7 +224,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-sanity.md contract: ultrafuzz/nonempty-markdown@1 @@ -239,7 +236,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/aviggiano.md contract: ultrafuzz/nonempty-markdown@1 @@ -252,7 +248,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/rounding.md contract: ultrafuzz/nonempty-markdown@1 @@ -265,7 +260,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/crytic.md contract: ultrafuzz/nonempty-markdown@1 @@ -278,7 +272,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/runtime-verification.md contract: ultrafuzz/nonempty-markdown@1 @@ -291,7 +284,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/a16z-erc4626.md contract: ultrafuzz/nonempty-markdown@1 @@ -304,7 +296,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/recon.md contract: ultrafuzz/nonempty-markdown@1 diff --git a/packages/config/topologies/exhaustive.yml b/packages/config/topologies/exhaustive.yml index 18b69c454..3dafaa7ec 100644 --- a/packages/config/topologies/exhaustive.yml +++ b/packages/config/topologies/exhaustive.yml @@ -200,7 +200,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/0kn0t.md contract: ultrafuzz/nonempty-markdown@1 @@ -213,7 +212,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-thinking.md contract: ultrafuzz/nonempty-markdown@1 @@ -226,7 +224,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-sanity.md contract: ultrafuzz/nonempty-markdown@1 @@ -239,7 +236,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/aviggiano.md contract: ultrafuzz/nonempty-markdown@1 @@ -252,7 +248,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/rounding.md contract: ultrafuzz/nonempty-markdown@1 @@ -265,7 +260,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/crytic.md contract: ultrafuzz/nonempty-markdown@1 @@ -278,7 +272,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/runtime-verification.md contract: ultrafuzz/nonempty-markdown@1 @@ -291,7 +284,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/a16z-erc4626.md contract: ultrafuzz/nonempty-markdown@1 @@ -304,7 +296,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/recon.md contract: ultrafuzz/nonempty-markdown@1 diff --git a/packages/config/topologies/invariant-only.yml b/packages/config/topologies/invariant-only.yml index c69fd4974..efc7b0d7f 100644 --- a/packages/config/topologies/invariant-only.yml +++ b/packages/config/topologies/invariant-only.yml @@ -91,7 +91,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/0kn0t.md contract: ultrafuzz/nonempty-markdown@1 @@ -104,7 +103,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-thinking.md contract: ultrafuzz/nonempty-markdown@1 @@ -117,7 +115,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/certora-sanity.md contract: ultrafuzz/nonempty-markdown@1 @@ -130,7 +127,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/aviggiano.md contract: ultrafuzz/nonempty-markdown@1 @@ -143,7 +139,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/rounding.md contract: ultrafuzz/nonempty-markdown@1 @@ -156,7 +151,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/crytic.md contract: ultrafuzz/nonempty-markdown@1 @@ -169,7 +163,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/runtime-verification.md contract: ultrafuzz/nonempty-markdown@1 @@ -182,7 +175,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/a16z-erc4626.md contract: ultrafuzz/nonempty-markdown@1 @@ -195,7 +187,6 @@ nodes: group: references depends_on: - __start__ - timeout_seconds: 300 outputs: - path: references/recon.md contract: ultrafuzz/nonempty-markdown@1 From 0ae8e2c8ecf692bd56ab869ffca87632f458269e Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:49:47 +0000 Subject: [PATCH 101/206] fix(prompts): stop inviting production interface edits the handoff rejects The stateful-invariant setup and implement-properties prompts told the agent not to edit production src/ or contracts/ "except for interfaces if they are genuinely required by the harness". Both nodes publish a workspace patch, and captureWorkspacePatch rejects every change under the configured production source roots (src and contracts by default), added files included: source-snapshot violation: workspace patch modifies protected production source: src/interfaces/IVault.sol So with the default roots, following the exception failed the handoff, and a retry starts a fresh session from the same prompt. Remove the exception and tell the agent to declare harness interfaces in the test tree, which the same capture accepts. Co-Authored-By: Claude Opus 5.5 --- .../prompts/strategies/invariants/implement-properties.md | 6 ++++-- .ultrafuzz/prompts/strategies/invariants/setup.md | 2 +- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/.ultrafuzz/prompts/strategies/invariants/implement-properties.md b/.ultrafuzz/prompts/strategies/invariants/implement-properties.md index 4e7aa6759..4d2180b4c 100644 --- a/.ultrafuzz/prompts/strategies/invariants/implement-properties.md +++ b/.ultrafuzz/prompts/strategies/invariants/implement-properties.md @@ -124,8 +124,10 @@ an unselected expected check is not implemented or fulfilled. guard, or other precondition that prevents the backend from observing the violating post-state. Preconditions may admit valid actions; they may not assume the property under test. - - Do not edit production contracts except interfaces that are genuinely - required by the test harness. + - Do not edit production contracts, not even to add an interface: the + workspace handoff rejects every change under the production source roots + (by default `src/` and `contracts/`). Declare any interface the harness + needs in the test tree instead. - Keep generated or changed invariant files in the test tree and include every changed `*.t.sol` test/reproducer in `generated-tests.json`. - Preserve Recon constructor deployment if property work changes `Setup`, diff --git a/.ultrafuzz/prompts/strategies/invariants/setup.md b/.ultrafuzz/prompts/strategies/invariants/setup.md index 1ec988aee..eb0b641a5 100644 --- a/.ultrafuzz/prompts/strategies/invariants/setup.md +++ b/.ultrafuzz/prompts/strategies/invariants/setup.md @@ -71,7 +71,7 @@ project's pinned revision. Use these Recon/Chimera rules while making decisions: - Read `AGENTS.md` and obey all repository-specific rules before editing. -- Do not edit production `src/` or `contracts/` except for interfaces if they are genuinely required by the harness. +- Do not edit production `src/` or `contracts/`, not even to add an interface: the workspace handoff rejects every change under the production source roots. Declare any interface the harness needs in the test tree instead. - [Chimera](https://github.com/Recon-Fuzz/create-chimera-app) is the write-once, run-everywhere scaffold for Foundry, Echidna, Medusa, Halmos, and Kontrol style runs. - The create-chimera-app layout under the repository's test root is: `/recon/Setup.sol`, `BeforeAfter.sol`, `Properties.sol`, From 5080b5a68b4d56fe0741e852fbf3cc5ed3085920 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:49:55 +0000 Subject: [PATCH 102/206] fix(runtime): warn at validate time when a project prompt differs from its built-in Runs use the project copy of every prompt, and ultrafuzz init without --force keeps an existing copy. After an upgrade a scaffolded copy therefore keeps an older release's text while the artifact gates move on, and nothing said so until a node failed its gate at the end of an attempt. projectPromptsDifferingFromBuiltIns lists the project prompts whose bytes differ from the built-in prompt at the same path. validate reports each one as a PROMPT_DIFFERS_FROM_BUILT_IN warning, which sets the prompts posture to warn without failing validation, planning or the dashboard's prompt save. A freshly scaffolded project and prompts the project adds at new paths produce no warning. The new runtime test fails on the previous validate.ts (prompts posture 'pass' where 'warn' is expected) and also checks the remedy the warning names: deleting the copy and rerunning init restores a passing posture. Co-Authored-By: Claude Opus 5.5 --- docs/how-to/edit-prompts-topology.md | 7 ++++ docs/reference/cli.md | 5 ++- packages/prompts/src/catalog.ts | 14 +++++++ packages/runtime/src/validate.ts | 17 +++++++- .../test/prompt-catalog-validation.test.ts | 40 +++++++++++++++++++ 5 files changed, 80 insertions(+), 3 deletions(-) create mode 100644 packages/runtime/test/prompt-catalog-validation.test.ts diff --git a/docs/how-to/edit-prompts-topology.md b/docs/how-to/edit-prompts-topology.md index 99ca17a04..2eaaaf9e7 100644 --- a/docs/how-to/edit-prompts-topology.md +++ b/docs/how-to/edit-prompts-topology.md @@ -14,6 +14,13 @@ After `ultrafuzz init`, editable prompts live under: Use the [prompt catalog](../reference/prompt-catalog.md) to find every shipped workflow prompt and its role before choosing what to customize. +Runs use these project copies, and `ultrafuzz init` keeps them unless you pass +`--force`, which also overwrites `ultrafuzz.toml` and the topology. After an +upgrade a copy therefore keeps the text of the release that scaffolded it. +`ultrafuzz validate` warns about every project prompt that differs from the +built-in prompt at the same path. To take the built-in version of a prompt, +delete your copy and rerun `ultrafuzz init`. + Prompt frontmatter may include only identity and display metadata: ```md diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 0e42102b2..6b21e979e 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -124,7 +124,10 @@ ultrafuzz validate [--project ] [--audit-profile ] \ Validation covers typed TOML config, `.ultrafuzz/topology.yml`, project prompt copies, safe paths, reference nodes, agent references, and trusted local -execution posture. It does not launch agents. +execution posture. It does not launch agents. A project prompt that differs +from the built-in prompt at the same path sets the prompts posture to `warn` +with one `PROMPT_DIFFERS_FROM_BUILT_IN` warning per file; warnings do not fail +validation. ## JSON Validate diff --git a/packages/prompts/src/catalog.ts b/packages/prompts/src/catalog.ts index 576a92ca9..a9e45f700 100644 --- a/packages/prompts/src/catalog.ts +++ b/packages/prompts/src/catalog.ts @@ -88,6 +88,20 @@ export function loadPromptCatalog(options: LoadPromptCatalogOptions = {}): Promp }; } +/** + * Relative paths of project prompts whose bytes differ from the built-in prompt at the same path. Runs + * use the project copy and `ultrafuzz init` keeps it unless run with `--force`, so after an upgrade such + * a copy is either a deliberate edit or an older release's prompt. + */ +export function projectPromptsDifferingFromBuiltIns(catalog: PromptCatalog): string[] { + const builtInMarkdown = new Map(loadBuiltInPromptAssets().map((asset) => [asset.relativePath, asset.markdown])); + return [...catalog.entries.values()] + .filter((entry) => entry.source === "project" && builtInMarkdown.has(entry.relativePath)) + .filter((entry) => builtInMarkdown.get(entry.relativePath) !== entry.markdown) + .map((entry) => entry.relativePath) + .sort(); +} + export function getPrompt(catalog: PromptCatalog, id: string): PromptCatalogEntry { const entry = catalog.entries.get(id); if (!entry) { diff --git a/packages/runtime/src/validate.ts b/packages/runtime/src/validate.ts index 8e509bb8f..5bbabe9e7 100644 --- a/packages/runtime/src/validate.ts +++ b/packages/runtime/src/validate.ts @@ -11,7 +11,7 @@ import { validateModelProfiles, type ResolvedConfig } from "@ultrafuzz/config"; -import { loadPromptCatalog, projectPromptDir } from "@ultrafuzz/prompts"; +import { loadPromptCatalog, projectPromptDir, projectPromptsDifferingFromBuiltIns } from "@ultrafuzz/prompts"; import { expandTopology, loadTopology, resolveTopologyPath, type ModelProfileSelection } from "@ultrafuzz/topology"; import type { @@ -202,8 +202,21 @@ function validatePrompts(projectRoot: string): { } try { const catalog = loadPromptCatalog({ projectRoot }); + const differing = projectPromptsDifferingFromBuiltIns(catalog); return { - posture: postureFromDiagnostics("prompts", "project prompt catalog loads and variables are strict", []), + posture: postureFromDiagnostics( + "prompts", + differing.length === 0 + ? "project prompt catalog loads and variables are strict" + : `project prompt catalog loads, but ${String(differing.length)} project prompt(s) differ from the built-in prompt at the same path; ultrafuzz validate --json lists them`, + differing.map((relativePath) => ({ + code: "PROMPT_DIFFERS_FROM_BUILT_IN", + message: `.ultrafuzz/prompts/${relativePath} differs from the built-in prompt at the same path; runs use the project copy, and ultrafuzz init without --force keeps it, so to take the built-in version, delete the file and rerun ultrafuzz init`, + severity: "warning" as const, + source: "prompts", + path: `.ultrafuzz/prompts/${relativePath}` + })) + ), summary: { prompt_dir: promptDir, prompt_count: catalog.orderedIds.length diff --git a/packages/runtime/test/prompt-catalog-validation.test.ts b/packages/runtime/test/prompt-catalog-validation.test.ts new file mode 100644 index 000000000..6adc19c36 --- /dev/null +++ b/packages/runtime/test/prompt-catalog-validation.test.ts @@ -0,0 +1,40 @@ +import assert from "node:assert/strict"; +import fs from "node:fs"; +import path from "node:path"; +import test from "node:test"; + +import { initProject, validateProject } from "../src/index.js"; +import { temporaryRoot } from "./temporary-root.js"; + +// Runs use the project copy of a prompt, and `ultrafuzz init` without `--force` keeps it. After an +// upgrade a scaffolded copy therefore keeps an older release's text while the gates move on, and +// nothing said so until a node failed its artifact gate. +test("validate warns about a project prompt that differs from the built-in prompt at its path", async () => { + const project = temporaryRoot("ufz-prompt-drift-"); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const promptPath = (relativePath: string) => path.join(project, ".ultrafuzz", "prompts", ...relativePath.split("/")); + const driftWarnings = (result: Awaited>) => + result.diagnostics + .filter((diagnostic) => diagnostic.code === "PROMPT_DIFFERS_FROM_BUILT_IN") + .map((diagnostic) => [diagnostic.severity, diagnostic.path]); + + const scaffolded = await validateProject({ projectRoot: project, env: {} }); + assert.equal(scaffolded.ok, true, JSON.stringify(scaffolded.diagnostics)); + assert.equal(scaffolded.value?.policy_posture.prompts.status, "pass"); + assert.deepEqual(driftWarnings(scaffolded), []); + + fs.appendFileSync(promptPath("review/triage.md"), "\nA local edit.\n", "utf8"); + // A prompt the project adds has no built-in counterpart, so it is not drift. + fs.writeFileSync(promptPath("strategies/project-only.md"), "A project-only prompt.\n", "utf8"); + const edited = await validateProject({ projectRoot: project, env: {} }); + assert.equal(edited.ok, true, JSON.stringify(edited.diagnostics)); + assert.equal(edited.value?.policy_posture.prompts.status, "warn"); + assert.deepEqual(driftWarnings(edited), [["warning", ".ultrafuzz/prompts/review/triage.md"]]); + + // The remedy the warning names restores the built-in copy. + fs.rmSync(promptPath("review/triage.md")); + assert.equal(initProject({ projectRoot: project }).ok, true); + const restored = await validateProject({ projectRoot: project, env: {} }); + assert.equal(restored.value?.policy_posture.prompts.status, "pass"); + assert.deepEqual(driftWarnings(restored), []); +}); From 0333333b0af767341313e213d55e476d48e9b077 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 22:51:33 +0000 Subject: [PATCH 103/206] docs: changelog for the review timeout and prompt/gate fixes Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..ae4ab7f1b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[config] [prompts] [runtime] [docs]** The shipped topologies pin the `review` group to the same 7,200-second timeout as the goal, strategy and specialist groups. The group had no pin, so dedupe, triage, severity classification, test aggregation and the final report ran on `run.default_timeout_seconds` (3,600 by default), although the triage and severity prompts promise an extended timeout; like the other group pins, it takes precedence over the run and model-profile timeouts (#675), and an existing `.ultrafuzz/topology.yml` keeps its old review budget until `defaults: { timeout_seconds: 7200 }` is added to its `review` group. The stateful-invariant setup and property prompts no longer permit interface edits under the production source roots (`src/` and `contracts/` by default), which the workspace handoff rejects, the unused `timeout_seconds: 300` on reference nodes is removed, and `ultrafuzz validate` warns (`PROMPT_DIFFERS_FROM_BUILT_IN`) about each project prompt that differs from the built-in prompt at the same path, because `ultrafuzz init` without `--force` keeps existing prompts (#1150). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 6af84e4efc7677540e0d4ad9fd0201128409987c Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:44:24 +0000 Subject: [PATCH 104/206] docs: let CHANGELOG entries exceed the markdownlint line length Super-linter lints a changed Markdown file whole, and dozens of existing CHANGELOG.md entries are single lines longer than its 400-character MD013 limit. So a PR that adds an entry fails "External static analysis", which fails `release-gates` and skips every release validation lane, including the runtime lanes this PR's new test runs in. This is the same two-line hunk #1169 adds, so whichever merges second applies it cleanly. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 0cdf295d9..02ccbf7bd 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,7 @@ # Changelog + + ## Unreleased ### Breaking changes From 93c52595ba7c167f42edb0813e2e6ab0028701a3 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:44:55 +0000 Subject: [PATCH 105/206] fix(security): end the BIP39 word run at Anvil's public mnemonic The mnemonic exemption skipped a window equal to the public phrase with `continue`, so the word run was never reset. Every later BIP39 word was then checked in 12- to 24-word windows that overlap the phrase, windows origin/main never examined because it reset the run after redacting the phrase. Reproduced on the previous head: - The gate still rejected the phrase when a whitespace-separated BIP39 word followed it ("... junk one account per index"), and two copies of it on consecutive lines. - Redaction weakened next to a real mnemonic. The phrase followed by "abandon x11 about" on the next line became "test abandon ... about", leaving 11 of the real mnemonic's 12 words in the clear. origin/main redacted both. A valid window that is exactly the public phrase now ends the run like any other detected mnemonic; only its redaction range is dropped. The run therefore evolves as on origin/main, and the rule's output is origin/main's with those placeholders restored. Co-Authored-By: Claude Opus 5.5 --- packages/security/src/sensitive-redaction.ts | 8 +++++++- packages/security/test/sensitive-redaction.test.ts | 5 +++++ 2 files changed, 12 insertions(+), 1 deletion(-) diff --git a/packages/security/src/sensitive-redaction.ts b/packages/security/src/sensitive-redaction.ts index 73d7f6261..d5860c4ea 100644 --- a/packages/security/src/sensitive-redaction.ts +++ b/packages/security/src/sensitive-redaction.ts @@ -380,7 +380,13 @@ function redactBip39Mnemonics(value: string, placeholder: string): string { if (run.length < wordCount) continue; const window = run.slice(-wordCount); const phrase = window.map((token) => token.word).join(" "); - if (phrase === PUBLIC_DEVELOPMENT_MNEMONIC || !validateMnemonic(phrase, englishWordlist)) continue; + if (!validateMnemonic(phrase, englishWordlist)) continue; + // The public phrase is left in place but still ends the run, so the + // words after it are checked exactly as if it had been redacted. + if (phrase === PUBLIC_DEVELOPMENT_MNEMONIC) { + run = []; + break; + } ranges.push({ start: window[0]!.start, end: window.at(-1)!.end }); run = []; break; diff --git a/packages/security/test/sensitive-redaction.test.ts b/packages/security/test/sensitive-redaction.test.ts index 7216e8d36..c9f0ac2b0 100644 --- a/packages/security/test/sensitive-redaction.test.ts +++ b/packages/security/test/sensitive-redaction.test.ts @@ -232,6 +232,11 @@ test("positive-only scans publish Anvil's dev mnemonic and keys while other mnem redactSecretsInText(`mnemonic: ${anvilMnemonic.replace(/junk$/u, "absent")}`, undefined, [], "positive-only"), "mnemonic: " ); + // The phrase ends its word run: BIP39 words that follow it ("one") start a + // new run, so prose after it publishes and a real mnemonic on the next line + // is still redacted in full. + assert.equal(containsSensitiveSecrets(`${anvilMnemonic} one account per index`, [], "positive-only"), false); + assert.equal(redactSecretsInText(`${anvilMnemonic}\n${mnemonic}`), `${anvilMnemonic}\n`); const anvilKeys = [ "0xac0974bec39a17e36ba4a6b4d238ff944bacb478cbed5efcae784d7bf4f2ff80", From d101467d65b4f567e5e496b0d6e6cd3a68bb62bb Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:45:09 +0000 Subject: [PATCH 106/206] fix(security): constrain only the JWT header segment The payload (?=eyJ) lookahead removed none of the false positives this branch targets: the header lookahead alone already stops the rule matching each dotted identifier fixture. What the payload lookahead did remove was detection of compact JWE tokens, whose second segment is the encrypted key, and of JWS tokens whose payload is not compact JSON ('{ "sub": ...' encodes to "eyA"). origin/main detected both. The rule now requires "eyJ" on the header segment only, and the test pins an RSA-OAEP JWE and a space-padded payload as still detected. The CHANGELOG entry drops the claims that real JWTs, other keys and other mnemonics are all still rejected, and tells cloud operators what to do when resume recomputes the classification of an allowlisted variable that holds one of these values. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- packages/artifacts/test/artifacts.test.ts | 2 +- packages/security/src/sensitive-redaction.ts | 12 +++++++----- packages/security/test/sensitive-redaction.test.ts | 11 ++++++++++- 4 files changed, 19 insertions(+), 8 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b44e86fbb..2d8c496f1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[security]** The fail-closed publication secret gate no longer rejects Anvil's public development accounts or long dotted identifiers in agent output. The JWT rule now requires the `eyJ` prefix that a JSON header and payload encode to, so a qualified name such as `ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck` no longer reads as a token; the labeled-private-key and mnemonic rules skip the published `test test … junk` mnemonic and the 10 keys `anvil` prints, which Foundry tests sign with. Real JWTs, other labeled 64-hex keys, and other valid mnemonics are still rejected. +- **[security]** The publication secret gate's pattern rules skip Anvil's public `test test … junk` mnemonic and the 10 keys `anvil` prints, and its JWT rule requires an `eyJ` header, not just three long dotted segments. A cloud run started before this change that allowlists a variable holding that mnemonic or such a dotted value must drop it from `ULTRAFUZZ_AGENT_ENV_ALLOWLIST` to resume, replay or fork. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/packages/artifacts/test/artifacts.test.ts b/packages/artifacts/test/artifacts.test.ts index 7acc0cee3..b96c1cd4b 100644 --- a/packages/artifacts/test/artifacts.test.ts +++ b/packages/artifacts/test/artifacts.test.ts @@ -1121,7 +1121,7 @@ test("canonical publication secret gate fails closed without rewriting bytes", ( // The positive rules must not fire on ordinary contract output either: a // Foundry test deriving and signing with Anvil's published mnemonic and // account (0) key, and a qualified identifier with three long dotted - // segments (JWT-shaped until the rule required the eyJ header and payload). + // segments (JWT-shaped until the rule required an eyJ header). assert.doesNotThrow(() => assertArtifactPublicationsContainNoSecrets( new Map([ diff --git a/packages/security/src/sensitive-redaction.ts b/packages/security/src/sensitive-redaction.ts index d5860c4ea..ba72ed136 100644 --- a/packages/security/src/sensitive-redaction.ts +++ b/packages/security/src/sensitive-redaction.ts @@ -49,11 +49,13 @@ const SUPPLEMENTAL_SECRET_PATTERNS: readonly RegExp[] = [ /\b(?:ak|as)-[A-Za-z0-9_]{16,}\b/gu, // No Google OAuth access-token rule in the recommended preset. /\bya29\.[A-Za-z0-9._-]{20,}\b/gu, - // No JWT rule in the recommended preset. A JWT's header and payload are - // base64url JSON objects, which encode to "eyJ" when they open with '{"' and - // a letter; that prefix is the identification. Three dotted segments alone - // also match qualified identifiers such as Contract.function.check. - /\b(?=eyJ)[A-Za-z0-9_-]{20,}\.(?=eyJ)[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\b/gu, + // No JWT rule in the recommended preset. A JWT header is a base64url JSON + // object, which encodes to "eyJ" when it opens with '{"' and a letter; that + // prefix is the identification. Three long dotted segments alone also + // describe qualified names such as + // ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck. Later + // segments stay unconstrained: a JWE's second one is its encrypted key. + /\b(?=eyJ)[A-Za-z0-9_-]{20,}\.[A-Za-z0-9_-]{10,}\.[A-Za-z0-9_-]{10,}\b/gu, // Provider-keyed RPC URLs embed the credential in the path; no secretlint // rule covers Alchemy/Infura project keys. /\b(?:https?|wss?):\/\/[^\s"'`]*(?:alchemy\.com\/v2\/|infura\.io\/v3\/)[A-Za-z0-9_-]{16,}\b/giu diff --git a/packages/security/test/sensitive-redaction.test.ts b/packages/security/test/sensitive-redaction.test.ts index c9f0ac2b0..896a24872 100644 --- a/packages/security/test/sensitive-redaction.test.ts +++ b/packages/security/test/sensitive-redaction.test.ts @@ -272,9 +272,18 @@ test("positive-only scans publish Anvil's dev mnemonic and keys while other mnem ); }); -test("the JWT rule requires eyJ header and payload segments, so dotted identifiers publish", () => { +test("the JWT rule requires an eyJ header, so dotted identifiers publish", () => { const jwt = "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dBjftJeZ4CVPmB92K27uhbUJU1p1r_wW1gFWFOEjXk"; // gitleaks:allow -- fake credential fixture for the redaction tests assert.equal(redactSecretsInText(`session=${jwt};`, undefined, [], "positive-only"), "session=;"); + // Only the header is constrained. A JWE's second segment is its encrypted + // key, and a payload with a space after '{' encodes to "eyA". + const base64url = (json: string): string => Buffer.from(json).toString("base64url"); + for (const token of [ + `${base64url('{"alg":"RSA-OAEP","enc":"A256GCM"}')}.OKOawDo13gRp2ojaHV7LFpZcgV7T.48V1_ALb6US04U3b.5eym8TW_c8SuK0ltJ3rpYIzOeDQz.XFBoMYUZodetZdvTiFvSkQ`, + `${base64url('{"alg":"HS256"}')}.${base64url('{ "sub": "1234567890" }')}.dBjftJeZ4CVPmB92K27uhbUJU1p1r_wW1gFWFOEjXk` + ]) { + assert.equal(containsSensitiveSecrets(`session=${token};`, [], "positive-only"), true, token); + } for (const fixture of [ "ReentrancyGuardUpgradeable.nonReentrantModifier.lockedStateCheck", "IExampleLendingPoolCore.liquidationCall.healthFactorBefore", From b964ae2301582ecfdc5fa0484e2f80fd09d11dbd Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:46:47 +0000 Subject: [PATCH 107/206] fix(runtime): restate run-summary tokens, spend and partial pricing together withWholeRunSummary restated tokens_used from run.json's cumulative accounting but kept the agent's estimated_spend and partial_pricing whenever the cumulative spend was "unavailable". For a direct run, that agent spend can come from the live Smithers fallback taken when the report task started, which may price only part of the run. The runtime presentation then showed whole-run tokens beside a report-start spend, with no partial marker. Tokens, spend and partial pricing now move as one group: when the cumulative record carries a token count, all three come from it, and a whole-run spend recorded as "unavailable" stays "unavailable". Without cumulative accounting, all three keep the agent's copy as before. The sparse test case pinned the mixed summary, using a cumulative block (estimated_spend "unavailable" beside estimated_spend_usd 41.2) that the runtime's accounting validator rejects. It now uses a ledger that priced no event, and the fixture's pricing markers carry the component field the validator requires. The fixture passes storedAccountingDocument (via assertRunMetadataAccountingUsageAuthority, which stops only at the ledger comparison). Co-Authored-By: Claude Opus 5.5 --- docs/reference/artifacts-reports.md | 8 +-- .../runtime/src/terminal-report-projection.ts | 10 ++-- .../test/terminal-report-projection.test.ts | 53 +++++++++++-------- 3 files changed, 43 insertions(+), 28 deletions(-) diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index dff578046..7235995eb 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -766,9 +766,11 @@ that snapshot. Runtime presentations (the verified terminal publication and unchecked reports) restate the run summary instead: elapsed time from `run.json#created_at` to `state.json#finished_at`, and models, tokens, estimated spend, and `partial_pricing` from the current -`accounting.cumulative`. A value those records lack, or record as -`unavailable`, keeps the agent's copy. Use `ultrafuzz stats` for the full -accounting breakdown. +`accounting.cumulative`. Tokens, estimated spend, and `partial_pricing` are +restated together whenever `accounting.cumulative` records a token count, so a +whole-run spend recorded as `unavailable` stays `unavailable` instead of showing +the agent's report-start figure. Otherwise, a value those records lack keeps the +agent's copy. Use `ultrafuzz stats` for the full accounting breakdown. `accounting.segments` publishes one rollup per checkpoint generation, and `accounting.current` identifies the latest segment. Each segment retains every diff --git a/packages/runtime/src/terminal-report-projection.ts b/packages/runtime/src/terminal-report-projection.ts index ebc6094cd..a32da8800 100644 --- a/packages/runtime/src/terminal-report-projection.ts +++ b/packages/runtime/src/terminal-report-projection.ts @@ -61,8 +61,10 @@ export function projectTerminalReport(input: TerminalReportProjectionInput): Can /** * The report agent copies a run summary that the host captured when the report task started, so its * elapsed time and accounting miss the report task itself and anything that finished later. Runtime - * presentations restate them from run.json and the recorded finish time. A value those records do - * not provide keeps the agent's copy; malformed records are ignored, never thrown. + * presentations restate them from run.json and the recorded finish time. Tokens, spend, and partial + * pricing move together, because the agent's spend may price only part of the run; a spend the + * whole-run record calls unavailable stays unavailable. A value those records do not provide keeps + * the agent's copy; malformed records are ignored, never thrown. */ export function withWholeRunSummary( runMetadata: Record, @@ -76,9 +78,9 @@ export function withWholeRunSummary( const models = field(cumulative, "models"); if (isNonEmptyStringList(models)) summary.models_used = [...models]; const tokens = field(cumulative, "tokens_used"); - if (availableLabel(tokens)) summary.tokens_used = tokens; const spend = field(cumulative, "estimated_spend"); - if (availableLabel(spend)) { + if (availableLabel(tokens) && typeof spend === "string" && spend.trim() !== "") { + summary.tokens_used = tokens; summary.estimated_spend = spend; summary.partial_pricing = field(cumulative, "partial_pricing") === true; } diff --git a/packages/runtime/test/terminal-report-projection.test.ts b/packages/runtime/test/terminal-report-projection.test.ts index 43c7e5b3e..018cafce4 100644 --- a/packages/runtime/test/terminal-report-projection.test.ts +++ b/packages/runtime/test/terminal-report-projection.test.ts @@ -197,8 +197,12 @@ test("terminal projection enforces bounded and internally consistent completion assert.match(result.markdown, /identities omitted from this bounded census: `1`/u); }); -/** A valid run.json whose cumulative accounting already includes the report task's own usage. */ -function metadataWithAccounting(cumulative: { estimated_spend?: string; models?: string[] } = {}): RunMetadataDocument { +/** + * A valid run.json whose cumulative accounting already includes the report task's own usage. When + * `priced` is false the usage ledger priced no event, so the whole-run spend is unavailable. + */ +function metadataWithAccounting(priced = true): RunMetadataDocument { + const unpricedModels = priced ? ["model-b"] : ["model-a", "model-b"]; const summary = { uncached_input_tokens: 9_000_000, input_tokens: 9_000_000, @@ -210,18 +214,27 @@ function metadataWithAccounting(cumulative: { estimated_spend?: string; models?: billable_token_total: 12_345_678, total_tokens: 12_345_678, tokens_used: "12,345,678", - estimated_spend: "$41.20+", - estimated_spend_usd: 41.2, - component_costs_usd: { uncached_input: 30, cache_read: 0, cache_write: 0, output: 11.2, reasoning: 0 }, + ...(priced + ? { + estimated_spend: "$41.20+", + estimated_spend_usd: 41.2, + component_costs_usd: { uncached_input: 30, cache_read: 0, cache_write: 0, output: 11.2, reasoning: 0 } + } + : { + estimated_spend: "unavailable", + component_costs_usd: { uncached_input: 0, cache_read: 0, cache_write: 0, output: 0, reasoning: 0 } + }), usage_complete: true, usage_incomplete_reasons: [], pricing_complete: false, - pricing_incomplete_reasons: [{ code: "model-pricing-unavailable" as const, model: "model-b" }], + pricing_incomplete_reasons: (["output", "uncached_input"] as const).flatMap((component) => + unpricedModels.map((model) => ({ code: "model-pricing-unavailable" as const, component, model })) + ), partial_pricing: true, cache_read_pricing_estimated: false, event_count: 2, - priced_event_count: 1, - unpriced_event_count: 1, + priced_event_count: priced ? 1 : 0, + unpriced_event_count: priced ? 1 : 2, models: ["model-a", "model-b"], agents: ["agent-a"] }; @@ -260,7 +273,7 @@ function metadataWithAccounting(cumulative: { estimated_spend?: string; models?: workflow_run_id: "workflow-1", current: structuredClone(segment), segments: [structuredClone(segment)], - cumulative: { ...summary, source_run_ids: [], ...cumulative }, + cumulative: { ...summary, source_run_ids: [] }, checkpoint: { schema_version: "ultrafuzz.accounting-checkpoint.v1", ledger_event_count: 2, @@ -272,9 +285,9 @@ function metadataWithAccounting(cumulative: { estimated_spend?: string; models?: source: "configured-catalog", status: "available", fetched_at: CREATED_AT, - resolved_models: ["model-a"], - unresolved_models: ["model-b"], - model_prices: { "model-a": { inputUsdPerMillion: 1, outputUsdPerMillion: 2 } } + resolved_models: priced ? ["model-a"] : [], + unresolved_models: unpricedModels, + model_prices: priced ? { "model-a": { inputUsdPerMillion: 1, outputUsdPerMillion: 2 } } : {} }, updated_at: FINISHED_AT } @@ -304,16 +317,14 @@ test("terminal projection restates whole-run accounting instead of the report-st assert.equal((result.report.run_metadata as Record).partial_pricing, true); assert.deepEqual(result.report.issues, agentReport().issues); - // Values the run records do not have keep the agent's copy. - const sparse = projectTerminalReport({ - ...input, - metadata: metadataWithAccounting({ estimated_spend: "unavailable", models: [] }) - }); - assert.deepEqual(runSummaryLines(sparse.markdown), [ + // A ledger that priced nothing makes the whole-run spend unavailable. The agent's report-start + // spend covers only part of the run, so it is not shown next to whole-run tokens. + const unpriced = projectTerminalReport({ ...input, metadata: metadataWithAccounting(false) }); + assert.deepEqual(runSummaryLines(unpriced.markdown), [ "- Elapsed time: `6h 02m`", - "- Models used: `example-model`", + "- Models used: `model-a, model-b`", "- Tokens used: `12,345,678`", - "- Estimated spend: `$0.01`" + "- Estimated spend: `unavailable`" ]); - assert.equal((sparse.report.run_metadata as Record).partial_pricing, false); + assert.equal((unpriced.report.run_metadata as Record).partial_pricing, true); }); From dde848fb9b1abfe5368a22f5aae57e831930bd31 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:47:00 +0000 Subject: [PATCH 108/206] fix(runtime): drop the renderer's legacy prose rules and escape `](` in warning codes The canonical renderer re-checks its own output. Its remaining prose rules were written for Markdown that agents used to write by hand, and they still matched byte-preserved upstream content. Each of these made projectCanonicalFinalReport throw, on this branch before this commit and on main: - a family variant titled "Item 1" (legacy numbered-item field); - a family variant titled "Source Node Id" (legacy source identifier); - a blocker summary "Critical" (unsupported Critical severity); - a blocker summary "Strategy loops: 4" (legacy run metadata). The final-report gate requires those fields to equal upstream, so the agent cannot repair them and every retry fails the same way. The other three rules (#### Sources, legacy report sections, legacy issue subsections) could not match rendered output at all, because publicProse escapes `#`. The self-check now keeps only its raw-HTML rule. Removing checks does not change any rendered bytes. artifact_validation_warnings[].code is free-form (1-128 characters) and is the one free-form field rendered as plain text without publicProse. With the link and image rules gone, a code such as `[notice](https://example.com/x)` rendered as a live link, in the report and in the public warnings companion. Only `](` is escaped there, so real gate codes such as ARTIFACT_OPTIONAL_METADATA_MISSING keep their bytes (full publicProse would escape their underscores). The three hand-mutated Markdown tests that used the deleted rules as probes now probe the raw-HTML rule through the same fence handling. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/final-report-markdown.ts | 30 ++--- .../test/final-report-markdown.test.ts | 109 ++++++++++++------ 2 files changed, 83 insertions(+), 56 deletions(-) diff --git a/packages/runtime/src/final-report-markdown.ts b/packages/runtime/src/final-report-markdown.ts index 03c17f2b2..3a8909aed 100644 --- a/packages/runtime/src/final-report-markdown.ts +++ b/packages/runtime/src/final-report-markdown.ts @@ -258,9 +258,11 @@ function finalReportMarkdownDirectiveViolation(markdown: string, report: JsonRec if (!markdown.includes("\n## Property provenance\n")) { return "missing property provenance"; } + // Report prose is preserved byte-for-byte from upstream artifacts that the agent cannot repair, so + // the only rule left is one escaped prose cannot match: publicProse escapes `<`, and only + // unescaped inline-code values can still carry raw HTML. const prose = markdownOutsideFencedCode(markdown).replace(//giu, ""); - const proseViolation = finalReportProseDirectiveViolation(prose); - if (proseViolation !== undefined) return proseViolation; + if (/<[A-Za-z][^>]*>/u.test(prose)) return "contains raw HTML outside fenced code"; const rendered = renderedIssues(Array.isArray(report.issues) ? report.issues.filter(isRecord) : []); const expectedHeadings = rendered.map(renderedIssueHeading); const headings = markdown.split("\n").filter((line) => line.startsWith("## [")); @@ -345,26 +347,6 @@ function completionFindingsViolation( return undefined; } -function finalReportProseDirectiveViolation(prose: string): string | undefined { - // Critical is not a supported report severity, but the word remains valid in explanatory prose - // (for example, "a critical invariant"). Reject only a standalone severity-like label rather than - // rewriting or discarding the validated finding text. - const forbiddenPatterns: ReadonlyArray = [ - [/(?:^|\n)(?:#{1,6}\s+|-\s+)?(?:\*\*)?Critical(?:\*\*)?\s*$/imu, "contains the unsupported Critical severity"], - [/(?:^|\n)#### Sources\s*$/imu, "contains a legacy Sources section"], - [/\*\*Source (?:Node|Property) Id\*\*/iu, "contains a legacy source identifier field"], - [/(?:^|\n)- \*\*Item \d+\*\*/imu, "contains a legacy numbered-item field"], - [/(?:^|\n)## (?:Executive summary|Issue index|Additional report data)\s*$/imu, "contains a legacy report section"], - [/(?:^|\n)#{3,6} (?:Lifecycle|Strategy|Strategy provenance)\s*$/imu, "contains a legacy issue subsection"], - [ - /(?:^|\n)- (?:Strategy loops|Audit profile catalog digest|Topology digest|Prompt digest|Expanded graph fingerprint):/imu, - "contains legacy run metadata" - ], - [/<[A-Za-z][^>]*>/u, "contains raw HTML outside fenced code"] - ]; - return forbiddenPatterns.find(([pattern]) => pattern.test(prose))?.[1]; -} - function validateReport(report: unknown): JsonRecord { const serialized = `${JSON.stringify(report)}\n`; const validation = validateArtifactContract("ultrafuzz/report@3", serialized, "report.json"); @@ -830,8 +812,10 @@ function appendArtifactValidationWarnings(lines: string[], value: unknown): void "" ); for (const warning of value.filter(isRecord)) { + // Codes render as plain text. Only `](` is escaped, so the bytes of real gate codes do not change. + const code = inlineValue(warning.code).replaceAll("](", "]\\("); lines.push( - `- ${inlineValue(warning.code)} — \`${inlineValue(warning.artifact_path)}#${inlineValue(warning.field_path)}\`: ${publicProse(String(warning.message))}` + `- ${code} — \`${inlineValue(warning.artifact_path)}#${inlineValue(warning.field_path)}\`: ${publicProse(String(warning.message))}` ); if (warning.source_path !== undefined) lines.push(` - Available context: \`${inlineValue(warning.source_path)}\``); } diff --git a/packages/runtime/test/final-report-markdown.test.ts b/packages/runtime/test/final-report-markdown.test.ts index 215bc4cc7..fba566021 100644 --- a/packages/runtime/test/final-report-markdown.test.ts +++ b/packages/runtime/test/final-report-markdown.test.ts @@ -1082,7 +1082,7 @@ test("directive validation rejects injected or presentation-divergent Markdown", const projection = projectCanonicalFinalReport(renderableReport()); assert.equal( isDirectiveConformingFinalReportMarkdown( - `${projection.markdown}\n## Executive summary\n\nInjected presentation.\n`, + `${projection.markdown}\n\n`, projection.report ), false @@ -1094,17 +1094,13 @@ test("directive validation rejects injected or presentation-divergent Markdown", ), false ); - for (const obsoleteLine of [ - "- Strategy loops: `3`", - `- Prompt digest: \`${"a".repeat(64)}\``, - "### Strategy\n\n| Strategy | Detection rate |\n| --- | --- |\n| stateful-invariant | 2/3 |" - ]) { - assert.equal( - isDirectiveConformingFinalReportMarkdown(`${projection.markdown}\n${obsoleteLine}\n`, projection.report), - false, - obsoleteLine - ); - } + assert.equal( + isDirectiveConformingFinalReportMarkdown( + projection.markdown.replace("\n## Property provenance\n", "\n## Provenance\n"), + projection.report + ), + false + ); }); test("directive conformance requires both coverage headings without an opt-out", () => { @@ -1133,35 +1129,27 @@ test("directive validation treats fenced proof code as code while retaining pros const input = renderableReport(); const issue = (input.issues as Array>)[0]!; issue.proof_of_concept = { - scenario: ["Prepare the bounded state.", "Execute the transition and observe the mismatch."], + scenario: ["Prepare the state where balance exceeds .", "Execute the transition and observe the mismatch."], language: "solidity", code: [ - "contract CriticalStateProbe {", - ' string internal constant label = "#### Sources";', - " // **Source Node Id** and ### Strategy provenance are target identifiers here.", + "contract MarkupProbe {", + ' string internal constant label = "bold";', + " // is a target string here.", "}" ].join("\n") }; const projection = projectCanonicalFinalReport(input); - assert.match(projection.markdown, /contract CriticalStateProbe/u); - assert.match(projection.markdown, /#### Sources/u); + assert.match(projection.markdown, /", "~~~~ " @@ -1190,7 +1177,7 @@ test("directive validation recognizes CommonMark tilde fences and matching close isDirectiveConformingFinalReportMarkdown( tildeProof.replace( "~~~~ \n\n## Property implementation coverage", - "~~~~ \n\nCritical\n\n## Property implementation coverage" + "~~~~ \n\nbold\n\n## Property implementation coverage" ), projection.report ), @@ -1198,7 +1185,7 @@ test("directive validation recognizes CommonMark tilde fences and matching close ); assert.equal( isDirectiveConformingFinalReportMarkdown( - insertProofBlock([" ~~~solidity", "Critical", " ~~~"].join("\n")), + insertProofBlock([" ~~~solidity", "", " ~~~"].join("\n")), projection.report ), false, @@ -1449,6 +1436,62 @@ test("upstream prose with link or image syntax renders as literal text", () => { assert.ok(nodes.some((node) => node.type === "paragraph" && node.text === "Call handlers[id](payload).")); }); +test("artifact validation warning codes with link or image syntax render as literal text", () => { + const report = renderableReport(); + const codes = ["[notice](https://example.com/x)", "![t](https://example.com/p.png)"]; + const warnings = codes.map((code) => ({ + code, + artifact_path: "artifacts/dedupe/strategy-detections.json", + field_path: "$[5].family_id", + message: "Optional metadata is missing", + gate: "strategy-detection-review-stage-reconciliation" + })); + (report.run_metadata as Record).artifact_validation_warnings = warnings; + + for (const markdown of [ + projectCanonicalFinalReport(report).markdown, + projectPublicArtifactValidationWarnings(warnings).markdown + ]) { + const nodes = markdownNodes(markdown); + assert.deepEqual( + nodes.filter((node) => node.type === "image" || (node.type === "link" && !node.url?.startsWith("#"))), + [] + ); + for (const code of codes) { + assert.ok( + nodes.some((node) => node.type === "paragraph" && node.text.startsWith(`${code} — `)), + code + ); + } + } +}); + +test("upstream prose that reads like a legacy report label still renders", () => { + const report = renderableReport(); + const [issue] = report.issues as Array>; + if (issue === undefined) throw new Error("missing issue fixture"); + issue.family_variants = [ + { id: "variant-1", title: "Item 1", summary: "The first sibling path.", dedupe_key: "variant-1" }, + { id: "variant-2", title: "Source Node Id", summary: "Mislabelled identifiers.", dedupe_key: "variant-2" } + ]; + (report.property_implementation_coverage as Record).blocker_summaries = [ + "Critical", + "Strategy loops: 4" + ]; + + const items = markdownNodes(projectCanonicalFinalReport(report).markdown) + .filter((node) => node.type === "listItem") + .map((node) => node.text); + for (const text of [ + "Item 1: The first sibling path.", + "Source Node Id: Mislabelled identifiers.", + "Critical", + "Strategy loops: 4" + ]) { + assert.ok(items.includes(text), text); + } +}); + test("issue titles with non-ASCII letters render with index anchors that resolve to their headings", () => { // GitHub heading slugs: lowercase, keep letters, marks, digits, spaces, "-" and "_", then spaces become "-". const slug = (text: string): string => From 561a8815e29d3dbd151a271d314aa18f953200eb Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:47:00 +0000 Subject: [PATCH 109/206] docs(changelog): file the terminal-publication relabel under breaking changes Terminal report publications written before this release no longer verify, because readers re-render them and the restated run summary differs from the stored bytes. `ultrafuzz report` and the dashboard then serve a complete run as the unchecked PARTIAL presentation. `report bundle` falls back to a report-only archive. `--require-verified` reads and the Modal public worker fail. Sync does not re-publish, because a publication status already exists for that terminal state. This is a user-visible break, so it moves to "Breaking changes". The old entry also named `](#anchor)` among the prose forms main accepted. main rejected it, because publicProse escapes `#`. The forms main accepted were `\](` and `](` before an escape-free `../` path. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 732c71a9d..e1e34b60b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,7 @@ ### Breaking changes +- **[runtime] [cli] [dashboard] [modal]** Terminal report publications written before this release no longer verify, because readers re-render them and the restated run summary (at least its elapsed time) differs from the stored bytes. `ultrafuzz report` and the dashboard then show a complete run as `# Ultrafuzz report — PARTIAL` with verification `not-checked`, `report bundle` falls back to a report-only archive, and `--require-verified` reads and the Modal public worker fail for that run; nothing re-publishes it automatically. Agent reports whose prose contains `\](` or `](` before an escape-free `../` relative path also render differently and no longer re-verify (#1151). - **[runtime] [cli] [docs]** Restores native Smithers continuation as the normal `resume` behavior. Ordinary resume now sends the persisted workflow and same Smithers run ID directly to Smithers with workflow-change acceptance instead of requiring Ultrafuzz control seals, link journals, controller generations, current schema bindings, graph identity, or metadata projections. `--refresh-controller` renders the current controller beside the historical source and continues that same Smithers run; it no longer publishes an authenticated historical generation. Completed Smithers rows and artifact bytes are not reset, replayed, migrated, or rewritten, and missing optional final-report metadata renders as `unavailable`. Replay determinism after accepting changed workflow source is therefore Smithers' caller-visible responsibility. The controller re-finalization CLI option is removed; use native resume or the existing explicit reset/retry operations (#939). - **[config] [docs] [evals]** Removes the `full` audit profile and renames its packaged topology to `topologies/exhaustive.yml` (byte-identical, topology digest unchanged); the `exhaustive` profile now runs that complete specialist workflow at its existing maximum settings, and no longer follows an edited `.ultrafuzz/topology.yml` (a project `topology_path` or `--topology-path` still overrides it). Configurations that select `full` now fail with the standard unknown-profile diagnostic. The public `full` and `threat-model` benchmark lanes launch with the `exhaustive` profile; the only effective policy change on those lanes is `same_agent_attempts` 3 → 5, since lane and target pins outrank the remaining profile settings (#792). - **[config] [docs]** Renames the unmodified built-in audit profile from `balanced` to `default` and removes the separate default pointer. Configurations that select `balanced` now fail with the standard unknown-profile diagnostic and must select `default` or omit `audit_profile`. Historical run metadata retains its exact recorded id; eval-history and benchmark reporting do not key cohorts by audit-profile id, so otherwise matching runs remain in the same reporting cohort across the migration boundary (#534). @@ -12,7 +13,7 @@ ### Other changes -- **[runtime] [prompts] [docs]** Runtime report presentations (the verified terminal publication and unchecked reports) now restate elapsed time, models, tokens, and estimated spend from `run.json` and the recorded finish time instead of the snapshot the report agent received when its task started, and unchecked reports render the run's goal-search census instead of always stating unknown coverage. The canonical renderer escapes `](` in prose, so it no longer rejects its own output when byte-preserved upstream prose contains link or image syntax, when an issue title contains a non-ASCII letter, or when the public projection redacts a path followed by `(`; the dead audit-context renderer and the prompt's contradictory `## Audit context` instructions are removed. Markdown for reports without `](` in prose is byte-identical, but terminal publications written by earlier versions read back as `not-checked` once the restated summary differs from their stored bytes (usually at least the elapsed time), and an agent report whose prose held a previously accepted `](` form (a `#anchor` or `../` link, or `\](`) no longer re-verifies (#1151). +- **[runtime] [prompts] [docs]** Runtime report presentations (the verified terminal publication and unchecked reports) restate elapsed time, models, tokens, and estimated spend from `run.json` and the recorded finish time instead of the report agent's task-start snapshot, and unchecked reports render the run's goal-search census instead of always stating unknown coverage. The canonical renderer escapes `](` in prose and in warning codes and keeps only its raw-HTML rule for prose, so byte-preserved upstream text with link or image syntax, a non-ASCII issue title, or a legacy-looking label such as a standalone `Critical` no longer makes it throw and fail the final-report gate on every retry. The dead audit-context renderer and the prompt's contradictory `## Audit context` instructions are removed (#1151). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 75cb1fcbf4995bbcb47591fe89e761d144691019 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:27:46 +0000 Subject: [PATCH 110/206] test(cli): poll the event log for the end of the resumed campaign `ultrafuzz status` synchronizes the run before it answers, unless the control evidence has diverged. `events` streams the engine's event log without synchronizing. The test now polls `events` until a terminal run event, calls `status` once, and reuses those events for the no-restart assertion. Locally, resume to run end fell from 165 s to 122 s. The wait for the held node also fails at once if the workflow stops before that node starts, instead of after the 15-minute bound. Co-Authored-By: Claude Opus 5.5 --- packages/cli/test/e2e/campaign-resume.test.ts | 21 ++++++++++++------- 1 file changed, 14 insertions(+), 7 deletions(-) diff --git a/packages/cli/test/e2e/campaign-resume.test.ts b/packages/cli/test/e2e/campaign-resume.test.ts index 3e777615b..42acb8c02 100644 --- a/packages/cli/test/e2e/campaign-resume.test.ts +++ b/packages/cli/test/e2e/campaign-resume.test.ts @@ -367,9 +367,13 @@ async function interruptMidRun(campaign: Campaign, mark: (phase: string) => void const [workflowRunId] = launched.workflow_ids; assert.ok(workflowRunId !== undefined, "run did not report its workflow run ID"); - const held = await waitFor(`${INTERRUPTED_NODE} to start`, 15 * MINUTE, () => - agentCalls(campaign).find((call) => call.node === INTERRUPTED_NODE && call.event === "held") - ); + const held = await waitFor(`${INTERRUPTED_NODE} to start`, 15 * MINUTE, () => { + const call = agentCalls(campaign).find((entry) => entry.node === INTERRUPTED_NODE && entry.event === "held"); + if (call === undefined && processesMentioning(workflowRunId).length === 0) { + assert.fail(`the workflow stopped before ${INTERRUPTED_NODE} started`); + } + return call; + }); // A host crash takes down the detached engine and the supervisor that would otherwise restart it. const controller = processesMentioning(workflowRunId); assert.ok(controller.length > 0, "no detached controller process is running the workflow"); @@ -409,11 +413,15 @@ test( mark("resume submitted"); assert.equal(resumed.submitted, true); - const health = await waitFor("the resumed run to end", 15 * MINUTE, async () => { - const current = await ultrafuzz(campaign, ["status", runId]); - return current.ended ? current : undefined; + // `events` reads the engine's event log without synchronizing the run, so it is the cheaper poll. + const events = await waitFor("the resumed workflow to finish", 15 * MINUTE, async () => { + const current = await ultrafuzz<{ events: WorkflowEvent[]; truncated: boolean }>(campaign, ["events", runId]); + const ended = ["RunFinished", "RunFailed", "RunCancelled"]; + return current.events.some((event) => ended.includes(event.category)) ? current : undefined; }); + const health = await ultrafuzz(campaign, ["status", runId]); mark("run ended"); + assert.equal(health.ended, true); assert.equal(health.status, "succeeded"); assert.deepEqual( { @@ -439,7 +447,6 @@ test( AGENT_NODES.map((node) => [node, starts.filter((call) => call.node === node).length]), AGENT_NODES.map((node) => [node, node === INTERRUPTED_NODE ? 2 : 1]) ); - const events = await ultrafuzz<{ events: WorkflowEvent[]; truncated: boolean }>(campaign, ["events", runId]); assert.equal(events.truncated, false); assert.equal(events.events.filter((event) => event.category === "RunStarted").length, 2); assertNoFinishedTaskRestarted(events.events); From f7a5fb7356b3539bceaaf578c40572e93735d2f6 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:50:58 +0000 Subject: [PATCH 111/206] refactor(prompts): leave getPrompt to the catalog work item getPrompt still has no caller, but #1164 adds projectPromptsDifferingFromBuiltIns directly above it in catalog.ts, so deleting it here made the two PRs conflict textually (git merge-tree reported CONFLICT (content) in catalog.ts). catalog.ts is outside this change's files, so restore it to origin/main and leave the deletion for a follow-up once #1164 lands. Co-Authored-By: Claude Opus 5.5 --- packages/prompts/src/catalog.ts | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/packages/prompts/src/catalog.ts b/packages/prompts/src/catalog.ts index b339be718..576a92ca9 100644 --- a/packages/prompts/src/catalog.ts +++ b/packages/prompts/src/catalog.ts @@ -88,6 +88,14 @@ export function loadPromptCatalog(options: LoadPromptCatalogOptions = {}): Promp }; } +export function getPrompt(catalog: PromptCatalog, id: string): PromptCatalogEntry { + const entry = catalog.entries.get(id); + if (!entry) { + throw new PromptError("missing-template-variable", `prompt id \`${id}\` was not found`); + } + return entry; +} + export function builtInPromptRelativePaths(): string[] { return discoverBuiltInPromptRelativePaths(builtInPromptRoot()); } From d7704ae79a53a3d2c8bfcc0dba461647d8f04ac3 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:50:58 +0000 Subject: [PATCH 112/206] docs(changelog): state the dead-code entry only as strongly as the evidence Drop getPrompt, which this change no longer deletes. Name the actual callers of the two CI scripts (eval-benchmarks.yml and the watch-modal-benchmark.sh it ran, both deleted in #1131) instead of "the workflows #1131 deleted": the other two workflows #1131 removed never invoked either script. Stop describing the config prompt-metadata layer as code no production path calls: resolveConfig ran it on every resolution, only ever with an empty layer. Mention the run-tests.mjs selectors and the docs/config.md change, and add the [runtime] and [docs] tags for them. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c347b82f3..7cae95fc1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[prompts] [config] [modal] [dashboard] [ci]** Deletes code that no production path calls: the prompt rename feature (`renamePromptId`, `renamePromptArtifactReferences`, `diffPromptIdentity`, `serializePromptDocument`, `getPrompt`), the config prompt-metadata layer, `restoreRedactedConfig`, `assertNoRedactionPlaceholders`, `applyDefaultProfileOverrides` and `assertResolvedConfigZod`, unreferenced or test-only Modal exports, and `scripts/ci/validate-threat-model-benchmark-gate.mjs` and `scripts/ci/modal-benchmark-control-window.mjs`, whose only callers were the workflows #1131 deleted. Two observable changes remain: output-contract templates now resolve through the same packaged-first lookup as the other prompt assets, so the renderer no longer probes a `dist/prompts` layout that no build produces, and the dashboard's `/api/flow` `capabilities` object drops seven always-false commands the frontend never read (#462). +- **[prompts] [config] [modal] [dashboard] [runtime] [ci] [docs]** Deletes code that no production path calls: the prompt rename feature (`renamePromptId`, `renamePromptArtifactReferences`, `diffPromptIdentity`, `serializePromptDocument`), `restoreRedactedConfig`, `assertNoRedactionPlaceholders`, `applyDefaultProfileOverrides` and `assertResolvedConfigZod`, unreferenced or test-only Modal exports, three `packages/runtime/scripts/run-tests.mjs` selectors, and `scripts/ci/validate-threat-model-benchmark-gate.mjs` and `scripts/ci/modal-benchmark-control-window.mjs`, whose only callers outside tests were `eval-benchmarks.yml` and the `watch-modal-benchmark.sh` it ran, both deleted in #1131. It also deletes the config prompt-metadata layer, which `resolveConfig` ran on every resolution but only ever with an empty layer, since no production caller passed `promptMetadata`; `docs/config.md` no longer describes the redaction restore step and placeholder launch guard that no code runs. Two observable changes remain: output-contract templates now resolve through the same packaged-first lookup as the other prompt assets, so the renderer no longer probes the `dist/prompts` layout that builds stopped producing in #796, and the dashboard's `/api/flow` `capabilities` object drops seven always-false commands the frontend never read (#462). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 36b70679906ef3abdd23dcb1fd9f14f461da96a4 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:00 +0000 Subject: [PATCH 113/206] docs(changelog): say the restated summary keeps values run.json lacks The entry said runtime presentations restate the run summary from run.json, but not that a value those records do not carry keeps the agent's copy. docs/reference/artifacts-reports.md already describes that fallback; the entry now says so too. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index e1e34b60b..7f2d61415 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,7 +13,7 @@ ### Other changes -- **[runtime] [prompts] [docs]** Runtime report presentations (the verified terminal publication and unchecked reports) restate elapsed time, models, tokens, and estimated spend from `run.json` and the recorded finish time instead of the report agent's task-start snapshot, and unchecked reports render the run's goal-search census instead of always stating unknown coverage. The canonical renderer escapes `](` in prose and in warning codes and keeps only its raw-HTML rule for prose, so byte-preserved upstream text with link or image syntax, a non-ASCII issue title, or a legacy-looking label such as a standalone `Critical` no longer makes it throw and fail the final-report gate on every retry. The dead audit-context renderer and the prompt's contradictory `## Audit context` instructions are removed (#1151). +- **[runtime] [prompts] [docs]** Runtime report presentations (the verified terminal publication and unchecked reports) restate elapsed time, models, tokens, and estimated spend from `run.json` and the recorded finish time, where those records carry them, instead of the report agent's task-start snapshot, and unchecked reports render the run's goal-search census instead of always stating unknown coverage. The canonical renderer escapes `](` in prose and in warning codes and keeps only its raw-HTML rule for prose, so byte-preserved upstream text with link or image syntax, a non-ASCII issue title, or a legacy-looking label such as a standalone `Critical` no longer makes it throw and fail the final-report gate on every retry. The dead audit-context renderer and the prompt's contradictory `## Audit context` instructions are removed (#1151). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From cd145af0dcb24b499b23826cd97606ba2d45a905 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:21 +0000 Subject: [PATCH 114/206] fix(evals): re-read run state after the watch loop before classifying a row Deleting the telemetry pump also deleted the state.json re-read that followed the final drain. When the watch deadline passed during the last poll sleep, the row was then classified from the state read before that sleep. The watch is not the only writer of state.json: `ultrafuzz status`, `inspect` and the dashboard synchronize the run too. A run one of them made terminal during that sleep was recorded as timed-out, with an EVAL_ROW_WATCH_TIMEOUT diagnostic and a `running` workflow lifecycle, and counted as incomplete. origin/main recorded its real terminal status. Restore the single read, as origin/main had it. The new watchEvalRow test polls once with a 1 s deadline and a 1.5 s sleep, and a timer writes the terminal run 300 ms after that poll. It passes on origin/main (test ported with `reporters: []`), fails on the previous branch head with final_status "timed-out", and passes with this change. Co-Authored-By: Claude Opus 5.5 --- docs/reference/evals.md | 7 +++- packages/evals/src/runner.ts | 3 ++ packages/evals/test/runner-publish.test.ts | 49 +++++++++++++++++++++- 3 files changed, 56 insertions(+), 3 deletions(-) diff --git a/docs/reference/evals.md b/docs/reference/evals.md index abeca5d22..0f7d612db 100644 --- a/docs/reference/evals.md +++ b/docs/reference/evals.md @@ -225,8 +225,11 @@ The eval driver owns the local loop and writes `matrix.json`, `runs.jsonl`, `ultrafuzz eval run` launches each row as a detached Ultrafuzz run. When it watches a row, it repeats three steps until the run's durable `state.json` is terminal or the watch deadline passes: synchronize the run, read `state.json`, -and sleep for the poll interval. A failed synchronization is counted and -recorded on the row as `EVAL_ROW_SYNC_FAILED`; it does not end the watch. +and sleep for the poll interval. It reads `state.json` once more before +recording the row, so a run that another command synchronized to a terminal +state during the last sleep is not recorded as timed out. A failed +synchronization is counted and recorded on the row as `EVAL_ROW_SYNC_FAILED`; +it does not end the watch. There is no reporter or telemetry-export interface. ## Versioned lineage diff --git a/packages/evals/src/runner.ts b/packages/evals/src/runner.ts index f9264799f..a2edac233 100644 --- a/packages/evals/src/runner.ts +++ b/packages/evals/src/runner.ts @@ -617,6 +617,9 @@ export async function watchEvalRow( await sleep(pollIntervalMs); } + // Other syncers (`ultrafuzz status`, the dashboard) also write state.json; a + // run they finished during the final sleep is terminal, not timed out. + state = readStateSafe(runRoot, input.record.ultrafuzz_run_id); const watchTimedOut = Date.now() >= deadline && (state === undefined || !isTerminalRunStatus(state.status)); const syncFailureDiagnostic: RuntimeDiagnostic | undefined = syncFailureCount === 0 diff --git a/packages/evals/test/runner-publish.test.ts b/packages/evals/test/runner-publish.test.ts index c0c74a124..3a43b33ea 100644 --- a/packages/evals/test/runner-publish.test.ts +++ b/packages/evals/test/runner-publish.test.ts @@ -3,7 +3,12 @@ import { mkdtempSync } from "node:fs"; import { tmpdir } from "node:os"; import path from "node:path"; -import { EVENT_SCHEMA_VERSION, assertEventRecord, createNodeAttemptLedgerEntry } from "@ultrafuzz/artifacts"; +import { + EVENT_SCHEMA_VERSION, + assertEventRecord, + createNodeAttemptLedgerEntry, + readRunState +} from "@ultrafuzz/artifacts"; import { initProject } from "@ultrafuzz/runtime"; import { describe, expect, it, vi } from "vitest"; @@ -1342,6 +1347,48 @@ Write a neutral fixture message to {{artifact_path}}/fixture.md. expect(watched.record.diagnostics.map((diagnostic) => diagnostic.code)).toContain("EVAL_ROW_WATCH_TIMEOUT"); }); + it("records the terminal status another writer reached while the watch slept past its deadline", async () => { + const base = mkdtempSync(path.join(fs.realpathSync(tmpdir()), "ufz-evals-watch-final-sleep-")); + const suite = testSuite(path.join(base, "gt")); + const row = testRow(suite); + const runRoot = path.join(base, "target", ".ultrafuzz", "runs", "run-1"); + writeRunFixture({ + runRoot, + events: [], + state: currentRunState({ + runId: "run-1", + status: "running", + nodes: {}, + overrides: { created_at: T0, started_at: T0, last_transition_at: T0 } + }) + }); + // Compile the state validator first so the one poll starts well inside the 1 s deadline. + readRunState(path.join(runRoot, "state.json")); + const evalRunRoot = path.join(base, "eval-run"); + let syncCalls = 0; + + const watched = await watchEvalRow({ + plan: { suite_path: "suite.yml", project_root: base, suite, matrix: [row] }, + row, + record: materializeLaunchedJournal(row, runRoot, evalRunRoot), + evalRunRoot, + sync: async () => { + syncCalls += 1; + // Another syncer (`ultrafuzz status`, the dashboard) finishes the run during the 1.5 s sleep. + setTimeout(() => terminalRunFixture(runRoot), 300); + }, + pollIntervalMs: 1_500, + timeoutSeconds: 1 + }); + + expect(syncCalls).toBe(1); + expect(watched.record).toMatchObject({ + final_status: "succeeded", + workflow: { status: "succeeded", terminal: true } + }); + expect(watched.diagnostics.map((diagnostic) => diagnostic.code)).not.toContain("EVAL_ROW_WATCH_TIMEOUT"); + }, 20_000); + it("coalesces and persists workflow synchronization failures", async () => { const base = mkdtempSync(path.join(fs.realpathSync(tmpdir()), "ufz-evals-watch-sync-failure-")); const suite = testSuite(path.join(base, "gt")); From 262852506e4f2d571490aa375155611f403a45f0 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:28 +0000 Subject: [PATCH 115/206] docs(evals): stop describing eval runs as keeping telemetry With the pump gone, eval runs write no telemetry, and the suite's `reporting.artifacts` policy has no reporter to stream to. Two docs still said eval runs retain telemetry, and the EvalArtifactPolicy and EvalTarget.sensitivity comments still described payload streaming and manifest-only artifact reporting. The docs and the default suite comment also said "nothing reads" the artifacts policy. That overstated it: normalizeReporting still computes it, and `eval plan --json` and the Modal worker still carry it. They now say no eval behaviour depends on it. The `sensitivity` comment now names what `private` does enforce: ground truth bound to the target's repo and ref, and ULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA before an LLM judge receives the data. Co-Authored-By: Claude Opus 5.5 --- .ultrafuzz/evals/bug-finding.yml | 2 +- docs/how-to/run-evals-on-modal.md | 4 ++-- docs/reference/cli.md | 4 ++-- docs/reference/evals.md | 2 +- packages/evals/src/types.ts | 17 +++++++++-------- 5 files changed, 15 insertions(+), 14 deletions(-) diff --git a/.ultrafuzz/evals/bug-finding.yml b/.ultrafuzz/evals/bug-finding.yml index 66ed203e8..7302e11ab 100644 --- a/.ultrafuzz/evals/bug-finding.yml +++ b/.ultrafuzz/evals/bug-finding.yml @@ -45,7 +45,7 @@ reporting: # policy only — provider selection/credentials live in ultrafuzz.to node_telemetry: true heartbeat_interval_seconds: 60 experiment_prefix: bug-finding - artifacts: # still validated, but nothing reads the artifacts policy + artifacts: # still validated, but no eval behaviour depends on it mode: manifest-only include: ["report.md", "report.json"] max_file_bytes: 5000000 diff --git a/docs/how-to/run-evals-on-modal.md b/docs/how-to/run-evals-on-modal.md index 32e28f7a2..89165747f 100644 --- a/docs/how-to/run-evals-on-modal.md +++ b/docs/how-to/run-evals-on-modal.md @@ -1,8 +1,8 @@ # Run evals on Modal Ultrafuzz can run long evaluation rows in Modal sandboxes while retaining -telemetry, scores, and reports in the run workspace. The runner uses the Modal -TypeScript SDK; Python is not required. +scores and reports in the run workspace. The runner uses the Modal TypeScript +SDK; Python is not required. ## Security and storage boundaries diff --git a/docs/reference/cli.md b/docs/reference/cli.md index e6cd7ac03..ee5ba4cf0 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -744,8 +744,8 @@ directory per target id under `--target-root`) against the pinned git refs; `--skip-target-validation` skips that check. `run` launches Ultrafuzz runs for matrix rows (all rows, or a `--row` -selection) and polls them to a terminal state. Reports and telemetry remain -in local run artifacts. `--no-watch` launches detached without polling. +selection) and polls them to a terminal state. Reports remain in local run +artifacts. `--no-watch` launches detached without polling. `status` reads the eval matrix, its latest `runs.jsonl` records, and each linked durable `state.json` without synchronizing or changing workflow state. diff --git a/docs/reference/evals.md b/docs/reference/evals.md index 0f7d612db..919e90e39 100644 --- a/docs/reference/evals.md +++ b/docs/reference/evals.md @@ -42,7 +42,7 @@ provider, an endpoint, or an env var. launched rows by default; `--no-watch` overrides it. It and `reporting.heartbeat_interval_seconds` remain inputs to the execution-policy fingerprint. `reporting.experiment_prefix` and `reporting.artifacts` are still -validated, but nothing reads them. +validated, but no eval behaviour depends on them. ### Suite contract and workflow input diff --git a/packages/evals/src/types.ts b/packages/evals/src/types.ts index 95c7112bf..fcf8d09fb 100644 --- a/packages/evals/src/types.ts +++ b/packages/evals/src/types.ts @@ -94,7 +94,10 @@ export interface EvalTarget { signal_profile?: string; /** Relative ground-truth file resolved strictly under the operator-supplied `[eval].ground_truth_root`. */ ground_truth: string; - /** `private` forces manifest-only artifact reporting unless the suite explicitly opts into `upload`. */ + /** + * `private` requires the ground truth to be bound to this repo and ref, and + * `ULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA=true` before an LLM judge receives it. + */ sensitivity?: "public" | "private"; /** * Benchmark paths the run must never read, such as a reference solution the @@ -165,17 +168,15 @@ export interface EvalRecoveryEquivalence { export type EvalArtifactMode = "manifest-only" | "upload"; +/** + * The suite's `reporting.artifacts` block: validated and recorded, but no eval + * behaviour depends on it. `mode` defaults to `manifest-only`. + */ export interface EvalArtifactPolicy { - /** - * `manifest-only` publishes file names/sizes/hashes only; `upload` also streams payloads. - * `manifest-only` is the default and is always forced for `sensitivity: private` - * targets unless the suite explicitly sets `upload` (the opt-in). - */ mode: EvalArtifactMode; - /** Allowlist of artifact file names eligible for streaming to the provider. */ include: string[]; max_file_bytes: number; - /** True when the suite YAML explicitly set `mode` (the privacy opt-in signal). */ + /** True when the suite YAML explicitly set `mode`. */ mode_explicit: boolean; } From 6a17f729179e2f08869ba14b54a74a9c8c5c68ef Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:59 +0000 Subject: [PATCH 116/206] fix(prompts): tell Vyper setup agents to declare interfaces in the test tree The setup-foundry and base-test-setup prompts ask for Solidity interfaces to the Vyper contracts' ABI but never say where to put them. Both nodes publish a workspace patch, and captureWorkspacePatch rejects any change under the production source roots (src and contracts by default), so an interface written under contracts/interfaces/ fails the setup handoff, and the setup group halts the run on failure. Say to declare new interfaces in the test tree, as the stateful-invariant prompts now do. Whether agents actually chose a production location here was not observed; this closes the same gap the invariant prompts had. Co-Authored-By: Claude Opus 5.5 --- .ultrafuzz/prompts/setup/discover-base-test.md | 6 ++++-- .ultrafuzz/prompts/setup/prepare-foundry-harness.md | 4 +++- 2 files changed, 7 insertions(+), 3 deletions(-) diff --git a/.ultrafuzz/prompts/setup/discover-base-test.md b/.ultrafuzz/prompts/setup/discover-base-test.md index 64ccf3ecc..1d9563348 100644 --- a/.ultrafuzz/prompts/setup/discover-base-test.md +++ b/.ultrafuzz/prompts/setup/discover-base-test.md @@ -39,8 +39,10 @@ If the setup handoffs identify Vyper-only or mixed Solidity/Vyper production contracts, make the reusable Foundry fixture Vyper-aware while keeping the tests Solidity-based. Define Solidity interfaces for the Vyper contracts' ABI-visible public/external functions and events, or reuse ABI-derived interfaces generated -by the target repository. Do not require Foundry to compile `.vy` files as -Solidity sources. +by the target repository. Declare any new interface in the test tree: the +workspace handoff rejects every change under the production source roots (by +default `src/` and `contracts/`). Do not require Foundry to compile `.vy` files +as Solidity sources. For Vyper deployment helpers, prefer one reusable path that compiles creation bytecode with the target project's pinned compiler/tooling from the project diff --git a/.ultrafuzz/prompts/setup/prepare-foundry-harness.md b/.ultrafuzz/prompts/setup/prepare-foundry-harness.md index 0cc1e14ac..226ab1644 100644 --- a/.ultrafuzz/prompts/setup/prepare-foundry-harness.md +++ b/.ultrafuzz/prompts/setup/prepare-foundry-harness.md @@ -22,7 +22,9 @@ production contracts, keep Foundry as the test harness and do not ask Foundry to compile `.vy` files as Solidity sources. Configure the harness so generated `.t.sol` tests interact with Vyper contracts through Solidity interfaces that match the contracts' public/external ABI, or through ABI-derived Solidity -interfaces when the target repository already generates them. +interfaces when the target repository already generates them. Declare any new +interface in the test tree: the workspace handoff rejects every change under the +production source roots (by default `src/` and `contracts/`). Create only the minimal harness layout needed by later fuzzing agents. From ff229f9c6fa508a57b833617fa69ac9dedc5558f Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:59 +0000 Subject: [PATCH 117/206] fix(runtime): doctor no longer summarizes a validation with warnings as a pass doctor reported the validate check as `warning` but kept the summary "config, topology, prompts, paths, agents, and trust posture pass". Now that validate warns about project prompts that differ from their built-in, every upgraded project with such a copy would see that contradiction. The summary now says validation passed with warnings and points to `ultrafuzz validate --json`, since text-mode output does not print warnings. The new test fails on the previous doctor.ts with actual 'config, topology, prompts, paths, agents, and trust posture pass'. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/doctor.ts | 9 ++++++--- .../runtime/test/prompt-catalog-validation.test.ts | 13 ++++++++++++- 2 files changed, 18 insertions(+), 4 deletions(-) diff --git a/packages/runtime/src/doctor.ts b/packages/runtime/src/doctor.ts index 0353a8151..ce64db2d6 100644 --- a/packages/runtime/src/doctor.ts +++ b/packages/runtime/src/doctor.ts @@ -58,9 +58,12 @@ export async function diagnoseProject(input: DoctorInput) { checks.push({ name: "validate", status: validationStatus, - summary: validation.ok - ? "config, topology, prompts, paths, agents, and trust posture pass" - : "configuration validation reported errors; run ultrafuzz validate for detail" + summary: + validationStatus === "ok" + ? "config, topology, prompts, paths, agents, and trust posture pass" + : validationStatus === "warning" + ? "configuration validation passed with warnings; run ultrafuzz validate --json for detail" + : "configuration validation reported errors; run ultrafuzz validate for detail" }); diagnostics.push(...validation.diagnostics); diff --git a/packages/runtime/test/prompt-catalog-validation.test.ts b/packages/runtime/test/prompt-catalog-validation.test.ts index 6adc19c36..02b7c1fb3 100644 --- a/packages/runtime/test/prompt-catalog-validation.test.ts +++ b/packages/runtime/test/prompt-catalog-validation.test.ts @@ -3,7 +3,7 @@ import fs from "node:fs"; import path from "node:path"; import test from "node:test"; -import { initProject, validateProject } from "../src/index.js"; +import { diagnoseProject, initProject, validateProject } from "../src/index.js"; import { temporaryRoot } from "./temporary-root.js"; // Runs use the project copy of a prompt, and `ultrafuzz init` without `--force` keeps it. After an @@ -38,3 +38,14 @@ test("validate warns about a project prompt that differs from the built-in promp assert.equal(restored.value?.policy_posture.prompts.status, "pass"); assert.deepEqual(driftWarnings(restored), []); }); + +test("doctor does not summarize a validation that only warned as a pass", async () => { + const project = temporaryRoot("ufz-prompt-drift-doctor-"); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + fs.appendFileSync(path.join(project, ".ultrafuzz", "prompts", "review", "triage.md"), "\nA local edit.\n", "utf8"); + + const doctor = await diagnoseProject({ projectRoot: project, env: {}, offline: true }); + const validation = doctor.value?.checks.find((check) => check.name === "validate"); + assert.equal(validation?.status, "warning"); + assert.match(validation?.summary ?? "", /with warnings/u); +}); From 1193208798c2c3ef7ea9d2a88cf9c2aff30df8ba Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:59 +0000 Subject: [PATCH 118/206] chore: move the prompt drift helper and state what the anchor test covers Move projectPromptsDifferingFromBuiltIns below builtInPromptRelativePaths so it no longer sits in the hunk where #1175 deletes getPrompt; with that placement the two branches merge without a conflict in catalog.ts. No behaviour change. The anchor test comment claimed a prompt and its topology "cannot drift apart again". The test checks only nodes whose built-in prompt promises an extended timeout (triage and severity-classification) in the packaged topologies, so say that, and that dedupe-findings is not covered. Co-Authored-By: Claude Opus 5.5 --- packages/prompts/src/catalog.ts | 24 +++++++++---------- .../topology/test/packaged-topologies.test.ts | 5 ++-- 2 files changed, 15 insertions(+), 14 deletions(-) diff --git a/packages/prompts/src/catalog.ts b/packages/prompts/src/catalog.ts index a9e45f700..d9a39199f 100644 --- a/packages/prompts/src/catalog.ts +++ b/packages/prompts/src/catalog.ts @@ -88,6 +88,18 @@ export function loadPromptCatalog(options: LoadPromptCatalogOptions = {}): Promp }; } +export function getPrompt(catalog: PromptCatalog, id: string): PromptCatalogEntry { + const entry = catalog.entries.get(id); + if (!entry) { + throw new PromptError("missing-template-variable", `prompt id \`${id}\` was not found`); + } + return entry; +} + +export function builtInPromptRelativePaths(): string[] { + return discoverBuiltInPromptRelativePaths(builtInPromptRoot()); +} + /** * Relative paths of project prompts whose bytes differ from the built-in prompt at the same path. Runs * use the project copy and `ultrafuzz init` keeps it unless run with `--force`, so after an upgrade such @@ -102,18 +114,6 @@ export function projectPromptsDifferingFromBuiltIns(catalog: PromptCatalog): str .sort(); } -export function getPrompt(catalog: PromptCatalog, id: string): PromptCatalogEntry { - const entry = catalog.entries.get(id); - if (!entry) { - throw new PromptError("missing-template-variable", `prompt id \`${id}\` was not found`); - } - return entry; -} - -export function builtInPromptRelativePaths(): string[] { - return discoverBuiltInPromptRelativePaths(builtInPromptRoot()); -} - function discoverBuiltInPromptRelativePaths(root: string): string[] { return discoverPromptFiles(root) .map((absolutePath) => normalizePromptRelativePath(path.relative(root, absolutePath))) diff --git a/packages/topology/test/packaged-topologies.test.ts b/packages/topology/test/packaged-topologies.test.ts index 08c2ccdeb..26b8cb7d9 100644 --- a/packages/topology/test/packaged-topologies.test.ts +++ b/packages/topology/test/packaged-topologies.test.ts @@ -493,8 +493,9 @@ describe("packaged topology collection", () => { // #1150. The triage and severity-classification prompts tell the agent that "the topology gives // this review node an extended timeout". No shipped topology did: the review group lost its pin in // v0.0.2, so the widest fan-in and the per-finding panel stages ran on the run default while every - // strategy lane had the reviewed window. Each promise is checked against the expanded graph, so a - // prompt and its topology cannot drift apart again. + // strategy lane had the reviewed window. Every built-in prompt that makes this promise is checked + // against each packaged topology's expanded graph; nodes whose prompt promises nothing, such as + // dedupe-findings, are not covered here. it("gives every node whose prompt promises an extended timeout more than the default", { timeout: 30_000 }, () => { const shadowed = largestShadowedDefault(); const promisingPrompts = new Set( From ba7bd6c911e8e814fdf72914385b7f1a1725f264 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:51:59 +0000 Subject: [PATCH 119/206] docs: record the review pin's precedence break and the topology refresh A group pin takes precedence over run.default_timeout_seconds and model-profile timeout_seconds (#675), so a project that raised either above 7200 now gets 7200 on review nodes of the packaged exhaustive and invariant-only topologies and of new scaffolds, and a cloud config whose explicit global resource timeout is below 7200 needs per-node overrides for the review nodes as well. Record that under Breaking changes with the escape hatch. Tighten the Other changes entry: name the three pinned topologies instead of "the shipped topologies" (smoke stays unpinned), and give the one-command refresh for an uncustomized project topology, which init keeps. The how-to now says the same about the topology as it already did about prompts. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 3 ++- docs/how-to/edit-prompts-topology.md | 7 +++++++ 2 files changed, 9 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ae4ab7f1b..cfd2b7c4a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,7 @@ ### Breaking changes +- **[config] [docs]** Because a group pin takes precedence over `run.default_timeout_seconds` and model-profile `timeout_seconds` (#675), review nodes in the packaged `exhaustive` and `invariant-only` topologies and in new project scaffolds now get 7,200 seconds even where either setting is higher. To keep a larger review budget on a packaged topology, `ultrafuzz topology copy` it into the project, raise the `review` group's `timeout_seconds` there, and point `topology_path` at the copy. In cloud mode, an explicit `execution.resources.timeout_seconds` below 7,200 now needs per-node `execution.nodes..resources.timeout_seconds` overrides for the review nodes too, or the run fails to launch with `CLOUD_TASK_TIMEOUT_BUDGET_EXCEEDED` (#1150). - **[runtime] [cli] [docs]** Restores native Smithers continuation as the normal `resume` behavior. Ordinary resume now sends the persisted workflow and same Smithers run ID directly to Smithers with workflow-change acceptance instead of requiring Ultrafuzz control seals, link journals, controller generations, current schema bindings, graph identity, or metadata projections. `--refresh-controller` renders the current controller beside the historical source and continues that same Smithers run; it no longer publishes an authenticated historical generation. Completed Smithers rows and artifact bytes are not reset, replayed, migrated, or rewritten, and missing optional final-report metadata renders as `unavailable`. Replay determinism after accepting changed workflow source is therefore Smithers' caller-visible responsibility. The controller re-finalization CLI option is removed; use native resume or the existing explicit reset/retry operations (#939). - **[config] [docs] [evals]** Removes the `full` audit profile and renames its packaged topology to `topologies/exhaustive.yml` (byte-identical, topology digest unchanged); the `exhaustive` profile now runs that complete specialist workflow at its existing maximum settings, and no longer follows an edited `.ultrafuzz/topology.yml` (a project `topology_path` or `--topology-path` still overrides it). Configurations that select `full` now fail with the standard unknown-profile diagnostic. The public `full` and `threat-model` benchmark lanes launch with the `exhaustive` profile; the only effective policy change on those lanes is `same_agent_attempts` 3 → 5, since lane and target pins outrank the remaining profile settings (#792). - **[config] [docs]** Renames the unmodified built-in audit profile from `balanced` to `default` and removes the separate default pointer. Configurations that select `balanced` now fail with the standard unknown-profile diagnostic and must select `default` or omit `audit_profile`. Historical run metadata retains its exact recorded id; eval-history and benchmark reporting do not key cohorts by audit-profile id, so otherwise matching runs remain in the same reporting cohort across the migration boundary (#534). @@ -12,7 +13,7 @@ ### Other changes -- **[config] [prompts] [runtime] [docs]** The shipped topologies pin the `review` group to the same 7,200-second timeout as the goal, strategy and specialist groups. The group had no pin, so dedupe, triage, severity classification, test aggregation and the final report ran on `run.default_timeout_seconds` (3,600 by default), although the triage and severity prompts promise an extended timeout; like the other group pins, it takes precedence over the run and model-profile timeouts (#675), and an existing `.ultrafuzz/topology.yml` keeps its old review budget until `defaults: { timeout_seconds: 7200 }` is added to its `review` group. The stateful-invariant setup and property prompts no longer permit interface edits under the production source roots (`src/` and `contracts/` by default), which the workspace handoff rejects, the unused `timeout_seconds: 300` on reference nodes is removed, and `ultrafuzz validate` warns (`PROMPT_DIFFERS_FROM_BUILT_IN`) about each project prompt that differs from the built-in prompt at the same path, because `ultrafuzz init` without `--force` keeps existing prompts (#1150). +- **[config] [prompts] [runtime] [docs]** The default, exhaustive and invariant-only topologies pin the `review` group to 7,200 seconds (previously `run.default_timeout_seconds`, 3,600 by default), as the triage and severity prompts already promised, and drop the unused `timeout_seconds: 300` on reference nodes. An existing `.ultrafuzz/topology.yml` keeps its old review budget until you run `ultrafuzz topology copy default .ultrafuzz/topology.yml --force` (if uncustomized) or add `defaults: { timeout_seconds: 7200 }` to its `review` group. The stateful-invariant prompts no longer invite interface edits under the production source roots, which the workspace handoff rejects; they and the Vyper setup prompts point agents at the test tree instead. `ultrafuzz validate` warns (`PROMPT_DIFFERS_FROM_BUILT_IN`) about each project prompt that differs from the built-in prompt at its path, and `doctor` no longer summarizes such a validation as a pass (#1150). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/how-to/edit-prompts-topology.md b/docs/how-to/edit-prompts-topology.md index 2eaaaf9e7..d2c3fc106 100644 --- a/docs/how-to/edit-prompts-topology.md +++ b/docs/how-to/edit-prompts-topology.md @@ -91,6 +91,13 @@ nodes: Agentic nodes run through the configured workflow adapter. Topology does not define arbitrary shell runners. +`ultrafuzz init` also keeps an existing `.ultrafuzz/topology.yml`, so after an +upgrade it still has the topology of the release that scaffolded it. If you +have not customized it, +`ultrafuzz topology copy default .ultrafuzz/topology.yml --force` replaces it +with the current packaged default and leaves `ultrafuzz.toml` and the prompts +alone; otherwise, merge the release's topology changes by hand. + ## Declare Durable Handoffs Use `outputs` for files a node must write under its artifact directory. Every From 2b9ac052c70f1d8d194ee8bb8368201945388971 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:52:10 +0000 Subject: [PATCH 120/206] fix(runtime): report a launch still publishing its snapshot as incomplete A launch seals its workflow controls, publishes the execution snapshot (its longest step), and only then writes the workflow-link journal, with the run still `pending` throughout. A reader that takes no control lock meets that window: `status` already did, and `pause` and `cancel` now do too since they read evidence observe-only. All of them answered WORKFLOW_RUN_LINK_JOURNAL_MISSING, whose text says the run "cannot be safely upgraded in place" and tells the operator to start a new run with a new run ID, which is the wrong advice for a healthy launch. A pending run whose seal exists but whose journal does not now gets the same incomplete-launch diagnostic as a pending run with no seal (WORKFLOW_CONTROL_SEAL_PENDING, which `status` renders as `launch-incomplete`), naming the journal. A run that is no longer pending and lacks its journal still gets WORKFLOW_RUN_LINK_JOURNAL_MISSING. The new test reproduces that on-disk state from a finished fake launch (state set back to pending, journal removed) and checks that pause and cancel report the incomplete launch without invoking the runner and that status reports launch-incomplete. It fails on origin/main and on the previous branch head with WORKFLOW_RUN_LINK_JOURNAL_MISSING. The CHANGELOG entry also stops saying that pause and cancel refuse a run whose "published execution snapshot files" changed: the generated workflow in that snapshot is a control document, and a hand-patched copy is tolerated; only the sealed execution files are still refused. Refs #674, #921 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- packages/runtime/src/start-run.ts | 14 ++++++++++- packages/runtime/test/runtime.test.ts | 36 +++++++++++++++++++++++++++ 3 files changed, 50 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c9f6685fc..465fd0120 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime] [docs]** `pause` and `cancel` can stop a run whose sealed control evidence diverged. Both read evidence observe-only and divergence-tolerant, as `status` does, instead of requiring execution-grade evidence, which refused with `WORKFLOW_CONTROL_EVIDENCE_INVALID` after, for example, a rebuild changed the validator build identity or an operator hot-patched the published workflow, leaving the run executing with no way to stop it through Ultrafuzz. A confirmed cancellation still persists `canceled`, and a run whose published execution snapshot files changed is still refused (#674, #921). +- **[runtime] [docs]** `pause` and `cancel` can stop a run whose sealed control documents diverged, for example after a rebuild changed the validator build identity or an operator hot-patched the published workflow; both used to refuse with `WORKFLOW_CONTROL_EVIDENCE_INVALID` while the run kept executing. They now read evidence the way `status` does (observe-only, divergence-tolerant), still persist `canceled` only when the runner confirms it, and still refuse a run whose sealed execution files changed. A launch that is still publishing its execution snapshot is now reported to `status`, `pause` and `cancel` as incomplete launch preparation instead of as a run to abandon for a new run ID (#674, #921). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index f8cf4d22c..98b3fa735 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -1476,7 +1476,7 @@ function invalidLinkedWorkflowEvidence(metadataPath: string, error: unknown): Li /** * A pending state records incomplete launch preparation, not launcher liveness. * It also survives an interrupted launch. Unreadable state retains the strict - * missing-seal diagnostic. + * missing-seal or missing-journal diagnostic. */ function runStatusIsPreSubmission(layout: RunLayout): boolean { try { @@ -1511,6 +1511,18 @@ function missingLinkedWorkflowEvidenceDiagnostic( } const linkJournalPath = workflowRunLinkJournalPath(layout); if (pathIsMissing(linkJournalPath)) { + // Launch writes this journal only after publishing the execution snapshot, its longest step, so a + // lock-free reader such as `status`, `pause` or `cancel` can meet a launch still in progress here. + // Report it with the seal's pending code, which `status` already renders as an incomplete launch. + if (runStatusIsPreSubmission(layout)) { + return { + code: "WORKFLOW_CONTROL_SEAL_PENDING", + message: `run ${layout.runId} has incomplete launch preparation: its workflow-link journal has not been written. Launcher liveness is unknown. If the original launch is still active, wait for it to finish; otherwise inspect its error before retrying`, + severity: "warning", + source: "workflow", + path: linkJournalPath + }; + } return { code: "WORKFLOW_RUN_LINK_JOURNAL_MISSING", message: `run ${layout.runId} lacks the required authenticated workflow-link journal; it may predate authenticated lifecycle links or be incomplete and cannot be safely upgraded in place. Preserve its stored artifacts and start a new run with a new run ID`, diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index 18a4995cd..5c835aee5 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -25981,6 +25981,42 @@ test("launch observation rechecks a seal published while state remains pending", } }); +test("a launch that has sealed its controls but not written its link journal reads as incomplete, not legacy", async () => { + const project = tempProject(); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + writeSmallTopology(project); + const runId = "launch-link-pending"; + const env = fakeSmithersEnv(project); + const run = await startRun({ projectRoot: project, runId, env }); + assert.equal(run.ok, true, JSON.stringify(run.diagnostics)); + assert.ok(run.value); + const layout = layoutForRunRoot(run.value.run_root, runId); + // Launch publishes the execution snapshot after sealing its controls and before writing the link + // journal, with the run still pending. `pause`, `cancel` and `status` take no control lock, so they + // can read exactly this state while that publication runs. + writeRunState(layout, { ...readRunState(layout), status: "pending" }); + const journalPath = path.join(layout.root, "smithers", "workflow-run-link-journal.json"); + fs.unlinkSync(journalPath); + const commandLog = path.join(project, "smithers-commands.log"); + fs.writeFileSync(commandLog, "", "utf8"); + + for (const result of [ + await pauseRun({ projectRoot: project, runId, env }), + await cancelRun({ projectRoot: project, runId, env }) + ]) { + assert.equal(result.ok, false); + assert.equal(result.diagnostics[0]?.code, "WORKFLOW_CONTROL_SEAL_PENDING", JSON.stringify(result.diagnostics)); + assert.equal(result.diagnostics[0]?.path, journalPath); + assert.match(result.diagnostics[0]?.message ?? "", /Launcher liveness is unknown/u); + assert.doesNotMatch(result.diagnostics[0]?.message ?? "", /new run ID|cannot be safely upgraded/u); + } + assert.doesNotMatch(fs.readFileSync(commandLog, "utf8"), /^(?:pause|cancel) /mu); + const health = await getRunHealth({ projectRoot: project, runId, env }); + assert.equal(health.ok, true, JSON.stringify(health.diagnostics)); + assert.equal(health.value?.verdict, "launch-incomplete"); + assert.equal(readRunState(layout).status, "pending"); +}); + test("native continuation restores only authenticated snapshot governance before launch", async () => { const project = tempProject(); assert.equal(initProject({ projectRoot: project, force: true }).ok, true); From 777a841993e499fab28eaaefb7f8a026a05d172a Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:52:23 +0000 Subject: [PATCH 121/206] docs(runtime): say what pause and cancel still refuse, and when they act on the wrong run The cli.md paragraph said pause and cancel still refuse a run whose "published execution snapshot files changed", which contradicted its own hand-patched-published-workflow example: the generated workflow in the published snapshot is a control document whose divergence is tolerated. What is still refused is a change to the sealed execution files the control seal lists in that snapshot (run plan, prompts, agent adapters, runtime packages and dependencies), because the runner is started from them. It now also says what taking no lock costs: - issued while a launch is still preparing, pause and cancel fail without changing the run and should be retried once `ultrafuzz run` returns; - they do not reconcile a replay or fork that was interrupted while linking its new workflow run, so they act on the workflow run it replaced, or refuse. `ultrafuzz why` reconciles that link (it does so before checking the control documents, so also on a diverged run), so the docs tell the operator to run it first. The readLinkedWorkflowEvidence comment no longer claims pause and cancel report divergences as warnings; they only proceed. Refs #674, #921 Co-Authored-By: Claude Opus 5.5 --- docs/reference/cli.md | 15 +++++++++++++-- packages/runtime/src/start-run.ts | 7 ++++--- 2 files changed, 17 insertions(+), 5 deletions(-) diff --git a/docs/reference/cli.md b/docs/reference/cli.md index 5efd717db..4830ac307 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -484,8 +484,19 @@ diagnostic. workflow control lock. A run whose sealed control documents diverged, for example a hand-patched published workflow or a planned graph that no longer matches the current build's artifact contracts after a rebuild, can therefore -still be paused or cancelled; `status` reports the divergence. Like `status`, -they still refuse a run whose published execution snapshot files changed. +still be paused or cancelled; `status` reports the divergence. They start the +workflow runner from the run's published execution snapshot, so, like +`status`, they still refuse a run whose sealed execution files changed: the +files the control seal lists in that snapshot, such as the run plan, prompts, +agent adapters, and the runtime packages and their dependencies. + +Because they take no lock, `pause` and `cancel` issued while a launch is still +preparing fail without changing the run; retry once `ultrafuzz run` has +returned. They also do not reconcile a `replay` or `fork` that was interrupted +while linking its new workflow run, so they act on the workflow run it +replaced, or refuse. Run `ultrafuzz why ` before pausing or cancelling +such a run: it reconciles that link, even when it then reports a diverged +control document. `why` returns a deterministic diagnosis: a summary, the current node, and typed blockers with `kind`, `node_id`, `iteration`, `reason`, `unblocker`, diff --git a/packages/runtime/src/start-run.ts b/packages/runtime/src/start-run.ts index 98b3fa735..0449d0f05 100644 --- a/packages/runtime/src/start-run.ts +++ b/packages/runtime/src/start-run.ts @@ -1328,9 +1328,10 @@ export async function readLinkedWorkflowEvidence( throw new Error("run metadata workflow IDs do not exactly match the active workflow run"); } - // Observers, `pause` and `cancel` pass `tolerateControlDivergence` so a divergent control file - // downgrades to a reported warning instead of hiding a live run entirely, or leaving it unstoppable - // (issue #674). Execution callers omit it and keep failing closed. + // Observers, `pause` and `cancel` pass `tolerateControlDivergence` so a divergent control file is + // collected in `divergences` instead of failing the read: observers report it as a warning, and + // `pause` and `cancel` still reach the runner (issue #674). Execution callers omit it and keep + // failing closed. const tolerateDivergence = options.tolerateControlDivergence === true; const verifiedControl = verifyWorkflowControlSnapshot(resolvedProjectRoot, layout, { tolerateDivergence }); // A live document replaced under the completeness re-derivation is a transient race, not a From 0df3af510da3ce8ea2eeacd880ae0dcb48579a0b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 02:59:12 +0000 Subject: [PATCH 122/206] docs: keep the security CHANGELOG entry within the 400-character line limit markdownlint's MD013 in the external static analysis job allows 400 characters per line. The rewritten entry was 408; it is now 398 with the same claims. The job still fails on the file's pre-existing long lines. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2d8c496f1..4724a514a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[security]** The publication secret gate's pattern rules skip Anvil's public `test test … junk` mnemonic and the 10 keys `anvil` prints, and its JWT rule requires an `eyJ` header, not just three long dotted segments. A cloud run started before this change that allowlists a variable holding that mnemonic or such a dotted value must drop it from `ULTRAFUZZ_AGENT_ENV_ALLOWLIST` to resume, replay or fork. +- **[security]** The publication secret gate's patterns skip Anvil's public `test test … junk` mnemonic and the 10 keys `anvil` prints, and its JWT rule requires an `eyJ` header, not just three long dotted segments. A cloud run from before this change whose allowlisted variable holds that mnemonic or such a dotted value must drop it from `ULTRAFUZZ_AGENT_ENV_ALLOWLIST` to resume, replay or fork. - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From dde0c7b07aab639386785f5f214a0d0cf023ec49 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:16:03 +0000 Subject: [PATCH 123/206] fix(runtime): stop retrying deterministic admission failures and label timeouts by code Two defects from #1144. Retries. A consumer's dependency-admission failure re-reads the same producer bytes, so every retry fails the same way; each consumer still spent its whole retry budget and agent fallback chain. The admission entry points (preparation's assertTaskDependencyInputs and the agent's assertDependencyArtifactAdmissionCurrent rechecks) now throw with details.failureRetryable=false, which Smithers honors, and preparationStep keeps that flag when it rewraps the error. Everything else stays retryable, including the #672 engine-boundary TypeError. A real Smithers run of the template's preparationStep and helper pins this: the admission failure runs once and its agent is skipped, while a transient TypeError is retried and finishes. With the flag removed the same preparation runs three times and ends `stalled`. Timeout labels. errorLooksLikeTimeout ran /timeout|timed out|heartbeat/ over every string in the NodeFailed error, including the stack and causes. That is a failure taxonomy from free text, which #572 rules out, and since #1027 every JSON-validator preflight failure mentions "timeout", so a 2ms `spawnSync ultrafuzz ENOENT` was recorded as a timed-out provider interruption. A node now times out only on Smithers' typed deadline codes (TASK_TIMEOUT, TASK_HEARTBEAT_TIMEOUT, PROCESS_TIMEOUT, PROCESS_IDLE_TIMEOUT) or a TaskHeartbeatTimeout event. The agent-failure normalizer used to strip PROCESS_TIMEOUT and PROCESS_IDLE_TIMEOUT, which left the regex as the only label for real agent CLI deadlines, so it now keeps them. The #1144 analysis also proposed guarding generate()'s catch-path recheck. It is not included: resetTaskArtifactsForRetry admits dependencies before that try block on every first generation in a process, so the catch path never runs without an admission. Refs #1144 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- docs/config.md | 7 +- .../templates/smithers/workflows/workflow.tsx | 43 +++++- packages/runtime/src/workflow-sync.ts | 32 +++-- .../test/generated-workflow-verifier.test.ts | 18 ++- packages/runtime/test/runtime.test.ts | 32 ++++- ...ithers-dependency-skip.integration.test.ts | 126 ++++++++++++++---- 7 files changed, 204 insertions(+), 56 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 0b2766f60..0fcdfe9a5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime] [docs]** Agent retries now wait one minute, doubling up to Smithers' five-minute cap, instead of 1s and 2s, which spent a three-attempt budget in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can take; Smithers' three-identical-failures stall verdict no longer ends the planned chain early, so later same-agent attempts and `[retry].agents` fallback profiles run (#1084). +- **[runtime] [docs]** Agent retries now wait one minute, doubling up to Smithers' five-minute cap, instead of 1s and 2s, which spent a three-attempt budget in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can take; Smithers' three-identical-failures stall verdict no longer ends the planned chain early, so later same-agent attempts and `[retry].agents` fallback profiles run. A dependency-admission failure is no longer retried, and synchronization labels a node timed out only from Smithers' typed deadline codes, not from any error text that mentions a timeout or heartbeat (#1084, #1144). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/config.md b/docs/config.md index fbe9f9df0..af776ddb2 100644 --- a/docs/config.md +++ b/docs/config.md @@ -54,9 +54,10 @@ A retry waits one minute, then two, then four, and at most five minutes (the Smithers cap), and uses a fresh session and the same effective task prompt, including Smithers' safety contracts; Ultrafuzz does not inspect provider error text. The planned chain is the whole budget: repeated identical failures do not -end it before later attempts or fallback profiles run. The planned chain and -actual producer are recorded in the task manifest, attempt ledger, and final -report. +end it before later attempts or fallback profiles run. A failure to admit a +dependency's verified artifacts is not retried, because a retry would re-read +the same producer files. The planned chain and actual producer are recorded in +the task manifest, attempt ledger, and final report. Retry chains currently require local execution. Cloud planning accepts one effective attempt, and local fallback across different agent implementations diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index f24ef08b8..c70c0dcae 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -3346,7 +3346,10 @@ function artifactAwareAgent( "AGENT_CONFIG_INVALID", "AGENT_SESSION_LOST", "AGENT_CHECKPOINT_INVALID", - "TASK_ABORTED" + "TASK_ABORTED", + // Agent CLI deadlines. Run synchronization labels timeouts by code only. + "PROCESS_TIMEOUT", + "PROCESS_IDLE_TIMEOUT" ]); // #677: a routed-gateway HTTP 402 (provider credit exhausted) reaches this // normalizer as an anonymous CLI failure because the subprocess boundary @@ -3951,10 +3954,11 @@ function preparationStep(attemptId: string, step: string, run: () => T): T { try { return run(); } catch (error) { - throw new Error( + const wrapped = new Error( `prepare:${attemptId} failed at step ${step}: ${error instanceof Error ? error.message : String(error)}`, { cause: error } ); + throw isNonRetryableFailure(error) ? nonRetryableFailure(wrapped) : wrapped; } } @@ -5578,6 +5582,14 @@ function assertTaskInputs(task: (typeof taskSpecs)[number], workspaceRoot: strin } function assertTaskDependencyInputs(task: (typeof taskSpecs)[number]): void { + try { + admitTaskDependencyInputs(task); + } catch (error) { + throw nonRetryableFailure(error); + } +} + +function admitTaskDependencyInputs(task: (typeof taskSpecs)[number]): void { const existingAdmission = dependencyArtifactAdmissionsByTask.get(task.attemptId); if (existingAdmission !== undefined) { assertDependencyArtifactAdmissionCurrent(task, existingAdmission); @@ -5704,6 +5716,17 @@ function admittedDependencyArtifactDirs(task: (typeof taskSpecs)[number]): reado function assertDependencyArtifactAdmissionCurrent( task: (typeof taskSpecs)[number], expected: DependencyArtifactAdmission = dependencyArtifactAdmission(task) +): DependencyArtifactAdmission { + try { + return checkDependencyArtifactAdmissionCurrent(task, expected); + } catch (error) { + throw nonRetryableFailure(error); + } +} + +function checkDependencyArtifactAdmissionCurrent( + task: (typeof taskSpecs)[number], + expected: DependencyArtifactAdmission ): DependencyArtifactAdmission { if (expected.task !== task || dependencyArtifactAdmissionsByTask.get(task.attemptId) !== expected) { throw new Error(`artifact-contract failure: dependency admission identity changed ${task.attemptId}`); @@ -5749,6 +5772,22 @@ function assertDependencyArtifactAdmissionCurrent( return expected; } +/** + * A dependency admission failure is deterministic for its consumer: the + * producer's verified artifacts are missing, changed, or unsafe, and a retry + * re-reads the same bytes. Smithers fails an attempt whose error carries + * `details.failureRetryable: false` once, instead of spending the consumer's + * retry budget and agent fallback chain on it (#1144). + */ +function nonRetryableFailure(error: unknown): Error { + const failure = error instanceof Error ? error : new Error(String(error)); + return Object.assign(failure, { details: { failureRetryable: false } }); +} + +function isNonRetryableFailure(error: unknown): boolean { + return (error as { details?: { failureRetryable?: unknown } } | undefined)?.details?.failureRetryable === false; +} + function assertVerifiedDependency( task: (typeof taskSpecs)[number], dependency: string, diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index b8810c84e..8f8c4f1f2 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -5683,7 +5683,7 @@ function terminalOutcomeForEvent(event: WorkflowEvent): return { outcome: "succeeded" }; case "NodeFailed": { const failureMessage = errorText(event.payload.error); - return errorLooksLikeTimeout(event.payload.error) + return workflowErrorIsTimeout(event.payload.error) ? { outcome: "timed-out", failureCategory: "timeout", ...(failureMessage ? { failureMessage } : {}) } : { outcome: "failed", failureCategory: "executor-error", ...(failureMessage ? { failureMessage } : {}) }; } @@ -5929,8 +5929,7 @@ function evidenceFromEvents(events: WorkflowEvent[]): NodeWorkflowEvidence | und break; case "NodeFailed": { const error = errorText(payload.error); - const timedOut = - evidence?.timedOut === true || errorLooksLikeTimeout(payload.error) || errorLooksLikeTimeout(error); + const timedOut = evidence?.timedOut === true || workflowErrorIsTimeout(payload.error); evidence = { ...evidence, status: timedOut ? "timed-out" : "failed", @@ -6844,17 +6843,22 @@ function errorText(value: unknown): string | undefined { return stringField(value, "message") ?? stringField(value, "code") ?? stringField(value, "_tag"); } -function errorLooksLikeTimeout(value: unknown): boolean { - if (value === undefined) { - return false; - } - if (typeof value === "string") { - return /timeout|timed out|heartbeat/iu.test(value); - } - if (!isRecord(value)) { - return false; - } - return Object.values(value).some((entry) => errorLooksLikeTimeout(entry)); +/** + * Smithers reports each deadline it enforces with a typed code: the engine's + * task and heartbeat watchdogs, and the process driver's total and idle timers + * for agent CLIs. Message, stack, and cause text are not classification input: + * a validator preflight that failed in 2ms mentions "timeout" in its message, + * and node ids can contain the word (#1144). + */ +const WORKFLOW_TIMEOUT_ERROR_CODES: ReadonlySet = new Set([ + "TASK_TIMEOUT", + "TASK_HEARTBEAT_TIMEOUT", + "PROCESS_TIMEOUT", + "PROCESS_IDLE_TIMEOUT" +]); + +function workflowErrorIsTimeout(error: unknown): boolean { + return isRecord(error) && typeof error.code === "string" && WORKFLOW_TIMEOUT_ERROR_CODES.has(error.code); } /** diff --git a/packages/runtime/test/generated-workflow-verifier.test.ts b/packages/runtime/test/generated-workflow-verifier.test.ts index 839f7f256..6880eb0fb 100644 --- a/packages/runtime/test/generated-workflow-verifier.test.ts +++ b/packages/runtime/test/generated-workflow-verifier.test.ts @@ -6175,7 +6175,12 @@ test("generated optional admission rejects a present malformed marker before pub assert.throws( () => harness.prepareAndBuild(consumer, runRoot, { agentRef: "codex" }), - /verification marker is schema-invalid optional-producer/u + (error: Error & { details?: unknown }) => { + assert.match(error.message, /verification marker is schema-invalid optional-producer/u); + // A retry would re-read the same producer bytes, so Smithers must not spend the budget (#1144). + assert.deepEqual(error.details, { failureRetryable: false }); + return true; + } ); assert.equal(dependencyAuthenticationCount, 1, "a present optional marker must be authenticated"); assert.equal(harness.hasAdmission(consumer.attemptId), false, "failed authentication must not publish admission"); @@ -6619,7 +6624,11 @@ test("generated dependency admission retains one exact snapshot epoch and never }; assert.throws( () => harness.current(consumer), - /dependency authority changed after admission patch-producer/u, + (error: Error & { details?: unknown }) => { + assert.match(error.message, /dependency authority changed after admission patch-producer/u); + assert.deepEqual(error.details, { failureRetryable: false }, "the agent's admission recheck is deterministic"); + return true; + }, "an identical-byte marker replacement must not inherit the admitted identity" ); @@ -8113,6 +8122,11 @@ test("agent failure normalization preserves only validated Smithers recovery con assert.equal(normalized.code, code); assert.deepEqual(normalized.details, details); } + // Agent CLI deadlines keep their code: run synchronization labels timeouts by code alone (#1144). + for (const code of ["PROCESS_TIMEOUT", "PROCESS_IDLE_TIMEOUT"]) { + const deadline = await captureAgentFailure(Object.assign(new Error("CLI timed out after 1800000ms"), { code })); + assert.equal(deadline.code, code); + } const abort = new Error("operation aborted") as Error & { code: string }; abort.name = "AbortError"; diff --git a/packages/runtime/test/runtime.test.ts b/packages/runtime/test/runtime.test.ts index eceaa6746..4d57c3e19 100644 --- a/packages/runtime/test/runtime.test.ts +++ b/packages/runtime/test/runtime.test.ts @@ -22178,13 +22178,28 @@ test("syncRun maps failed workflow nodes into durable failed run state", async ( "utf8" ); const workflowRunId = "ultrafuzz-sync-failed-node"; + // Only Smithers' typed deadline codes make a timeout. The second failure is + // the #1144 report: 2ms old, yet its text mentions a timeout. + const heartbeatTimeout = "Task node:project-discovery has not heartbeated in 1800250ms (timeout: 1800000ms)."; + const timeoutWordedFailure = + "artifact-contract failure: JSON validator preflight failed after 2ms against a 180000ms budget (only the launcher is signalled on timeout, and its grandchild CLI holds the inherited pipes open, so elapsed can overrun the budget): spawnSync ultrafuzz ENOENT"; const failedEvents = [ { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 1 }, { type: "TaskHeartbeatTimeout", nodeId: "node:project-discovery", attempt: 1 }, - { type: "NodeFailed", nodeId: "node:project-discovery", attempt: 1, error: { message: "agent failed" } }, + { + type: "NodeFailed", + nodeId: "node:project-discovery", + attempt: 1, + error: { code: "TASK_HEARTBEAT_TIMEOUT", message: heartbeatTimeout } + }, { type: "NodeRetrying", nodeId: "node:project-discovery", attempt: 2 }, { type: "NodeStarted", nodeId: "node:project-discovery", attempt: 2 }, - { type: "NodeFailed", nodeId: "node:project-discovery", attempt: 2, error: { message: "agent failed again" } } + { + type: "NodeFailed", + nodeId: "node:project-discovery", + attempt: 2, + error: { message: timeoutWordedFailure } + } ]; const env = fakeLifecycleSmithersEnv(project, { inspect: workflowInspect({ @@ -22223,7 +22238,7 @@ test("syncRun maps failed workflow nodes into durable failed run state", async ( }; assert.equal(state.nodes?.["project-discovery"]?.status, "failed"); assert.equal(state.nodes?.["project-discovery"]?.retry_count, 1); - assert.equal(state.nodes?.["project-discovery"]?.last_error, "agent failed again"); + assert.equal(state.nodes?.["project-discovery"]?.last_error, timeoutWordedFailure); assert.deepEqual(state.nodes?.["project-discovery"]?.provenance?.failure, { category: "agent-failure", causal_task_id: "node:project-discovery", @@ -22240,12 +22255,15 @@ test("syncRun maps failed workflow nodes into durable failed run state", async ( .split("\n") .map((line) => JSON.parse(line) as Record); assert.deepEqual( - failedLedger.map((entry) => entry.outcome), - ["failed", "failed"] + failedLedger.map((entry) => [entry.outcome, entry.failure_category]), + [ + ["timed-out", "timeout"], + ["failed", "executor-error"] + ] ); assert.deepEqual( failedLedger.map((entry) => entry.failure_message), - ["agent failed", "agent failed again"] + [heartbeatTimeout, timeoutWordedFailure] ); assert.deepEqual( failedLedger.map((entry) => entry.agent), @@ -22740,7 +22758,7 @@ test("syncRun keeps reset workflow nodes pending while the workflow is running", type: "NodeFailed", nodeId: "node:project-discovery", attempt: 1, - error: { message: "CLI timed out after 1800000ms" } + error: { code: "PROCESS_TIMEOUT", message: "CLI timed out after 1800000ms" } } ]) }); diff --git a/packages/runtime/test/smithers-dependency-skip.integration.test.ts b/packages/runtime/test/smithers-dependency-skip.integration.test.ts index 85a32cee1..5392e5b5e 100644 --- a/packages/runtime/test/smithers-dependency-skip.integration.test.ts +++ b/packages/runtime/test/smithers-dependency-skip.integration.test.ts @@ -14,16 +14,55 @@ interface Inspection { } test("failed and stalled required inputs skip dependent attempts while independent work and review finish", async () => { - const root = temporaryRoot("ufz-skip-required-"); + const { root, state } = await runSyntheticWorkflow("required-inputs", syntheticWorkflowSource); + assert.equal(state("verify:producer"), "failed"); + assert.equal(state("prepare:dependent"), "skipped"); + assert.equal(state("node:dependent"), "skipped"); + assert.equal(state("verify:dependent"), "skipped"); + assert.equal(state("prepare:own-failure"), "stalled"); + assert.equal(state("node:own-failure"), "skipped"); + assert.equal(state("verify:own-failure"), "skipped"); + assert.equal(state("independent"), "finished"); + assert.equal(state("review"), "finished"); + assert.deepEqual(fs.readFileSync(path.join(root, "executed.log"), "utf8").trim().split("\n"), [ + "independent", + "review" + ]); +}); + +test("a dependency admission failure fails its preparation once while a transient failure is retried", async () => { + const { root, state } = await runSyntheticWorkflow("non-retryable", nonRetryableWorkflowSource); + // #1144: the default three-attempt budget used to re-read the same producer + // bytes three times and then stall; the attempt now fails once. + assert.equal(state("prepare:consumer"), "failed"); + assert.equal(state("node:consumer"), "skipped"); + // #672: an engine-boundary TypeError still spends the preparation retry. + assert.equal(state("prepare:transient"), "finished"); + assert.deepEqual(fs.readFileSync(path.join(root, "executed.log"), "utf8").trim().split("\n").sort(), [ + "consumer", + "transient", + "transient" + ]); +}); + +async function runSyntheticWorkflow( + name: string, + workflowSource: (root: string, template: string) => string +): Promise<{ root: string; state: (nodeId: string) => string | undefined }> { + const root = temporaryRoot(`ufz-${name}-`); const packageRoot = runtimePackageRoot(); const smithers = path.join(packageRoot, "node_modules", ".bin", "smithers"); - const runId = `skip-required-${process.pid}-${Date.now()}`; - const workflow = path.join(root, ".smithers", "workflows", "required-inputs.tsx"); + const runId = `${name}-${process.pid}-${Date.now()}`; + const workflow = path.join(root, ".smithers", "workflows", `${name}.tsx`); fs.mkdirSync(path.dirname(workflow), { recursive: true }); const dependencyRoot = path.dirname(fs.realpathSync(path.join(packageRoot, "node_modules", "smthrs"))); fs.symlinkSync(dependencyRoot, path.join(root, ".smithers", "node_modules"), "dir"); execFileSync("git", ["init", "--quiet", "--initial-branch=main"], { cwd: root }); - fs.writeFileSync(workflow, syntheticWorkflowSource(root, packageRoot)); + const template = fs.readFileSync( + path.join(packageRoot, "src", "templates", "smithers", "workflows", "workflow.tsx"), + "utf8" + ); + fs.writeFileSync(workflow, workflowSource(root, template)); execFileSync( smithers, ["up", workflow, "--detach", "--run-id", runId, "--root", root, "--input", "{}", "--format", "json"], @@ -41,35 +80,68 @@ test("failed and stalled required inputs skip dependent attempts while independe await new Promise((resolve) => setTimeout(resolve, 100)); } assert.equal(inspected.status ?? inspected.run?.status, "finished", JSON.stringify(inspected)); - const state = (nodeId: string) => inspected.steps?.find((step) => (step.id ?? step.nodeId) === nodeId)?.state; - assert.equal(state("verify:producer"), "failed"); - assert.equal(state("prepare:dependent"), "skipped"); - assert.equal(state("node:dependent"), "skipped"); - assert.equal(state("verify:dependent"), "skipped"); - assert.equal(state("prepare:own-failure"), "stalled"); - assert.equal(state("node:own-failure"), "skipped"); - assert.equal(state("verify:own-failure"), "skipped"); - assert.equal(state("independent"), "finished"); - assert.equal(state("review"), "finished"); - assert.deepEqual(fs.readFileSync(path.join(root, "executed.log"), "utf8").trim().split("\n"), [ - "independent", - "review" - ]); + return { + root, + state: (nodeId) => inspected.steps?.find((step) => (step.id ?? step.nodeId) === nodeId)?.state + }; +} + +function templateSlice(template: string, startMarker: string, endMarker: string): string { + const start = template.indexOf(startMarker); + const end = template.indexOf(endMarker, start); + assert.ok(start >= 0 && end > start, `${startMarker} is missing from the workflow template`); + return template.slice(start, end); +} + +function nonRetryableWorkflowSource(root: string, template: string): string { + return `/** @jsxImportSource smthrs */ +import fs from "node:fs"; +import { createSmithers } from "smthrs"; +import { z } from "zod/v4"; +${templateSlice(template, "type WorkflowTaskStateContext =", "type DependencyVerificationProducer =")} +${templateSlice(template, "function preparationStep", "\n\nfunction prepareArtifactMirror")} +${templateSlice(template, "function nonRetryableFailure", "\n\nfunction assertVerifiedDependency")} +const evidence = ${JSON.stringify(path.join(root, "executed.log"))}; +const { Workflow, Parallel, Task, smithers, outputs } = createSmithers({ + input: z.object({}), + result: z.object({ value: z.string() }) +}); +const executions = (value: string) => + fs.existsSync(evidence) ? fs.readFileSync(evidence, "utf8").split("\\n").filter((line) => line === value).length : 0; +const record = (value: string) => { fs.appendFileSync(evidence, value + "\\n"); return { value }; }; +export default smithers((ctx) => { + const preparation = failedWorkflowPrerequisites(ctx, ["prepare:consumer"]); + return + + {() => preparationStep("consumer", "assert-task-inputs", () => { + record("consumer"); + throw nonRetryableFailure(new Error("artifact-contract failure: artifact dependency has not passed verification producer")); + })} + + + {() => record("forbidden-consumer")} + + + {() => preparationStep("transient", "resolve-workspace-root", () => { + if (executions("transient") === 0) { + record("transient"); + throw new TypeError("undefined is not an object (evaluating 'get')"); + } + return record("transient"); + })} + + ; }); +`; +} -function syntheticWorkflowSource(root: string, packageRoot: string): string { - const source = fs.readFileSync( - path.join(packageRoot, "src", "templates", "smithers", "workflows", "workflow.tsx"), - "utf8" - ); - const start = source.indexOf("type WorkflowTaskStateContext ="); - const end = source.indexOf("type DependencyVerificationProducer =", start); - assert.ok(start >= 0 && end > start); +function syntheticWorkflowSource(root: string, template: string): string { return `/** @jsxImportSource smthrs */ import fs from "node:fs"; import { createSmithers } from "smthrs"; import { z } from "zod/v4"; -${source.slice(start, end)} +${templateSlice(template, "type WorkflowTaskStateContext =", "type DependencyVerificationProducer =")} const evidence = ${JSON.stringify(path.join(root, "executed.log"))}; const { Workflow, Parallel, Task, smithers, outputs } = createSmithers({ input: z.object({}), From 52cf3390eadfabd5617cda53899a81bd776ada9b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:01:15 +0000 Subject: [PATCH 124/206] docs: correct the OpenCode profile claim and scope the guardrail advice Review follow-up for the OpenRouter guardrails section. - The section said the OpenCode openrouter/ model came from "the shipped profile". No OpenCode model profile ships: packages/config/defaults.toml and ultrafuzz.toml carry [agents.OpenCodeAgent] but no [models.opencode]. Drop the parenthetical, and correct the OpenCode agent section's "The default root config includes an opt-in OpenCode profile", the claim it repeated. The doctor paragraph's copy of the claim is left to #1179, which rewrites that paragraph. - OpenRouter's scan_scope for the prompt-injection builtin defaults to all_messages and can be set to user_only, so say that every message is scanned by default rather than always. - Say that the workspace default and member guardrails also cover other keys, that a workspace used only for the Ultrafuzz key confines the workspace-default change to that key, and that in an organization account only an organization admin can change guardrails. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- docs/config.md | 21 ++++++++++++++------- 2 files changed, 15 insertions(+), 8 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 195754c2f..9bae88f54 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[docs]** `docs/config.md` now explains the HTTP 403 `Request blocked: prompt injection patterns detected` that `OpenRouterAgent`, `OpenCodeAgent`, and `PiAgent` get when an OpenRouter guardrail on their key sets prompt-injection detection to Block: set it to Flag, or turn it off, on every guardrail that covers the key, and do not use Redact, which forwards altered content. Runtime behavior is unchanged; the rejection is an ordinary agent failure under the `[retry]` policy (#1149). +- **[docs]** `docs/config.md` now explains the HTTP 403 `Request blocked: prompt injection patterns detected` that `OpenRouterAgent`, `OpenCodeAgent`, and `PiAgent` get when an OpenRouter guardrail covering their key sets prompt-injection detection to Block: set it to Flag, or turn it off, on every guardrail that covers the key, and do not use Redact, which forwards altered content. Runtime behavior is unchanged; the rejection is an ordinary agent failure under the `[retry]` policy. The OpenCode section no longer claims the default config ships an OpenCode model profile (#1149). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/config.md b/docs/config.md index 1387a1fc2..bd236fcb4 100644 --- a/docs/config.md +++ b/docs/config.md @@ -260,7 +260,8 @@ million cache-hit input tokens, and $0.87 per million output tokens. `ultrafuzz init` also generates a dedicated `OpenCodeAgent`. It is opt-in and non-default: nothing selects it until a topology group or `--agent` names it. -The default root config includes an opt-in OpenCode profile: +The default root config has the `[agents.OpenCodeAgent]` block below but no +OpenCode model profile, so add one such as `[models.opencode]` to use it: ```toml [models.opencode] @@ -379,17 +380,17 @@ forwards that final level without degrading it. ## OpenRouter guardrails `OpenRouterAgent`, `PiAgent`, and `OpenCodeAgent` with an `openrouter/` model -(as in the shipped profile) send their requests through OpenRouter with the key -in `OPENROUTER_API_KEY`. If any guardrail covering that key sets +send their requests through OpenRouter with the key in `OPENROUTER_API_KEY`. If +any guardrail covering that key sets [prompt-injection detection](https://openrouter.ai/docs/guides/features/guardrails/prompt-injection) to **Block**, OpenRouter rejects each request its detector matches with HTTP 403 `Request blocked: prompt injection patterns detected` before it reaches a model. -The match need not be in Ultrafuzz's task prompt. OpenRouter scans every -message in a request, including base64- and hex-decoded text, and these -harnesses also send their own system prompts and the target source and test -output the agent reads. For example, the default system prompt of OpenCode +The match need not be in Ultrafuzz's task prompt. These harnesses also send +their own system prompts and the target source and test output the agent reads, +and by default OpenRouter scans every message in a request, including base64- +and hex-decoded text. For example, the default system prompt of OpenCode 1.18.18, used for models without a model-specific prompt such as DeepSeek, Qwen, or GLM, has an `assistant: [...]` line followed by a `user:` line, which matches OpenRouter's documented `role_delimiter_injection` pattern. @@ -402,6 +403,12 @@ apply. Do not use **Redact** either: it replaces each match with `[PROMPT_INJECTION]` and forwards the request, so the model can work from altered source or tool output with no error for Ultrafuzz to report. +The workspace default covers every key in its workspace and a member guardrail +every key of that member, so relaxing either can affect more than Ultrafuzz. +Creating the Ultrafuzz key in a workspace of its own confines the +workspace-default change to that key. In an organization account, only an +organization admin can change guardrails. + Ultrafuzz has no special handling for this rejection: the attempt fails like any other agent error and follows the `[retry]` policy above. A retry on the same profile sends the same task prompt with the same key, so it is rejected again From af55c862b9b49dc28114451ecb4ad4b5f33b6ff3 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:10:47 +0000 Subject: [PATCH 125/206] perf(runtime): cut per-process and per-render cost of the generated workflow agent-registry.ts and smithers.ts imported the TypeScript compiler at module load, and both are reachable from @ultrafuzz/runtime's index. Only registry analysis (validate and init) and the controller-refresh helper that only tests call ever use it. Require it at first use instead. The saving lands in processes that import the runtime without smthrs: the ultrafuzz CLI, including every agent `ultrafuzz json validate` call. Smithers engine processes still load the compiler: the generated workflow imports smthrs, whose index re-exports @smthrs/scorers, and two scorer modules import typescript at module load. Importing the built runtime index alone on this host went from 528-547 ms / 248-250 MB RSS to 409-437 ms / 196-198 MB under Node 24, and from 449-483 ms / 252-261 MB to 336-351 ms / 209-213 MB under Bun 1.3.14 (5 runs each). Changing the parameter type on topLevelNameIsBound's signature makes the strict diff lint report that function's existing complexity, so its import-clause check moves into a helper with the same conditions. materializeDynamicRuntime runs on every render and rewrote tasks.json and graph.json, with fsyncs, even when it re-derived the bytes already on disk. Skip the write in that case. taskSpecsFromCompiled looked every task up with serializedTaskSpecs.find; index the specs by id once per call instead. At the default topology's 65 static tasks the difference is not measurable. Add a test that compiles the packaged default topology and holds the generated workflow under 2.3 MB, measured with a fixed-length project root (about 2.06-2.07 MB today, depending on the checkout path), so the size #1146 describes cannot grow silently. Refs #1146 Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/agent-registry.ts | 69 ++++++++-------- packages/runtime/src/dynamic-runtime.ts | 17 ++-- packages/runtime/src/smithers.ts | 6 +- .../templates/smithers/workflows/workflow.tsx | 3 +- .../runtime/test/dynamic-expansion.test.ts | 79 ++++++++++++++++++ .../test/generated-workflow-footprint.test.ts | 81 +++++++++++++++++++ 6 files changed, 212 insertions(+), 43 deletions(-) create mode 100644 packages/runtime/test/generated-workflow-footprint.test.ts diff --git a/packages/runtime/src/agent-registry.ts b/packages/runtime/src/agent-registry.ts index d1d655818..b04b6f20f 100644 --- a/packages/runtime/src/agent-registry.ts +++ b/packages/runtime/src/agent-registry.ts @@ -1,9 +1,14 @@ +import { createRequire } from "node:module"; import path from "node:path"; import { readSinglyLinkedRegularFileSnapshotInside } from "@ultrafuzz/artifacts"; -import * as ts from "typescript"; +import type * as TypeScript from "typescript"; import { errorMessage } from "@ultrafuzz/artifacts"; +// Required by analyzeAgentRegistry rather than imported: every Smithers engine process imports this +// package's index through the generated workflow, and only `validate` and `init` analyze a registry. +let ts: typeof TypeScript; + export const AGENT_REGISTRY_RELATIVE_PATH = ".smithers/agents/index.ts"; const MAX_AGENT_REGISTRY_BYTES = 256 * 1024; const MAX_ANALYSIS_DEPTH = 64; @@ -50,6 +55,7 @@ export function agentRegistryRegisters(inspection: AgentRegistryInspection, agen const SAFE_AGENT_REF_PATTERN = /^(?!.*\.\.)[A-Za-z_][A-Za-z0-9_.:-]{0,127}$/u; function analyzeAgentRegistry(sourceText: string): ReadonlySet { + ts = createRequire(import.meta.url)("typescript") as typeof TypeScript; const source = ts.createSourceFile( AGENT_REGISTRY_RELATIVE_PATH, sourceText, @@ -57,7 +63,8 @@ function analyzeAgentRegistry(sourceText: string): ReadonlySet { false, ts.ScriptKind.TS ); - const parseDiagnostics = (source as ts.SourceFile & { parseDiagnostics: readonly ts.Diagnostic[] }).parseDiagnostics; + const parseDiagnostics = (source as TypeScript.SourceFile & { parseDiagnostics: readonly TypeScript.Diagnostic[] }) + .parseDiagnostics; if (parseDiagnostics.length > 0) throw new Error("agent registry contains invalid TypeScript syntax"); const exported = exportedAgentFactoriesLocalNames(source); if (exported.size !== 1) { @@ -65,10 +72,10 @@ function analyzeAgentRegistry(sourceText: string): ReadonlySet { } const bindings = topLevelConstBindings(source); const objectIsShadowed = topLevelNameIsBound(source, "Object"); - const memo = new Map | null>(); + const memo = new Map | null>(); let steps = 0; - const evaluate = (expression: ts.Expression, depth: number): ReadonlyMap | undefined => { + const evaluate = (expression: TypeScript.Expression, depth: number): ReadonlyMap | undefined => { steps += 1; if (steps > MAX_ANALYSIS_STEPS) throw new Error("agent registry static analysis exceeded its step budget"); if (depth > MAX_ANALYSIS_DEPTH) throw new Error("agent registry static analysis exceeded its depth budget"); @@ -127,14 +134,14 @@ function analyzeAgentRegistry(sourceText: string): ReadonlySet { } function objectMemberIsUsableFactory( - property: ts.ObjectLiteralElementLike, - bindings: ReadonlyMap, + property: TypeScript.ObjectLiteralElementLike, + bindings: ReadonlyMap, visiting: ReadonlySet, depth: number ): boolean { if (depth > MAX_ANALYSIS_DEPTH) throw new Error("agent registry factory analysis exceeded its depth budget"); if (ts.isMethodDeclaration(property)) return true; - let expression: ts.Expression | undefined; + let expression: TypeScript.Expression | undefined; if (ts.isPropertyAssignment(property)) expression = property.initializer; else if (ts.isShorthandPropertyAssignment(property)) expression = property.objectAssignmentInitializer ?? property.name; @@ -149,8 +156,8 @@ function objectMemberIsUsableFactory( } function factoryExpressionIsNonNullish( - expression: ts.Expression, - bindings: ReadonlyMap, + expression: TypeScript.Expression, + bindings: ReadonlyMap, visiting: ReadonlySet, depth: number ): boolean { @@ -164,7 +171,7 @@ function factoryExpressionIsNonNullish( return factoryExpressionIsNonNullish(initializer, bindings, new Set([...visiting, current.text]), depth + 1); } -function exportedAgentFactoriesLocalNames(source: ts.SourceFile): ReadonlySet { +function exportedAgentFactoriesLocalNames(source: TypeScript.SourceFile): ReadonlySet { const localNames = new Set(); for (const statement of source.statements) { if (ts.isVariableStatement(statement) && hasExportModifier(statement)) { @@ -191,8 +198,8 @@ function exportedAgentFactoriesLocalNames(source: ts.SourceFile): ReadonlySet { - const bindings = new Map(); +function topLevelConstBindings(source: TypeScript.SourceFile): ReadonlyMap { + const bindings = new Map(); for (const statement of source.statements) { if (!ts.isVariableStatement(statement) || (statement.declarationList.flags & ts.NodeFlags.Const) === 0) continue; for (const declaration of statement.declarationList.declarations) { @@ -203,25 +210,10 @@ function topLevelConstBindings(source: ts.SourceFile): ReadonlyMap entry.name.text === name) - ) - return true; - } + if (ts.isImportDeclaration(statement) && importClauseBindsName(statement.importClause, name)) return true; if ( ts.isVariableStatement(statement) && statement.declarationList.declarations.some((declaration) => bindingNameContains(declaration.name, name)) @@ -235,7 +227,14 @@ function topLevelNameIsBound(source: ts.SourceFile, name: string): boolean { return false; } -function bindingNameContains(binding: ts.BindingName, name: string): boolean { +function importClauseBindsName(clause: TypeScript.ImportClause | undefined, name: string): boolean { + if (clause?.name?.text === name) return true; + const bindings = clause?.namedBindings; + if (bindings !== undefined && ts.isNamespaceImport(bindings)) return bindings.name.text === name; + return bindings !== undefined && bindings.elements.some((entry) => entry.name.text === name); +} + +function bindingNameContains(binding: TypeScript.BindingName, name: string): boolean { if (ts.isIdentifier(binding)) return binding.text === name; return binding.elements.some( (element) => !ts.isOmittedExpression(element) && bindingNameContains(element.name, name) @@ -243,9 +242,9 @@ function bindingNameContains(binding: ts.BindingName, name: string): boolean { } function isUnshadowedObjectFreeze( - expression: ts.Expression, + expression: TypeScript.Expression, objectIsShadowed: boolean -): expression is ts.CallExpression { +): expression is TypeScript.CallExpression { return ( !objectIsShadowed && ts.isCallExpression(expression) && @@ -257,18 +256,18 @@ function isUnshadowedObjectFreeze( ); } -function hasExportModifier(node: ts.VariableStatement): boolean { +function hasExportModifier(node: TypeScript.VariableStatement): boolean { return ts.getModifiers(node)?.some((modifier) => modifier.kind === ts.SyntaxKind.ExportKeyword) ?? false; } -function unwrapTypeExpressions(expression: ts.Expression): ts.Expression { +function unwrapTypeExpressions(expression: TypeScript.Expression): TypeScript.Expression { let current = expression; while (ts.isAsExpression(current) || ts.isSatisfiesExpression(current) || ts.isParenthesizedExpression(current)) current = current.expression; return current; } -function objectMemberName(property: ts.ObjectLiteralElementLike): string | undefined { +function objectMemberName(property: TypeScript.ObjectLiteralElementLike): string | undefined { if ( !ts.isPropertyAssignment(property) && !ts.isShorthandPropertyAssignment(property) && diff --git a/packages/runtime/src/dynamic-runtime.ts b/packages/runtime/src/dynamic-runtime.ts index c71cbc8ae..9ac3b2315 100644 --- a/packages/runtime/src/dynamic-runtime.ts +++ b/packages/runtime/src/dynamic-runtime.ts @@ -12,8 +12,7 @@ import { sha256Bytes, validateNodeReference, validateSafeId, - writeFileDurable, - writeJsonDurable + writeFileDurable } from "@ultrafuzz/artifacts"; import { renderPrompt, type PromptConcreteNode, type PromptGraphNode } from "@ultrafuzz/prompts"; @@ -183,9 +182,11 @@ function deriveDynamicRuntime( if (mode === "publish") { // Publish tasks first. A concurrent synchronizer may temporarily skip an // unknown graph node, while the reverse ordering could finalize a graph node - // against stale task identity. Both files are themselves atomically replaced. - writeJsonDurable(tasksPath, runtimeTaskDocument); - writeJsonDurable(graphPath, runtimeGraph); + // against stale task identity. Both files are themselves atomically replaced, + // and only when their bytes change: this runs on every render, and most + // renders re-derive exactly the documents already on disk. + writeJsonDurableIfChanged(tasksPath, runtimeTaskDocument); + writeJsonDurableIfChanged(graphPath, runtimeGraph); } else { const observedTasks = readRecord(tasksPath); const observedGraph = readPlannedGraph(graphPath); @@ -638,6 +639,12 @@ function readPlannedGraph(graphPath: string): PlannedGraph { return value as PlannedGraph; } +/** Same bytes as `writeJsonDurable`, skipping the replace and fsyncs when the file already holds them. */ +function writeJsonDurableIfChanged(filePath: string, value: unknown): void { + const bytes = `${JSON.stringify(value, null, 2)}\n`; + if (fs.readFileSync(filePath, "utf8") !== bytes) writeFileDurable(filePath, bytes); +} + function readRecord(filePath: string): Record { const value = JSON.parse(fs.readFileSync(filePath, "utf8")) as unknown; if (typeof value !== "object" || value === null || Array.isArray(value)) { diff --git a/packages/runtime/src/smithers.ts b/packages/runtime/src/smithers.ts index 51c96f889..4add29265 100644 --- a/packages/runtime/src/smithers.ts +++ b/packages/runtime/src/smithers.ts @@ -1,7 +1,7 @@ import { execFile, spawn } from "node:child_process"; import crypto from "node:crypto"; import fs from "node:fs"; -import { isBuiltin } from "node:module"; +import { createRequire, isBuiltin } from "node:module"; import os from "node:os"; import path from "node:path"; import { createInterface } from "node:readline"; @@ -60,7 +60,7 @@ import { type ExpandedNode, type ModelFanoutProvenance } from "@ultrafuzz/topology"; -import * as ts from "typescript"; +import type * as TypeScript from "typescript"; import { DATA_GOVERNANCE_PROVENANCE_PATH, @@ -3746,6 +3746,8 @@ function assertRefreshedModuleAuthority( } const dependencies = new Set(Object.keys(issuers[0]!.dependencies)); const executablePaths = new Set(dependencyMap.executable_paths); + // Required here rather than imported: the generated workflow imports this module in every engine process. + const ts = createRequire(import.meta.url)("typescript") as typeof TypeScript; for (const file of files) { if (file.executable !== executablePaths.has(file.snapshotPath)) { throw new Error(`controller module ${moduleName} changed executable authority for ${file.snapshotPath}`); diff --git a/packages/runtime/src/templates/smithers/workflows/workflow.tsx b/packages/runtime/src/templates/smithers/workflows/workflow.tsx index 44d9634f8..87e562dd9 100644 --- a/packages/runtime/src/templates/smithers/workflows/workflow.tsx +++ b/packages/runtime/src/templates/smithers/workflows/workflow.tsx @@ -478,9 +478,10 @@ function compiledTaskSourceIdentity(task: (typeof compiledBaseTasks)[number]) { } function taskSpecsFromCompiled(tasks: typeof compiledBaseTasks) { + const compiledById = new Map(serializedTaskSpecs.map((candidate) => [candidate.id, candidate])); return tasks.map((task) => { const controlPaths = taskWorkflowControlPaths(task.execution.mode, admittedWorkflowControls); - const compiled = serializedTaskSpecs.find((candidate) => candidate.id === task.smithersNodeId); + const compiled = compiledById.get(task.smithersNodeId); const runtimePromptPath = task.renderedPromptPath === undefined ? undefined diff --git a/packages/runtime/test/dynamic-expansion.test.ts b/packages/runtime/test/dynamic-expansion.test.ts index 0d9190a11..43034850d 100644 --- a/packages/runtime/test/dynamic-expansion.test.ts +++ b/packages/runtime/test/dynamic-expansion.test.ts @@ -681,6 +681,85 @@ test("100 generated attempts remain queued under the ordinary concurrency projec assert.equal(Object.values(projection.state.nodes).filter((node) => node.wait_reason === "capacity").length, 96); }); +test("re-rendering an unchanged dynamic runtime does not replace its published task plan or graph", () => { + const runId = "unchanged-publication"; + const projectRoot = tempDirectory(); + const runRoot = path.join(projectRoot, "runs", runId); + const sourceArtifactPath = path.join(runRoot, "artifacts", "planner", "plan.json"); + const templatePath = path.join(runRoot, "templates", "worker.md"); + const graphPath = path.join(runRoot, "graph.json"); + const tasksPath = path.join(runRoot, "smithers", "tasks.json"); + const baseGraphPath = path.join(runRoot, "smithers", "runtime-base-graph.json"); + const baseTasksPath = path.join(runRoot, "smithers", "runtime-base-tasks.json"); + fs.mkdirSync(path.dirname(sourceArtifactPath), { recursive: true }); + fs.mkdirSync(path.dirname(templatePath), { recursive: true }); + fs.mkdirSync(path.dirname(tasksPath), { recursive: true }); + fs.writeFileSync(sourceArtifactPath, `${JSON.stringify({ goals: [item(0), item(1)] })}\n`, "utf8"); + fs.writeFileSync(templatePath, "Investigate {{item.goal_prompt}}.\n", "utf8"); + const templateTask = compiledTask(projectRoot, runRoot, "fanout", "fanout", templatePath); + const joinTask = compiledTask(projectRoot, runRoot, "join", "join", undefined, ["fanout"]); + const group: CompiledSmithersDynamicGroup = { + groupNodeId: "fanout", + logicalNodeId: "fanout", + source: { + concreteNodeId: "planner", + attemptId: "planner", + verifierSmithersNodeId: "verify:planner", + artifactPath: sourceArtifactPath + }, + sourcePath: "$.goals", + keyPath: "id", + nodeIdTemplate: "dynamic:item:{{ item.id }}", + templatePath, + templateDigest: digest(fs.readFileSync(templatePath)), + templateFingerprint: digest("template-fingerprint"), + continueOnFail: true, + maxDynamicNodes: 100, + reservedNodeIds: ["planner", "fanout", "join"], + taskTemplates: [templateTask], + promptContext: promptContext(projectRoot, runRoot) + }; + fs.writeFileSync(graphPath, `${JSON.stringify(plannedGraph(runId))}\n`, "utf8"); + fs.writeFileSync( + tasksPath, + `${JSON.stringify({ schema_version: "1.0", run_id: runId, tasks: [joinTask], dynamic_groups: [group] })}\n`, + "utf8" + ); + fs.copyFileSync(graphPath, baseGraphPath); + fs.copyFileSync(tasksPath, baseTasksPath); + const controls = { + runId, + projectRoot, + runRoot, + graphPath, + tasksPath, + baseGraphPath, + baseTasksPath, + baseTasks: [joinTask], + groups: [group], + readyGroupIds: ["fanout"] + }; + // A durable write replaces the file through a rename, so an unchanged inode means no write. + const snapshot = (filePath: string) => ({ + ino: fs.statSync(filePath, { bigint: true }).ino, + bytes: fs.readFileSync(filePath, "utf8") + }); + const published = () => ({ tasks: snapshot(tasksPath), graph: snapshot(graphPath) }); + + const seeded = published(); + materializeDynamicRuntime(controls); + const expanded = published(); + assert.notEqual(expanded.tasks.bytes, seeded.tasks.bytes, "the first expansion must publish its generated tasks"); + assert.notEqual(expanded.graph.bytes, seeded.graph.bytes, "the first expansion must publish its generated nodes"); + + // Smithers calls this on every render and resume. Compare after each call: a replaced file frees + // its old inode, which a later replacement could reuse. + materializeDynamicRuntime(controls); + assert.deepEqual(published(), expanded); + materializeDynamicRuntime({ ...controls, readyGroupIds: [] }); + assert.deepEqual(published(), expanded); +}); + function plannedGraph(_runId: string): PlannedGraph { const base = { display_name: "Node", diff --git a/packages/runtime/test/generated-workflow-footprint.test.ts b/packages/runtime/test/generated-workflow-footprint.test.ts new file mode 100644 index 000000000..e7799f5ee --- /dev/null +++ b/packages/runtime/test/generated-workflow-footprint.test.ts @@ -0,0 +1,81 @@ +import assert from "node:assert/strict"; +import { execFileSync } from "node:child_process"; +import fs from "node:fs"; +import path from "node:path"; +import test from "node:test"; + +import { loadReferenceCatalog } from "@ultrafuzz/references"; + +import { compileSmithersWorkflow, initProject, planRun } from "../src/index.js"; +import { writeShippedDocumentReferenceCaches, writeShippedVulnerabilityDatabaseCache } from "./reference-fixtures.js"; +import { temporaryRoot } from "./temporary-root.js"; + +/** + * Every Smithers engine process parses and transpiles the whole generated workflow (#1146). With a + * fixed-length project root the packaged default topology compiles to 2,070,382 bytes, 1.59 MB of it + * the same 65 tasks serialized twice (as compiled tasks and as task specs). Raise this only + * deliberately: #1146 asks for the file to shrink. + */ +const GENERATED_DEFAULT_WORKFLOW_BUDGET_BYTES = 2_300_000; + +test("the packaged default topology compiles to a generated workflow under its byte budget", async () => { + const project = temporaryRoot("ufz-footprint-"); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const xdgCacheHome = path.join(project, "xdg-cache"); + writeShippedDocumentReferenceCaches(xdgCacheHome, loadReferenceCatalog(project)); + writeShippedVulnerabilityDatabaseCache(xdgCacheHome); + const previousXdgCacheHome = process.env.XDG_CACHE_HOME; + process.env.XDG_CACHE_HOME = xdgCacheHome; + let plan: Awaited>; + try { + plan = await planRun({ projectRoot: project, runId: "footprint", env: {} }); + } finally { + if (previousXdgCacheHome === undefined) delete process.env.XDG_CACHE_HOME; + else process.env.XDG_CACHE_HOME = previousXdgCacheHome; + } + assert.ok(plan.ok && plan.value, JSON.stringify(plan.diagnostics)); + const compiled = compileSmithersWorkflow({ + projectRoot: project, + config: plan.value.resolved_config, + graph: plan.value.expanded_graph, + runLayout: plan.value.layout, + workflowName: "ultrafuzz-footprint", + renderedPrompts: plan.value.rendered_prompts + }); + + // The absolute project root appears thousands of times, so measure with a fixed-length root. + // Otherwise the budget would depend on how long the machine's temporary directory path is. + const source = fs.readFileSync(compiled.workflowPath, "utf8"); + const bytes = Buffer.byteLength(source.replaceAll(project, "/project")); + assert.ok( + bytes <= GENERATED_DEFAULT_WORKFLOW_BUDGET_BYTES, + `generated default workflow is ${bytes} bytes, over its ${GENERATED_DEFAULT_WORKFLOW_BUDGET_BYTES}-byte budget` + ); +}); + +test("importing the runtime package does not load the TypeScript compiler", () => { + // The generated workflow imports this package's index in every engine process. Run the import in a + // fresh process so this test file's own TypeScript import cannot mask it. + const project = temporaryRoot("ufz-footprint-registry-"); + assert.equal(initProject({ projectRoot: project, force: true }).ok, true); + const runtimeIndex = new URL("../src/index.js", import.meta.url).href; + const agentRegistry = new URL("../src/agent-registry.js", import.meta.url).href; + const probe = ` + import { createRequire } from "node:module"; + await import(${JSON.stringify(runtimeIndex)}); + const require = createRequire(${JSON.stringify(runtimeIndex)}); + const compiler = require.resolve("typescript"); + const afterImport = compiler in require.cache; + const registry = await import(${JSON.stringify(agentRegistry)}); + const inspection = registry.inspectAgentRegistry(${JSON.stringify(project)}); + console.log(JSON.stringify({ + afterImport, + afterInspection: compiler in require.cache, + registersCodex: registry.agentRegistryRegisters(inspection, "CodexAgent") + })); + `; + const observed = JSON.parse( + execFileSync(process.execPath, ["--input-type=module", "--eval", probe], { encoding: "utf8" }) + ) as { afterImport: boolean; afterInspection: boolean; registersCodex: boolean }; + assert.deepEqual(observed, { afterImport: false, afterInspection: true, registersCodex: true }); +}); From 75c4f2b96804064ebe7572c4f344736f3b47712b Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:02:11 +0000 Subject: [PATCH 126/206] docs: state what still ends the retry chain, and that text-only deadlines are labelled failed Review of this change found the docs, CHANGELOG and one comment stronger than the code: - docs/config.md said the planned chain is the whole budget. Smithers still stops a chain at a failure it classifies as non-retryable, such as a CLI auth or configuration error, and pauses the run on a quota limit. - The admission wrappers mark every failure inside admission non-retryable, including a file-system or validator error while reading the producer files, not only missing or changed bytes. - Timeout labelling also keeps the TaskHeartbeatTimeout event rule, and a deadline reported only as text, such as the Modal provider's cloud-node deadline, is now labelled failed. docs/reference/artifacts-reports.md states the rule next to the node statuses. - The workflow-sync comment read as if it listed every Smithers deadline code. Refs #1144 Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- docs/config.md | 13 ++++++++----- docs/reference/artifacts-reports.md | 6 ++++++ packages/runtime/src/workflow-sync.ts | 12 +++++++----- 4 files changed, 22 insertions(+), 11 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 0fcdfe9a5..df9050e72 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime] [docs]** Agent retries now wait one minute, doubling up to Smithers' five-minute cap, instead of 1s and 2s, which spent a three-attempt budget in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can take; Smithers' three-identical-failures stall verdict no longer ends the planned chain early, so later same-agent attempts and `[retry].agents` fallback profiles run. A dependency-admission failure is no longer retried, and synchronization labels a node timed out only from Smithers' typed deadline codes, not from any error text that mentions a timeout or heartbeat (#1084, #1144). +- **[runtime] [docs]** Agent retries now wait one minute, doubling up to Smithers' five-minute cap, instead of 1s and 2s, which spent a three-attempt budget in about 25 seconds, inside the minute a contended Claude Code OAuth refresh can take; Smithers' three-identical-failures stall verdict no longer ends the planned chain early, so later same-agent attempts and `[retry].agents` fallback profiles run. A dependency-admission failure is no longer retried, including a file-system or validator error while admission reads the producer's files. Synchronization labels a node timed out only from Smithers' typed deadline codes and heartbeat-timeout events, not from error text that mentions a timeout or heartbeat, so a deadline reported only as text, such as a Modal cloud-node deadline, is now labelled failed (#1084, #1144). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). diff --git a/docs/config.md b/docs/config.md index af776ddb2..bb14c2bf1 100644 --- a/docs/config.md +++ b/docs/config.md @@ -53,11 +53,14 @@ attempts. Omitting `agents`, or leaving it empty, keeps model fallback disabled. A retry waits one minute, then two, then four, and at most five minutes (the Smithers cap), and uses a fresh session and the same effective task prompt, including Smithers' safety contracts; Ultrafuzz does not inspect provider error -text. The planned chain is the whole budget: repeated identical failures do not -end it before later attempts or fallback profiles run. A failure to admit a -dependency's verified artifacts is not retried, because a retry would re-read -the same producer files. The planned chain and actual producer are recorded in -the task manifest, attempt ledger, and final report. +text. Repeated identical failures do not end the planned chain early; Smithers +still stops it at a failure it classifies as non-retryable, such as a CLI +configuration or authentication error, and pauses the run on a provider quota +limit. Dependency admission is not retried: it re-reads the same producer files, +so an admission failure, including a file-system or validator error while +reading them, fails the task without a retry or fallback. The planned chain and +actual producer are recorded in the task manifest, attempt ledger, and final +report. Retry chains currently require local execution. Cloud planning accepts one effective attempt, and local fallback across different agent implementations diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index 2683bb70e..593f0df7b 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -173,6 +173,12 @@ Node statuses are: - `reused-from-prior-run` - `invalidated` +Synchronization records a node as `timed-out` from Smithers' typed deadline +codes (`TASK_TIMEOUT`, `TASK_HEARTBEAT_TIMEOUT`, `PROCESS_TIMEOUT`, +`PROCESS_IDLE_TIMEOUT`) and heartbeat-timeout events, not from error text: a +failure whose message mentions a timeout, or a deadline reported only as text +such as a Modal cloud-node deadline, is `failed`. + Every nonterminal node records `wait_since`, a typed `wait_reason`, and a typed `next_eligible_action`. Wait reasons distinguish ready work, capacity and dependency waits, retry backoff, external gates, controller loss, and active diff --git a/packages/runtime/src/workflow-sync.ts b/packages/runtime/src/workflow-sync.ts index 8f8c4f1f2..e08c38b39 100644 --- a/packages/runtime/src/workflow-sync.ts +++ b/packages/runtime/src/workflow-sync.ts @@ -6844,11 +6844,13 @@ function errorText(value: unknown): string | undefined { } /** - * Smithers reports each deadline it enforces with a typed code: the engine's - * task and heartbeat watchdogs, and the process driver's total and idle timers - * for agent CLIs. Message, stack, and cause text are not classification input: - * a validator preflight that failed in 2ms mentions "timeout" in its message, - * and node ids can contain the word (#1144). + * Smithers reports its deadlines that can fail an Ultrafuzz task with typed + * codes: the engine's task and heartbeat watchdogs, and the process driver's + * total and idle timers for agent CLIs. Message, stack, and cause text are not + * classification input: a validator preflight that failed in 2ms mentions + * "timeout" in its message, and node ids can contain the word (#1144). A + * deadline reported only as text, such as the Modal provider's cloud-node + * deadline, is therefore labelled failed. */ const WORKFLOW_TIMEOUT_ERROR_CODES: ReadonlySet = new Set([ "TASK_TIMEOUT", From b03a83a16ee04c4c741611997cc800a7ab6bb435 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Mon, 28 Sep 2026 23:10:47 +0000 Subject: [PATCH 127/206] docs: changelog for render robustness and runtime footprint Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 1 + 1 file changed, 1 insertion(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index 960ba9394..d3b900dc9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,6 +12,7 @@ ### Other changes +- **[runtime] [artifacts] [docs]** A literal `{{word}}` in a task or operator prompt no longer fails every workflow render: the agent-prompt renderer now checks only its own template's placeholders. `ultrafuzz/goal-plan@1` now rejects `{{` in replacement values, so a label the dynamic renderer cannot bind fails `goal-plan`'s own verify instead of every render of the run. Each engine process runs the JSON validator CLI preflight once instead of in every prepare, attempt reset and verify, no longer loads the TypeScript compiler when it imports `@ultrafuzz/runtime`, and skips rewriting unchanged dynamic `tasks.json` and `graph.json` on each render; a new test holds the generated default-topology workflow under 2.3 MB (#1146). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 668ad911144a07d51fe078bc6bdc7cbf2b9a2c5d Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:02:11 +0000 Subject: [PATCH 128/206] ci: disable markdownlint MD013 for CHANGELOG.md Every CHANGELOG entry is one long line, so super-linter's markdownlint fails MD013 (line length 400) on any pull request that touches the file. release-gates then fails and every full release-validation lane is skipped, so the runtime suite never runs on the pull request. This is the same directive #1179 adds, byte for byte, so whichever lands first merges cleanly with the other. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index df9050e72..589e681e4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -127,3 +127,5 @@ ## v0.0.1 - First external private release. + + From c71b8671c530b51657e5701d8e5ef1606acb8b17 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:07:13 +0000 Subject: [PATCH 129/206] test(runtime): make the campaign join tests fail only for the rule they name Three rewritten campaign-join tests ("no finding covers", "unrelated finding", and the dropped-property half of "ambiguous and mismatched same-ID findings") and the property-combination test still used findings with no partition metadata. property-campaign-context-joins then reports every such finding as having no contributions and every failure as unclaimed, so each assertion held whether or not the fixture broke the rule in the test's title. Delete the three redundant tests. Partition-less findings are already covered by "requires partition metadata and accepts an explicit complete partition", the finding-ID rule by "requires each finding ID to represent one of its contributions", a dropped property by "accepts non-property findings and validates property-derived failures", and duplicate finding IDs by the host document-gate test. Convert the combination test to accounted findings for failure-1 and failure-2: it now requires exactly one join issue, for failure-3, and passes once failure-3 is claimed. Also pin, at host level, that a campaign result omitting the optional sequence_length is accepted. On origin/main this test fails with CAMPAIGN_TIMEOUT_EVIDENCE_INVALID and CAMPAIGN_SEQUENCE_LENGTH_MISMATCH from the deleted host copy; the only earlier assertion was registry-level and passed on main too. Drop the run-state mutation from the fan-in lens test, which the sealed-authority gate never reads, and retitle it for the catalog omission it actually tests. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/test/artifact-gates.test.ts | 166 ++++--------------- 1 file changed, 32 insertions(+), 134 deletions(-) diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 52becf417..d7308eecf 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -6820,8 +6820,8 @@ test("property fan-in selects a declared lens contract without relying on the pr assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); }); -test("property fan-in cannot hide a planned lens by omitting its state declaration and catalog rows", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-planned-lens-state-omission" }); +test("property fan-in cannot hide a planned lens by omitting its catalog rows", () => { + const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-planned-lens-catalog-omission" }); const node = writeMinimalPropertyFaninFixture(layout, { sourceNodeId: "project-discovery", sourcePropertyId: "evidence-1", @@ -6843,10 +6843,6 @@ test("property fan-in cannot hide a planned lens by omitting its state declarati ] }) ); - // Run state is not a declaration source: the planned lens stays bound. - const state = readRunState(layout); - delete state.nodes["property-specification-recon"]!.outputs; - writeRunState(layout, state); const result = verifyRequiredArtifactsForAttempt(layout, node, node.id); @@ -10245,7 +10241,7 @@ interface CampaignTimeoutResultFixture extends Record { schema_version: "ultrafuzz.property-campaign.v3"; fuzzer_backend: "recon"; configured_timeout_seconds: number; - sequence_length: number; + sequence_length?: number; exact_command: string; start_timestamp: string; end_timestamp: string; @@ -10745,6 +10741,15 @@ test("current campaign timeout gate accepts ordinary shell framing around the re } }); +// The result schema makes sequence_length optional and the campaign prompt does +// not ask for it; the deleted host copy required it after the budget was spent. +test("current campaign timeout gate accepts a result that omits the optional sequence_length", () => { + const result = runCampaignTimeoutGate((fixture) => { + delete fixture.backend.sequence_length; + }); + assert.equal(result.ok, true, JSON.stringify(result.diagnostics)); +}); + test("current campaign timeout gate names only the missing flag when the command also redirects output", () => { const result = runCampaignTimeoutGate((fixture) => withCampaignCommand( @@ -11460,124 +11465,6 @@ test("campaign gate conditionally reconciles the R55 summary failure counts with assert.deepEqual(fs.readFileSync(incompletePath), incompleteBytes); }); -test("campaign gate still rejects a property-derived failure no finding covers", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-campaign-uncovered" }); - campaignPropertyCatalog(layout, ["property-1", "property-2"]); - const campaignId = "stateful-invariant-campaign"; - writeArtifact( - layout, - campaignId, - "recon-fuzzer-results.json", - JSON.stringify( - currentCampaign( - ["property-1", "property-2"], - [ - { id: "failure-1", property_ids: ["property-1"] }, - { id: "failure-2", property_ids: ["property-2"] } - ] - ) - ) - ); - writeArtifact(layout, campaignId, "findings.json", JSON.stringify([campaignFinding("failure-1", ["property-1"])])); - writeCampaignSummary(layout, campaignId, 2, 1); - const node = { - ...currentCampaignNode(["recon-fuzzer-results.json", "findings.json"]), - id: campaignId, - logical_id: campaignId - }; - - const result = verifyRequiredArtifactsForAttempt(layout, node, campaignId); - assert.equal(result.ok, false); - assert.ok( - campaignJoinIssues(result).some((entry) => entry.includes('"failure-2" must be claimed by exactly one finding')), - JSON.stringify(result.diagnostics) - ); -}); - -test("campaign gate does not let an unrelated finding cover a campaign failure", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-campaign-unrelated-coverage" }); - campaignPropertyCatalog(layout, ["property-1"]); - const campaignId = "stateful-invariant-campaign"; - writeArtifact( - layout, - campaignId, - "recon-fuzzer-results.json", - JSON.stringify(currentCampaign(["property-1"], [{ id: "failure-1", property_ids: ["property-1"] }])) - ); - writeArtifact( - layout, - campaignId, - "findings.json", - JSON.stringify([campaignFinding("unrelated-finding", ["property-1"])]) - ); - writeCampaignSummary(layout, campaignId, 1, 1); - const node = { - ...currentCampaignNode(["recon-fuzzer-results.json", "findings.json"]), - id: campaignId, - logical_id: campaignId - }; - - const result = verifyRequiredArtifactsForAttempt(layout, node, campaignId); - assert.equal(result.ok, false); - assert.ok( - campaignJoinIssues(result).some((entry) => entry.includes('"failure-1" must be claimed by exactly one finding')), - JSON.stringify(result.diagnostics) - ); -}); - -test("campaign gate keeps flagging ambiguous and mismatched same-ID findings", () => { - const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-campaign-ambiguous" }); - campaignPropertyCatalog(layout, ["property-1"]); - const campaignId = "stateful-invariant-campaign"; - writeArtifact( - layout, - campaignId, - "recon-fuzzer-results.json", - JSON.stringify( - currentCampaign( - ["property-1"], - [ - { id: "failure-1", property_ids: ["property-1"] }, - { id: "failure-2", property_ids: ["property-1"] } - ] - ) - ) - ); - writeArtifact( - layout, - campaignId, - "findings.json", - JSON.stringify([campaignFinding("failure-1", ["property-1"]), campaignFinding("failure-1", ["property-1"])]) - ); - writeCampaignSummary(layout, campaignId, 2, 2); - const node = { - ...currentCampaignNode(["recon-fuzzer-results.json", "findings.json"]), - id: campaignId, - logical_id: campaignId - }; - const ambiguous = verifyRequiredArtifactsForAttempt(layout, node, campaignId); - assert.equal(ambiguous.ok, false); - assert.deepEqual(gateIssuePaths(ambiguous, "findings-id-uniqueness"), ["$[1]"]); - - // A finding that claims a failure's ID must still carry that failure's properties, - // even though other failures may now be covered by a different finding. - writeArtifact( - layout, - campaignId, - "findings.json", - JSON.stringify([campaignFinding("failure-1", []), campaignFinding("failure-2", ["property-1"])]) - ); - writeCampaignSummary(layout, campaignId, 2, 2); - const mismatched = verifyRequiredArtifactsForAttempt(layout, node, campaignId); - assert.equal(mismatched.ok, false); - assert.ok( - campaignJoinIssues(mismatched).some((entry) => - entry.includes('"failure-1" must be claimed by exactly one finding') - ), - JSON.stringify(mismatched.diagnostics) - ); -}); - test("campaign gate rejects a failure whose property combination no single finding claims", () => { const layout = createRunLayout({ projectRoot: tempProject(), runId: "run-campaign-combination" }); campaignPropertyCatalog(layout, ["property-1", "property-2"]); @@ -11599,12 +11486,11 @@ test("campaign gate rejects a failure whose property combination no single findi ) ) ); - writeArtifact( - layout, - campaignId, - "findings.json", - JSON.stringify([campaignFinding("failure-1", ["property-1"]), campaignFinding("failure-2", ["property-2"])]) - ); + const singlePropertyFindings = [ + accountedCampaignFinding("failure-1", ["property-1"], ["failure-1"]), + accountedCampaignFinding("failure-2", ["property-2"], ["failure-2"]) + ]; + writeArtifact(layout, campaignId, "findings.json", JSON.stringify(singlePropertyFindings)); writeCampaignSummary(layout, campaignId, 3, 2); const node = { ...currentCampaignNode(["recon-fuzzer-results.json", "findings.json"]), @@ -11614,10 +11500,22 @@ test("campaign gate rejects a failure whose property combination no single findi const result = verifyRequiredArtifactsForAttempt(layout, node, campaignId); assert.equal(result.ok, false); - assert.ok( - campaignJoinIssues(result).some((entry) => entry.includes('"failure-3" must be claimed by exactly one finding')), - JSON.stringify(result.diagnostics) + assert.deepEqual(campaignJoinIssues(result), [ + '$.failures: Semantic gate property-campaign-context-joins failed: Property-derived campaign failure "failure-3" must be claimed by exactly one finding' + ]); + + writeArtifact( + layout, + campaignId, + "findings.json", + JSON.stringify([ + ...singlePropertyFindings, + accountedCampaignFinding("failure-3", ["property-1", "property-2"], ["failure-3"]) + ]) ); + writeCampaignSummary(layout, campaignId, 3, 3); + const claimed = verifyRequiredArtifactsForAttempt(layout, node, campaignId); + assert.equal(claimed.ok, true, JSON.stringify(claimed.diagnostics)); }); test("campaign gate accepts a deduplicated finding that unions the properties of the failures it covers", () => { From b375f43da87ebeaa7401e867c77e40df06563a11 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:07:14 +0000 Subject: [PATCH 130/206] test(artifacts): share one campaign timeout evidence fixture The sequence-length registry test copied the 45-line plan and result fixture from the adjacent forged-deadline test. Both now build it with one helper that returns the gate's issues and asserts the gate had its context, so a missing context cannot read as a pass. Co-Authored-By: Claude Opus 5.5 --- .../artifacts/test/semantic-gates.test.ts | 166 +++++++----------- 1 file changed, 65 insertions(+), 101 deletions(-) diff --git a/packages/artifacts/test/semantic-gates.test.ts b/packages/artifacts/test/semantic-gates.test.ts index 2d487271e..d190e3a49 100644 --- a/packages/artifacts/test/semantic-gates.test.ts +++ b/packages/artifacts/test/semantic-gates.test.ts @@ -6446,31 +6446,43 @@ test("campaign evidence-file closure names duplicated, missing, and unreferenced ); }); -test("campaign timeout evidence reports the expected final artifact deadline next to a forged execution deadline", () => { - const command = - "timeout --preserve-status --signal=INT --kill-after=300s 60s recon fuzz . --workers 1 " + - "--timeout 60 --test-limit 18446744073709551615 --seq-len 100"; - const plan = { - configured_fuzzer_timeout_seconds: 60, - recon_internal_timeout_seconds: 60, - host_soft_timeout_seconds: 60, - host_force_kill_grace_seconds: 300, - artifact_finalization_reserve_seconds: 100, - finalization_reserve_seconds: 100, - configured_budget_seconds: 460, - recon_test_limit: "18446744073709551615", - recon_sequence_length: 100, - backend_started_at: "2026-01-01T00:00:00.000Z", - fuzzing_deadline_utc: "2026-01-01T00:01:00.000Z", - force_kill_deadline_utc: "2026-01-01T00:06:00.000Z", - final_artifact_deadline_utc: "2026-01-01T00:07:40.000Z", - deadline: "2026-01-01T00:07:40.000Z", - backend: { exact_shell_escaped_command: command }, - command_plan: [{ phase: "campaign", command }] - }; - const result = executeSemanticGate("property-campaign-timeout-evidence", { +const CAMPAIGN_TIMEOUT_EVIDENCE_COMMAND = + "timeout --preserve-status --signal=INT --kill-after=300s 60s recon fuzz . --workers 1 " + + "--timeout 60 --test-limit 18446744073709551615 --seq-len 100"; + +interface CampaignTimeoutEvidenceFields { + plan: Record; + document: Record; + summary: Record; +} + +/** property-campaign-timeout-evidence issues for a consistent 60-second Recon campaign after `mutate`. */ +function campaignTimeoutEvidenceIssues( + mutate: (fields: CampaignTimeoutEvidenceFields) => void = () => undefined +): { path: string; message: string }[] { + const command = CAMPAIGN_TIMEOUT_EVIDENCE_COMMAND; + const fields: CampaignTimeoutEvidenceFields = { + plan: { + configured_fuzzer_timeout_seconds: 60, + recon_internal_timeout_seconds: 60, + host_soft_timeout_seconds: 60, + host_force_kill_grace_seconds: 300, + artifact_finalization_reserve_seconds: 100, + finalization_reserve_seconds: 100, + configured_budget_seconds: 460, + recon_test_limit: "18446744073709551615", + recon_sequence_length: 100, + backend_started_at: "2026-01-01T00:00:00.000Z", + fuzzing_deadline_utc: "2026-01-01T00:01:00.000Z", + force_kill_deadline_utc: "2026-01-01T00:06:00.000Z", + final_artifact_deadline_utc: "2026-01-01T00:07:40.000Z", + deadline: "2026-01-01T00:07:40.000Z", + backend: { exact_shell_escaped_command: command }, + command_plan: [{ phase: "campaign", command }] + }, document: { configured_timeout_seconds: 60, + sequence_length: 100, exact_command: command, start_timestamp: "2026-01-01T00:00:00.000Z", end_timestamp: "2026-01-01T00:01:00.000Z", @@ -6482,13 +6494,16 @@ test("campaign timeout evidence reports the expected final artifact deadline nex usable_results: true, started_at: "2026-01-01T00:00:00.000Z", finished_at: "2026-01-01T00:01:00.000Z", - // The incident's wrong-field bug: copied from the plan's - // fuzzing_deadline_utc instead of final_artifact_deadline_utc. - deadline: "2026-01-01T00:01:00.000Z" + deadline: "2026-01-01T00:07:40.000Z" } }, + summary: { outcome: "complete", sequence_length: 100 } + }; + mutate(fields); + const result = executeSemanticGate("property-campaign-timeout-evidence", { + document: fields.document, context: { - artifactSet: { campaignPlan: plan, campaignSummary: { outcome: "complete", sequence_length: 100 } }, + artifactSet: { campaignPlan: fields.plan, campaignSummary: fields.summary }, propertyCampaignTimeout: { configuredFuzzerTimeoutSeconds: 60, plannedTimeoutSeconds: 600, @@ -6496,11 +6511,17 @@ test("campaign timeout evidence reports the expected final artifact deadline nex } } }); + assert.notEqual(result.status, "requires-context"); + return result.status === "failed" ? result.issues.map((entry) => ({ path: entry.path, message: entry.message })) : []; +} - assert.equal(result.status, "failed"); - const issues = result.status === "failed" ? result.issues : []; +test("campaign timeout evidence reports the expected final artifact deadline next to a forged execution deadline", () => { assert.deepEqual( - issues.map((entry) => ({ path: entry.path, message: entry.message })), + campaignTimeoutEvidenceIssues(({ document }) => { + // The incident's wrong-field bug: copied from the plan's + // fuzzing_deadline_utc instead of final_artifact_deadline_utc. + (document.execution as Record).deadline = "2026-01-01T00:01:00.000Z"; + }), [ { path: "$.execution.deadline", @@ -6513,106 +6534,49 @@ test("campaign timeout evidence reports the expected final artifact deadline nex }); test("campaign timeout evidence requires the stateful Recon sequence length wherever it is recorded", () => { - const command = - "timeout --preserve-status --signal=INT --kill-after=300s 60s recon fuzz . --workers 1 " + - "--timeout 60 --test-limit 18446744073709551615 --seq-len 100"; - const gate = ( - mutate: (fields: { - plan: Record; - document: Record; - summary: Record; - }) => void = () => undefined - ) => { - const plan: Record = { - configured_fuzzer_timeout_seconds: 60, - recon_internal_timeout_seconds: 60, - host_soft_timeout_seconds: 60, - host_force_kill_grace_seconds: 300, - artifact_finalization_reserve_seconds: 100, - finalization_reserve_seconds: 100, - configured_budget_seconds: 460, - recon_test_limit: "18446744073709551615", - recon_sequence_length: 100, - backend_started_at: "2026-01-01T00:00:00.000Z", - fuzzing_deadline_utc: "2026-01-01T00:01:00.000Z", - force_kill_deadline_utc: "2026-01-01T00:06:00.000Z", - final_artifact_deadline_utc: "2026-01-01T00:07:40.000Z", - deadline: "2026-01-01T00:07:40.000Z", - backend: { exact_shell_escaped_command: command }, - command_plan: [{ phase: "campaign", command }] - }; - const document: Record = { - configured_timeout_seconds: 60, - sequence_length: 100, - exact_command: command, - start_timestamp: "2026-01-01T00:00:00.000Z", - end_timestamp: "2026-01-01T00:01:00.000Z", - termination_reason: "configured-timeout", - campaign_outcome: "complete", - usable_results: true, - execution: { - command, - usable_results: true, - started_at: "2026-01-01T00:00:00.000Z", - finished_at: "2026-01-01T00:01:00.000Z", - deadline: "2026-01-01T00:07:40.000Z" - } - }; - const summary: Record = { outcome: "complete", sequence_length: 100 }; - mutate({ plan, document, summary }); - const result = executeSemanticGate("property-campaign-timeout-evidence", { - document, - context: { - artifactSet: { campaignPlan: plan, campaignSummary: summary }, - propertyCampaignTimeout: { - configuredFuzzerTimeoutSeconds: 60, - plannedTimeoutSeconds: 600, - finalizationReserveSeconds: 100 - } - } - }); - return result.status === "failed" ? result.issues.map((entry) => entry.path) : []; - }; - const withCommand = (fields: { plan: Record; document: Record }, next: string) => { - fields.plan.backend = { exact_shell_escaped_command: next }; - fields.plan.command_plan = [{ phase: "campaign", command: next }]; - fields.document.exact_command = next; - (fields.document.execution as Record).command = next; + const issuePaths = (mutate?: (fields: CampaignTimeoutEvidenceFields) => void) => + campaignTimeoutEvidenceIssues(mutate).map((entry) => entry.path); + const withCommand = ({ plan, document }: CampaignTimeoutEvidenceFields, next: string) => { + plan.backend = { exact_shell_escaped_command: next }; + plan.command_plan = [{ phase: "campaign", command: next }]; + document.exact_command = next; + (document.execution as Record).command = next; }; - assert.deepEqual(gate(), []); + assert.deepEqual(issuePaths(), []); // The result record's sequence_length is optional in its schema. assert.deepEqual( - gate(({ document }) => { + issuePaths(({ document }) => { delete document.sequence_length; }), [] ); assert.deepEqual( - gate(({ plan }) => { + issuePaths(({ plan }) => { plan.recon_sequence_length = 1; }), ["$.campaign_plan_ref#recon_sequence_length"] ); assert.deepEqual( - gate(({ summary }) => { + issuePaths(({ summary }) => { summary.sequence_length = 1; }), ["$.campaign_summary_ref#sequence_length"] ); assert.deepEqual( - gate(({ document }) => { + issuePaths(({ document }) => { document.sequence_length = 1; }), ["$.sequence_length"] ); + const command = CAMPAIGN_TIMEOUT_EVIDENCE_COMMAND; for (const next of [ command.replace("--seq-len 100", "--seq-len 1"), command.replace(" --seq-len 100", ""), `${command} --seq-len 100` ]) { assert.deepEqual( - gate((fields) => withCommand(fields, next)), + issuePaths((fields) => withCommand(fields, next)), ["$.exact_command"], next ); From e3ae14504fc80a27de6ed331f3ac163e1a25aecb Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:07:14 +0000 Subject: [PATCH 131/206] refactor(artifacts): make the finding backend provenance parser module-private findingFuzzerBackendProvenance was exported for the host campaign partition joins deleted in this branch. Its only remaining caller is resolveCampaignFindingBackends in the same module, and no generated workflow template imports it. property-provenance.ts is not one of the modules hashed into the validator build identity. Co-Authored-By: Claude Opus 5.5 --- packages/artifacts/src/property-provenance.ts | 12 +++--------- 1 file changed, 3 insertions(+), 9 deletions(-) diff --git a/packages/artifacts/src/property-provenance.ts b/packages/artifacts/src/property-provenance.ts index 25aaf7954..571d0c200 100644 --- a/packages/artifacts/src/property-provenance.ts +++ b/packages/artifacts/src/property-provenance.ts @@ -275,19 +275,13 @@ export interface PropertyCampaignArtifact { failures: PropertyCampaignFailure[]; } -export type FindingFuzzerBackendProvenance = +type FindingFuzzerBackendProvenance = | { present: false; valid: true; backends: readonly [] } | { present: true; valid: false; backends: readonly [] } | { present: true; valid: true; backends: readonly string[] }; -/** - * Read backend provenance owned by a deduplicated campaign finding. Keeping - * this parser beside the campaign artifact contract gives the runtime gate and - * final-report verification one interpretation of the singular/plural fields. - */ -export function findingFuzzerBackendProvenance( - finding: Readonly> -): FindingFuzzerBackendProvenance { +/** Read backend provenance owned by a deduplicated campaign finding: its singular or plural field. */ +function findingFuzzerBackendProvenance(finding: Readonly>): FindingFuzzerBackendProvenance { const hasBackend = Object.prototype.hasOwnProperty.call(finding, "fuzzer_backend"); const hasBackends = Object.prototype.hasOwnProperty.call(finding, "fuzzer_backends"); if (!hasBackend && !hasBackends) { From cf00e20c337cbed944744924b429e69373e07a10 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:07:15 +0000 Subject: [PATCH 132/206] docs(changelog): scope the host gate de-duplication entry The entry opened with "no longer runs its own copies of registry gates", which the code contradicts: for example, final-report property references are still checked on the host as well as by report-property-provenance-join. Name only the three re-implementations this branch deletes and the three host-only rejections it removes, including the omitted result sequence_length, and keep the line under the 400-character markdownlint limit. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 212d923b2..24ce5eca9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ ### Other changes -- **[runtime] [artifacts]** Host artifact finalization no longer runs its own copies of registry gates. The host re-implementations of campaign timeout evidence, the campaign failure-partition joins, and invariant-ledger checks that the ledger schema already enforces are deleted, so for these checks the host runs only the shared `property-campaign-timeout-evidence` and `property-campaign-context-joins` gates that the in-attempt verifier also runs; a recorded Recon command framed as `cd && …`, `… | tee ` or `…; echo …`, or a property failure in a campaign declared at a nested path, is no longer rejected after the verifier passed. The `--seq-len 100` / `sequence_length` rules moved into the shared gate and now fail inside the attempt, `verifyRequiredArtifactsForAttempt` requires the sealed attempt authority every production caller already passed, and the implementation-selection gate reads `config.resolved.toml` with the config parser instead of line regexes (#998). +- **[runtime] [artifacts]** Host finalization no longer re-implements campaign timeout evidence, partition joins, or invariant-ledger checks its schema enforces. Recon commands with `cd`/`tee`/`echo` framing, results without the optional `sequence_length`, and nested campaign result paths no longer fail after the verifier passed; a sequence length other than 100 now fails in the attempt (#998). - **[runtime]** `resume --refresh-controller` on a dynamically-expanded run now publishes a task manifest that still re-derives from its sealed base, closing a data-integrity regression. Current-controller rendering keeps binding each plan-time prompt to its authenticated retained snapshot (#980), but that execution-only binding now reaches the task-spec literal alone instead of the compiled task manifest, whose `renderedPromptPath` is normalized back to the sealed launch path recorded in `plan.json`. Previously the refreshed controller published the rebound `prompt-snapshots/.md` path into `smithers/tasks.json`, while `verifyDynamicRuntimeMaterialization` re-derived the same document from the sealed `controls/runtime-base-tasks.json` carrying the launch path; the two fingerprinted differently and every gated command (`sync`, `pause`, replay, fork) failed `WORKFLOW_CONTROL_EVIDENCE_INVALID` for the rest of the run's life, with no self-heal on a repeat refresh. Because resume is ungated, a run already damaged this way is recoverable: the normalization scrubs a rebound manifest on the next refresh, including when the sealed base manifest is absent and the live document is the compile input. Prompt authentication is unchanged -- a retained snapshot whose bytes no longer match `plan.json`'s recorded digest still aborts the refresh, and a task with no plan row still fails closed rather than falling back to the cleanup-owned launch path (#1002). - Pi model profiles now accept and forward the CLI's `max` thinking level, in addition to the previously supported levels through `xhigh`. - Upgrades the pinned workflow engine to Smithers 0.35.0. The Effect 4 tree is unchanged at `4.0.0-beta.105`, so the `@effect/*` override family and the `pnpm-workspace.yaml` mirror keep their exact pins; the deterministic npm resolution cutoff moves to `2026-08-18T06:00:00Z`, past `smthrs@0.35.0`'s publish instant. Six compatibility patches are re-derived against upstream's new `smithersRuntimeSpawn`/`smithersRuntimeReentry` indirection and the re-nested `adapter.insertRun` call; none is retired, because every workaround they encode is still absent upstream. Ultrafuzz's Smithers state mirrors gain the new `succeeded-with-failures` run state and `stalled` node state, and the inspect contract admits `tokenUsage`, `run.cancellationSource` and `runState.warnings`. Existing `0.34.0` dependency manifests migrate forward in place at `init`, so an already-scaffolded project converges onto the new pin without `--force` regenerating the run. Native resume never read the target project's `.smithers/package.json` -- launch admission asserts the manifest only against the operator controller's own freshly rendered temporary root -- so a stopped run already resumed under its own run ID and continues to. Ultrafuzz's strict readers for `why`, `status`, `node` and the `TokenUsageReported` event stream also admit the release's new `warnings`, `counts.stalled`, `tokenUsage.freshInputTokens`, `freshInputTokens`/`costUsd` and `stalled` blocker fields, and its two new `NodeStalled`/`RunConcurrencySaturated` event types; a stalled node now resets under `resume --retry-failed` alongside failed ones, and the evmbench adapter no longer aborts a converged run on the release's widened `degraded` verdict. The runner adds four forward-only SQLite migrations (0041-0044) (#955). From 99c2f7ca9e88e2abb8bf11481167b0d97ed833ff Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:09:33 +0000 Subject: [PATCH 133/206] fix(runtime): keep unscoped report prose advisory when no coverage producer was admitted The previous commit made UNSCOPED_COVERAGE_* warnings, but verifyFinalReportCoverageEvidence still failed the final report when no coverage-evidence producer was planned or admitted. It treated any coverage-context score in report.md as invented coverage evidence and emitted REPORT_COVERAGE_EVIDENCE_UNPLANNED as an error. That path covers every smoke run, which plans no coverage producer, and default runs whose continue-policy stateful-invariant-coverage failed, so a sentence such as "Recon reached 85% line coverage on Vault.sol." still failed final-report there. Only a score bound to an exact declaration-completeness scope now counts as a Markdown coverage claim. Without an admitted producer, a typed report.json.coverage_evidence or a scoped report.md score still fails, so the existing invented-coverage case keeps failing. The producer-free and planned-but-unadmitted final-report tests now also publish unscoped coverage sentences. Both fail with REPORT_COVERAGE_EVIDENCE_UNPLANNED without this change. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/src/artifact-gates.ts | 16 +++++++------ packages/runtime/test/artifact-gates.test.ts | 25 ++++++++++++++++++++ 2 files changed, 34 insertions(+), 7 deletions(-) diff --git a/packages/runtime/src/artifact-gates.ts b/packages/runtime/src/artifact-gates.ts index 165bb268d..83b0e25c8 100644 --- a/packages/runtime/src/artifact-gates.ts +++ b/packages/runtime/src/artifact-gates.ts @@ -7252,8 +7252,10 @@ function verifyFinalReportCoverageEvidence( const producerStatus = plannedContractProducerStatus(layout, node, "ultrafuzz/coverage-evidence@1", attemptAuthority); if (producerStatus === "absent") { - const markdownClaimsCoverage = markdownCoverageScoreOccurrences(markdown).some((occurrence) => - coverageScoreHasContext(occurrence, false) + // A score that names an exact declaration-completeness scope claims + // coverage evidence; unscoped prose scores stay advisory warnings. + const markdownClaimsCoverage = markdownCoverageScoreOccurrences(markdown).some( + (occurrence) => occurrence.scopes.length > 0 ); if (report.coverage_evidence !== undefined || markdownClaimsCoverage) { diagnostics.push({ @@ -7512,11 +7514,11 @@ function reportCoverageScoreDiagnostic( path: `${artifactPath}:${occurrence.line}` }; } - // Prose scores are advisory. This natural-language scan runs only on the - // host, after the in-workflow verifier accepted the attempt, and it matches - // ordinary sentences such as "Recon reached 85% line coverage" or "Handlers - // reachable: 7/9". The typed coverage evidence and its canonical Markdown - // section remain errors when they disagree. + // Unscoped prose scores are advisory. This natural-language scan runs only on + // the host, after the in-workflow verifier accepted the attempt, and it + // matches ordinary sentences such as "Recon reached 85% line coverage" or + // "Handlers reachable: 7/9". The typed coverage evidence and scores that name + // an exact scope are checked separately. const percentage = occurrence.kind === "percentage"; return { code: percentage ? "UNSCOPED_COVERAGE_PERCENTAGE" : "UNSCOPED_COVERAGE_FRACTION", diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 226522f22..0c02125ec 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -9246,6 +9246,19 @@ test("producer-free final reports require the exact not-planned implementation c const valid = verifyRuntimeRequiredArtifactsForAttempt(layout, node, node.id); assert.equal(valid.ok, true, JSON.stringify(valid.diagnostics)); + for (const [prose, code] of [ + ["Recon reached 85% line coverage on Vault.sol.", "UNSCOPED_COVERAGE_PERCENTAGE"], + ["Coverage of withdraw() was 3/4 branches in the replay.", "UNSCOPED_COVERAGE_FRACTION"] + ] as const) { + writeDeclaredArtifactNode(layout, node.id, outputs, { + "deliverables/report.md": `# Ultrafuzz report\n\n${prose}\n`, + "deliverables/report.json": JSON.stringify(currentReport(layout.runId)) + }); + const unscopedProse = verifyRuntimeRequiredArtifactsForAttempt(layout, node, node.id); + assert.equal(unscopedProse.ok, true, `${prose}: ${JSON.stringify(unscopedProse.diagnostics)}`); + assertAdvisoryCoverageScore(unscopedProse, code, prose); + } + writeDeclaredArtifactNode(layout, node.id, outputs, { "deliverables/report.md": "# Ultrafuzz report\n\nrecon-selected-declaration-completeness: `1/1`\n", "deliverables/report.json": JSON.stringify(currentReport(layout.runId)) @@ -9350,6 +9363,18 @@ test("final reports disclose planned but omitted implementation coverage without fs.existsSync(path.join(layout.artifactsDir, optionalTask.attemptId, "implemented-properties.json")), false ); + + const prose = "Recon reached 85% line coverage on Vault.sol."; + writeArtifactFile(layout, reportTask.attemptId, "report.md", `# Ultrafuzz report\n\n${prose}\n`); + const unadmittedCoverageProse = verifyRuntimeRequiredArtifactsForAttempt( + layout, + reportNode, + reportTask.attemptId, + { task: reportTask, tasks, admittedDependencyAttemptIds: [] }, + authenticatedSnapshotsForNode(layout, reportNode, reportTask.attemptId) + ); + assert.equal(unadmittedCoverageProse.ok, true, JSON.stringify(unadmittedCoverageProse.diagnostics)); + assertAdvisoryCoverageScore(unadmittedCoverageProse, "UNSCOPED_COVERAGE_PERCENTAGE", prose); }); test("final report gate joins the default recon-only campaign backend", () => { From e5bf30c51171b8cff2fa68e130a16e4e5355081d Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:09:47 +0000 Subject: [PATCH 134/206] fix(prompts): keep the coverage projection partial identical to its reference docs 84b4a799 reworded the bare-score sentence only in docs/reference/artifacts-reports.md. The prompts semantic-anchors test requires that page to contain the coverage-evidence-markdown.mdx partial verbatim, so it failed. The partial also still told agents that runtime publication rejects bare coverage scores, which stopped being true when those became warnings. The partial now says publication rejects missing, duplicated, reordered, or contradicting scoped scores and warns about scores that name no exact scope. The docs page carries it verbatim again, followed by a docs-only paragraph on the scan limit and on final reports without an admitted coverage producer. The partial is rendered into the stateful-invariant coverage and final-report prompts; it is not a validator-build input. Co-Authored-By: Claude Opus 5.5 --- .../coverage-evidence-markdown.mdx | 5 +++-- docs/reference/artifacts-reports.md | 17 +++++++++++------ 2 files changed, 14 insertions(+), 8 deletions(-) diff --git a/.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx b/.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx index f963a7f16..050533c23 100644 --- a/.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx +++ b/.ultrafuzz/prompts/_templates/output-contract/coverage-evidence-markdown.mdx @@ -37,5 +37,6 @@ Blockers: Repeat blocker and evidence rows in artifact order. Runtime publication compares this section with the typed handoff and rejects missing, duplicated, reordered, -or bare coverage scores. Raw `covg-eval` output is for iteration only and -defines neither published declaration-completeness view. +or contradicting scoped scores, and it warns about coverage scores that name no +exact scope. Raw `covg-eval` output is for iteration only and defines neither +published declaration-completeness view. diff --git a/docs/reference/artifacts-reports.md b/docs/reference/artifacts-reports.md index b1a62f68c..97e692a59 100644 --- a/docs/reference/artifacts-reports.md +++ b/docs/reference/artifacts-reports.md @@ -737,12 +737,17 @@ Blockers: ``` Repeat blocker and evidence rows in artifact order. Runtime publication compares -this section with the typed handoff and rejects missing, duplicated, or -reordered scores. A coverage score elsewhere in the Markdown or in `report.json` -text that names no exact declaration-completeness scope is reported as a -warning rather than failing publication, although a document with more than -2,048 score candidates still fails the scan limit. Raw `covg-eval` output is -for iteration only and defines neither published declaration-completeness view. +this section with the typed handoff and rejects missing, duplicated, reordered, +or contradicting scoped scores, and it warns about coverage scores that name no +exact scope. Raw `covg-eval` output is for iteration only and defines neither +published declaration-completeness view. + +A coverage score that names no exact declaration-completeness scope, whether in +`report.md`, the coverage producer's Markdown, or `report.json` text, does not +fail publication, although text that exceeds the 2,048-candidate scan limit +still does. When no coverage producer was planned or admitted, +`report.json.coverage_evidence` or a `report.md` score that names an exact +scope fails the final report. Current-run `report.md` contains concise links to `THREAT_MODEL.md`, `threat-model.json`, and `goal-plan.json`, plus source-node provenance for each From d0bdb3b6e6292e0b935169a9056a06c138a9f524 Mon Sep 17 00:00:00 2001 From: Antonio Viggiano Date: Tue, 29 Sep 2026 03:09:48 +0000 Subject: [PATCH 135/206] test(runtime): pin aggregation finalization and implicit-visibility gate results The deleted aggregation-semantic-context tests covered an admitted producer without controller finalization. The failed-optional aggregation test now also removes the admitted producer's verification marker and expects REQUIRED_ARTIFACT_INVALID, which pins that the canonical finalized reader still guards aggregation intake. The old authentication chain also rejected a producer without a marker (that is how the failed optional producer fails on origin/main), so this is a regression guard rather than a behavioural change. The implicit-visibility coverage loop lost its result assertion when the prose findings became warnings. It now asserts that only the two stylesheet cases fail. Co-Authored-By: Claude Opus 5.5 --- packages/runtime/test/artifact-gates.test.ts | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) diff --git a/packages/runtime/test/artifact-gates.test.ts b/packages/runtime/test/artifact-gates.test.ts index 0c02125ec..54516e8df 100644 --- a/packages/runtime/test/artifact-gates.test.ts +++ b/packages/runtime/test/artifact-gates.test.ts @@ -3916,6 +3916,20 @@ test("host aggregation intake skips a failed optional generated-tests producer t omitted.diagnostics.some((diagnostic) => diagnostic.message.includes("omits authenticated source bundle")), JSON.stringify(omitted.diagnostics) ); + + // An admitted producer is read only once the controller finalized it. + writeAggregationManifestCopying(layout, aggregationNode, [source]); + fs.unlinkSync(path.join(layout.root, ".ultrafuzz-verification", `${verifiedTask.attemptId}.json`)); + const unfinalized = verifyAggregationAttempt(layout, aggregationTask, tasks, [verifiedTask.attemptId]); + assert.equal(unfinalized.ok, false); + assert.ok( + unfinalized.diagnostics.some( + (diagnostic) => + diagnostic.code === "REQUIRED_ARTIFACT_INVALID" && + diagnostic.message.includes(`authority is invalid for ${verifiedTask.attemptId}`) + ), + JSON.stringify(unfinalized.diagnostics) + ); }); test("host aggregation intake keys a directly consumed dynamic producer by its storage attempt ID", () => { @@ -14768,6 +14782,12 @@ test("final report preserves typed coverage evidence and its canonical Markdown ]) { writeArtifact(layout, reportNode.id, "report.md", `${scopedMarkdown}\n## Notes\n\n${implicitlyVisibleScore}\n`); const visibleAfterImplicitClose = verifyRequiredArtifactsForAttempt(layout, reportNode, reportNode.id); + // Only the stylesheet cases fail; the prose score itself stays advisory. + assert.equal( + visibleAfterImplicitClose.ok, + !implicitlyVisibleScore.startsWith("