Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion docs/reference/artifacts-reports.md
Original file line number Diff line number Diff line change
Expand Up @@ -227,7 +227,16 @@ status changes afterwards. Finished, failed, timed-out, and cancelled attempts
are recorded; a cancellation has outcome and category `canceled` and the
Smithers cancellation reason as its message. A terminal event with no started
attempt in the same Smithers activation, or stamped before that attempt
started, is skipped. A finished attempt whose node then fails, for example
started, is skipped. An attempt that was still running when its controller
stopped, for example because the controller was killed, has no terminal event:
when the run resumes, Smithers marks it cancelled without emitting one. Like a
failure, it is recorded unless it stopped before Smithers selected a rung: as
`canceled`, with the message `abandoned: the controller stopped during this
attempt, and the resumed run cancelled it`, and it counts toward the node's
`retry_count`. Its terminal is the next event of the same task, normally the
replacement attempt's start, because one resume can abandon several attempts;
it is not recorded before the task has such an event. A finished attempt whose
node then fails, for example
because the verifier or artifact gates reject its output, is recorded as failed
with category `invalid-output` for findings validation and
`artifact-validation` otherwise.
Expand Down
6 changes: 5 additions & 1 deletion docs/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -406,7 +406,11 @@ completed and currently elapsed execution time, token components, estimated
spend, model, and attempt count. JSON output uses `ultrafuzz.stats.v1` inside
the normal CLI envelope and additionally exposes retries, executed/reused
counts, outcomes, failure categories, completeness, unattributed usage, and
cumulative run accounting.
cumulative run accounting. Run state records a task that Smithers cancelled,
for example by `ultrafuzz cancel`, as failed. When Smithers cancelled every
failed task of a node, `stats` reports the node with status `canceled` and
counts it under `canceled` rather than `failed`; `status` likewise counts
cancelled tasks apart from failures.
The closed stats v1 field `pricing_complete` continues to mean complete cost
coverage: it is the inverse of run accounting `partial_pricing`. Accounting v4's
own `pricing_complete` field independently reports whether each event had
Expand Down
12 changes: 7 additions & 5 deletions docs/reference/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -74,11 +74,13 @@ pnpm --filter @ultrafuzz/modal test
`report`, and `events` as separate CLI processes, and the generated workflow
runs on the pinned Smithers engine under Bun, with a stub `codex` executable in
place of the model. It SIGKILLs the detached controller while one node is
running, resumes the run, and checks that it succeeds with a verified report and
that no finished task started again. It needs Linux, Bun, Git, and access to
the npm registry, because `run` and `resume` install the pinned engine from npm
as they do for any campaign. On SIGINT or SIGTERM, the test kills the detached
campaign and deletes its fixture, about 1 GB, before it exits.
running, resumes the run, and checks that it succeeds with a verified report,
that no finished task started again, and that `status` and `stats` count the
same agent attempts, including the one the kill interrupted. It needs Linux,
Bun, Git, and access to the npm registry, because `run` and `resume` install the
pinned engine from npm as they do for any campaign. On SIGINT or SIGTERM, the
test kills the detached campaign and deletes its fixture, about 1 GB, before it
exits.

Package-local `typecheck` and `test` scripts may build direct workspace
dependencies first because package exports point at `dist/**`.
Expand Down
3 changes: 3 additions & 0 deletions packages/cli/schema/cli-result.schema.json
Original file line number Diff line number Diff line change
Expand Up @@ -1561,6 +1561,7 @@
"running",
"succeeded",
"failed",
"canceled",
"skipped",
"timed-out",
"reused-from-prior-run",
Expand Down Expand Up @@ -1612,6 +1613,7 @@
"running",
"succeeded",
"failed",
"canceled",
"skipped",
"timed-out",
"reused-from-prior-run",
Expand All @@ -1625,6 +1627,7 @@
"running": { "$ref": "#/$defs/nonNegativeInteger" },
"succeeded": { "$ref": "#/$defs/nonNegativeInteger" },
"failed": { "$ref": "#/$defs/nonNegativeInteger" },
"canceled": { "$ref": "#/$defs/nonNegativeInteger" },
"skipped": { "$ref": "#/$defs/nonNegativeInteger" },
"timed-out": { "$ref": "#/$defs/nonNegativeInteger" },
"reused-from-prior-run": { "$ref": "#/$defs/nonNegativeInteger" },
Expand Down
36 changes: 31 additions & 5 deletions packages/cli/src/run-statistics.ts
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ export interface TokenStatistics {
models: string[];
}

export type NodeStatisticsStatus = NodeStatus | "unknown";
export type NodeStatisticsStatus = NodeStatus | "canceled" | "unknown";

export interface NodeStatistics {
node_id: string;
Expand Down Expand Up @@ -484,6 +484,17 @@ function taskWorkflowAgentId(nodeState: NodeState | undefined): string | undefin
return "agent_task_id" in provenance.workflow ? provenance.workflow.agent_task_id : undefined;
}

/** Whether Smithers ended every failed task of the node cancelled. */
function workflowCancelled(states: readonly NodeState[]): boolean {
const failedTaskStates = states.flatMap((nodeState) => {
const provenance = nodeState.provenance;
if (nodeState.status !== "failed" || provenance === undefined || !("workflow" in provenance)) return [];
const workflow = provenance.workflow;
return workflow !== undefined && "state" in workflow ? [workflow.state] : [];
});
return failedTaskStates.length > 0 && failedTaskStates.every((state) => state === "cancelled");
Comment on lines +491 to +495

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Failed tasks can disappear When a fan-out node has one task marked cancelled and another failed task without a recorded workflow state, this check drops the second task. It then reports the whole node as canceled, hiding the failure in totals.status_counts. A failed task needs cancellation evidence before it can be counted as canceled.

Knowledge Base Used: CLI workflows

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/cli/src/run-statistics.ts
Line: 491-495

Comment:
**Failed tasks can disappear** When a fan-out node has one task marked `cancelled` and another failed task without a recorded workflow state, this check drops the second task. It then reports the whole node as `canceled`, hiding the failure in `totals.status_counts`. A failed task needs cancellation evidence before it can be counted as canceled.

**Knowledge Base Used:** [CLI workflows](https://app.greptile.com/monad-foudnation/-/custom-context/knowledge-base/monad-developers/ultrafuzz/-/docs/cli-workflows.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

}

function usageNodeAliases(
descriptors: readonly NodeDescriptor[],
diagnostics: RuntimeDiagnostic[]
Expand Down Expand Up @@ -548,7 +559,10 @@ function nodeStatistics(

const canonicalState =
descriptor.states.find((nodeState) => nodeState.node_id === descriptor.nodeId) ?? descriptor.states[0];
const status = canonicalState?.status ?? aggregateNodeStatus(descriptor.states);
const stateStatus = canonicalState?.status ?? aggregateNodeStatus(descriptor.states);
// Run state records a task that Smithers cancelled as failed; `status` does
// not count it as a failure (#1087).
const status = stateStatus === "failed" && workflowCancelled(descriptor.states) ? "canceled" : stateStatus;
const strategyStates = descriptor.states.filter((nodeState) => nodeState.node_id !== descriptor.nodeId);
const timedStates = strategyStates.length > 0 ? strategyStates : descriptor.states;
const currentElapsedMs = timedStates.reduce<number | null>((total, nodeState) => {
Expand Down Expand Up @@ -803,9 +817,20 @@ function aggregateNodeStatus(states: readonly NodeState[]): NodeStatisticsStatus
}

function aggregateAttemptOutcome(attempts: readonly NodeAttemptLedgerEntry[]): NodeAttemptOutcome | "mixed" | null {
const latestByStrategy = new Map<string, NodeAttemptOutcome>();
for (const attempt of attempts) latestByStrategy.set(attempt.strategy_attempt_id, attempt.outcome);
const outcomes = [...new Set(latestByStrategy.values())];
const latestByStrategy = new Map<string, NodeAttemptLedgerEntry>();
for (const attempt of attempts) {
const previous = latestByStrategy.get(attempt.strategy_attempt_id);
// Rows are appended as they are recorded, so a version that records more
// attempts can append an earlier attempt of a workflow run after a later one.
if (
previous?.workflow_run_id === attempt.workflow_run_id &&
previous.source_event_sequence > attempt.source_event_sequence
) {
continue;
}
latestByStrategy.set(attempt.strategy_attempt_id, attempt);
}
const outcomes = [...new Set([...latestByStrategy.values()].map((attempt) => attempt.outcome))];
if (outcomes.length === 0) return null;
return outcomes.length === 1 ? outcomes[0]! : "mixed";
}
Expand All @@ -818,6 +843,7 @@ function emptyStatusCounts(): Record<NodeStatisticsStatus, number> {
running: 0,
succeeded: 0,
failed: 0,
canceled: 0,
skipped: 0,
"timed-out": 0,
"reused-from-prior-run": 0,
Expand Down
1 change: 1 addition & 0 deletions packages/cli/test/cli-contracts.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -234,6 +234,7 @@ test("stats has an exact command discriminator and a fully closed result shape",
running: 0,
succeeded: 0,
failed: 0,
canceled: 0,
skipped: 0,
"timed-out": 0,
"reused-from-prior-run": 0,
Expand Down
20 changes: 8 additions & 12 deletions packages/cli/test/e2e/campaign-resume.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -204,6 +204,7 @@ interface StatsValue {
status: string;
attempt_count: number | null;
executed_attempt_count: number | null;
failure_categories: string[] | null;
}>;
totals: { node_count: number; status_counts: Record<string, number> };
}
Expand Down Expand Up @@ -498,21 +499,16 @@ test(
assert.equal(stats.totals.node_count, AGENT_NODES.length);
assert.equal(stats.totals.status_counts.succeeded, AGENT_NODES.length);
assert.equal(stats.nodes.find((node) => node.node_id === "project-discovery")?.executed_attempt_count, 1);
// status counts every agent invocation, including the one the kill interrupted.
// status and stats count every agent invocation, including the one the kill interrupted, which
// stats records as canceled.
const statusAttempts = health.model_mix.reduce((total, entry) => total + entry.attempts, 0);
assert.equal(statusAttempts, starts.length);
mark("checks done");

await t.test(
"stats counts the agent attempt the controller crash interrupted",
{
todo: "Smithers emits no terminal event for the attempt it abandons when the resumed run starts, so attempts.jsonl never records it (#1187)"
},
() => {
const statsAttempts = stats.nodes.reduce((total, node) => total + (node.attempt_count ?? 0), 0);
assert.equal(statsAttempts, statusAttempts);
}
assert.equal(
stats.nodes.reduce((total, node) => total + (node.attempt_count ?? 0), 0),
statusAttempts
);
assert.deepEqual(stats.nodes.find((node) => node.node_id === INTERRUPTED_NODE)?.failure_categories, ["canceled"]);
mark("checks done");
} finally {
process.removeListener("SIGINT", interrupted);
process.removeListener("SIGTERM", interrupted);
Expand Down
76 changes: 74 additions & 2 deletions packages/cli/test/run-statistics.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ import {
manifestDigest,
type NodeAttemptLedgerEntry,
type NodeState,
type NodeWorkflowProvenance,
type PlannedGraphDocument,
type RunAccountingSummary,
type RunMetadataDocument,
Expand All @@ -20,6 +21,8 @@ import {
} from "@ultrafuzz/artifacts";
import { assertRunMetadataAccountingUsageAuthority } from "@ultrafuzz/runtime";

import { envelope } from "../src/command-shared.js";
import { buildStatisticsCommandResult } from "../src/commands/stats.js";
import { deriveRunStatistics, type StatisticsEvidence } from "../src/run-statistics.js";

const RUN_ID = "stats-unit";
Expand All @@ -40,7 +43,12 @@ function attempt(
nodeId: string,
strategyAttemptId: string,
sourceEventSequence: number,
options: { runId?: string; startedAt?: string; finishedAt?: string; outcome?: "succeeded" | "failed" } = {}
options: {
runId?: string;
startedAt?: string;
finishedAt?: string;
outcome?: "succeeded" | "failed" | "canceled";
} = {}
): NodeAttemptLedgerEntry {
const outcome = options.outcome ?? "succeeded";
return createNodeAttemptLedgerEntry(
Expand All @@ -59,7 +67,7 @@ function attempt(
outcome,
inputManifestDigest: manifestDigest("input"),
outputManifestDigest: outcome === "succeeded" ? manifestDigest("output") : null,
...(outcome === "failed" ? { failureCategory: "executor-error" as const } : {})
...(outcome === "succeeded" ? {} : { failureCategory: outcome === "failed" ? "executor-error" : "canceled" })
}
);
}
Expand Down Expand Up @@ -168,6 +176,7 @@ test("stats derives closed per-node timing, usage, cost, and status totals", ()
running: 0,
succeeded: 1,
failed: 0,
canceled: 0,
skipped: 0,
"timed-out": 0,
"reused-from-prior-run": 0,
Expand All @@ -192,6 +201,69 @@ test("stats derives closed per-node timing, usage, cost, and status totals", ()
assert.deepEqual(derived.diagnostics, []);
});

test("stats counts a failed node whose task Smithers cancelled as canceled", () => {
type WorkflowState = "cancelled" | "failed";
const failed = (nodeId: string, logicalNodeId: string, workflow: NodeWorkflowProvenance): NodeState => ({
...terminalNodeState(nodeId, logicalNodeId),
status: "failed",
outputs: [],
provenance: { workflow }
});
const task = (nodeId: string, logicalNodeId: string, state: WorkflowState): NodeState =>
failed(nodeId, logicalNodeId, {
run_id: WORKFLOW_RUN_ID,
task_id: `node:${nodeId}`,
agent_task_id: `node:${nodeId}`,
verifier_task_id: `verify:${nodeId}`,
state,
attempt: 1
});
const status = (states: NodeState[], graph?: PlannedGraphDocument) => {
// Smithers cancelled the tasks before it selected an agent, so no attempt was recorded.
const { value } = deriveRunStatistics(
evidence({
...(graph === undefined ? {} : { graph, usage: [] }),
state: runState(states, "canceled"),
attempts: []
}),
Date.parse(FINISHED_AT)
);
// The JSON envelope is validated against the closed CLI result schema.
envelope("stats", buildStatisticsCommandResult(value, []));
return [value.nodes[0]?.status, value.totals.status_counts.failed, value.totals.status_counts.canceled];
};

// Run state records a cancelled task as failed; `status` counts it apart from failures (#1087).
assert.deepEqual(status([task("node", "node", "cancelled")]), ["canceled", 0, 1]);
assert.deepEqual(status([task("node", "node", "failed")]), ["failed", 1, 0]);
// A fan-out node's canonical state aggregates its strategy tasks and has no Smithers state of its own.
const fanOut = (...states: WorkflowState[]) =>
status(
[
failed("fan", "fan", { run_id: WORKFLOW_RUN_ID, aggregate_attempt_statuses: ["failed", "failed"] }),
...states.map((state, index) => task(`fan__model_${index}__attempt_0`, "fan", state))
],
graphDocument(
"fan",
states.map((_, index) => `node:fan__model_${index}__attempt_0`),
["gpt-test", "gpt-other"]
)
);
assert.deepEqual(fanOut("cancelled", "cancelled"), ["canceled", 0, 1]);
assert.deepEqual(fanOut("cancelled", "failed"), ["failed", 1, 0]);
});

test("stats reports a node's latest attempt outcome in event order, not ledger row order", () => {
const outcome = (attempts: NodeAttemptLedgerEntry[]) =>
deriveRunStatistics(evidence({ attempts }), Date.parse(FINISHED_AT)).value.nodes[0]?.outcome;
const abandoned = attempt("node", "node", 1, { outcome: "canceled" });
const replacement = attempt("node", "node", 2);

assert.equal(outcome([abandoned, replacement]), "succeeded");
// A run synchronized by an earlier version records the abandoned attempt after its replacement.
assert.equal(outcome([replacement, abandoned]), "succeeded");
});

test("stats counts only the latest cumulative usage snapshot for each attempt", () => {
const snapshots = [
usage(1, "node:node", {
Expand Down
1 change: 1 addition & 0 deletions packages/cli/test/stats-command.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@ const statistics: RunStatisticsValue = {
running: 0,
succeeded: 0,
failed: 0,
canceled: 0,
skipped: 0,
"timed-out": 0,
"reused-from-prior-run": 0,
Expand Down
Loading
Loading