Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 0 additions & 19 deletions .ultrafuzz/prompts/review/final-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -385,25 +385,8 @@ Ultrafuzz is an automated smart-contract fuzzing campaign assistant. Issues belo
- Tokens used: `<token usage, or unavailable>`
- Estimated spend: `<cost estimate such as $123 or $123+ when pricing is partial, or unavailable>`
- Audit profile: `<effective audit profile, or unavailable>`

## Audit context

- Threat model: [THREAT_MODEL.md](<relative path to THREAT_MODEL.md>); [threat-model.json](<relative path to threat-model.json>)
- Goal plan: [goal-plan.json](<relative path to goal-plan.json>)
```

Render `## Audit context` with exactly this heading, bullet order, and link
text, immediately after `## Run summary`. Use repository-relative or
report-relative paths to the run's own `threat-model` and `goal-plan` artifacts;
never absolute paths or external URLs. Omit an individual link whose artifact
the run did not produce, omit the `Goal plan` bullet when there is no goal plan,
and omit the whole section when the run produced none of them. Do not invent a
different heading, ordering, or link text: `ultrafuzz report` regenerates this
exact section deterministically from the run's own artifacts and overwrites
anything else.
Keep detailed threat content in those dedicated artifacts; do not duplicate it
in `report.md`.

Each production issue entry must use exactly this Markdown section order. The
following example is structural only; replace the title, actor names, actions,
outcomes, explanations, code, variants, and strategy IDs with issue-specific
Expand Down Expand Up @@ -825,8 +808,6 @@ Before finishing, verify that:
- `report.md` contains `## Property provenance`, including every
property-derived finding and no invented property IDs for non-property
findings.
- `report.md` renders the fixed `## Audit context` section for every artifact
the run produced, without copying their detailed analysis.
- `report.md` contains `## Property implementation coverage` rendered from the
exact runtime-authoritative coverage object.
- `report.md` contains `## Goal search coverage` with counts recomputed from the
Expand Down
37 changes: 23 additions & 14 deletions docs/reference/artifacts-reports.md
Original file line number Diff line number Diff line change
Expand Up @@ -537,8 +537,9 @@ artifacts/final-report/report.json

`report.json` must satisfy `ultrafuzz/report@3` with the exact
`ultrafuzz.report.v3` version literal. After a run stops, the runtime can format
that report and attach whole-run completion information without changing the
agent's files. Verified publications use:
that report, attach whole-run completion information, and restate the run
summary's elapsed time and accounting without changing the agent's files.
Verified publications use:

```text
review/runtime-report/<authority-digest>/report.json
Expand Down Expand Up @@ -741,11 +742,12 @@ this section with the typed handoff and rejects missing, duplicated, reordered,
or bare coverage scores. Raw `covg-eval` output is for iteration only and
defines neither published declaration-completeness view.

Current-run `report.md` contains concise links to `THREAT_MODEL.md`,
`threat-model.json`, and `goal-plan.json`, plus source-node provenance for each
production issue. Detailed threat analysis stays in the dedicated threat-model
artifacts and is not duplicated into the report. `report.json` preserves the
same `source_nodes` arrays.
Current-run `report.md` contains source-node provenance for each production
issue and does not link to other run files. Detailed threat analysis stays in
the dedicated threat-model artifacts and is not duplicated into the report.
`report.json` preserves the same `source_nodes` arrays. Inline link and image
syntax inside report prose, including prose preserved byte-for-byte from
upstream findings, renders as literal text.

When workflow usage data is available, run metadata includes
`accounting.cumulative.tokens_used` and
Expand All @@ -755,13 +757,20 @@ available cumulative values into the markdown run summary and into
persisted estimate is partial because some token usage did not have pricing
data.

If cumulative metadata has not synchronized when the final-report producer
starts, its live Smithers fallback is a snapshot through that producer's start.
It includes earlier attempts but cannot include the producer's own eventual
duration, model fallback, tokens, or cost. A terminal presentation of an existing
verified agent report preserves those accounting values. Report v3 has no
metric-scope field, so use
`ultrafuzz stats` after terminal synchronization for closed-run accounting.
The final-report producer receives its run summary when its task starts: from
cumulative metadata when it has synchronized, otherwise from a live Smithers
fallback. Either way it is a snapshot through that producer's start. It includes
earlier attempts but cannot include the producer's own eventual duration, model
fallback, tokens, or cost, and the agent's `report.json` and `report.md` keep
that snapshot. Runtime presentations (the verified terminal publication and
unchecked reports) restate the run summary instead: elapsed time from
`run.json#created_at` to `state.json#finished_at`, and models, tokens,
estimated spend, and `partial_pricing` from the current
`accounting.cumulative`. Tokens, estimated spend, and `partial_pricing` are
restated together whenever `accounting.cumulative` records a token count, so a
whole-run spend recorded as `unavailable` stays `unavailable` instead of showing
the agent's report-start figure. Otherwise, a value those records lack keeps the
agent's copy. Use `ultrafuzz stats` for the full accounting breakdown.

`accounting.segments` publishes one rollup per checkpoint generation, and
`accounting.current` identifies the latest segment. Each segment retains every
Expand Down
69 changes: 12 additions & 57 deletions packages/runtime/src/final-report-markdown.ts
Original file line number Diff line number Diff line change
Expand Up @@ -21,18 +21,16 @@ import { redactSecretsInText, type SecretScanMode } from "@ultrafuzz/security";

export const MAX_FINAL_REPORT_JSON_BYTES = 64 * 1024 * 1024;
export const MAX_FINAL_REPORT_MARKDOWN_BYTES = 16 * 1024 * 1024;
const SAFE_REPORT_RELATIVE_LINK_PATTERN = /^\.\.\/(?:(?!\.\.?\/)[A-Za-z0-9._-]+\/)+(?!\.\.?$)[A-Za-z0-9._-]+$/u;
/**
* Secret placeholder for the public projection only. Two constraints pick it:
*
* - The public bundle re-projects the published report.json and requires the same bytes back, so
* the placeholder must be a fixed point of the redaction pass. The key-name assignment rule's
* unquoted value class stops at whitespace, `,`, `;`, `]`, and `}`, so a placeholder containing
* any of those is re-redacted on the next pass (`token=[redacted]` becomes `token=[redacted]]`).
* - The final-review Markdown gate rejects raw HTML, images, and links outside fenced code, and
* redacted values land unescaped in inline code (run summary values, coverage paths, source
* nodes), so the placeholder must not read as HTML (`<redacted>`), a link (`[redacted](`),
* an image, or emphasis (`*`, `_`).
* - The final-review Markdown gate rejects raw HTML outside fenced code, including inside inline
* code, where redacted values land unescaped (run summary values, coverage paths, source nodes),
* so the placeholder must not read as HTML (`<redacted>`).
*
* A bare uppercase word satisfies both. Every other redaction keeps the security package's default.
*/
Expand Down Expand Up @@ -260,9 +258,11 @@ function finalReportMarkdownDirectiveViolation(markdown: string, report: JsonRec
if (!markdown.includes("\n## Property provenance\n")) {
return "missing property provenance";
}
// Report prose is preserved byte-for-byte from upstream artifacts that the agent cannot repair, so
// the only rule left is one escaped prose cannot match: publicProse escapes `<`, and only
// unescaped inline-code values can still carry raw HTML.
const prose = markdownOutsideFencedCode(markdown).replace(/<br\s*\/?\s*>/giu, "");
const proseViolation = finalReportProseDirectiveViolation(prose);
if (proseViolation !== undefined) return proseViolation;
if (/<[A-Za-z][^>]*>/u.test(prose)) return "contains raw HTML outside fenced code";
const rendered = renderedIssues(Array.isArray(report.issues) ? report.issues.filter(isRecord) : []);
const expectedHeadings = rendered.map(renderedIssueHeading);
const headings = markdown.split("\n").filter((line) => line.startsWith("## ["));
Expand Down Expand Up @@ -347,31 +347,6 @@ function completionFindingsViolation(
return undefined;
}

function finalReportProseDirectiveViolation(prose: string): string | undefined {
// Critical is not a supported report severity, but the word remains valid in explanatory prose
// (for example, "a critical invariant"). Reject only a standalone severity-like label rather than
// rewriting or discarding the validated finding text.
const forbiddenPatterns: ReadonlyArray<readonly [RegExp, string]> = [
[/(?:^|\n)(?:#{1,6}\s+|-\s+)?(?:\*\*)?Critical(?:\*\*)?\s*$/imu, "contains the unsupported Critical severity"],
[/(?:^|\n)#### Sources\s*$/imu, "contains a legacy Sources section"],
[/\*\*Source (?:Node|Property) Id\*\*/iu, "contains a legacy source identifier field"],
[/(?:^|\n)- \*\*Item \d+\*\*/imu, "contains a legacy numbered-item field"],
[/(?:^|\n)## (?:Executive summary|Issue index|Additional report data)\s*$/imu, "contains a legacy report section"],
[/(?:^|\n)#{3,6} (?:Lifecycle|Strategy|Strategy provenance)\s*$/imu, "contains a legacy issue subsection"],
[
/(?:^|\n)- (?:Strategy loops|Audit profile catalog digest|Topology digest|Prompt digest|Expanded graph fingerprint):/imu,
"contains legacy run metadata"
],
[/<[A-Za-z][^>]*>/u, "contains raw HTML outside fenced code"],
[/!\[[^\]]*\]\(/u, "contains an embedded image outside fenced code"],
[
/(?<!\\)\]\((?!(?:#[a-z0-9-]+|\.\.\/(?:(?!\.\.?\/)[A-Za-z0-9._-]+\/)+(?!\.\.?\))[A-Za-z0-9._-]+)\))/iu,
"contains a disallowed Markdown link outside fenced code"
]
];
return forbiddenPatterns.find(([pattern]) => pattern.test(prose))?.[1];
}

function validateReport(report: unknown): JsonRecord {
const serialized = `${JSON.stringify(report)}\n`;
const validation = validateArtifactContract("ultrafuzz/report@3", serialized, "report.json");
Expand Down Expand Up @@ -695,7 +670,6 @@ function renderCanonicalReport(report: JsonRecord, goalSearchCoverage: unknown):
lines,
isRecord(report.run_metadata) ? report.run_metadata.artifact_validation_warnings : undefined
);
appendAuditContext(lines, report.audit_context);
const campaignDidNotRun = appendCampaignOutcome(lines, report.campaign_outcome);
appendCoverageEvidence(lines, report.coverage_evidence);
const goalCoverage = summarizeGoalSearchCoverage(goalSearchCoverage);
Expand Down Expand Up @@ -838,8 +812,10 @@ function appendArtifactValidationWarnings(lines: string[], value: unknown): void
""
);
for (const warning of value.filter(isRecord)) {
// Codes render as plain text. Only `](` is escaped, so the bytes of real gate codes do not change.
const code = inlineValue(warning.code).replaceAll("](", "]\\(");
lines.push(
`- ${inlineValue(warning.code)} — \`${inlineValue(warning.artifact_path)}#${inlineValue(warning.field_path)}\`: ${publicProse(String(warning.message))}`
`- ${code} — \`${inlineValue(warning.artifact_path)}#${inlineValue(warning.field_path)}\`: ${publicProse(String(warning.message))}`
);
if (warning.source_path !== undefined) lines.push(` - Available context: \`${inlineValue(warning.source_path)}\``);
}
Expand Down Expand Up @@ -961,29 +937,6 @@ function appendRunSummary(lines: string[], metadata: JsonRecord): void {
}
}

function appendAuditContext(lines: string[], value: unknown): void {
if (!isRecord(value)) return;
const threat = recordField(value, "threat_model");
const goalPlan = recordField(value, "goal_plan");
const threatMarkdown = safeReportLink(threat?.markdown);
const threatJson = safeReportLink(threat?.json);
const goalPlanJson = safeReportLink(goalPlan?.json);
if (threatMarkdown === undefined && threatJson === undefined && goalPlanJson === undefined) return;
lines.push("", "## Audit context", "");
if (threatMarkdown !== undefined || threatJson !== undefined) {
const links = [
threatMarkdown === undefined ? undefined : `[THREAT_MODEL.md](${threatMarkdown})`,
threatJson === undefined ? undefined : `[threat-model.json](${threatJson})`
].filter((entry): entry is string => entry !== undefined);
lines.push(`- Threat model: ${links.join("; ")}`);
}
if (goalPlanJson !== undefined) lines.push(`- Goal plan: [goal-plan.json](${goalPlanJson})`);
}

function safeReportLink(value: unknown): string | undefined {
return typeof value === "string" && SAFE_REPORT_RELATIVE_LINK_PATTERN.test(value) ? value : undefined;
}

function appendProductionIssue(lines: string[], rendered: RenderedIssue): void {
const { issue } = rendered;
lines.push("", renderedIssueHeading(rendered), "", publicProse(issueDescription(issue)), "", "### Severity", "");
Expand Down Expand Up @@ -1522,6 +1475,7 @@ function recordTitle(record: JsonRecord, fallback: string): string {
return typeof record.title === "string" && record.title.trim().length > 0 ? record.title.trim() : fallback;
}

/** Escaping `(` after every `]` keeps byte-preserved prose from forming an inline link or image. */
function publicProse(value: string): string {
return value
.replace(/\s+/gu, " ")
Expand All @@ -1533,6 +1487,7 @@ function publicProse(value: string): string {
.replaceAll("!", "\\!")
.replaceAll("#", "\\#")
.replaceAll("~", "\\~")
.replaceAll("](", "]\\(")
.replaceAll("<", "&lt;")
.replaceAll(">", "&gt;");
}
Expand Down
53 changes: 53 additions & 0 deletions packages/runtime/src/terminal-report-projection.ts
Original file line number Diff line number Diff line change
Expand Up @@ -51,8 +51,61 @@ export function projectTerminalReport(input: TerminalReportProjectionInput): Can
if (sourceRunId !== undefined && Reflect.get(reportMetadata, "source_run_id") !== sourceRunId) {
throw new Error("Verified final report has a different source run identity");
}
report.run_metadata = withWholeRunSummary(reportMetadata as Record<string, unknown>, metadata, state.finished_at);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Older verified reports stop verifying For a terminal publication written before this change, the restated summary can differ from its stored bytes. Reads regenerate the report and reject that publication on a byte mismatch, so an unchanged run falls back to an unchecked report and --require-verified access fails.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/terminal-report-projection.ts
Line: 54

Comment:
**Older verified reports stop verifying** For a terminal publication written before this change, the restated summary can differ from its stored bytes. Reads regenerate the report and reject that publication on a byte mismatch, so an unchanged run falls back to an unchecked report and `--require-verified` access fails.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

report.completion = completion;
return projectCanonicalFinalReport(report, context);
}
throw new ReportUnavailableError("no successful report-agent output is available");
}

/**
* The report agent copies a run summary that the host captured when the report task started, so its
* elapsed time and accounting miss the report task itself and anything that finished later. Runtime
* presentations restate them from run.json and the recorded finish time. Tokens, spend, and partial
* pricing move together, because the agent's spend may price only part of the run; a spend the
* whole-run record calls unavailable stays unavailable. A value those records do not provide keeps
* the agent's copy; malformed records are ignored, never thrown.
*/
export function withWholeRunSummary(
runMetadata: Record<string, unknown>,
metadata: unknown,
finishedAt: unknown
): Record<string, unknown> {
const summary = { ...runMetadata };
const elapsed = elapsedTime(field(metadata, "created_at"), finishedAt);
if (elapsed !== undefined) summary.elapsed_time = elapsed;
const cumulative = field(field(metadata, "accounting"), "cumulative");
const models = field(cumulative, "models");
if (isNonEmptyStringList(models)) summary.models_used = [...models];
const tokens = field(cumulative, "tokens_used");
const spend = field(cumulative, "estimated_spend");
if (availableLabel(tokens) && typeof spend === "string" && spend.trim() !== "") {
summary.tokens_used = tokens;
summary.estimated_spend = spend;
summary.partial_pricing = field(cumulative, "partial_pricing") === true;
}
Comment thread
greptile-apps[bot] marked this conversation as resolved.
return summary;
}

function field(value: unknown, key: string): unknown {
return typeof value === "object" && value !== null && !Array.isArray(value) ? Reflect.get(value, key) : undefined;
}

function isNonEmptyStringList(value: unknown): value is string[] {
return Array.isArray(value) && value.length > 0 && value.every((entry) => typeof entry === "string" && entry !== "");
}

function availableLabel(value: unknown): value is string {
return typeof value === "string" && value.trim() !== "" && value.trim().toLowerCase() !== "unavailable";
}

/** Same format as the host's report-start projection: `42.0s`, `5m 07s`, or `6h 02m`. */
function elapsedTime(createdAt: unknown, finishedAt: unknown): string | undefined {
if (typeof createdAt !== "string" || typeof finishedAt !== "string") return undefined;
const totalSeconds = (Date.parse(finishedAt) - Date.parse(createdAt)) / 1_000;
if (!Number.isFinite(totalSeconds) || totalSeconds < 0) return undefined;
if (totalSeconds < 60) return `${totalSeconds.toFixed(1)}s`;
const totalMinutes = Math.floor(totalSeconds / 60);
if (totalMinutes < 60) return `${String(totalMinutes)}m ${String(Math.floor(totalSeconds % 60)).padStart(2, "0")}s`;
return `${String(Math.floor(totalMinutes / 60))}h ${String(totalMinutes % 60).padStart(2, "0")}m`;
}
12 changes: 12 additions & 0 deletions packages/runtime/src/unverified-report-inputs.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ import {
reportSchema,
type ReportVerification
} from "@ultrafuzz/artifacts";
import { loadGoalSearchCoverageSnapshot } from "./final-report-markdown.js";
import { ReportUnavailableError } from "./report-unavailable.js";

type JsonRecord = Record<string, unknown>;
Expand All @@ -31,6 +32,7 @@ export interface UnverifiedReportInputs {
observed: ObservedReportCompletion;
verification: ReportVerification;
agentReport: JsonRecord;
goalSearchCoverage?: unknown;
sources_sha256: string;
}

Expand All @@ -53,10 +55,20 @@ export function readUnverifiedReportInputs(root: string): UnverifiedReportInputs
observed,
verification: { status: "not-checked", reason_codes: [...reader.reasons].sort() },
agentReport,
goalSearchCoverage: readGoalSearchCoverage(root),
sources_sha256: sha256Bytes(Buffer.from(JSON.stringify(reader.sources)))
};
}

/** An unreadable census renders as unknown goal-search coverage instead of hiding the report. */
function readGoalSearchCoverage(root: string): unknown {
try {
return loadGoalSearchCoverageSnapshot(root);
} catch {
return undefined;
}
}

class ReportInputReader {
readonly reasons = new Set<Reason>(["verification-unavailable"]);
readonly sources: [string, string][] = [];
Expand Down
19 changes: 14 additions & 5 deletions packages/runtime/src/unverified-report.ts
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ import {
} from "@ultrafuzz/artifacts";

import { projectCanonicalFinalReport } from "./final-report-markdown.js";
import { withWholeRunSummary } from "./terminal-report-projection.js";
import {
loadCurrentFinalReportSnapshot,
publishTerminalReport,
Expand Down Expand Up @@ -111,11 +112,19 @@ function publishUncheckedPresentation(root: string, file: string, bytes: Buffer)

function captureUnverifiedReport(inputs: UnverifiedReportInputs): ReportSnapshot {
validateSafeId(inputs.runId, "run ID");
const projection = projectCanonicalFinalReport({
...inputs.agentReport,
verification: inputs.verification,
observed_completion: inputs.observed
});
const projection = projectCanonicalFinalReport(
{
...inputs.agentReport,
run_metadata: withWholeRunSummary(
inputs.agentReport.run_metadata as Record<string, unknown>,
inputs.metadata,
inputs.state?.finished_at
),
verification: inputs.verification,
observed_completion: inputs.observed
},
{ goalSearchCoverage: inputs.goalSearchCoverage }
);
const jsonBytes = Buffer.from(`${JSON.stringify(projection.report, null, 2)}\n`, "utf8");
const markdownBytes = Buffer.from(projection.markdown, "utf8");
const generation = sha256Bytes(Buffer.concat([jsonBytes, markdownBytes]));
Expand Down
Loading
Loading