Skip to content

fix(runtime): reports use whole-run metadata and the renderer accepts link syntax and legacy labels in prose - #1167

Merged
aviggiano merged 9 commits into
mainfrom
claude/w11-report-accuracy
Sep 29, 2026
Merged

aviggiano merged 9 commits into
mainfrom
claude/w11-report-accuracy

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Issue #1151 lists several report problems. Two of them are real, and one of those can stop a run from producing any report.

  1. The run summary under-reports the run. The Run summary in a published report is the accounting the host captured when the report task started. Here is a local smoke run (baseline-main-1, tiny-vault):

    Source Elapsed Tokens Spend
    Agent report and runtime publication 20m 17s 2,363,772 $5.64
    Final run.json / state.json 36m 23s 4,591,556 $10.17
  2. The canonical renderer rejects its own output. Ordinary upstream content makes projectCanonicalFinalReport throw. Examples, all checked on origin/main:

    • [spec](https://…), ![x](…) or handlers[id](payload) in a description (disallowed Markdown link / embedded image);
    • an issue title with a non-ASCII letter (Δ, ï), because the link rule's anchor allowlist does not match the renderer's own \p{L} anchors;
    • the public projection of /srv/…/Repro.t.sol(line 12), which redaction turns into [redacted-path](line 12);
    • legacy-looking labels: a family variant titled Item 1 or Source Node Id, a blocker summary or description of exactly Critical, or a blocker summary Strategy loops: 4.

    The review gates require those fields to be byte-identical to upstream, so the agent cannot repair them. The final-report gate calls the renderer on the agent's report.json, so every retry fails the same way and the run ends with no report.

  3. The prompt contradicts the renderer. final-report.md tells the agent to write a ## Audit context section of relative links and says ultrafuzz report regenerates it. The renderer code for that section is dead: report@3 has no audit_context field. report.md must also equal the canonical render byte for byte, so the agent can only ignore the instruction or fail by following it.

  4. Unchecked fallback reports always print "Goal search coverage is unknown", even when goal-search-coverage.json exists.

Root cause

  • (1) run_metadata is owned by the host but passes through the LLM when the report task starts. The runtime presentations (the verified terminal publication and the unchecked report) then reused that stale copy. Both already load the final run.json and state.json.
  • (2) The renderer re-checks its own deterministic output with regex rules written for Markdown that agents used to write by hand. publicProse did not neutralize ](, and the legacy-label rules match escaped prose that the renderer wraps in list items or bold.
  • (4) captureUnverifiedReport never passed the census to the renderer.

Change

  • withWholeRunSummary() in terminal-report-projection.ts, applied in projectTerminalReport and in captureUnverifiedReport. It restates:

    • elapsed_time, from run.json#created_at to state.json#finished_at, in the host's format;
    • models_used, from accounting.cumulative.models when that list is non-empty;
    • tokens_used, estimated_spend and partial_pricing as one group, whenever accounting.cumulative records a token count. A whole-run spend recorded as unavailable stays unavailable. It is not replaced by the agent's copy, which can come from the report-start Smithers fallback and price only part of the run.

    Without those records the agent's copy is kept. The function never throws. The agent's own report.json and report.md are not touched.

  • publicProse escapes ( after ] (]( becomes ]\(), so prose can no longer form an inline link or image (! was already escaped).

    • I used ]\( rather than \]( because the inserted backslash then never sits next to another backslash.
    • With \](, a prose backslash right before ]( becomes \\\](. The public projection's UNC-path pattern flags that as a private path.
    • With \](, the key-name secret rule's unquoted value stops at ], so token=REDACTED\]( would redact again on the next pass.
  • The renderer's self-check keeps only its raw-HTML rule for prose. I deleted finalReportProseDirectiveViolation: the image and link rules, and the legacy-label rules (Critical, Item N, Source Node Id, legacy run-metadata lines). Three of the deleted rules (#### Sources, legacy report sections, legacy issue subsections) could never match rendered output, because publicProse escapes #. The structural checks (title, completion census, coverage headings, issue order) are unchanged. Removing checks changes no rendered bytes.

  • Warning codes escape ](. artifact_validation_warnings[].code is free-form and was the one free-form field rendered as plain text without publicProse. Once the link rule was gone, a code like [notice](https://…) rendered as a live link, in the report and in the public warnings companion. Only ]( is escaped there, so real gate codes keep their bytes. Full publicProse would escape the _ in ARTIFACT_OPTIONAL_METADATA_MISSING.

  • Deleted dead code: appendAuditContext, safeReportLink, SAFE_REPORT_RELATIVE_LINK_PATTERN, and the prompt's ## Audit context example, paragraph and checklist line. The schema rejects unknown keys, so this code could never produce output.

  • Unchecked reports read goal-search-coverage.json through loadGoalSearchCoverageSnapshot. An unreadable census still renders as unknown coverage instead of failing the report.

  • Docs and CHANGELOG.

    • artifacts-reports.md: removed the false "concise links to THREAT_MODEL.md…" claim and documented what runtime presentations restate.
    • CHANGELOG: an Other-changes entry, plus a Breaking-changes entry for the upgrade effect (see Risk).

Source diff: +91 / −62 lines in packages/runtime/src; the prompt loses 19 lines.

Deliberately not built (and why)

  • No schema change, provenance vocabulary, context annexes or Markdown lint in CI. Each summary field has one durable source, and the + suffix already marks partial pricing.
  • No re-publication of terminal publications written by earlier versions (see Risk). Sync would have to re-verify and re-publish stored publications, which is another reconciliation layer, and VERIFIED_OUTPUT_CHANGED is also the signal for changed evidence. The root-level fix is to stop byte-comparing re-rendered Markdown on read (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921). That call is the owner's.
  • No fix for the index anchors of titles containing _ or < (the renderer drops _; GitHub keeps it). Fixing it changes canonical bytes, so older reports with such titles would stop re-verifying.
  • The host→LLM run_metadata copy in workflow.tsx is unchanged. Removing it needs a schema and prompt change, and several parallel items own that file. Follow-up.
  • The raw-HTML rule still applies inside inline code. Inline-code values (run summary values, paths, source nodes) are not escaped for <. A model name such as <model> in accounting.cumulative.models would still make the renderer throw. This is contrived; model names come from TokenUsageReported events.
  • The CLI reportAccountingDiagnostics is not deleted. It is now redundant for runtime presentations but still inspects agent-only reports of running runs. Follow-up.
  • Brackets in general are not escaped, nor ]:. That would change the bytes of every report with [, ] or arr[i]: in its prose. A prose line starting [x]: https://… therefore still forms a link reference definition, as it does on main.
  • Not addressed, same on main: the public projection still rejects prose with a space before a backslash ( \d), because publicProse doubles the backslash into the UNC private-path pattern.

Verification

Discriminating tests. I compiled this branch's test files against other source trees and ran them:

Source tree Result (final-report-markdown, terminal-report-projection, unverified-report: 60 tests)
origin/main (b6dd1da9) 9 fail: every new or updated test below
previous PR head (628f4c43) 3 fail: the warning-code, legacy-label and unpriced-spend tests added in this revision
this branch 60 pass
Test Result on main
final-report-markdown: upstream prose with link or image syntax renders as literal text throws contains an embedded image outside fenced code
final-report-markdown: artifact validation warning codes with link or image syntax render as literal text throws embedded image (on 628f4c43: renders a live link and image)
final-report-markdown: upstream prose that reads like a legacy report label still renders throws contains the unsupported Critical severity. Probed one at a time, Item 1, Source Node Id and Strategy loops: 4 each throw their own rule.
final-report-markdown: issue titles with non-ASCII letters render with index anchors that resolve to their headings throws contains a disallowed Markdown link outside fenced code
final-report-markdown: public projection keeps a redacted path followed by a parenthesis as literal text throws the link violation
terminal-report-projection: terminal projection restates whole-run accounting instead of the report-start snapshot renders 2m, example-model, 100, $0.01. On 628f4c43, the unpriced case renders $0.01 next to whole-run tokens.
unverified-report: unchecked reports restate whole-run accounting from run.json renders 1m and unavailable
unverified-report: unchecked reports render the run's goal-search census renders "Goal search coverage is unknown"

The renderer assertions parse the output with the existing mdast-util-from-markdown dependency, so they check what a reader sees rather than escape bytes. The public-projection test also covers a prose backslash before ]( and token=…](x). I checked that it fails when the escape is temporarily changed to \](, the form the ]\( choice avoids.

Fixture realism. Before this revision, the sparse accounting case used a cumulative block that the runtime's accounting validator rejects (cumulative.estimated_spend does not match estimated_spend_usd). I checked both current fixtures, priced and unpriced, through assertRunMetadataAccountingUsageAuthority. Both pass storedAccountingDocument and stop only at the usage-ledger comparison, which needs real ledger entries.

Updated tests that pinned the report-start snapshot. Each now accepts the runtime elapsed_time and still deep-equals the agent report everywhere else:

  • terminal-report-projection, preserves all verified review data…;
  • verified-output, terminal presentation discloses tolerated failures…;
  • runtime.test, syncRun publishes a report after recovery of a failed report agent.

The three hand-mutated-Markdown tests that used the deleted rules as probes (directive validation rejects injected…, …treats fenced proof code as code…, …recognizes CommonMark tilde fences…) now probe the raw-HTML rule through the same fence handling. They pass on all three source trees.

Real-run probes on baseline-main-1. These were read-only: every file's size and mtime was unchanged afterwards.

  • The restated summary is 36m 23s, 4,591,556 tokens and $10.17, which matches run.json and state.json.
  • loadCurrentFinalReportSnapshot reads the stored publication as verified-runtime-report with main's source and throws VERIFIED_OUTPUT_CHANGED on this branch (see Risk).
  • The earlier revision also checked that the branch renderer reproduces the agent's report.md byte for byte.

Suites run on this revision:

  • runtime:
    • whole files: final-report-markdown, terminal-report-projection, unverified-report, verified-output, terminal-report-completion, report-publication-status (139);
    • filtered by name: artifact-gates final-report tests (11), generated-workflow-verifier report (19), runtime.test syncRun publishes a report|syncRun distinguishes available agent reports (4).
  • CLI: 11 report, render and bundle tests from cli.test and cli-contracts, including report accepts populated accounting snapshots… and agent-owned bytes stay identical… (11 pass).
  • dashboard: tests matching report (8).
  • modal:
    • public-bundle: 27 pass.
    • public-worker: 60 of 62 pass in the full-file run under load average ~18. The other two (recognizes and cleans the legacy persistent public workspace… and rejects a present dangling public bundle…) hit vitest's 30 s timeout, then passed when run alone with a longer timeout.
  • checks: prettier, eslint, lint:strict:ci (diff-limited), pnpm --filter @ultrafuzz/runtime typecheck, pnpm -w knip, docs-check, prompt-catalog-docs --check, audit-profile-docs --check, git diff --check.
  • The earlier revision also ran prompts render, semantic-anchors and scaffold-catalog (93), and the CLI run, ps, status, inspect, report, … lifecycle test. This revision does not touch prompts.

Not run: the full runtime and CLI suites, the Bun adapter contracts (untouched), and a live campaign.

Risk / compatibility

  • Breaking: terminal publications written by earlier versions no longer verify after upgrading. Readers re-render and compare bytes (Re-evaluate sealed execution snapshots: cost/benefit after repeated campaign losses #921), and the restated summary (in practice at least the elapsed time) differs from the stored bytes. For such a run:

    • ultrafuzz report and the dashboard serve the unchecked presentation. A complete run shows as # Ultrafuzz report — PARTIAL, not-checked.
    • report bundle falls back to a report-only archive of that pair (read from the code), and report --require-verified, report bundle --require-verified and the Modal public worker fail.
    • Nothing re-publishes it: sync publishes only while no publication status exists for that terminal state.

    Publications written by this version verify normally. This is filed under Breaking changes in the CHANGELOG.

  • Agent reports: Markdown bytes change only where prose contains ](. Of those forms, main accepted only \]( and ]( before an escape-free ../ path; it rejected ](#anchor) and ../ paths containing _, because publicProse escapes # and _. Existing agent reports with an accepted form no longer re-verify. I checked that prose without ]( renders byte-identically on main and on this branch. Warning codes change bytes only if a code contains ](; real codes are host gate identifiers. Eval scoring reads the agent snapshot, so only those rare reports are affected there.

  • Prompt edit: new runs get a new prompt digest. The earlier revision states that in-flight runs keep their sealed prompt snapshots; I did not re-verify that.

  • No schema, contract-description or validator-build change.

Closes #1151

🤖 Generated with Claude Code

RetriggerConfidence Score: 4/5

The PR is not safe to merge while older terminal publications lose verified access after upgrade.

Fix All in Claude CodeFindings

  1. P1 Older verified reports stop verifying ▶
Fix with agent prompt
### Issue 1
packages/runtime/src/terminal-report-projection.ts:undefined-54
For a terminal publication written before this change, the restated summary can differ from its stored bytes. Reads regenerate the report and reject that publication on a byte mismatch, so an unchanged run falls back to an unchecked report and `--require-verified` access fails.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR restates whole-run metadata in runtime report presentations, makes the canonical renderer accept more upstream prose, and supplies goal-search coverage to unchecked reports. The latest revision removes both corresponding changelog entries.

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

aviggiano and others added 4 commits September 28, 2026 22:58
…tions

The final-report agent copies a run summary that the host captures when the
report task starts, and both runtime presentations (the verified terminal
publication and the unchecked fallback report) republished that copy
unchanged. Elapsed time, models, tokens and spend therefore excluded the
report task itself and everything that finished after it started. On a local
smoke run the published report showed 2,363,772 tokens, $5.64 and 20m 17s,
while run.json recorded 4,591,556 tokens and $10.17 and the run finished
after 36m 23s.

projectTerminalReport and captureUnverifiedReport already read the final
run.json and state.json, so withWholeRunSummary() restates elapsed time
(run.json created_at to state.json finished_at) and models, tokens, spend
and partial_pricing from accounting.cumulative. A missing, malformed or
"unavailable" value keeps the agent's copy, and the helper never throws.
The agent's own report.json and report.md are untouched.

Unchecked reports also never received the run-root goal-search census, so
they always said "Goal search coverage is unknown". They now read
goal-search-coverage.json the same way the verified path does; an unreadable
census still renders as unknown coverage instead of failing the report.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…tput

projectCanonicalFinalReport re-validates the Markdown it just rendered with
regex rules written for hand-authored reports. publicProse did not escape
`](`, so upstream prose that the review gates require byte-for-byte (a spec
link, an image, `handlers[id](payload)`, a relative link whose `_` gets
escaped) tripped the link or image rule. The link rule also rejected the
renderer's own index anchors for any issue title with a non-ASCII letter,
and the public projection threw when a redacted path was followed by `(`
(`[redacted-path](line 12)`). The agent cannot repair byte-preserved fields,
so every retry failed the same way and the run ended without a report.

publicProse now escapes `(` after every `]`, so prose cannot form an inline
link or image, and the image and link rules are deleted; the raw-HTML rule
stays. Rendering of prose without `](` is unchanged, so existing reports
keep verifying.

The audit-context renderer (appendAuditContext, safeReportLink,
SAFE_REPORT_RELATIVE_LINK_PATTERN) is deleted: report@3 has no audit_context
field and rejects unknown keys, so it could never run. The prompt's
`## Audit context` instructions, which asked the agent to write a section
that the byte-exact canonical render never contains, are removed.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The artifacts reference claimed report.md links to THREAT_MODEL.md,
threat-model.json and goal-plan.json (only dead code ever did) and that
terminal presentations keep the report-start accounting snapshot. Document
what runtime presentations now restate, that prose link syntax renders as
literal text, and the upgrade effect on existing terminal publications.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…bracket

Escaping `](` as `\](` would break the public projection on two shapes:
after publicProse doubles a prose backslash, `\\\](` matches the UNC
private-path pattern, and `token=REDACTED\](` extends the redacted
assignment value so the fixed-point re-scan redacts it again. Both public
projections succeed with the `]\(` escape; with `\](` this test fails.

Refs #1151

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 23:10
Comment thread packages/runtime/src/terminal-report-projection.ts
if (sourceRunId !== undefined && Reflect.get(reportMetadata, "source_run_id") !== sourceRunId) {
throw new Error("Verified final report has a different source run identity");
}
report.run_metadata = withWholeRunSummary(reportMetadata as Record<string, unknown>, metadata, state.finished_at);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Older verified reports stop verifying For a terminal publication written before this change, the restated summary can differ from its stored bytes. Reads regenerate the report and reject that publication on a byte mismatch, so an unchanged run falls back to an unchecked report and --require-verified access fails.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/src/terminal-report-projection.ts
Line: 54

Comment:
**Older verified reports stop verifying** For a terminal publication written before this change, the restated summary can differ from its stored bytes. Reads regenerate the report and reject that publication on a byte mismatch, so an unchanged run falls back to an unchecked report and `--require-verified` access fails.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

Comment thread packages/runtime/src/final-report-markdown.ts Outdated
aviggiano and others added 3 commits September 29, 2026 02:46
…ogether

withWholeRunSummary restated tokens_used from run.json's cumulative
accounting but kept the agent's estimated_spend and partial_pricing
whenever the cumulative spend was "unavailable". For a direct run, that
agent spend can come from the live Smithers fallback taken when the
report task started, which may price only part of the run. The runtime
presentation then showed whole-run tokens beside a report-start spend,
with no partial marker.

Tokens, spend and partial pricing now move as one group: when the
cumulative record carries a token count, all three come from it, and a
whole-run spend recorded as "unavailable" stays "unavailable". Without
cumulative accounting, all three keep the agent's copy as before.

The sparse test case pinned the mixed summary, using a cumulative block
(estimated_spend "unavailable" beside estimated_spend_usd 41.2) that the
runtime's accounting validator rejects. It now uses a ledger that priced
no event, and the fixture's pricing markers carry the component field
the validator requires. The fixture passes storedAccountingDocument (via
assertRunMetadataAccountingUsageAuthority, which stops only at the
ledger comparison).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…in warning codes

The canonical renderer re-checks its own output. Its remaining prose
rules were written for Markdown that agents used to write by hand, and
they still matched byte-preserved upstream content. Each of these made
projectCanonicalFinalReport throw, on this branch before this commit and
on main:
- a family variant titled "Item 1" (legacy numbered-item field);
- a family variant titled "Source Node Id" (legacy source identifier);
- a blocker summary "Critical" (unsupported Critical severity);
- a blocker summary "Strategy loops: 4" (legacy run metadata).
The final-report gate requires those fields to equal upstream, so the
agent cannot repair them and every retry fails the same way. The other
three rules (#### Sources, legacy report sections, legacy issue
subsections) could not match rendered output at all, because
publicProse escapes `#`. The self-check now keeps only its raw-HTML
rule. Removing checks does not change any rendered bytes.

artifact_validation_warnings[].code is free-form (1-128 characters) and
is the one free-form field rendered as plain text without publicProse.
With the link and image rules gone, a code such as
`[notice](https://example.com/x)` rendered as a live link, in the report
and in the public warnings companion. Only `](` is escaped there, so
real gate codes such as ARTIFACT_OPTIONAL_METADATA_MISSING keep their
bytes (full publicProse would escape their underscores).

The three hand-mutated Markdown tests that used the deleted rules as
probes now probe the raw-HTML rule through the same fence handling.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
… changes

Terminal report publications written before this release no longer
verify, because readers re-render them and the restated run summary
differs from the stored bytes. `ultrafuzz report` and the dashboard then
serve a complete run as the unchecked PARTIAL presentation. `report
bundle` falls back to a report-only archive. `--require-verified` reads
and the Modal public worker fail. Sync does not re-publish, because a
publication status already exists for that terminal state. This is a
user-visible break, so it moves to "Breaking changes".

The old entry also named `](#anchor)` among the prose forms main
accepted. main rejected it, because publicProse escapes `#`. The forms
main accepted were `\](` and `](` before an escape-free `../` path.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano aviggiano changed the title fix(runtime): reports use whole-run metadata and the renderer stops rejecting its own output fix(runtime): reports use whole-run metadata and the renderer accepts link syntax and legacy labels in prose Sep 29, 2026
aviggiano and others added 2 commits September 29, 2026 02:51
The entry said runtime presentations restate the run summary from
run.json, but not that a value those records do not carry keeps the
agent's copy. docs/reference/artifacts-reports.md already describes that
fallback; the entry now says so too.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Generate accurate self-contained partial reports from durable metadata

1 participant