refactor(evals): delete the dead reporter/telemetry pipeline and publication leftovers - #1174
Conversation
… interface runEvalSuite always built an empty reporter list and resolveEvalProvider accepts only "none", so NodeTelemetryPump translated and "delivered" journal events to nobody. It still had to succeed on every poll: a drain that threw aborted the whole suite before run-summary.json was written. One schema-valid trigger is a node-synced transition without payload.attempt, which the event contract allows and sync emits when the runner reported no attempt; the pump rejected it with EVAL_TELEMETRY_EVENT_UNUSABLE. watchEvalRow is now: sync, read state.json, sleep. Deleted with the pump: EvalReporter and its graph/row types, guardReporter, reporterForReliableDelivery, the telemetry cursor (reader/writer, schema and common-schema definition), the eval telemetry/ directory, the graph.json read the reporter needed, the proper-lockfile dependency, and utils helpers only the pump used. boundedResponseText moves into scoring.ts next to the LLM judge, its only caller. The execution-policy fingerprint inputs are unchanged, so published history stays comparable. Co-Authored-By: Claude Opus 5.5 <[email protected]>
#1131 deleted the eval-benchmarks and eval-history-publication workflows, and docs/how-to/run-evals.md already says there is no automatic history publisher. Their producer and handoff checks stayed behind: history-publication.ts, its plan and generation schemas and semantic gates, and the `automatic` / `plan-rows` / `policy` commands of scripts/ci/prepare-eval-history-publication.mjs. Nothing runs them. The script keeps only the exports that validate-modal-benchmark-launch.mjs and prepare-modal-benchmark-cleanup.mjs import (manifest, pair-config and policy-checkout validation) and the helpers they reach; its CLI entry, the bundle-summary and plan-row code, and the TypeScript source-constant parser go. Tests of the deleted paths go with them; the strict-manifest reader test now targets readBenchmarkControlManifest, which still uses that reader. Co-Authored-By: Claude Opus 5.5 <[email protected]>
The flag and assertEvalHistoryRecency existed for the scheduled benchmark-history freshness monitor, which #1112 removed. No workflow, script or doc passes the flag any more. Co-Authored-By: Claude Opus 5.5 <[email protected]>
README.md embeds only the three overview charts (quality, latest-summary, performance-cost). The precision, recall, F1, cumulative-true-positive, wall-clock and cost SVGs were still generated and byte-checked by `eval history --check`, but nothing links to them. Their renderer and assets go. The three kept charts render byte-identically: `eval history --check` passes against the checked-in SVGs. Tests that used a per-metric chart to observe supersession or completeness now read the same commit links and completeness attributes from the overview charts; the tests of per-metric layout details go. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- read/write/parseEvalPublicationState, the publication-state schema and its common-schema definitions, and the EvalPublicationState types: left over from the retired `eval publish`; only tests used them. The CLI `json validate` eval-schema test now uses the ground-truth schema. - assertEvalReportAuthorityRemainedCurrent, isEvalNodeStatus, EVAL_EXPANSION_SOURCE_NODE_KEY and evmbenchSchemaPath: no caller in any package or script. - validateBenchmarkControlManifest and the testReportingPolicy test helper are only used inside their own modules, so they are no longer exported. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
… a row Deleting the telemetry pump also deleted the state.json re-read that followed the final drain. When the watch deadline passed during the last poll sleep, the row was then classified from the state read before that sleep. The watch is not the only writer of state.json: `ultrafuzz status`, `inspect` and the dashboard synchronize the run too. A run one of them made terminal during that sleep was recorded as timed-out, with an EVAL_ROW_WATCH_TIMEOUT diagnostic and a `running` workflow lifecycle, and counted as incomplete. origin/main recorded its real terminal status. Restore the single read, as origin/main had it. The new watchEvalRow test polls once with a 1 s deadline and a 1.5 s sleep, and a timer writes the terminal run 300 ms after that poll. It passes on origin/main (test ported with `reporters: []`), fails on the previous branch head with final_status "timed-out", and passes with this change. Co-Authored-By: Claude Opus 5.5 <[email protected]>
With the pump gone, eval runs write no telemetry, and the suite's `reporting.artifacts` policy has no reporter to stream to. Two docs still said eval runs retain telemetry, and the EvalArtifactPolicy and EvalTarget.sensitivity comments still described payload streaming and manifest-only artifact reporting. The docs and the default suite comment also said "nothing reads" the artifacts policy. That overstated it: normalizeReporting still computes it, and `eval plan --json` and the Modal worker still carry it. They now say no eval behaviour depends on it. The `sensitivity` comment now names what `private` does enforce: ground truth bound to the target's repo and ref, and ULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA before an LLM judge receives the data. Co-Authored-By: Claude Opus 5.5 <[email protected]>
| diagnostics.push(...finalDrain.warnings); | ||
| // Other syncers (`ultrafuzz status`, the dashboard) also write state.json; a | ||
| // run they finished during the final sleep is terminal, not timed out. | ||
| state = readStateSafe(runRoot, input.record.ultrafuzz_run_id); |
There was a problem hiding this comment.
When a poll has already observed a terminal run, this unconditional second read can overlap with another synchronizer replacing state.json. The snapshot reader then throws, so the watched row—and potentially the suite—fails instead of recording the terminal result. Keep the state already observed when it is terminal.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/evals/src/runner.ts
Line: 622
Comment:
**Terminal state can be lost**
When a poll has already observed a terminal run, this unconditional second read can overlap with another synchronizer replacing `state.json`. The snapshot reader then throws, so the watched row—and potentially the suite—fails instead of recording the terminal result. Keep the state already observed when it is terminal.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.| await startRowIfReady(true); | ||
| const finalDrain = await pump.drain(); | ||
| diagnostics.push(...finalDrain.warnings); | ||
| // Other syncers (`ultrafuzz status`, the dashboard) also write state.json; a |
There was a problem hiding this comment.
This comment names the dashboard as a writer of state.json, but the dashboard only reads run state. The new test repeats that example. It gives maintainers the wrong reason for the final read and makes the behavior harder to understand; use an actual state writer as the example.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/evals/src/runner.ts
Line: 620
Comment:
**Dashboard is not a syncer**
This comment names the dashboard as a writer of `state.json`, but the dashboard only reads run state. The new test repeats that example. It gives maintainers the wrong reason for the final read and makes the behavior harder to understand; use an actual state writer as the example.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
Every pull request in this batch inserts its entry at the same place in CHANGELOG.md, so each merge would conflict with the next. The entries are collected into one changelog update instead. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Problem
@ultrafuzz/evalsstill carried three pipelines that nothing uses. One of them could still end an eval run early.runEvalSuitealways built an empty reporter list (const reporters: EvalReporter[] = []), andresolveEvalProvideraccepts onlynone. SoNodeTelemetryPumptranslated and "delivered" journal events to nobody. It still had to succeed on every poll, though, because a faileddrain()rejectedrunEvalSuitebeforerun-summary.jsonwas written. One schema-valid trigger is anode-syncedtransition withoutpayload.attempt. The event contract makes that field optional (nodeSyncedPayloadSchema), and sync omits it when the runner reported no attempt (workflow-sync.tsnode-synced payload). The pump rejected such a transition withEVAL_TELEMETRY_EVENT_UNUSABLE.history-publication.ts, its plan and generation schemas, their semantic gates, and theautomatic/plan-rows/policycommands ofscripts/ci/prepare-eval-history-publication.mjsremained after Remove paid benchmark CI and repair CI gates #1131 deleted the workflows that called them.eval history --max-age-days(withassertEvalHistoryRecency) remained after ci: remove benchmark history freshness monitor #1112 removed the freshness monitor that used it. Six per-metric history SVGs were rendered and byte-checked even thoughREADME.mdembeds only the three overview charts.Root cause
When the external reporter providers and the paid benchmark workflows were removed, the extension points and handoff code they needed were left in place. The pump in particular kept running its fail-closed journal, cursor and lock checks on every watch tick, with no consumer at all.
Change
Telemetry pump and reporter interface deleted.
watchEvalRownow repeats three steps: sync, readstate.json, sleep. After the loop it readsstate.jsononce more, as main did, so a run that another command (ultrafuzz status,inspect, the dashboard) synchronized to a terminal state during the last sleep is recorded as terminal, not timed out. Deleted with the pump:NodeTelemetryPump,EvalReporterand its row/graph/event types,guardReporterandreporterForReliableDelivery.telemetry/.graph.jsonread (for reporters'onRowStart), theproper-lockfiledependency, and the utils helpers only the pump used.boundedResponseTextmoves intoscoring.tsnext to the LLM judge, its only caller. It now takes the realResponsetype; one scoring test double changed from a plain object tonew Response(...).Automatic history publication deleted.
history-publication.ts, its two schemas and gates, and its test are gone.prepare-eval-history-publication.mjskeeps only the exports thatvalidate-modal-benchmark-launch.mjsandprepare-modal-benchmark-cleanup.mjsimport, plus the helpers those reach. Its CLI entry, the bundle-summary and plan-row code, and the TypeScript source-constant parser are gone.eval history --max-age-daysandassertEvalHistoryRecencydeleted, with their tests.The six per-metric charts and their renderer deleted. The three README charts render byte-identically. History tests that used a per-metric chart to observe supersession now read the same commit links from
quality.svg.Unused exports deleted:
read/write/parseEvalPublicationState, plus the publication-state schema and types (leftovers of the retiredeval publish). The CLIjson validateeval-schema test now uses the ground-truth schema.assertEvalReportAuthorityRemainedCurrent,isEvalNodeStatus,EVAL_EXPANSION_SOURCE_NODE_KEYandevmbenchSchemaPath.validateBenchmarkControlManifestand thetestReportingPolicytest helper are no longer exported; they are used only inside their own modules.Docs, the default suite comment and the
eval runhelp text no longer promise telemetry. Thereporting.artifactsdocs, suite comment and type comments now say the block is validated and recorded but that no eval behaviour depends on it; theEvalTarget.sensitivitycomment names whatprivatedoes enforce (ground truth bound to the target's repo and ref, andULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA=truebefore an LLM judge receives it). Net: +249 / −5,717 lines outside the SVGs, plus 4,643 deleted SVG lines.The execution-policy fingerprint inputs are unchanged.
reporting.node_telemetryandreporting.heartbeat_interval_secondsstill feed it, andEVAL_EXECUTION_POLICY_REVISIONis untouched, so published history cohorts stay comparable.Deliberately not built (and why)
watchEvalRowalready records the failure count, first and last failure times, and the last message on the row asEVAL_ROW_SYNC_FAILED, and--watch-timeout-secondsbounds the watch. An early stop would end the watch while the detached run may still be progressing. In the Modal public worker,eval runreturning leads straight to scoring and bundling (or to a failure), and the run root is removed with the sandbox (see thepublic-worker.tscomment on retained artifacts). So an early stop on a repeated but transient error would end a healthy campaign, which works against the stability priority. It would also make the decision depend on comparing error-message text. The permanent-failure cases come from sync failing on bookkeeping, and the fix for that belongs in sync itself.graph.jsonstill fails the watch; this PR does not change that. After polling,evalRunExpansion(readStaticNodeIds→assertPlannedGraph) validates the graph for every row that has astate.json, andclassifyRecoveryEquivalencedoes so for terminal rows. Either throw still rejectsrunEvalSuitebeforerun-summary.jsonis written. On main the pump's graph read rejected it earlier, during polling; now it happens after polling. Making those readers tolerant is a separate stability change. The deleted watch-level test rejects a present graph that would require telemetry repair is not restored: its subject (telemetry repair) is gone, and the readers' rejection is covered byexpansion.test.ts› rejects absent or malformed state and graph evidence andrecovery-equivalence.test.ts› rejects contradictory graph kind and model fanout.reporting:block is kept. The suite schema is closed, so removing fields would reject existing suite YAML.node_telemetrystill selects the watch default, and two of the fields feed the execution-policy fingerprint.experiment_prefixandartifactsare now documented as validated, with no eval behaviour depending on them.--provider/resolveEvalProvidershim is kept. Whether to remove it is a separate, contested decision.prepare-eval-history-publication.mjsis not renamed. This keeps the diff to deletions. Its remaining exports validate Modal benchmark control manifests.buildPairwiseChartstays exported, although it was listed as unused. It is called insidecharts.ts, and droppingexportpulls its existing 142-line body into the diff-limited strict lint.Verification
Discriminating test 1:
runner-publish.test.ts› publishes the run summary for a watched row whose skipped node synced without an attempt. It uses a terminal run whose journal has a contract-validnode-syncedskippedevent with nopayload.attempt.packages/evals/srcandpackages/evals/schemareset toorigin/mainand the new test in place,runEvalSuiterejects:EvalError: node transition evt-4444… must carry a positive canonical payload.attempt, raised fromNodeTelemetryPump.draininsidewatchEvalRow.launched: 1, incomplete: 0, andrun-summary.jsonis written.Regression test 2 (added after review):
runner-publish.test.ts› records the terminal status another writer reached while the watch slept past its deadline. One poll runs against arunningstate with a 1 s deadline and a 1.5 s poll interval, and a timer writes the terminal run 300 ms after that poll, during the sleep.origin/main(b6dd1da), with the test ported (main'swatchEvalRowalso needsreporters: []): passes.final_status: "timed-out"andworkflow: { status: "running", terminal: false }.syncCallswould be 0 and the test would fail loudly. If the read after the poll took more than 300 ms, the test would still pass but would stop discriminating.Also checked by a throwaway probe (not committed): on this branch an invalid
graph.json(a node withoutlogical_id) rejectswatchEvalRowwithplanned graph is schema-invalid, both for a terminal row and for a running row whose watch times out.What I ran, after
pnpm -w build:--testTimeout=180000 --maxWorkers=4because the host is shared; an earlier round under load 30–40 with the default 5 s timeout hit 9 timeouts and no assertion failures. The new test sets its own 20 s timeout because it sleeps 1.5 s.watchEvalRow, add a test, and edit comments and docs):eval-history.test.ts: pass.json-validate.test.ts› json validate recognizes the pinned eval schema…: pass.cli.test.ts› agent-owned bytes stay identical across validation, sync, aggregation, report, dashboard, and bundle reads: pass, 139 s.contracts.test.tspasses (32), first round.pnpm -w test:ci-scripts, first round: 94 pass, 1 fail. The failure issafe-archive-bun.test.ts› repeatedly extracts a synthetic archive without stalling Bun callbacks: 2,000 extractions inside its own 15 s budget. It imports onlypackages/modal/src/safe-archive.ts, which this PR does not touch. The same test also timed out when I ran it in a pristineorigin/mainworktree on this host.pnpm -w benchmark:check:prebuilt(eval history --check), first round: passes against the checked-inquality.svg,latest-summary.svgandperformance-cost.svg(192 observations). The kept charts are therefore byte-identical.tsc --noEmitfor evals (cli, evmbench and modal in the first round).npx prettier --checkandnpx eslinton every changed file.CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci.pnpm -w knip.node scripts/docs-check.mjs.pnpm install --frozen-lockfile --offlinewith the updated lockfile (first round).Risk / compatibility
ultrafuzz eval history --max-age-daysis gone, and oclif rejects the flag.eval historyrenders 3 SVGs instead of 9. Render mode does not delete old SVGs in other checkouts; append mode already replaces the whole charts directory.telemetry-cursor:1,publication-state:1,history-automatic-publication-plan:1andhistory-publication-generation:1.json validateno longer recognizes them, and the eval schema bundle digest changes. Artifact schemas andvalidatorBuildIdentity()are untouched: that identity hashes only@ultrafuzz/artifactsvalidator modules and ajv versions. So in-flight campaign runs are unaffected; I checked this by reading the code, not by running an old run.graph.jsonstill rejects the watch, now after polling instead of during it (see Deliberately not built).state.jsonre-read after the loop is kept, so the terminal-versus-timed-out classification matches main.state.jsonswapped in mid-watch now fails withreadStateSafe's message ("run state identity does not match run …") instead of the pump's.<eval-run>/telemetry/*.cursor.jsonfiles are ignored.prepare-eval-history-publication.mjsno longer has a command-line entry. No workflow, script or doc invoked it.packages/cli/test/cli.test.ts(~2582) as this PR's test retitle. Whichever merges second keeps w21's(t)/tempProject(t)and this title.Refs #462
🤖 Generated with Claude Code
The PR does not appear safe to merge while the outstanding watch-state race can prevent a terminal row from being recorded.
Fix with agent prompt
Summary
The PR removes the unused eval telemetry/reporter pipeline, automatic history-publication code, retired schemas, the history recency flag, and six per-metric charts. It also restores a final state read after watch polling. The latest change removes the changelog entry for these compatibility changes.
Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."