You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
ci(bench): grid benchmark reports false regressions, compare build time over repeated runs #1132
#1131 did not change ResultGridView / ResultsTab; the base itself went from 11.4 to 15.8 ms between two runs.
Causes
The compared metric is FrameTiming.totalSpan, which includes vsync wait and the raster queue. The Dart side is build (p50 0.42-0.47 ms in every run, no base/PR difference). raster is software rendering (Mesa llvmpipe under xvfb) on a shared VM.
One run per side, always PR first, then base: any runner spike lands entirely on one side.
Threshold too tight: >5% and >0.3 ms. p99 of ~300 frames is 3 frames.
stutters(>12.5ms) assumes 120 Hz, but xvfb presents at 60 Hz, so ~every frame (290/291) counts.
Proposal (benchmark.yml, scripts/ci/bench_compare.py; no app code)
Flag regressions on build p50 / p90 only; keep raster, total and p99 in the table as information.
Problem
.github/workflows/benchmark.ymlcomments "regression" on PRs that did not touch the measured code. Evidence:#1131 did not change
ResultGridView/ResultsTab; the base itself went from 11.4 to 15.8 ms between two runs.Causes
FrameTiming.totalSpan, which includes vsync wait and the raster queue. The Dart side isbuild(p50 0.42-0.47 ms in every run, no base/PR difference).rasteris software rendering (Mesa llvmpipe under xvfb) on a shared VM.stutters(>12.5ms)assumes 120 Hz, but xvfb presents at 60 Hz, so ~every frame (290/291) counts.Proposal (benchmark.yml, scripts/ci/bench_compare.py; no app code)
buildp50 / p90 only; keep raster, total and p99 in the table as information.Cost: roughly 4-5 extra minutes per run on the benchmark job.
Related: #1065 (introduced the workflow), #987.