Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
7da4250 to
1d882e8
Compare
8392432 to
6d7729f
Compare
This comment has been minimized.
This comment has been minimized.
6d7729f to
7734516
Compare
d314dae to
f320e18
Compare
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit f320e18. Configure here.
| }, | ||
| runInfo: data.runInfo, | ||
| artifacts: data.artifacts.map((a) => a.name), | ||
| artifacts: artifactNames, |
There was a problem hiding this comment.
Correlation defaults to missing metric
Medium Severity
The public views route still defaults corrYMetric to temperature. buildCorrelationData now drops samples where that metric is absent, so power-only DCGM series return an empty correlation plot. The explorer already falls back to another collected metric for this case.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit f320e18. Configure here.
Register the migrate, verify and telemetry backfill admin scripts and pull in the PGlite and Blob fixture dependencies the new DB and API tests use. 中文:注册 migrate、verify 与遥测 backfill 管理脚本,并引入新 DB/API 测试所需的 PGlite 与 Blob fixture 依赖。
Migration 016 stores one series per (run, artifact, CSV) with full-resolution samples, per-GPU statistics and benchmark point links; 017 adds stats_version so digests can be upgraded without rescanning samples. verify-db and the shared table constants cover the new tables. 中文:迁移 016 按 (run, artifact, CSV) 存储序列、全分辨率样本、每 GPU 统计及 benchmark 点链接;017 增加 stats_version 以便升级摘要而无需重扫样本。verify-db 与共享表名常量覆盖新表。
Parse NVIDIA/AMD CSVs and multinode power_audit bundles into series, samples and statistics, replace a changed series atomically under a row lock, and count telemetry failures separately so they never fail the benchmark ingest. 中文:在 ingest 时把 NVIDIA/AMD CSV 与多节点 power_audit 包解析为序列、样本和统计,在行锁内原子替换变更的序列,遥测失败单独计数、不影响 benchmark 入库。
Write a per-run telemetry receipt (expected, produced, stored, API-readable points), widen required-power publication to 1K/1K, and verify sweep manifests, point evidence and curve preservation before a run is treated as published. 中文:为每个 run 写遥测回执(预期/产出/入库/API 可读点),required-power 发布范围扩展到 1K/1K,并在视为已发布前校验 sweep manifest、点证据与曲线保留。
Backfill historical runs from GitHub artifacts before retention expires, skipping purged points; purge telemetry alongside run overrides with explicit counts; refresh recovered power_audit fields and invalidate the cache after a backfill. 中文:在 artifact 过期前从 GitHub 回填历史 run 并跳过已清理点;run overrides 清理时同步删除遥测并显式计数;回填后刷新恢复的 power_audit 字段并失效缓存。
Read a run or point with prefix/source/host scoping pushed into SQL, page samples in 50k-row keyset pages under the Neon HTTP cap, re-check the series version after loading so a concurrent re-ingest is retried, and expose a revision hash for the point cache. 中文:按 prefix/source/host 在 SQL 侧限定读取范围,样本以 5 万行 keyset 分页控制在 Neon HTTP 上限内,读取后复核序列版本以重试并发重入库,并为 point 缓存暴露 revision 哈希。
Run admin:db:migrate before ingest and overrides so writers never meet a missing column; let agentic ingest target a Neon branch; only invalidate the cache after overrides actually applied. 中文:ingest 与 overrides 前先执行 admin:db:migrate,避免写入端遇到缺列;agentic ingest 可指向 Neon 分支;仅在 overrides 实际生效后失效缓存。
Serve /api/gpu-metrics from the database first and fall back to live GitHub artifacts only when a run is not ingested; add /api/v1/gpu-metrics-point behind a revision-keyed Blob cache; keep the public views route shape identical for stored and live runs and normalize live NVIDIA timestamps to ISO UTC. 中文:/api/gpu-metrics 优先读数据库,仅在 run 未入库时回退到 GitHub artifact;新增以 revision 为键的 Blob 缓存 /api/v1/gpu-metrics-point;公开 views 路由对入库与 live run 保持同一形状,live NVIDIA 时间戳归一为 ISO UTC。
The /gpu-metrics explorer reads stored series, shows full-record statistics from the ingest digest, adds points/rolling display modes with a mean-of-chips line, and refetches on focus so a re-ingest shows up in an open tab. 中文:/gpu-metrics explorer 读取入库序列,展示 ingest 摘要的全记录统计,新增点/滑动平均显示模式与芯片均值线,并在窗口聚焦时刷新以反映重入库。
Document the telemetry digest, receipts, migration prerequisites, targeted repair and rollback, the gpu-metrics views route, and the new API route entries. 中文:记录遥测摘要、回执、迁移前置条件、定向修复与回滚、gpu-metrics views 路由及新增 API 路由条目。
中文:将博客列表页的访问放入每个用例的准备阶段,使 Cypress 能在页面加载超时时重试。
中文:博客列表页的 E2E 用例统一替换图片优化请求,避免 Firefox 等待缩略图时页面加载超时。
f320e18 to
0181958
Compare
…h ruler #1167 was rebased onto master twice today; #1220 had merged the 13:32 rebase (f320e18), which still carried the date-comparison Perf Ruler, while the 14:19 rebase (0181958) removed it after #1232 landed on master. Merge the new parent with the earlier rebase as the base, so only GPUGraph.tsx and its component spec conflict, and resolve them by removing the ruler: the #1229 run-specific ruler share binding, the legend toggle, the hit strokes and drag handlers, and the ruler share-link spec. `ScatterGraph` keeps `i_rulers`; docs and the store comment now say the date comparison draws no rulers. 中文:#1167 今日两次 rebase 到 master;#1220 已合入 13:32 的 rebase (f320e183),其中仍含日期对比图的性能标尺,而 14:19 的 rebase(0181958d) 在 #1232 合入 master 后移除了它。以较早的 rebase 为 base 合并新父分支, 仅 GPUGraph.tsx 及其组件用例冲突,并按移除标尺解决:删除 #1229 的运行级 标尺分享绑定、图例开关、命中描边与拖拽处理,以及标尺分享链接用例。 `ScatterGraph` 保留 `i_rulers`;文档与 store 注释改为日期对比图不绘制标尺。


Summary
GitHub deletes benchmark telemetry artifacts after 90 days and the dashboard re-downloads and re-parses them on every view, so older PowerX curves are gone and recent ones are slow. This PR ingests the telemetry into the database once, serves it DB-first with GitHub artifacts as the fallback and backfill source. The dashboard views built on that stored data follow in #1237, stacked on this branch.
flowchart TB subgraph Before["Before"] direction LR B1["Page"] --> B2["GitHub artifacts"] --> B3["Parse CSV"] --> B4["Display"] B2 -.-> B5["Expired: unavailable"] end subgraph After["After"] direction LR A1["Artifacts"] --> A2["Ingest"] --> A3["Persistent DB"] A4["Page"] --> A5["DB-first API"] A3 --> A5 --> A6["Display"] A5 -. "Missing / incomplete" .-> A7["GitHub fallback"] A7 --> A6 endDB failures return errors. Failed background refreshes preserve cached telemetry; public statistics omit the storage-only metric field. Live artifact reads normalize NVIDIA timestamps to ISO UTC like stored reads.
AI model disclosure
GPT-6: implementation/delegated verification/description; exact variant unavailable. claude-opus-5: review. Earlier contributors' exact versions unavailable.
Validation
I have completed the AI model disclosure and kept it current
Current head 7734516: the DB/ETL/API/explorer half of the former 6d7729f tree as 10 feature commits (read bottom-up from schema to docs); the dashboard views are [PowerX] power boundary metrics, comparison overlays, point view and timeline / 功耗边界指标、对比叠加、逐点视图与时间线 #1237. Local on this head: typecheck, lint, format and typography pass; app 6,001 unit tests pass; DB 881 pass with one pre-existing OperatorX PGlite timeout.
Test pruning (8c36086, perf-ruler part reverted in ad05dce): removed 38 redundant PowerX unit cases (types tests here; measured-metric-config and chart-utils tests in [PowerX] power boundary metrics, comparison overlays, point view and timeline / 功耗边界指标、对比叠加、逐点视图与时间线 #1237); 22 added regressions. On unchanged production code across a fixed 91-file scope, aggregate coverage lost at most 0.11 percentage points. The 27 perf-ruler deletions moved to [Tests] remove redundant perf-ruler cases / 删除冗余的 perf-ruler 测试用例 #1235.
Earlier recorded checks: local DB/browser 12/12, context checks, mixed-source browser 2/2, and official/unofficial fixtures passed on earlier revisions.
Limits: hosted acceptance and production repair remain unexecuted. The local audit reported 23 dependency advisories; dependency manifests and lockfile were unchanged. Real Neon response size/latency was not measured.
中文说明
GitHub 在 90 天后删除 benchmark 遥测 artifact,而仪表板每次查看都要重新下载、解析,导致较早的 PowerX 曲线消失、较新的曲线加载缓慢。本 PR 把遥测一次性入库,API 优先读数据库、GitHub artifact 退为回退与回填来源。建立在入库数据之上的仪表板视图见堆叠于本分支的 #1237。上图对比旧版按页面请求下载、解析 artifact,以及新版入库后优先读取 DB 的路径。缺失或不完整的数据可尝试 GitHub fallback;DB 查询故障会明确报错。后台刷新失败时保留已缓存的遥测图表,公开统计不再返回数据库内部的 metric 字段。live artifact 读取与入库路径一样将 NVIDIA 时间戳归一为 ISO UTC。
GPT-6 负责实现、委派验证及描述修改,精确变体无法核实;claude-opus-5 负责审查。早期贡献者的精确模型版本无法核实。
Note
High Risk
Touches production ingest, schema migrations, benchmark publication preflight, and CDN invalidation workflows; incorrect telemetry or manifest gating could block sweeps or serve stale dashboard data.
Overview
PowerX telemetry is ingested into the database and served DB-first, with GitHub artifacts kept as fallback/backfill so curves survive 90-day retention and repeat views avoid re-parsing zips.
Ingest and CI: New
gpu_metric_*storage (migrations 016/017) digests CSVs and multinodepower_auditbundles at ingest, links series to benchmark points, and versions full-record per-GPU statistics. Workflows gain optionalrequire-power/INGEST_REQUIRE_POWER, stricter required-power manifest v2 preflight (shared fixture), telemetry receipts, andadmin:db:backfill-gpu-metrics(including--stats-only). apply-run-overrides runs migrations before verify and still invalidates/warms the CDN when overrides apply successfully.APIs:
/api/gpu-metricsreads stored telemetry first (series=powerTimeline via read-only POST +sourceCoverage), returns503for DB outage or known incomplete storage, and merges healthy DB windows with artifact recovery when allowed./api/v1/gpu-metrics-pointexposes per-point samples/stats with revision-checked Blob cache andno-store. The public gpu-metrics view prefers stored digests for the full-record stats table (startup/warmup included).UI: PowerX explorer adds rolling-average / mean-of-chips display controls, overlay axis support, refetch-on-focus, and stats scoped to stored digests when present. Docs, fixtures, PGlite acceptance tests, and Cypress tweaks support recovery and contract checks.
Reviewed by Cursor Bugbot for commit 0181958. Bugbot is set up for automated code reviews on this repo. Configure here.