Skip to content

[PowerX] persist telemetry and recover historical reads / 持久化遥测并恢复历史数据读取 - #1167

Open
cquil11 wants to merge 12 commits into
masterfrom
feat/powerx-db-ingest
Open

cquil11 wants to merge 12 commits into
masterfrom
feat/powerx-db-ingest

Conversation

@cquil11

@cquil11 cquil11 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

GitHub deletes benchmark telemetry artifacts after 90 days and the dashboard re-downloads and re-parses them on every view, so older PowerX curves are gone and recent ones are slow. This PR ingests the telemetry into the database once, serves it DB-first with GitHub artifacts as the fallback and backfill source. The dashboard views built on that stored data follow in #1237, stacked on this branch.

flowchart TB
    subgraph Before["Before"]
        direction LR
        B1["Page"] --> B2["GitHub artifacts"] --> B3["Parse CSV"] --> B4["Display"]
        B2 -.-> B5["Expired: unavailable"]
    end
    subgraph After["After"]
        direction LR
        A1["Artifacts"] --> A2["Ingest"] --> A3["Persistent DB"]
        A4["Page"] --> A5["DB-first API"]
        A3 --> A5 --> A6["Display"]
        A5 -. "Missing / incomplete" .-> A7["GitHub fallback"]
        A7 --> A6
    end
Loading
Area Before After
Storage Artifact-dependent Samples, provenance, statistics, benchmark links
Sources CSV viewer NVIDIA/AMD CSVs + multinode bundle windows
Inspection Run-level charts Stored series, source coverage; Timeline and per-point View PowerX are in #1237
Display Existing controls/statistics Smoothing controls; stored full-record statistics
Consistency Artifact reads Atomic series replacement; revision-aware point cache
Validation Existing ingest checks Completeness receipts; required-power manifest checks
Recovery Existing ingest paths Purge-aware telemetry backfill; statistics upgrades; migration-first overrides
Queries Artifact downloads and parsing SQL selectors; 50,000-row pages; selected-host views

DB failures return errors. Failed background refreshes preserve cached telemetry; public statistics omit the storage-only metric field. Live artifact reads normalize NVIDIA timestamps to ISO UTC like stored reads.

  • Follow-ups: unified default collection, automatic repair, route/component decomposition.
  • Operations: repair/rollback runbook. Apply migrations 016+017 before ingest or backfill.

AI model disclosure

GPT-6: implementation/delegated verification/description; exact variant unavailable. claude-opus-5: review. Earlier contributors' exact versions unavailable.

Validation

中文说明

GitHub 在 90 天后删除 benchmark 遥测 artifact,而仪表板每次查看都要重新下载、解析,导致较早的 PowerX 曲线消失、较新的曲线加载缓慢。本 PR 把遥测一次性入库,API 优先读数据库、GitHub artifact 退为回退与回填来源。建立在入库数据之上的仪表板视图见堆叠于本分支的 #1237。上图对比旧版按页面请求下载、解析 artifact,以及新版入库后优先读取 DB 的路径。缺失或不完整的数据可尝试 GitHub fallback;DB 查询故障会明确报错。后台刷新失败时保留已缓存的遥测图表,公开统计不再返回数据库内部的 metric 字段。live artifact 读取与入库路径一样将 NVIDIA 时间戳归一为 ISO UTC。

范围 改动前 改动后
存储 依赖 artifact 可用 保存采样、来源、统计及 benchmark 关联
来源 CSV 查看器 支持 NVIDIA/AMD CSV 及多节点 bundle 的测量窗口
查看 Run 级图表 已存储序列与来源覆盖情况;Timeline 与单点 View PowerX 见 #1237
展示 已有控件和统计 增加平滑控件,使用已存储的全采集周期统计
一致性 从 artifact 读取 单个 series 原子更新,单点缓存随数据版本刷新
验证 已有入库检查 增加完整性回执及 required-power manifest 检查
恢复 已有入库路径 遥测回填跳过已清理点,支持统计升级;覆盖操作前先执行 migration
查询 下载并解析产物 SQL 筛选、每页最多 50,000 行;公开视图只加载所选主机序列
  • 后续工作: 统一默认采集入口、自动修复、拆分过大的路由和组件。
  • 运维: 定向修复与回滚见上方 runbook;运行 ingest 或 backfill 前先执行迁移 016+017。

GPT-6 负责实现、委派验证及描述修改,精确变体无法核实;claude-opus-5 负责审查。早期贡献者的精确模型版本无法核实。


Note

High Risk
Touches production ingest, schema migrations, benchmark publication preflight, and CDN invalidation workflows; incorrect telemetry or manifest gating could block sweeps or serve stale dashboard data.

Overview
PowerX telemetry is ingested into the database and served DB-first, with GitHub artifacts kept as fallback/backfill so curves survive 90-day retention and repeat views avoid re-parsing zips.

Ingest and CI: New gpu_metric_* storage (migrations 016/017) digests CSVs and multinode power_audit bundles at ingest, links series to benchmark points, and versions full-record per-GPU statistics. Workflows gain optional require-power / INGEST_REQUIRE_POWER, stricter required-power manifest v2 preflight (shared fixture), telemetry receipts, and admin:db:backfill-gpu-metrics (including --stats-only). apply-run-overrides runs migrations before verify and still invalidates/warms the CDN when overrides apply successfully.

APIs: /api/gpu-metrics reads stored telemetry first (series=power Timeline via read-only POST + sourceCoverage), returns 503 for DB outage or known incomplete storage, and merges healthy DB windows with artifact recovery when allowed. /api/v1/gpu-metrics-point exposes per-point samples/stats with revision-checked Blob cache and no-store. The public gpu-metrics view prefers stored digests for the full-record stats table (startup/warmup included).

UI: PowerX explorer adds rolling-average / mean-of-chips display controls, overlay axis support, refetch-on-focus, and stats scoped to stored digests when present. Docs, fixtures, PGlite acceptance tests, and Cypress tweaks support recovery and contract checks.

Reviewed by Cursor Bugbot for commit 0181958. Bugbot is set up for automated code reviews on this repo. Configure here.

@vercel

vercel Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
inferencemax-app Ready Ready Preview Sep 30, 2026 9:20pm UTC

Request Review

@edwingao28 edwingao28 changed the title [PowerX] Digest gpu_metrics telemetry into the database at ingest time [PowerX] persist telemetry and recover historical reads / 持久化遥测并恢复历史数据读取 Sep 22, 2026
@edwingao28
edwingao28 force-pushed the feat/powerx-db-ingest branch from 7da4250 to 1d882e8 Compare September 22, 2026 17:22
@blacksmith-sh

This comment has been minimized.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit f320e18. Configure here.

},
runInfo: data.runInfo,
artifacts: data.artifacts.map((a) => a.name),
artifacts: artifactNames,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correlation defaults to missing metric

Medium Severity

The public views route still defaults corrYMetric to temperature. buildCorrelationData now drops samples where that metric is absent, so power-only DCGM series return an empty correlation plot. The explorer already falls back to another collected metric for this case.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit f320e18. Configure here.

Register the migrate, verify and telemetry backfill admin scripts and pull in the
PGlite and Blob fixture dependencies the new DB and API tests use.

中文:注册 migrate、verify 与遥测 backfill 管理脚本,并引入新 DB/API 测试所需的 PGlite 与 Blob fixture 依赖。
Migration 016 stores one series per (run, artifact, CSV) with full-resolution samples,
per-GPU statistics and benchmark point links; 017 adds stats_version so digests can be
upgraded without rescanning samples. verify-db and the shared table constants cover the new tables.

中文:迁移 016 按 (run, artifact, CSV) 存储序列、全分辨率样本、每 GPU 统计及 benchmark 点链接;017 增加 stats_version 以便升级摘要而无需重扫样本。verify-db 与共享表名常量覆盖新表。
Parse NVIDIA/AMD CSVs and multinode power_audit bundles into series, samples and
statistics, replace a changed series atomically under a row lock, and count telemetry
failures separately so they never fail the benchmark ingest.

中文:在 ingest 时把 NVIDIA/AMD CSV 与多节点 power_audit 包解析为序列、样本和统计,在行锁内原子替换变更的序列,遥测失败单独计数、不影响 benchmark 入库。
Write a per-run telemetry receipt (expected, produced, stored, API-readable points),
widen required-power publication to 1K/1K, and verify sweep manifests, point evidence
and curve preservation before a run is treated as published.

中文:为每个 run 写遥测回执(预期/产出/入库/API 可读点),required-power 发布范围扩展到 1K/1K,并在视为已发布前校验 sweep manifest、点证据与曲线保留。
Backfill historical runs from GitHub artifacts before retention expires, skipping
purged points; purge telemetry alongside run overrides with explicit counts; refresh
recovered power_audit fields and invalidate the cache after a backfill.

中文:在 artifact 过期前从 GitHub 回填历史 run 并跳过已清理点;run overrides 清理时同步删除遥测并显式计数;回填后刷新恢复的 power_audit 字段并失效缓存。
Read a run or point with prefix/source/host scoping pushed into SQL, page samples in
50k-row keyset pages under the Neon HTTP cap, re-check the series version after loading
so a concurrent re-ingest is retried, and expose a revision hash for the point cache.

中文:按 prefix/source/host 在 SQL 侧限定读取范围,样本以 5 万行 keyset 分页控制在 Neon HTTP 上限内,读取后复核序列版本以重试并发重入库,并为 point 缓存暴露 revision 哈希。
Run admin:db:migrate before ingest and overrides so writers never meet a missing
column; let agentic ingest target a Neon branch; only invalidate the cache after
overrides actually applied.

中文:ingest 与 overrides 前先执行 admin:db:migrate,避免写入端遇到缺列;agentic ingest 可指向 Neon 分支;仅在 overrides 实际生效后失效缓存。
Serve /api/gpu-metrics from the database first and fall back to live GitHub artifacts
only when a run is not ingested; add /api/v1/gpu-metrics-point behind a revision-keyed
Blob cache; keep the public views route shape identical for stored and live runs and
normalize live NVIDIA timestamps to ISO UTC.

中文:/api/gpu-metrics 优先读数据库,仅在 run 未入库时回退到 GitHub artifact;新增以 revision 为键的 Blob 缓存 /api/v1/gpu-metrics-point;公开 views 路由对入库与 live run 保持同一形状,live NVIDIA 时间戳归一为 ISO UTC。
The /gpu-metrics explorer reads stored series, shows full-record statistics from the
ingest digest, adds points/rolling display modes with a mean-of-chips line, and refetches
on focus so a re-ingest shows up in an open tab.

中文:/gpu-metrics explorer 读取入库序列,展示 ingest 摘要的全记录统计,新增点/滑动平均显示模式与芯片均值线,并在窗口聚焦时刷新以反映重入库。
Document the telemetry digest, receipts, migration prerequisites, targeted repair and
rollback, the gpu-metrics views route, and the new API route entries.

中文:记录遥测摘要、回执、迁移前置条件、定向修复与回滚、gpu-metrics views 路由及新增 API 路由条目。
中文:将博客列表页的访问放入每个用例的准备阶段,使 Cypress 能在页面加载超时时重试。
中文:博客列表页的 E2E 用例统一替换图片优化请求,避免 Firefox 等待缩略图时页面加载超时。
@edwingao28
edwingao28 force-pushed the feat/powerx-db-ingest branch from f320e18 to 0181958 Compare September 30, 2026 21:19
edwingao28 added a commit that referenced this pull request Sep 30, 2026
…h ruler

#1167 was rebased onto master twice today; #1220 had merged the 13:32 rebase
(f320e18), which still carried the date-comparison Perf Ruler, while the
14:19 rebase (0181958) removed it after #1232 landed on master. Merge the
new parent with the earlier rebase as the base, so only GPUGraph.tsx and its
component spec conflict, and resolve them by removing the ruler: the #1229
run-specific ruler share binding, the legend toggle, the hit strokes and
drag handlers, and the ruler share-link spec. `ScatterGraph` keeps
`i_rulers`; docs and the store comment now say the date comparison draws no
rulers.

中文:#1167 今日两次 rebase 到 master;#1220 已合入 13:32 的 rebase
(f320e183),其中仍含日期对比图的性能标尺,而 14:19 的 rebase(0181958d)
在 #1232 合入 master 后移除了它。以较早的 rebase 为 base 合并新父分支,
仅 GPUGraph.tsx 及其组件用例冲突,并按移除标尺解决:删除 #1229 的运行级
标尺分享绑定、图例开关、命中描边与拖拽处理,以及标尺分享链接用例。
`ScatterGraph` 保留 `i_rulers`;文档与 store 注释改为日期对比图不绘制标尺。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
中文:同步 #1220 合入 rebase 后的 #1167 与移除日期对比图标尺的结果到 PowerX 导航分支。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
中文:同步 #1220 合入 rebase 后的 #1167 与移除日期对比图标尺的结果到 NVL72 建模分支。

This branch was successfully deployed

1 active deployment
Preview — 0181958d Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants