Skip to content

[PowerX] unify telemetry inspection, comparisons and power labels / 统一遥测查看、对比视图与功耗标签 - #1220

Open
edwingao28 wants to merge 119 commits into
feat/powerx-db-ingestfrom
feat/powerx-article-parity
Open

edwingao28 wants to merge 119 commits into
feat/powerx-db-ingestfrom
feat/powerx-article-parity

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Unify point inspection, Timeline, comparisons, mobile tooltips and four power boundaries. Keep stitched curves together, pin comparison selections, support P75/P90 role panels, and simplify profit controls.

Based on #1167; #1228 adds public entry points, #1190 adds NVL72 modeling. Production approval remains separate. Saved preview keys containing audit hashes require reselection.

AI model disclosure

  • Model/version: claude-opus-5-5[1m]; claude-sonnet-5; GPT-6 (exact variant unverified); claude-fable-5-1. Earlier versions unverified.
  • Role: implementation; review; integration/fixes/verification; comparison fixes/Chinese review, respectively.

Validation

Local: 6,010 app tests, 190 smoke cases and desktop/mobile official/overlay checks pass. Three other-workspace timeouts passed in isolated retries. CI: current checks.

  • I have completed the AI model disclosure and kept it current

中文说明

整合逐点查看、Timeline、对比、移动端提示框及四个功耗边界。拼接后的曲线保持为同一来源,固定对比选择,支持 P75/P90 角色面板,并简化利润功耗控件。

基于 #1167;#1228 开放入口,#1190 增加 NVL72 建模。生产发布仍需独立确认;此前保存且含审计 hash 的预览来源键需重新选择。

claude-opus-5-5[1m] 负责实现,claude-sonnet-5 负责审阅,GPT-6 负责整合、修复与验证(精确变体无法核实),claude-fable-5-1 负责对比修复与中文审核;更早版本无法核实。

本地 6,010 项 app 单元测试、190 项浏览器 smoke 用例及桌面/手机的官方与叠加数据检查通过。其他 workspace 的三项超时用例分别重跑通过。CI 结果见上方链接。


Note

Medium Risk
Large inference/PowerX surface area (charts, read-only API panels, telemetry, share URLs) with extensive doc and test updates; behavioral risk is mitigated by new Cypress coverage but production paths span power boundaries and service comparison math.

Overview
This PR consolidates PowerX on /inference: four power boundary metrics (GPU measured, TDP, all-in provisioned, all-in modeled/PUE) ride on existing i_metric and Measured controls, with i_pcompare=boundaries|roles, a Power Timeline display (y_measuredPowerTimeline / i_pt*), and article-style analysis panels (equal-service / same-concurrency compare, role energy share, power–throughput fit, frontier points table). Boundary labels are renamed for clarity (e.g. profit views now use All in Provisioned / All in Measured).

Chart behavior gains i_xmode=concurrency (observed load sweeps, no Pareto/ruler/replay), mean vs median service stats (i_mstat → API xstat), and date-comparison GPUGraph support for ?unofficialrun= overlays and concurrency—aligned with expanded /api/v1/views/inference query keys (topologies, serviceCompare, roleShare, powerFit, etc.). Perf Ruler share state (i_rulers) is documented for the primary scatter chart; the date-comparison ruler is removed per product direction.

The old power metric availability panel and legacy-power-ring / measured-power-summary UI are dropped; quick filters and targeted empty states (e.g. All in Measured) replace them. Docs (powerx-permanent-view.md, dashboard read-only contract, data transforms, state ownership) and Cypress specs are updated to match.

Reviewed by Cursor Bugbot for commit 75b6f61. Bugbot is set up for automated code reviews on this repo. Configure here.

cquil11 and others added 30 commits September 17, 2026 12:29
…t time

PowerX read chip telemetry by downloading and parsing gpu_metrics GitHub
artifacts on every page view, and lost the data once GitHub's 90-day
artifact retention expired. This moves that work to ingest time.

- Migration 016 adds gpu_metric_series (one row per artifact CSV),
  gpu_metric_samples (full-resolution per-GPU samples), gpu_metric_gpu_stats
  (per-GPU min/max/mean/median/p95/p99/stddev digest), and
  benchmark_result_gpu_metrics (point <-> series links).
- A pure nvidia-smi/amd-smi CSV parser plus artifact discovery and an
  idempotent upsert (same CSV hash refreshes links only; a changed CSV
  replaces samples and digest in one transaction).
- CI ingest links gpu_metrics_<suffix> next to bmk_<suffix> using the same
  pairing rule as server logs.
- New backfill CLI: bun run admin:db:backfill-gpu-metrics --all --yes
  (bounded by GitHub retention; the GCS backup does not mirror gpu_metrics).
- /api/gpu-metrics serves the stored digest first and falls back to live
  GitHub artifacts for runs that are not ingested yet.
- New /api/v1/gpu-metrics-point?id=N and a PowerX tab on the per-point
  detail page showing the telemetry recorded while that point ran.
- Shared benchmark-result lookup extracted from the server-log backfill.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
…lemetry charts

Add a Points / Rolling average control to the PowerX telemetry charts on both
the explorer page and the per-point PowerX tab, with a 10/30/60/300 s window
selector. The average is a centered time-window mean per chip (pure helper in
telemetry-smoothing.ts, unit-tested for empty input, single sample, irregular
timestamps and inclusive window edges), so it follows elapsed time rather than
sample count. Averaged mode draws smooth lines and keeps the sample circles as
invisible hover targets so the tooltip and crosshair still work.

The per-point PowerX tab gains the same chip legend as the explorer so
individual chips can be hidden and restored, scoped to the selected series.
The chart's t=0 now comes from the whole series rather than the visible chips
so hiding a chip does not shift the time axis.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Add a Lines control (Per chip / Mean of chips / Both) next to the display-mode
toggle on both PowerX surfaces. The mean line averages the currently visible
chips at each timestamp of the longest series, aligning the other chips by
nearest sample within one estimated poll interval (median gap), so the few ms
of skew between nvidia-smi rows and the occasional dropped row do not
interpolate or misalign. It is drawn in the foreground color at a heavier
stroke, labelled in a key row under the caption, and its tooltip reports how
many chips contributed. The rolling-average mode smooths the mean line too.

The shared line layer gains an optional per-series getStrokeWidth so one
keyed join can hold both the chip lines and the heavier mean line.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Add an "Overlay decode throughput" switch (default off) to the per-point
PowerX tab. When on, the point's decodeTps series from the trace server
metrics is drawn over the telemetry in violet on its own right-hand axis,
with a key-row entry and a row in every tooltip giving the nearest decode
rate. The trace's startNs is an epoch-ns wall-clock timestamp, so the two
series are aligned by absolute time; when a trace carries no wall-clock
start the overlay falls back to the telemetry start and the note under the
switch says so. Points without server metrics get an "unavailable" note
instead of an empty overlay.

The chart takes the overlay as an optional prop and renders it in a custom
D3 layer that redraws on zoom and removes its own axis when toggled off. In
rolling-average mode the overlay is smoothed with the same time window as
the chip lines so both curves describe the same interval.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
At 1 s cadence the raw per-chip points read as noise, so both PowerX
surfaces now open in rolling-average mode (30 s window); Points remains
one click away in the Display control.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
…rlay

Replace the decode-only overlay switch on the per-point PowerX tab with a
single-select menu of server metrics. The menu lists only the series the
point's trace actually reports (decode/prefill throughput, KV and host KV
cache utilization, prefix cache hit rate and hits, queue depth), defaults
to "None", and is disabled when the trace has no server metrics.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
… digest

Register GET /api/v1/gpu-metrics-point as a page-owned BFF in the API route
catalog and refresh the /api/gpu-metrics digest after the DB-first digest read
landed in d5e90c4. Neither route is part of the public API reference, so the
OpenAPI document and human reference are unchanged; the exclusion reason for
/api/gpu-metrics now describes the stored-digest-then-live-artifact behavior.

中文:在 API 路由目录中登记 GET /api/v1/gpu-metrics-point(页面专用 BFF),并在
d5e90c4 引入 DB 优先读取后刷新 /api/gpu-metrics 的摘要。两条路由都不属于公开 API
参考,OpenAPI 文档与人工参考无需改动;/api/gpu-metrics 的排除说明改为描述
"已入库摘要优先、否则回退实时制品"的行为。
The pipeline doc blamed an artifact-name filter in the GCS backup. The backup
keeps every artifact name; it only mirrors scheduled and main-branch runs, so
PR sweeps and manual dispatches - nearly every telemetry-bearing run - are never
copied, and our own GCS reader ignores non-bmk_/server_logs_ objects anyway.
Also note that the multinode template uploads no gpu_metrics_ artifact, so
multinode and disaggregated points never get a telemetry series.

中文:修正数据管线文档中 gpu_metrics telemetry 回填受 90 天限制的原因说明。
GCS 备份并不按 artifact 名过滤,而是只镜像定时任务和 main 分支的运行,PR sweep
与手动触发的运行从不被复制;应用自身的 GCS reader 也只读取 bmk_/server_logs_
对象。同时注明多节点模板不上传 gpu_metrics_ artifact,多节点与 disagg 数据点
不会有 telemetry 序列。
30 PGlite-backed tests over the PowerX telemetry ingest code that had none.

- queries/gpu-metrics: null on unknown run and on a run with zero series,
  full run payload shape, includeSamples=false, latest-attempt selection,
  multinode point fan-out, availability map, and a characterization of the
  `?? 0` sample defaulting (a dropped reading reports 0 W).
- lib/benchmark-result-lookup: natural-key match scoped to run attempt and
  config, offload_mode exact match vs. unique fallback vs. ambiguous drop,
  agentic null isl/osl, and artifact JSON mapping.
- backfill-gpu-metrics: candidate selection driven through the real CLI main
  against a migrated PGlite database, with GitHub listing stubbed. Pins the
  90-day default --since, the resume hazard (runs with stored series are
  skipped without --force), --run bypassing both filters, --from-run,
  --limit, and that --shard-count/--shard-index do not partition the set.

No production source was changed.

中文:为此前完全没有测试的 PowerX 遥测 ingest 代码补充 30 个基于 PGlite 的单元
测试,覆盖 gpu-metrics 读取查询、backfill 候选筛选与 benchmark 点位匹配。要点:
读取查询在 run 不存在或没有 series 时返回 null,多节点点位返回全部关联 series,
并固化了 `?? 0` 默认值行为(采集中断的样本会显示为 0 W,与真实的 0 W 无法区分);
点位匹配按 run attempt 与完整 config 自然键限定,并区分 offload_mode 精确匹配、
唯一回退与歧义丢弃;backfill 测试通过真实 main 驱动候选筛选,固化默认 90 天
--since 窗口、--force 之前会跳过已有 series 的 run(断点续跑隐患)、--run 同时
绕过日期与 series 过滤,以及 --shard-count/--shard-index 实际不分片的现状。
未修改任何生产代码。
`gh api …/runs/<id>/artifacts` returns HTTP 404 once a run is deleted or
purged, and `retryArtifactOperation` spent the full backoff schedule on it
before `backfill-gpu-metrics --all --dry-run` aborted on the first such run
(34533943809, 18 runs into a 224-run sweep). Classify that 404 as
`WorkflowRunNotFoundError`, a `NonRetryableArtifactError` the retry helper
rethrows at once, and let both gpu-metrics and server-log backfills report
the run as gone and continue. The gpu-metrics header also drops the wrong
"GCS mirrors only bmk_/server_logs_" retention explanation.

中文:`gh api` 在 run 被删除或清理后返回 HTTP 404,之前 `retryArtifactOperation`
会把完整退避重试跑完,`backfill-gpu-metrics --all --dry-run` 在 224 个候选中
第 19 个 run(34533943809)处直接崩溃。现将该 404 归类为不可重试的
`WorkflowRunNotFoundError`,重试助手立即抛出,gpu-metrics 与 server-log 两个
backfill 记录该 run 已不存在并继续;同时修正 gpu-metrics 头注释中错误的
GCS 镜像保留说明。
…yment mean

Aggregate multinode producers (Kimi K3 B200 dynamo-vLLM TP8/PP2, H200 vLLM
TP16×2) emit no per-worker telemetry, so `modelSystemPower` rejected every
such row as `topology` and the smart-provision basis showed no NVIDIA K3 SKU.
Add a `uniform-hosts` topology basis for non-disaggregated multinode rows
without a worker array: the telemetry GPU count must fill whole eight-GPU
hosts and match TP×PP×PCP×replicas, and each chassis is modeled at the
deployment mean; rows that carry per-worker telemetry keep the exact
worker-hosts path and disaggregated rows still require it. `planningKwPerGpu`
now admits every fully measured multi-chassis estimate instead of only one
single-node chassis. Tooltip copy (en/zh) names the uniform-hosts assumption
and the system-power doc records the admission rule.

中文:聚合多节点的采集端(Kimi K3 B200 dynamo-vLLM TP8/PP2、H200 vLLM
TP16×2)不输出逐 worker 功耗,`modelSystemPower` 一律按 `topology` 拒绝,
smart provision 里因此没有任何 NVIDIA K3 SKU。新增 `uniform-hosts` 拓扑基准:
非 disagg 多节点且无 worker 数组时,要求 telemetry GPU 数填满整数个八卡主机并
等于 TP×PP×PCP×副本数,每个机箱按部署平均功耗建模;带逐 worker 功耗的行仍走
精确的 worker-hosts 路径,disagg 行仍需逐 worker 数据。`planningKwPerGpu`
改为接受所有机箱均完整实测的估算。tooltip 中英文说明该假设,系统功耗文档同步。
Multinode InferenceX jobs upload no `gpu_metrics_` artifact; their per-GPU
1 Hz power lives in `power_audit_<suffix>/LOGS/power/samples.csv`, one
deployment-wide CSV from srt-slurm's `dcgm-power` collector. Every Kimi K3
NVIDIA point therefore had an empty PowerX tab while the data sat in GitHub.

Accept the bundle as a fallback telemetry artifact: `gpuMetricsArtifactSuffix`
recognises `power_audit_`, discovery (CI ingest) and backfill pairing prefer a
`gpu_metrics_` sibling and use the bundle only when none exists, and
`prepareGpuMetricsArtifact` regroups `samples.csv` by hostname into one
power-only series per host (`file_name` `LOGS/power/samples.csv#<host>`,
manifest as context sidecar, GPU UUIDs as identity).

Power-only series made a latent reader defect visible: `toSampleRow` coerced
null clocks, temperature and utilization to 0, so the UI offered and plotted
fabricated flat-zero metrics. The five non-power core fields are now optional
end to end (reader, `GpuMetricRow`, anomaly checks, chart and correlation
points, tooltip, correlation default). The PowerX tab labels the collector
from the recorded producer instead of `nvidia-smi`, and takes the point's
hardware key for the TDP line because multinode artifact names are
hash-truncated. Docs record the adapter.

Backfilled into the branch DB: K3 B200 34674595026 (7 points, 14 series),
GB300 34873998796 (5 of 11; the six disaggregated jobs report
`multinode_power_contract_missing`), H200 34744300699 (10 points, 40 series).
Not covered here: the explorer's live GitHub fallback for not-yet-ingested
runs still reads `gpu_metrics_` only.

中文:多节点 InferenceX 任务不上传 `gpu_metrics_` artifact,每 GPU 1 Hz 功耗数据
在 `power_audit_<suffix>/LOGS/power/samples.csv`(srt-slurm `dcgm-power` 采集,
整个部署一个 CSV)里,Kimi K3 NVIDIA 数据点的 PowerX 标签页因此一直为空。
本次将该 bundle 作为回退 telemetry artifact:suffix 识别 `power_audit_`,
CI 入库发现与 backfill 配对优先使用 `gpu_metrics_`,仅在没有时使用 bundle;
`prepareGpuMetricsArtifact` 按 hostname 拆成每主机一条仅含功耗的序列。
同时修正读取端把空的时钟/温度/利用率读成 0 的问题(五个字段全链路改为可选),
PowerX 标签页的 Collector 改为显示真实采集器,TDP 参考线改用数据点的硬件 key。
已回填 branch DB:K3 B200、GB300(5/11,disagg 任务无功耗合约)、H200。
未覆盖:explorer 对未入库 run 的 GitHub 实时回退仍只读 `gpu_metrics_`。
…timeline / 新增功耗边界、标尺、对比序列与功耗时间线 (#1177)

* feat: share placed perf rulers through the i_rulers URL param

Rulers placed on the primary Inference chart now live in a provider-owned
store and serialize into the share link as `i_rulers`
(`isoX|curveA|curveB;...`). Opening such a link restores the rulers once both
curves have rendered, switches the tool on, clears them when the x-axis mode
changes the rendered chart, and ignores malformed values. Overlay-run curves
are addressable too.

中文:图表上放置的 Perf Ruler 现在保存在 provider 级的 store 中,并以
`i_rulers` 参数写入分享链接;打开链接后会在两条曲线都渲染完成时恢复标尺并自动
开启该工具,切换 X 轴模式导致图表更换时清除,格式错误的值将被忽略。非官方运行
的 overlay 曲线同样可被引用。

* feat: derive provisioned and modelled power boundary fields

Adds `lib/power-basis.ts` with the B2–B4 boundary maths and wires it into
`buildDerivedChartFields`: GPU provisioned (registry TDP), utility provisioned
(registry all-in kW per GPU) and utility modelled (measured GPU power carried
through the chassis model to the utility meter, PUE applied once). Energy
variants count every allocated GPU per output token; values are omitted, never
estimated, when a spec, throughput or telemetry input is missing. Registers the
six `y_*` metric keys with bilingual labels and axis explanations.

中文:新增 `lib/power-basis.ts`,实现 B2–B4 功耗边界的计算并接入
`buildDerivedChartFields`:GPU 额定(注册表 TDP)、全电源配置(注册表每 GPU
all-in kW)以及数据中心建模(GPU 实测功耗经机箱模型推算至市电侧,PUE 仅计入一
次)。能耗指标按每输出 token 计入全部已分配 GPU;缺少规格、吞吐量或遥测输入时
省略数值而不做估算。同时注册六个 `y_*` 指标键及中英文标签与坐标轴说明。

* feat: add a Boundary control to the gated measured power groups

The ↑↑↓↓-gated Measured Power / Measured Energy controls gain a Boundary select
(GPU measured, GPU provisioned, utility provisioned, utility modelled). The
boundary rides on `i_metric`, so existing share links keep working and a
shared boundary view renders even while the gate is locked. Boundary watt
gauges draw the same upper envelope as measured watts and stay envelope-locked
under Optimal Only; tied maxima now stay on the envelope so a flat TDP series
spans its tested range. The availability panel explains per point why a
boundary value is withheld, and the chart caption states each boundary's
formula and assumptions.

中文:↑↑↓↓ 门控的 Measured Power / Measured Energy 控件新增“功耗边界”下拉
(GPU 实测、GPU 额定、全电源配置、数据中心建模)。边界随 `i_metric` 传递,现有
分享链接不受影响,门控锁定时分享的边界视图仍可渲染。边界功率指标与实测功率一
样绘制上包络线,并在“仅最优”下保持包络锁定;包络线现在保留并列最大值,使平坦
的 TDP 序列覆盖完整测试范围。可用性面板逐点说明边界值缺失的原因,图表说明列出
各边界的公式与假设。

* Shorten chip config comparison label / 缩短芯片配置对比文案 (#1168)

* fix: shorten chip config comparison label

Shorten the comparison selector placeholder in the inference chart and profit estimator, with matching Simplified Chinese copy.\n\n中文:缩短推理图表和利润估算器中的芯片配置对比选择器占位文案,并同步更新简体中文。

* test: update chip config placeholder assertions

Update the profit estimator E2E assertions for the shortened English and Simplified Chinese placeholders.\n\n中文:更新利润估算器端到端测试断言,使其匹配缩短后的英文和简体中文占位文案。

* feat: add a measured power timeline display to the gated power group

Third value of the Measured Power Display control (`timeline`,
`y_measuredPowerTimeline`). The key aliases the measured average so the
point set, table view, availability panel and share link are unchanged;
ChartDisplay swaps ScatterGraph for the new PowerTimeline, which joins
each validated point to its `gpu_metrics_<RESULT_FILENAME>` artifact via
`power_audit.source`, fetches one prefixed request per workflow run, and
draws one-second per-GPU power over the whole job with the validated
window emphasized, TDP references per hardware, an opt-in all-in line,
wall-clock / elapsed axes, per-GPU lines, overlay-run colours, and an
explicit list of undrawn configs by reason.

`/api/gpu-metrics` gains `series=power` (compact one-second buckets,
~1 MB instead of ~27 MB raw for a 25-config run) and `prefix=` to skip
other models' artifacts; downloads run four at a time.

中文:在受门控的实测功耗组中新增「时间线」显示方式
(`y_measuredPowerTimeline`)。该指标键复用实测平均功耗,因此数据点集合、
表格视图、可用性面板和分享链接保持不变;ChartDisplay 在该模式下用新的
PowerTimeline 替换散点图:按 `power_audit.source` 将每个有效数据点关联到
对应的 `gpu_metrics_<RESULT_FILENAME>` 产物,按工作流运行各发起一次带前缀的
请求,绘制整个基准测试任务期间每 GPU 的逐秒功耗,突出显示有效测量窗口,
按硬件绘制 TDP 参考线(全电源配置线可选开启),支持实际时刻 / 相对起点两种
时间轴、每 GPU 单独曲线、非官方运行的配色,并按原因列出未绘制的配置。

`/api/gpu-metrics` 新增 `series=power`(一秒分桶的紧凑序列,25 个配置约
1 MB,而原始行约 27 MB)和 `prefix=` 参数以跳过其他模型的产物;下载并发数为 4。

* fix(landing): hide DeepSeek V4 Pro and Qwen3.8 Flash Next / 首页隐藏两个模型 (#1169)

* fix(landing): hide DeepSeek V4 Pro and Qwen3.8 Flash Next

中文:从首页隐藏 DeepSeek V4 Pro 和 Qwen3.8 Flash Next,保留比较页及仪表板行为,并添加双语回归测试。

* test(landing): update model row and badge expectations

中文:更新首页模型行数及 NEW 标记数量的测试预期。

* fix: temporarily label DSpark run 35166686551 as MoRI UMBP SGLang (#1170)

中文:为 DSpark 运行 35166686551 添加三周有效的 MoRI UMBP SGLang 显示名称,保留其他运行和原始数据,并测试到期边界。

Co-authored-by: Perplexity Computer <[email protected]>

* docs: require exact AI model disclosure (#1171)

中文:要求 PR 描述列出确切的 AI 模型名称/版本及工作内容,更新 agent 指引并新增双语 PR 模板。

* docs: expand bilingual glossary from Rubin AgentX article (#1172)

Add 12 terms, update 10 definitions, and test bilingual coverage and metric boundaries.

中文:根据 Rubin AgentX 文章扩充双语术语表,新增 12 个词条、更新 10 个定义,并补充双语覆盖与指标边界回归测试。

* feat: redirect /xvideo and /xxvideo aliases to /video (#1173)

Add a small alias table (`video-alias-redirects.ts`) wired into
`next.config.ts` so `/xvideo` and `/xxvideo` (plus the `/zh` siblings)
308-redirect to the canonical `/video` route. Subpaths and query strings
carry through via `:path*`. Unit test mirrors the inference-model alias
redirect test.

中文:新增 `video-alias-redirects.ts` 别名表并接入 `next.config.ts`,使
`/xvideo` 与 `/xxvideo`(含 `/zh` 版本)以 308 永久重定向到规范路由
`/video`,子路径和查询参数通过 `:path*` 原样透传。单元测试与推理模型
别名重定向测试保持一致。

* fix: refresh the gpu-metrics route digest after formatting

The pre-commit formatter rewrote src/app/api/gpu-metrics/route.ts after its
catalog digest was computed, so the API route catalog guard failed on a
clean checkout. Documentation and classification are unchanged.

中文:提交前的格式化工具在计算摘要之后重写了 gpu-metrics 路由文件,导致 API
路由目录校验在干净检出时失败;此处仅刷新 SHA-256 摘要,文档与分类不变。

* feat: overlay power boundary and role comparison series via i_pcompare

Add a Compare control to the gated Measured controls that overlays sibling
series on the selected metric without changing it: every power boundary
(GPU measured, GPU provisioned, utility provisioned, utility modeled) or
the prefill / decode worker pools. useChartData and the ?unofficialrun=
overlay processor both append one clone per sibling to every base point
(same x, y from the sibling's field, powerVariant set) through
utils/power-compare.ts; base points are untouched, so charts without a
comparison render byte-identically.

ScatterGraph keys series on hwKey + precision + variant, draws siblings in
the hardware colour (overlay runs keep the run colour) with a per-variant
dash and their own frontier, dims clone points, adds legend rows with line
swatches that toggle and highlight a series across hardware, and labels
the series in tooltips, the Table and the CSV export. Clones stay out of
best-per-SKU ranking, tier counts, the legend points table, the
availability panel and the date-comparison GPUGraph.

For PowerX Figure 7 the prefill pool's J per input token is carried onto
the output-token axis by the served input:output ratio
(utils/role-energy.ts, reconstructedPrefillJPerOutputToken) so it stacks
against the decode pool's J per output token. The comparison rides on the
new i_pcompare URL parameter (boundaries | roles; empty = off) and pauses
with a hint on metrics without a common axis instead of being cleared.

中文:在受门控的 Measured 控件中新增“对比”选项,可在所选指标上叠加同源系列而不改变
该指标:四种功耗边界(GPU 实测、GPU 额定、全电源配置、数据中心建模)或预填充 / 解码
worker 池。useChartData 与 ?unofficialrun= overlay 处理路径均通过 utils/power-compare.ts
为每个基础数据点追加对应系列的克隆点(x 相同、y 取自对应字段、标记 powerVariant),
基础点保持不变,未开启对比的图表渲染结果与之前完全一致。ScatterGraph 以
hwKey + 精度 + 变体作为系列键,同色不同虚线绘制,图例新增可切换、可高亮的线型行,
tooltip、表格与 CSV 标注系列;克隆点不参与 best-per-SKU、功耗等级计数、图例数据点表、
可用性面板及日期对比 GPUGraph。针对 PowerX 图 7,utils/role-energy.ts 按实际服务的
输入/输出 token 比将预填充池的每输入 token 能耗折算到每输出 token 轴
(reconstructedPrefillJPerOutputToken),与解码池并列。对比模式由新 URL 参数
i_pcompare 承载(boundaries | roles,空为关闭),在没有公共坐标轴的指标上暂停并提示,
而不是被清除。

* Add Qwen3.8-27B / 新增 Qwen3.8-27B (#1176)

Register the qwen3.827b and qwen3.827beager DB keys as the experimental
Qwen3.8-27B and Qwen3.8-27B-Eager dashboard models (InferenceX#3260, #3263):
constants, ETL normalizer path, Model enum and MODEL_CONFIG, compare slugs,
compare SSR known models, refreshed constants digest, AGENTS.md parameter row,
and the pinned registry counts.

Co-authored-by: Claude Fable 5.1 <[email protected]>

* [Cache Reuse] add prefix-cache tier share tab and agentic entry links / 新增前缀缓存复用页及智能体入口链接 (#1175)

* feat: add Prefix Cache Reuse tab with agentic entry links

Footer-only /cache-reuse (+ /zh) stacks per-concurrency shares of prompt
tokens served from HBM, the host tier, or recomputed, for one configuration.
Host tier is the CPU-offload rate, falling back to the external rate; never
summed. TensorRT-LLM offload rows draw one combined segment. Fixed sequences
keep the scenario selector and explain that no cache tiers are recorded.
Unofficial runs on the same hardware render as outlined series. c_cfg seeds
the configuration from a share link. Agentic chart footer and point summary
link into the tab.

中文:新增 footer 级 /cache-reuse(含 /zh)页面,按并发数堆叠展示单一配置的
prompt token 来源占比:HBM 缓存、主机层、未复用。主机层取 CPU offload 命中率,
缺失时改用外部缓存命中率,两者不相加;TensorRT-LLM offload 行绘制为单一合并段。
固定序列保留场景选择器并说明无缓存分层数据。同硬件的非官方运行以描边系列叠加。
c_cfg 参数用于分享链接定位配置。智能体图表页脚和数据点摘要新增入口链接。

* fix: pin point links to one precision and drop the empty official series

The point-detail link now writes i_prec from the point, so the tab keys its
groups by bare hardware key and c_cfg matches instead of falling back to
another sweep. buildCacheReuse lists an official series only when official
rows exist, so an overlay-only configuration draws full-width run bars.

中文:数据点详情页链接现在写入该点的 i_prec,页面按纯硬件键分组,c_cfg 能精确
匹配而不会回退到其他配置。buildCacheReuse 仅在存在官方数据行时列出官方系列,
仅存在于非官方运行中的配置以全宽柱形绘制。

* feat: draw worker-pool power traces from power-audit bundles on the timeline

PowerX Figure 1 on the gated Timeline display (y_measuredPowerTimeline):

- /api/gpu-metrics?series=power also downloads power_audit_<RESULT_FILENAME>
  bundles (Slurm / Dynamo DCGM collector, up to 256 MiB), cuts
  LOGS/power/samples.csv into one series per power_validation_*.json window
  (±60 s) and labels devices <hostname>/<GPU-uuid> with their prefill / decode
  role; a corrupt archive skips that artifact instead of failing the run.
- The timeline joins bundle-cut series by validation file name, keeps the
  gpu_metrics join for single-node rows, and gains a "Prefill / decode pools"
  line mode: summed pool watts with pool size × TDP references, pool tooltip,
  and an all-GPU total for traces without roles.
- Pinned scatter tooltips on measured-power metrics offer "View power trace",
  which opens the timeline focused on that config (official and
  ?unofficialrun= points); the anchor href carries the canonical chart state.
- Docs, catalog digest, unit / component / e2e coverage in en and zh.

中文:在受门控的功耗时间线视图上实现 PowerX 图 1。/api/gpu-metrics?series=power
现在也会下载 Slurm / Dynamo(DCGM)运行上传的 power_audit_* 产物包(上限 256 MiB),
按每个 power_validation_*.json 的测量窗口(前后各 60 s)切出独立序列,并按
prefill / decode 角色标注每个 <hostname>/<GPU-uuid> 设备;损坏的压缩包只跳过该产物,
不再让整个请求失败。时间线通过校验文件名关联这些序列(单节点行仍按 gpu_metrics
关联),新增"预填充 / 解码 GPU 池"线型:按池求和的功耗、池内 GPU 数 × TDP 的参考线、
池级提示框,无角色的曲线则显示全部 GPU 总功耗。实测功耗散点图的固定提示框新增
"查看功耗曲线"操作,可直接跳转到聚焦该配置的时间线(官方点与 ?unofficialrun= 叠加点
均支持),链接携带规范化的图表状态。同步更新文档、路由目录摘要及中英文单元 /
组件 / 端到端测试。

* [First-Token Limits] add measured TTFT-cap cost comparison tab / 新增首 token 延迟约束成本对比页 (#1174)

* feat: add First-Token Limits tab comparing cheapest measured configs per TTFT cap

New footer-only dashboard tab at /first-token and /zh/first-token. For each
time-to-first-token cap it picks the cheapest measured row per chip vendor
that also clears an interactivity floor, draws grouped bars with the
best-vs-runner-up gap under the axis, and lists every winner with its run
link. Unofficial runs render as their own series. TTFT is now carried on
GPUDataPoint and kept by the calculator API view.

中文:新增页脚入口的仪表板页面 /first-token 与 /zh/first-token。按每一档首 token
延迟(TTFT)上限,从满足交互性下限的实测数据行中逐厂商选出成本最低的配置,以分组
柱形展示并在坐标轴下方标注最优与次优的成本差距,表格列出各优胜配置及其运行链接。
非官方运行以独立系列呈现。GPUDataPoint 新增 ttft 字段,calculator API 视图保留
TTFT 指标。

* fix: count overlay rows in the first-token empty state

Overlay-only rows that carry a TTFT but miss the interactivity floor were
described as having no first-token measurement, because the empty state only
looked at official rows. Report overlay readable rows separately so the
caption stays official-only while the empty state sees both.

中文:仅有非官方运行数据、且带 TTFT 但未达交互性下限时,空状态误报为「未报告首
token 延迟」。现单独统计叠加运行的可读行,标题说明仍只计官方数据,空状态则同时
考虑两者。

* fix: blame the cap ladder when qualifying rows clear no first-token cap

The empty state told readers to lower the interactivity floor whenever no
bar drew, even when rows cleared the floor and only the TTFT caps were too
tight. It now distinguishes three cases: no TTFT reported, nothing reaches
the floor, and nothing finishes within the largest cap. Overlay rows count
toward the floor check through overlayQualifyingRows; formatCap moves to
the shared module so the message prints caps like the axis ticks.

中文:此前只要没有柱形可画,空状态就提示降低交互性下限,即使数据行已达到下限、
只是 TTFT 上限过紧。现在区分三种情况:未报告 TTFT、无配置达到下限、无配置落在
最大上限之内。非官方运行行通过 overlayQualifyingRows 参与下限判断;formatCap
移入共享模块,使提示中的上限格式与坐标轴刻度一致。

* fix: prevent measured power statistic label overflow

Size the control grid to its container and preserve statistic label widths. Add English and Chinese desktop/mobile geometry regression coverage.

中文:修复实测功耗统计量标签溢出。根据容器宽度排列控件,并保留按钮标签所需宽度;新增中英文桌面和移动端布局回归测试。

* fix: keep power envelope ties only for provisioned and modelled gauges

The tie-keeping added for flat boundary gauges applied to every power
curve, so the measured watt axes kept tied points that Optimal Only used
to collapse (certified-power-filter rendered 6 of 6 instead of 2 of 6).
`upperPowerEnvelope` takes `keepTies`; both charts pass it per series via
`isPowerGaugeSeries`, true for a boundary axis or a boundary comparison
clone, false for measured series, the measured clone and role clones.

中文:此前为平直的边界量表加入的"保留并列点"逻辑作用到了所有功耗曲线,
实测功率轴上 Optimal Only 本应收起的并列点被保留(certified-power-filter
显示 6/6 而非 2/6)。`upperPowerEnvelope` 新增 `keepTies` 参数,两张图按
序列通过 `isPowerGaugeSeries` 传入:边界轴或边界对比克隆为 true,实测序列、
实测克隆和角色克隆为 false。

* fix: key comparison legend rows by the base series identity

`isBase` was "no official point carries this variant", which is wrong when
only an `?unofficialrun=` overlay carries the comparison: every row became
the base, all swatches drew solid and every toggle hid the base key, so a
sibling could never be hidden on its own. Derive the base from
`powerCompareBase` for the selected metric instead. A component test
covers the overlay-only case.

中文:`isBase` 原来的判断是"没有官方点带这个 variant",当只有 `?unofficialrun=`
overlay 带对比序列时所有行都被当成 base,图例全画实线且每次切换都隐藏
base,兄弟序列无法单独隐藏。改为按所选指标通过 `powerCompareBase` 推导
base。新增组件测试覆盖仅 overlay 的情况。

* fix: rename the timeline axis so a "Measured Power" search stays unique

The Timeline display adds a fourteenth `y_measured*` option. Its label
began with "Measured Power", so the family search matched two options.
Lead with "Measured Average Power", as the %TDP display does, and update
the selector count test, the timeline specs and the Chinese label.

中文:Timeline 显示模式新增了第 14 个 `y_measured*` 选项,其标签以
"Measured Power" 开头,导致按系列名搜索时匹配到两项。改为与 %TDP 显示
一致地以"Measured Average Power"开头,同步更新选择器数量测试、时间线
测试和中文标签。

* fix: label each power comparison series on its own line

Under i_pcompare the line labels were grouped by hardware only, so one
label per hardware landed on whichever variant had the most points
(usually the TDP boundary) and multi-precision views drew identical
labels on every sibling. Group by hardware and variant, keep the plain
label on the base measured series, and suffix siblings with the short
variant name plus the flat watts of a provisioned boundary
(`H200 (SGLang) · TDP 700 W`, `… · Decode GPUs`). Label groups now carry
data-series-id / data-power-variant hooks; overlays get the same text.

中文:对比模式(i_pcompare)下曲线标签原先只按硬件分组,每个硬件只有一个
标签且落在点数最多的 variant 线上(通常是 TDP 边界线),多精度视图则给每条
兄弟线画出相同文字。现改为按硬件 + variant 分组:基线保留原标签,兄弟线追加
简短 variant 名及平直边界的瓦数(如 `H200 (SGLang) · TDP 700 W`、
`… · Decode GPUs`);标签组新增 data-series-id / data-power-variant 钩子,
叠加运行同样处理。

* fix: load overlay runs first in the power timeline run cap

`?unofficialrun=` telemetry sorted after official runs by run id and was the
run dropped by POWER_TIMELINE_MAX_RUNS; overlay runs now take the cap's slots
right after the deep-linked run. Adds prioritizeRuns with unit and component
coverage and records the order in docs/powerx-permanent-view.md.

中文:功耗时间线每张图最多加载 4 个 run,此前按 run id 排序,`?unofficialrun=`
叠加的 run 通常最新、排在最后而被丢弃;现在叠加 run 紧随深链 run 优先占位。
新增 prioritizeRuns 及单元 / 组件测试,并在 docs/powerx-permanent-view.md 记录顺序。

* fix: route timeline legend clicks through the unified overlay selection

With an overlay loaded the timeline reads localOfficialOverride, so the
context toggleHwType changed nothing visible. Official rows now solo through
setUnifiedOverlaySelection like ScatterGraph; component test added.

中文:加载 `?unofficialrun=` 叠加后,时间线读取 localOfficialOverride,图例点击
只改 activeHwTypes、界面无变化。现在官方硬件行与 ScatterGraph 一样通过
setUnifiedOverlaySelection 独显,并补充组件测试。

* fix: name the hardware instead of the branch on overlay line labels

Unofficial-run pills read `✕ <hardware>`, parsed like official pills so the
GPU stays bold; a short run tag (` · main`, ` · …<date>-<sha>`) is appended only
when several overlay runs draw the same hardware. Branch names run to 70+
characters, so two of them could not share a chart and one pill was hidden.
Legend rows keep the full branch. Unit and component coverage added.

中文:叠加曲线标签改为「✕ 硬件名」,与官方标签同样解析、GPU 名加粗;仅当多个
叠加 run 画同一硬件时追加短 run 标签(` · main`、` · …<日期>-<sha>`)。分支名
常超过 70 字符,两条标签放不下会被隐藏。图例行仍显示完整分支。补充单元与
组件测试。

* fix: keep overlay line labels visible when only their overlay row is active

The filter-sync opacity pass judged every line label by the official
hardware set, so soloing one official hardware hid the pills of every other
overlay hardware. Overlay labels now follow activeOverlayHwTypes; unit tests
cover the rule and the component test pins the rendered pill visibility.

中文:图例筛选后的标签透明度同步只看官方硬件集合,独显一个官方硬件会把其他
硬件的叠加曲线标签一并隐藏。叠加标签现在跟随 activeOverlayHwTypes;单元测试
覆盖规则,组件测试固定渲染后的可见性。

* fix: merge equal-size pool TDP references and stack coinciding labels

In pool mode a prefill pool and a decode pool of the same GPU count drew two
references at the same watts, printing their labels over each other. Pools
of one hardware that share a size now draw one line labelled
`prefill / decode ×16`, and labels of lines at equal watts stack upward.

中文:池模式下 prefill 池与 decode 池 GPU 数相同时,两条参考线重合、标签互相
覆盖。同一硬件、相同 GPU 数的池现在只画一条参考线,标签为
`prefill / decode ×16`;瓦数相同的参考线标签向上错开排列。

---------

Co-authored-by: Alec Ibarra <[email protected]>
Co-authored-by: functionstackx <[email protected]>
Co-authored-by: Perplexity Computer <[email protected]>
Co-authored-by: Claude Fable 5.1 <[email protected]>
…licts

Brings master (20 commits) into feat/powerx-db-ingest after #1177 landed.
Three conflicts, all "both sides inserted at the same anchor":

- AGENTS.md: kept master's Pareto cross-repo bullet; our gpu-metrics-point
  API line auto-merged elsewhere in the same file.
- packages/app/cypress/e2e/landing-dashboard-navigation.cy.ts: master's
  "landing model curation" block is a strict superset of ours (verified:
  zero lines exist on our side that master lacks), so master's file wins.
- packages/app/src/lib/glossary.ts: kept both. Our vera-rubin and
  extreme-co-design entries plus our KV/CPU/NVMe offload rewrites stay;
  master's 15 new Engram-article entries are appended.

packages/app/src/lib/api-route-catalog.ts auto-merged cleanly this time.

Local gate on the merge result: lint, fmt, typecheck, typography, and
6,872 unit tests (app 5,888 + 4 skipped, constants 63, db 921). The one
db failure is the known PGlite cold-start timeout in operatorx.test.ts and
passes on rerun.

中文:在 #1177 落地后把 master 的 20 个提交合入 feat/powerx-db-ingest。三处冲突
都是两侧在同一锚点插入内容:

- AGENTS.md:保留 master 的 Pareto 跨仓库条目;我们新增的 gpu-metrics-point
  API 说明在同文件其他位置已自动合并。
- landing-dashboard-navigation.cy.ts:master 的 landing model curation 块是
  我们版本的严格超集(已核对,我们没有任何一行是 master 缺失的),整体取 master。
- glossary.ts:两侧都保留。我们的 vera-rubin、extreme-co-design 词条以及
  KV/CPU/NVMe offload 相关改写保持不变,追加 master 新增的 15 条 Engram 词条。

api-route-catalog.ts 本次自动合并干净。

合并结果的本地检查:lint、fmt、typecheck、typography,以及 6,872 个单元测试
(app 5,888 通过 + 4 跳过,constants 63,db 921)。db 侧唯一失败是已知的
operatorx.test.ts PGlite 冷启动超时,重跑通过。
…证据并保护已发布曲线 (#1152)

* feat: enforce required PowerX publication coverage

中文:强制校验 PowerX 必需功耗点的入库与发布完整性。
按实际入库身份匹配并发点,保留过滤后的缺失检查,并核对 AgentX 数据库与 API 结果。

* fix: validate power evidence and preserve published curves

中文:校验必需功耗证据,并在入库前保护已发布曲线。版本化 manifest 绑定来源、物理 GPU、测量窗口和产物哈希;有意替换必须精确声明旧快照及移除点。
Migration 016 declares `on delete cascade` from gpu_metric_series to
workflow_runs and from benchmark_result_gpu_metrics to benchmark_results, so
a purge already removed telemetry — with no count in the preview and no line
in the transcript. Past GitHub's 90-day artifact retention the stored samples
are the only copy, which is the reason 016 exists, so that loss should not be
invisible at the moment the operator confirms it.

The same rows are removed as before; this is accounting, not a change of
behaviour. previewPurge now prints the series and sample counts alongside the
benchmark and server_log counts, a whole-run purge deletes the series itself
before dropping workflow_runs and logs what went, and a point purge drops only
the point-to-series links. The series stays with its run in that second case:
/api/gpu-metrics?runId= reads series by run rather than through the links, and
other points of the same run may still reference it.

New lib/telemetry-purge.ts holds the three queries, covered by 11 PGlite tests
against the real 001 and 016 migrations so the cascades under test are the ones
the migration declares.

中文:迁移 016 声明了 gpu_metric_series 到 workflow_runs、
benchmark_result_gpu_metrics 到 benchmark_results 的 `on delete cascade`,因此
purge 一直在删除遥测数据,但预览里没有计数、日志里没有记录。超过 GitHub 90 天
产物保留期后,库里的采样就是唯一副本,这正是 016 存在的理由,所以操作者确认的
那一刻不应该看不见这笔损失。

删除的行与此前完全相同,这次改的是账目而不是行为。previewPurge 现在会在
benchmark 与 server_log 计数旁一并打印 series 与采样数;整个 run 的 purge 会在
删除 workflow_runs 之前先显式删除 series 并记录;单点 purge 只解除点与 series 的
链接。后一种情况下 series 会随 run 保留,因为 /api/gpu-metrics?runId= 是按 run
读取 series 而不是走这些链接,而且同一个 run 的其他点可能仍在引用它。

新增的 lib/telemetry-purge.ts 收拢这三条查询,由 11 个 PGlite 测试覆盖,直接对
真实的 001 与 016 迁移运行,确保被测的正是迁移声明的那些级联。
This workflow fires on a push to run-overrides.ts, which can land before the
next ingest dispatch has applied a pending migration. It ran admin:db:verify
with no migrate step of its own, and verify-db counts every table it knows
about, so it would fail on a schema the checked-out ref expects but production
does not have yet. Added the same Run migrations step ingest-results.yml uses;
migrations are idempotent by filename.

The worse half was what a red verify did to the two steps after it. Neither
carried if: always(), so a verify failure skipped both the production cache
invalidation and the warmup while the overrides were already committed to the
database. The dashboard then served pre-override data with a red workflow as
the only signal. Both now run whenever the overrides step itself succeeded.

中文:该工作流在 run-overrides.ts 的 push 上触发,而这次 push 可能早于下一次
ingest 派发应用待处理的迁移。它自己没有 migrate 步骤就直接跑 admin:db:verify,
而 verify-db 会统计它已知的每一张表,于是会在一个"检出的 ref 期望、但生产还没有"
的 schema 上失败。现已加入与 ingest-results.yml 相同的 Run migrations 步骤;迁移
按文件名幂等。

更严重的一半是 verify 变红对其后两个步骤的影响。两者都没有 if: always(),所以
verify 一失败就会跳过生产缓存失效和预热,而此时 override 已经写进数据库了。仪表板
于是继续提供改写前的数据,唯一的信号只有一个变红的工作流。现在只要 override 步骤
本身成功,这两步就会执行。
A gpu_metrics digest failure was recorded through recordDbError, which feeds
tracker.skips.dbError, which the ingest writes into the publication manifest
as an ingestError, which verify-power-publication folds into errors and exits
non-zero on. That script is the required "Verify PowerX source, database and
public API" step with no continue-on-error, so one malformed telemetry CSV
would turn a whole production ingest red even though every benchmark row
landed correctly. This failure mode does not exist before migration 016.

Telemetry failures now have their own counter and their own recordTelemetryError,
with its own print budget so telemetry noise cannot suppress real DB errors.
The manifest reports them under telemetryWarnings, and the new
fatalPublicationErrors makes the fatal set explicit at the one place that
decides the exit code. What is lost on a digest failure is one point's PowerX
tab, and admin:db:backfill-gpu-metrics --run <id> can re-digest the artifact.

Coverage: the two counters are proven independent, and fatalPublicationErrors
is proven to ignore telemetry warnings while still failing on real ingest
errors and verification mismatches. The catch site itself has no direct test —
it sits inside the artifact loop of a script the suite only runs as a
subprocess against a purged run, which returns before reaching it.

中文:gpu_metrics 摘要失败此前经 recordDbError 记录,进入 tracker.skips.dbError,
再被 ingest 作为 ingestError 写入发布 manifest,verify-power-publication 把它折进
errors 并以非零码退出。该脚本是必需的 "Verify PowerX source, database and public
API" 步骤且没有 continue-on-error,因此一个格式错误的遥测 CSV 就能让整条生产
ingest 变红,哪怕每一行 benchmark 数据都正确落库。这个失败模式在迁移 016 之前
并不存在。

遥测失败现在有独立计数器和独立的 recordTelemetryError,并有自己的打印额度,避免
遥测噪声淹没真正的 DB 错误。manifest 将其归入 telemetryWarnings,新增的
fatalPublicationErrors 在决定退出码的唯一位置显式界定致命集合。摘要失败损失的只是
某一个点的 PowerX 标签页,用 admin:db:backfill-gpu-metrics --run <id> 可以重新摘要
该产物。

覆盖范围:已验证两个计数器互相独立,并验证 fatalPublicationErrors 忽略遥测警告、
同时仍然对真正的 ingest 错误和校验不一致判为失败。catch 处本身没有直接测试——它
位于一个脚本的产物循环内部,而测试套件只以子进程方式对一个已 purge 的 run 运行该
脚本,那条路径在到达此处之前就返回了。
The anchor pass in placeLineLabels tried four fractions along each line and, if
all of them collided, emitted visible:false. In the PowerX article's Fig 6 —
roles mode with two ?unofficialrun= overlays, so nine lines whose anchors
compete in one narrow band — that silently cost the GB300 overall series its
pill. Its curve was still drawn, next to two labelled siblings of the same
colour, so the reader could only identify it by elimination.

The drop was a layering mistake. It decided visibility from a crude nominal box
of collisionWidth/2 by 21px, while layoutPills runs right after with the real
measured boxes, mirrors, shifts rows and clamps into the plot, and says in its
own doc comment that an overlapped label beats a missing one. A rendered pill in
this figure measures 223-330px against that 120px model, so the pass that gave
up was the one least able to judge. The same function's pinned-anchor branch
already emits visible:true unconditionally.

A series with no clear slot is now deferred and placed after the clean ones, on
whichever of its slots carries the least overlap, and it no longer anchors on
points[0] — the axis-hugging index lineCandidates deliberately skips. Deferring
matters on its own: a doomed series used to be able to reserve a slot that a
later series could have had to itself. keepVisibleOnCollision is gone; it was
the single-point special case of the rule this generalises.

This ends the promise that line labels never overlap, stated in #132 and
restated in #434. It was worth less than it cost: a dropped pill is silent, and
with line labels on the PNG export omits the legend, so the series loses its
only identifier.

Verified at /inference?i_metric=y_measuredAvgPower&i_pcompare=roles&unofficialruns=
35319969159,35319956855 against the branch DB: 17 labels, all rendered, none
overlapping, including the pill the figure was missing. The 12 pills still
hidden there are the separate GH #470 de-duplication path, which is untouched.
Gates: lint, fmt, typecheck, typography, 6,949 unit tests, 521 Cypress
component tests. Three of the five new unit tests and the new component test
fail on the old code.

中文:placeLineLabels 的锚点阶段沿每条线尝试四个位置,若全部碰撞就发出
visible:false。在 PowerX 文章的 Fig 6 中(roles 模式加两个 ?unofficialrun= 叠加,
九条线的锚点挤在同一窄带里),这让 GB300 整体序列悄悄丢掉了标签。它的曲线仍然
画着,旁边是两条同色且有标签的兄弟线,读者只能靠排除法辨认。

这个丢弃是分层错误。它用 collisionWidth/2 乘 21px 的粗略估算盒来决定可见性,而
紧随其后的 layoutPills 拿着真实测量盒做镜像、错行和边界钳制,其文档注释明确写着
重叠的标签也好过消失的标签。该图里一个实际渲染的标签宽 223 到 330px,而模型只按
120px 估算,所以放弃的恰恰是最没有判断力的那一遍。同一函数的固定锚点分支本来就
无条件发出 visible:true。

现在没有空闲槽位的序列会被推迟到干净标签之后放置,落在重叠代价最小的槽位上,并且
不再锚定到 points[0]——lineCandidates 刻意跳过的贴轴位置。推迟本身也有意义:原先
一个注定重叠的序列可能占掉后面序列本可独享的槽位。keepVisibleOnCollision 已删除,
它只是本规则的单点特例。

这终结了「折线标签永不重叠」的承诺,该承诺由 #132 提出、#434 重申。它的价值抵不上
代价:标签被丢弃是无声的,而开启折线标签后 PNG 导出会省略图例,序列就失去了唯一的
标识。

验证:在指向分支数据库的 /inference?i_metric=y_measuredAvgPower&i_pcompare=roles
&unofficialruns=35319969159,35319956855 上,17 个标签全部渲染、互不覆盖,包括该图
原本缺失的那一个。页面上仍隐藏的 12 个标签来自独立的 GH #470 去重路径,未受影响。
闸门:lint、fmt、typecheck、typography、6,949 个单元测试、521 个 Cypress 组件测试。
五个新单元测试中的三个以及新增的组件测试在旧代码上失败。
Remove unused availability and sample-omission paths, identity wrappers,
test-only label/source helpers, and write-only telemetry fields. Keep label
formatting and nested-source coverage on the production helpers, preserve
artifact failure isolation, and refresh the unchanged API contract digest.

中文:删除未使用的 PowerX 遥测查询模式、包装函数与只写不读的字段;
将标签格式及嵌套路径测试迁到实际生产入口,保留逐产物错误隔离,
并同步 API 路由摘要。此次仅落实精简清单 1–8。
Deduplicate each GPU/timestamp before computing metadata and statistics,
keeping the first sample for both CSV and per-host power-audit series.
Include sidecars and normalized sample counts in the replay check so
corrected context is persisted and explicit re-ingest repairs old digests.

Add PGlite regressions for duplicate populations, legacy replay, timezone
and identity corrections, and JSON key-order no-ops. The focused suite
fails on the old implementation and passes all 12 tests with the fix.

中文:统一 PowerX 遥测明细与摘要的采样集合,在统计前按 GPU 和时间戳去重。
重新摄取时同时核对 sidecar 与去重后的样本数,使时区及身份修正能够落库,
并支持显式重跑修复旧摘要;补充真实 PGlite 回归测试及恢复说明。
…参考线 (#1196)

* fix: keep constant telemetry axes readable

中文:为恒定及近恒定遥测值保留最小坐标范围,避免显存时钟刻度坍缩;补充桌面、手机、全零值和 TDP 参考线的组件回归测试。

* fix: clear power reference when switching telemetry metrics

切换到显存频率时清除旧 TDP 参考线,并验证切回功耗后仅恢复一条参考线。
Use plotted role and boundary values in tables, sorting, and CSV exports. Reuse the locally verified 8dbab1ed table correction. Distinguish dates on multi-day UTC timelines and clarify per-chip power accounting. Cover regressions and exercise the mobile telemetry tooltip with a real click.

中文:修正 PowerX 对比数值与时间轴标签。表格、排序和 CSV 使用各角色或功耗边界的绘图数值,复用已验证的 8dbab1ed 修复;跨日 UTC 时间轴补充日期,明确每芯片功率与部署总能耗的区别,并补充回归测试和手机端点击提示框验证。
Space timeline ticks by plot width and wrap long curve-label segments without dropping hardware, framework, or role text. Add desktop/mobile geometry and mobile ruler-drag coverage.

中文:修复 PowerX 手机端时间轴与角色标签的可读性。按绘图区宽度安排时间刻度,长曲线标签换行且保留硬件、框架及角色信息;补充桌面和手机布局回归以及手机宽度的标尺拖动测试。
中文:合并 master 并解决 PowerX 入库分支冲突,保留多节点与部分 GPU 功耗估算、遥测图表和共享 API 数据转换。
中文:全记录统计复用数据库摘要,统一 NVIDIA、AMD 和多节点指标的单位、缺失值、重复样本与时间戳语义;包含统计界面、API、回归测试和文档,不改变 serving-window 功率或 J/token。
中文:历史 Timeline 优先读取数据库,保留主机、GPU、验证窗口及分桶语义;缺少存储时按来源回退,数据库故障明确报错。统一 context 时区与文件顺序,包含历史读取的测试和文档。
中文:点详情以数据库版本选择缓存,修正 sidecar、新增关联及共享序列更新后自动读取新数据;恢复页面聚焦刷新,包含缓存失败恢复、事务回滚和真实数据库浏览器验收。
中文:扩展现有 run/attempt 发布回执,独立追踪预期点、产物、序列、样本、关联与 API 可读性;保留未知分母、失败身份和定向恢复动作,包含 CI、backfill、幂等恢复及历史身份修复的测试和操作说明。
Reuse the stored point telemetry view in a lazy in-page dialog. Preserve the PowerX gate, measured share links, and unofficial trace fallback. Cover desktop, Chinese mobile, and collapsed comparison legends.

中文:从仪表板数据点按需打开 PowerX 遥测面板,复用已存储遥测。保留功能门控、实测指标分享链接及非正式运行曲线入口,并验证桌面、中文手机端和收起图例后的交互。
中文:格式化对比组件用例(oxfmt)。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
中文:同步 #1220 的格式化提交到 PowerX 导航分支。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
Brings ff989ca/fbf41938 (Compare pins its pair on toggle) and the #1236,
#1239, #1240 consolidation into the NVL72 branch. Conflict in
ProfitEstimatorDisplay: #1239 dropped the header power disclosure, so the
one-line note stays and the NVL72 basis notes (module sensor, modeled trays,
DLC PUE) move to where the methodology now lives, the Power Estimation help
and the CSV caption; the profit spec asserts them there.

中文:把 ff989ca/fbf41938(对比在开启时固定一对)以及 #1236、#1239、#1240
的整合合入 NVL72 分支。ProfitEstimatorDisplay 冲突:#1239 已移除图表头部的
功耗假设折叠块,因此保留一行功耗说明,NVL72 的边界说明(module 传感器、
建模的交换机托盘、DLC PUE)移到方法学现在所在的位置——Power Estimation 帮助
与 CSV 说明;利润用例改为在这两处断言。
Register the migrate, verify and telemetry backfill admin scripts and pull in the
PGlite and Blob fixture dependencies the new DB and API tests use.

中文:注册 migrate、verify 与遥测 backfill 管理脚本,并引入新 DB/API 测试所需的 PGlite 与 Blob fixture 依赖。
Migration 016 stores one series per (run, artifact, CSV) with full-resolution samples,
per-GPU statistics and benchmark point links; 017 adds stats_version so digests can be
upgraded without rescanning samples. verify-db and the shared table constants cover the new tables.

中文:迁移 016 按 (run, artifact, CSV) 存储序列、全分辨率样本、每 GPU 统计及 benchmark 点链接;017 增加 stats_version 以便升级摘要而无需重扫样本。verify-db 与共享表名常量覆盖新表。
Parse NVIDIA/AMD CSVs and multinode power_audit bundles into series, samples and
statistics, replace a changed series atomically under a row lock, and count telemetry
failures separately so they never fail the benchmark ingest.

中文:在 ingest 时把 NVIDIA/AMD CSV 与多节点 power_audit 包解析为序列、样本和统计,在行锁内原子替换变更的序列,遥测失败单独计数、不影响 benchmark 入库。
Write a per-run telemetry receipt (expected, produced, stored, API-readable points),
widen required-power publication to 1K/1K, and verify sweep manifests, point evidence
and curve preservation before a run is treated as published.

中文:为每个 run 写遥测回执(预期/产出/入库/API 可读点),required-power 发布范围扩展到 1K/1K,并在视为已发布前校验 sweep manifest、点证据与曲线保留。
Backfill historical runs from GitHub artifacts before retention expires, skipping
purged points; purge telemetry alongside run overrides with explicit counts; refresh
recovered power_audit fields and invalidate the cache after a backfill.

中文:在 artifact 过期前从 GitHub 回填历史 run 并跳过已清理点;run overrides 清理时同步删除遥测并显式计数;回填后刷新恢复的 power_audit 字段并失效缓存。
Read a run or point with prefix/source/host scoping pushed into SQL, page samples in
50k-row keyset pages under the Neon HTTP cap, re-check the series version after loading
so a concurrent re-ingest is retried, and expose a revision hash for the point cache.

中文:按 prefix/source/host 在 SQL 侧限定读取范围,样本以 5 万行 keyset 分页控制在 Neon HTTP 上限内,读取后复核序列版本以重试并发重入库,并为 point 缓存暴露 revision 哈希。
Run admin:db:migrate before ingest and overrides so writers never meet a missing
column; let agentic ingest target a Neon branch; only invalidate the cache after
overrides actually applied.

中文:ingest 与 overrides 前先执行 admin:db:migrate,避免写入端遇到缺列;agentic ingest 可指向 Neon 分支;仅在 overrides 实际生效后失效缓存。
Serve /api/gpu-metrics from the database first and fall back to live GitHub artifacts
only when a run is not ingested; add /api/v1/gpu-metrics-point behind a revision-keyed
Blob cache; keep the public views route shape identical for stored and live runs and
normalize live NVIDIA timestamps to ISO UTC.

中文:/api/gpu-metrics 优先读数据库,仅在 run 未入库时回退到 GitHub artifact;新增以 revision 为键的 Blob 缓存 /api/v1/gpu-metrics-point;公开 views 路由对入库与 live run 保持同一形状,live NVIDIA 时间戳归一为 ISO UTC。
The /gpu-metrics explorer reads stored series, shows full-record statistics from the
ingest digest, adds points/rolling display modes with a mean-of-chips line, and refetches
on focus so a re-ingest shows up in an open tab.

中文:/gpu-metrics explorer 读取入库序列,展示 ingest 摘要的全记录统计,新增点/滑动平均显示模式与芯片均值线,并在窗口聚焦时刷新以反映重入库。
Document the telemetry digest, receipts, migration prerequisites, targeted repair and
rollback, the gpu-metrics views route, and the new API route entries.

中文:记录遥测摘要、回执、迁移前置条件、定向修复与回滚、gpu-metrics views 路由及新增 API 路由条目。
中文:将博客列表页的访问放入每个用例的准备阶段,使 Cypress 能在页面加载超时时重试。
中文:博客列表页的 E2E 用例统一替换图片优化请求,避免 Firefox 等待缩略图时页面加载超时。
中文:将 PowerX 遥测父分支合入文章对比功能分支,并解决冲突。
…h ruler

#1167 was rebased onto master twice today; #1220 had merged the 13:32 rebase
(f320e18), which still carried the date-comparison Perf Ruler, while the
14:19 rebase (0181958) removed it after #1232 landed on master. Merge the
new parent with the earlier rebase as the base, so only GPUGraph.tsx and its
component spec conflict, and resolve them by removing the ruler: the #1229
run-specific ruler share binding, the legend toggle, the hit strokes and
drag handlers, and the ruler share-link spec. `ScatterGraph` keeps
`i_rulers`; docs and the store comment now say the date comparison draws no
rulers.

中文:#1167 今日两次 rebase 到 master;#1220 已合入 13:32 的 rebase
(f320e183),其中仍含日期对比图的性能标尺,而 14:19 的 rebase(0181958d)
在 #1232 合入 master 后移除了它。以较早的 rebase 为 base 合并新父分支,
仅 GPUGraph.tsx 及其组件用例冲突,并按移除标尺解决:删除 #1229 的运行级
标尺分享绑定、图例开关、命中描边与拖拽处理,以及标尺分享链接用例。
`ScatterGraph` 保留 `i_rulers`;文档与 store 注释改为日期对比图不绘制标尺。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
中文:同步 #1220 合入 rebase 后的 #1167 与移除日期对比图标尺的结果到 PowerX 导航分支。
edwingao28 added a commit that referenced this pull request Sep 30, 2026
中文:同步 #1220 合入 rebase 后的 #1167 与移除日期对比图标尺的结果到 NVL72 建模分支。
@edwingao28
edwingao28 marked this pull request as ready for review September 30, 2026 21:54
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 3m 25s —— View job


Re-review: PowerX article panels + same-load timeline

  • Gather context (new commits since last review, CI status)
  • Review net-new source since last review
  • Runtime verification
  • Post findings

Focused on the net-new PR-authored change since the last review at 2f982f8c: e73694e0 — remove Perf Ruler from date-comparison GPUGraph (#1232) (landed via the a900e0d7 merge). Everything else in 2f982f8c..HEAD is base-branch telemetry/DB/ETL work (feat(db)/feat(etl)/feat(api) — the powerx-db-ingest parent) pulled in through the re-integration merges and reviewed on its own PR, so it's out of scope here.

e73694e0 — clean removal. The commit deletes the entire Perf Ruler subsystem from GPUGraph only (layer, legend toggle, clear action, curve/point click + drag handlers, hit strokes, analytics) — 502 lines — while ScatterGraph keeps its ruler. I verified the deletion is complete and self-consistent:

  • No dangling references to perfRuler* / perf-ruler / lastRenderCtxRef / currentZoomRenderContext / ROOFLINE_CLASS_PREFIX / usePerfRulerAxisReset anywhere in GPUGraph.tsx, ChartDisplay.tsx, or gpu-graph.cy.tsx. ✅
  • Now-unused imports are dropped (useLayoutEffect, the @/lib/d3-chart/layers/perf-ruler symbols, usePerfRulerAxisReset); RenderContext/ContinuousScale are still used elsewhere (15 refs) so they correctly stay. ✅
  • ChartDisplay renders <GPUGraph> with no ruler-related props, so nothing is left passing into a removed surface. The perfRulerLayer is removed from the customLayers array and the onPointClick ruler-mode branch is gone — point clicks in the comparison chart now always pin the tooltip, matching the commit intent. ✅
  • perf-ruler-store.ts documents the new reality accurately ("the date-comparison GPUGraph draws no rulers"); chart-0 persistence stays bound to ScatterGraph, replay chart stays local — consistent. ✅
  • The GPUGraph Cypress spec drops its ruler case with no leftover helper references. ✅

CI: Tests (Unit) ✅ and Lint & Format ✅ are green on this sync (so typecheck passes with the cleaned imports); Tests (E2E) and Tests (Skills) are still in progress — confirm they finish green before merge.

Runtime: /inference (dsr1 · Measured Power) renders real data with 0 console errors. Deleting the ruler code introduces no new runtime surface, and the static check confirms no orphaned references remain; I wasn't able to force the browser into date-comparison mode via URL params (they get hydrated away without an i_gpus selection), but prior syncs verified the date-comparison GPUGraph renders, and this change only subtracts from it.

Prior findings (all resolved in earlier syncs): measured-metric-controls.cy.tsx open/click fixed at 2f982f8c; groupConcurrencySeries wired + tested; concurrency API short-circuits the e2e clip; localized source labels; export names comparison sources; overlay/date-toggle/concurrency preserved; official-point "View power trace" restored; run rulers persisted + mobile tooltip actions reachable. Two carried-over non-blocking Bugbot notes still stand (default comparison sources stay unpinned PowerServiceComparison.tsx; derived X axis empties role panels equal-service-comparison.ts) — both degrade gracefully; not re-raising.

🟢 LGTM — no blocking issues in the new change; the Perf Ruler removal is clean, complete, and self-consistent, with the Cypress spec updated to match.

Note: @functionstackx's earlier "address feedback from slack" change request references a discussion not visible to me — please confirm those items are handled before merge.
· feat/powerx-article-parity

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit a900e0d. Configure here.

Group snapshot-backed observations across telemetry producer and exporter
revisions while retaining serving-image, recipe and topology boundaries.
Cover complete comparison/fit ladders and preserve raw endpoint provenance.
Synchronize bilingual API documentation and the packaged skill contract.

中文:保留拼接后的 PowerX 曲线对比来源。曲线快照允许不同遥测采集版本,
同时保留服务镜像、测试配置和拓扑边界;补充完整对比与拟合回归测试,
并同步中英文 API 文档与技能包说明。
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 1m 40s —— View job


Re-review: preserve stitched PowerX comparison sources

  • Gather context (new commits since last review, CI status)
  • Review net-new source since last review (75b6f61a)
  • Runtime verification
  • Post findings

Focused on the only net-new PR commit since the last review at a900e0d7: 75b6f61a — fix: preserve stitched PowerX comparison sources. It resolves the open Cursor Bugbot finding "Stitched curve sources can split" (equal-service-comparison.ts:118).

75b6f61a — correct and complete. The prior equalServiceSourceKey still folded power_audit.producer_sha / exporter_image_sha256 into the identity even after keying by curve-${snapshot}, so an append-only stitch that kept one snapshot but re-ran telemetry with a different producer/exporter would split into several Compare sources. The fix nulls both hashes in the key when a snapshot is present, keeping stitched observations as one source, while recipe_fingerprint, image, and pointTopologyKey still distinguish genuinely different configurations, and rows without a snapshot keep their own run URL + hashes:

snapshot === null ? (point.power_audit?.producer_sha ?? null) : null,
snapshot === null ? (point.power_audit?.exporter_image_sha256 ?? null) : null,

This is symmetric with the existing snapshot handling of run_url (→ curve-${snapshot}) and attempt (→ null) at lines 100–103, so the whole identity behaves consistently under stitching.

Test coverage is thorough:

  • equal-service-comparison.test.ts — the stitched-curve test now gives the two points different producer/exporter hashes and still expects one source; a new it.each asserts distinct image/recipe/decode_tp/curve_workflow_run_id do split, and another asserts producer/exporter do split when no snapshot authorizes stitching. Both directions pinned.
  • route.test.ts — API path now stitches two producers under one snapshot, asserts serviceSources length 2, preserves both producer endpoints in baseline.endpoints[].point.runUrl, and checks matchedConcurrency pairs both loads.
  • powerx-compare.cy.ts — official B200 rows share a snapshot with differing run URLs + producer/exporter and assert the baseline select collapses to 2 options (one grouped official source + overlay).
  • power-fit.test.ts updated with a producer hash on the appended point.

Contract sync kept: api-route-catalog.ts digest bumped, views-api/docs/inference.ts (EN+ZH) and the inferencex-api skill (references/dashboard-views.md + integrity.json hash) all updated to describe the stitching rule. Read-only-coverage contract stays in sync.

Chinese copy: the new EN↔ZH doc/parameter strings read naturally and correctly keep established English infra terms (telemetry producer/exporter hash, append-only, curve_workflow_run_id). No findings.

Runtime: /inference (dsr1 · Measured Power, Compare on) renders real data with 0 console errors (1 benign warning). This commit only changes source-identity computation with no new UI surface; the stitched-grouping behavior is covered by the E2E assertion above.

Prior findings (all resolved in earlier syncs): groupConcurrencySeries wired + tested; concurrency API short-circuits the e2e clip; localized source labels; export names comparison sources; overlay/date-toggle/concurrency preserved; official-point "View power trace" restored; run rulers persisted + mobile tooltip actions reachable; Perf Ruler removed from date-comparison GPUGraph; measured-metric-controls.cy.tsx open/click fixed. Two carried-over non-blocking Bugbot notes still stand (default comparison sources stay unpinned in PowerServiceComparison.tsx; derived X axis empties role panels in equal-service-comparison.ts) — both degrade gracefully; not re-raising.

🟢 LGTM — no blocking issues in the new change; the stitched-source fix is correct, symmetric with the existing snapshot handling, and directly tested in unit, API, and E2E.

Notes: (1) At review time Lint & Format is green; Tests (Unit) / Tests (E2E) / Tests (Skills) are still in progress on this sync — confirm they finish green before merge. (2) @functionstackx's earlier "address feedback from slack" change request references a discussion not visible to me — please confirm those items are handled before merge.
· feat/powerx-article-parity

This branch was successfully deployed

1 active deployment
Preview — 75b6f61a Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants