diff --git a/docs/d3-charts.md b/docs/d3-charts.md index 21d290035..94881f945 100644 --- a/docs/d3-charts.md +++ b/docs/d3-charts.md @@ -81,7 +81,7 @@ For machine-readable frontier and hinterland observations, see the [Pareto boundary API](./pareto-api.md). The server and chart share `src/lib/pareto-frontier.ts`; their input-selection pipelines differ. -Every inference scatter chart exposes a **Pareto Frontier** switch under +Inference scatter charts with a performance-preference X axis expose a **Pareto Frontier** switch under **Advanced**, off by default. Collapsing Advanced does not turn off an enabled highlight. `i_frontier=1` preserves plain shading in share links; value `2` restores the scenic background. @@ -122,12 +122,24 @@ The six measured-power metrics (average, prefill, decode, P75, P90 and percentag Modeled chassis power retains its existing behavior: **Optimal Only** on shows the minimum-power Pareto frontier, which can legitimately contain one point. Turning it off draws the upper power boundary. In that mode, the separate **Show all measurements** switch (`i_allpoints=1`) reveals off-boundary points without changing the curve. -Dividing watts by one hardware's positive, constant TDP preserves its boundary membership. A lower percentage across different chips is not, by itself, an energy-efficiency comparison. Historical rings remain attached to visible historical points. +Dividing watts by one hardware's positive, constant TDP preserves its boundary membership. A lower percentage across different chips is not, by itself, an energy-efficiency comparison. Upper boundaries use monotone interpolation between unique-X vertices, including after zoom. Curves are grouped by hardware, precision and date, and additionally by run for unofficial overlays; unrelated dates and runs never share a curve. **Perf Ruler** is available on all six measured-power axes in the single-run scatter chart (`ScatterGraph`) only; the date-comparison chart (`GPUGraph`) has no ruler. It measures the ratio of the drawn upper-boundary values at the same X coordinate, using the rendered paths after zoom; clicking an off-boundary dot selects its series' boundary, clamped to the curves' shared X range. The ratio compares power or percentage of TDP, not energy efficiency. Optimal Only changes point visibility without moving the ruler. Hidden or removed curves cannot be measured, and changing either axis clears existing rulers. Modeled chassis power keeps its existing ruler restriction while showing an upper boundary; energy and other Pareto views retain their ruler behavior. +### Observed concurrency sweeps + +`i_xmode=concurrency` is an observed-load view, not an optimization axis. It uses the exact positive `conc` values and a linear X scale. Valid metric-bearing load points remain visible even when saved Optimal Only or Best per SKU preferences are enabled; Pareto Frontier, gradient strategy labels, Perf Ruler and Replay are unavailable in this mode. Y-axis units and measured/modelled boundaries do not change. + +`groupConcurrencySeries` groups straight-line segments by hardware, precision, topology (`pointTopologyKey`), recipe fingerprint, date, run and power-comparison variant. It never joins TP4 to TP8 or 4P/4D to 16P/16D. A group with repeated concurrency values, or points without run provenance, stays as markers instead of being reduced to an arbitrary average or envelope. Markers retain their original values and identities. These segments connect observations; they do not establish a hardware-controlled comparison or estimate untested loads. + +Official and unofficial paths share this behavior. Date comparisons stay on `GPUGraph` and split each compared (date, hardware) series and each unofficial run into the same segments, drawn with linear curves and without frontier, power envelope or Optimal Only; each series keeps one line label, on its longest segment. Exact topology quick filters retain that topology's entire load sweep. Share URLs, tables and CSV exports preserve the selected `conc` coordinates. + +### Unofficial runs in date comparisons + +`GPUGraph` plots `?unofficialrun=` rows next to the compared dates as their own (run, hardware) series, `overlay-run_`. They pass the same precision, quick-filter and overlay-hardware (`activeOverlayHwTypes`) gates as `ScatterGraph`, draw as X markers in `overlayRunColor(runIndex)`, and their curves take `overlayRooflineDasharray(runIndex)` through the roofline layer's per-curve `getDasharray`. They are not date series, so `activeDates` toggles leave them on; the legend lists them first, one `UNOFFICIAL: ` group per run, and dismissing the run removes them. Pinned tooltips, official and overlay alike, offer **View power trace**, which opens the comparison Timeline focused on that trace; official points also keep **View PowerX**. + ## Gradient Roofline Labels Parallelism strategies (TP4, TEP8, DPAEP4) color roofline paths with gradient stops and use one label per contiguous strategy segment. The reasoning: diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md index 760ec12fb..2cce12fe0 100644 --- a/docs/dashboard-readonly-views.md +++ b/docs/dashboard-readonly-views.md @@ -22,28 +22,28 @@ All endpoints are GET under /api/v1/views. Unsupported and repeated keys return This table records accepted query names; behavioral tests exercise representative combinations, not the full Cartesian product of all possible filter values. -| View | Accepted query keys | -| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `cache-reuse` | `config`, `date`, `gpus`, `model`, `percentile`, `precisions`, `recipe`, `runId`, `sequence`, `tcoBasis`, `unofficialrun` | -| `calculator` | `costProvider`, `costType`, `costcap`, `date`, `format`, `gpus`, `hideSkuAboveConfigLimit`, `mode`, `model`, `mw`, `percentile`, `precisions`, `runId`, `sequence`, `target`, `tcoBasis`, `unofficialrun` | -| `collectivex` | `activeSeries`, `kvSeries`, `swapSeries`, `backend`, `epSize`, `kvOp`, `kvX`, `kvY`, `modes`, `operation`, `overlapIsl`, `pageTokens`, `percentile`, `phase`, `precision`, `runs`, `sku`, `suite`, `swapDirection`, `swapLayout`, `swapMetric`, `swapPercentile`, `version`, `yAxis` | -| `compare` | `format`, `gpus`, `model`, `scenario`, `slug`, `tiers`, `variant` | -| `current-inferencex-image` | `asOf`, `frameworks`, `hardware`, `model`, `nodeType`, `precision`, `sequence`, `spec` | -| `evaluation` | `unofficialrun`, `benchmark`, `date`, `format`, `gpus`, `model`, `precisions` | -| `first-token` | `caps`, `costProvider`, `costType`, `date`, `gpus`, `minInteractivity`, `model`, `percentile`, `precisions`, `runId`, `sequence`, `tcoBasis`, `unofficialrun` | -| `fleet` | `cache`, `costProvider`, `costType`, `format`, `gpus`, `horizon`, `metric`, `model`, `mtbi`, `mw`, `oprice`, `percentile`, `precisions`, `price`, `ramp`, `recovery`, `sequence`, `target`, `tcoBasis` | -| `gpu-metrics` | `artifact`, `chartView`, `corrXMetric`, `corrYMetric`, `direction`, `downsample`, `gpus`, `metric`, `runId`, `sort` | -| `gpu-specs` | `format`, `metric` | -| `historical` | `deployment`, `end`, `extendToDate`, `format`, `frameworks`, `gpus`, `metric`, `model`, `precisions`, `priceSource`, `sequence`, `start`, `target`, `tcoBasis`, `vendors` | -| `inference` | `allPoints`, `best`, `date`, `dates`, `deployment`, `end`, `format`, `frameworks`, `gpus`, `metric`, `model`, `optimal`, `percentile`, `power`, `precisions`, `priceSource`, `runId`, `sequence`, `spec`, `start`, `tcoBasis`, `unofficialrun`, `userCosts`, `userPowers`, `vendors`, `xmetric`, `xmode` | -| `options` | `format` | -| `overview` | `compare`, `engine`, `format`, `hwrows`, `models`, `ref`, `rows`, `tier` | -| `profit-estimator` | `cachedInputPrice`, `costProvider`, `customCosts`, `date`, `dates`, `end`, `gpus`, `inputPrice`, `labCut`, `model`, `outputPrice`, `percentile`, `powerBasis`, `precisions`, `priceSource`, `runId`, `sequence`, `start`, `target`, `tcoBasis`, `unofficialrun`, `utilization` | -| `profit-estimator-per-gigawatt` | `cachedInputPrice`, `costProvider`, `customCosts`, `date`, `dates`, `end`, `gpus`, `inputPrice`, `labCut`, `model`, `outputPrice`, `percentile`, `powerBasis`, `precisions`, `priceSource`, `runId`, `sequence`, `start`, `target`, `tcoBasis`, `unofficialrun`, `utilization` | -| `rankings` | `format`, `kind`, `model`, `scenario` | -| `reliability` | `asOf`, `format`, `gpus`, `range` | -| `submissions` | `direction`, `limit`, `lines`, `mode`, `offset`, `onChangeOnly`, `search`, `sort` | -| `video` | `artifact`, `cell`, `compare`, `costs`, `gpuBasis`, `page`, `phase`, `run`, `selected`, `slot`, `source`, `view`, `workload`, `xAxis`, `yAxis` | +| View | Accepted query keys | +| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `cache-reuse` | `config`, `date`, `gpus`, `model`, `percentile`, `precisions`, `recipe`, `runId`, `sequence`, `tcoBasis`, `unofficialrun` | +| `calculator` | `costProvider`, `costType`, `costcap`, `date`, `format`, `gpus`, `hideSkuAboveConfigLimit`, `mode`, `model`, `mw`, `percentile`, `precisions`, `runId`, `sequence`, `target`, `tcoBasis`, `unofficialrun` | +| `collectivex` | `activeSeries`, `kvSeries`, `swapSeries`, `backend`, `epSize`, `kvOp`, `kvX`, `kvY`, `modes`, `operation`, `overlapIsl`, `pageTokens`, `percentile`, `phase`, `precision`, `runs`, `sku`, `suite`, `swapDirection`, `swapLayout`, `swapMetric`, `swapPercentile`, `version`, `yAxis` | +| `compare` | `format`, `gpus`, `model`, `scenario`, `slug`, `tiers`, `variant` | +| `current-inferencex-image` | `asOf`, `frameworks`, `hardware`, `model`, `nodeType`, `precision`, `sequence`, `spec` | +| `evaluation` | `unofficialrun`, `benchmark`, `date`, `format`, `gpus`, `model`, `precisions` | +| `first-token` | `caps`, `costProvider`, `costType`, `date`, `gpus`, `minInteractivity`, `model`, `percentile`, `precisions`, `runId`, `sequence`, `tcoBasis`, `unofficialrun` | +| `fleet` | `cache`, `costProvider`, `costType`, `format`, `gpus`, `horizon`, `metric`, `model`, `mtbi`, `mw`, `oprice`, `percentile`, `precisions`, `price`, `ramp`, `recovery`, `sequence`, `target`, `tcoBasis` | +| `gpu-metrics` | `artifact`, `chartView`, `corrXMetric`, `corrYMetric`, `direction`, `downsample`, `gpus`, `metric`, `runId`, `sort` | +| `gpu-specs` | `format`, `metric` | +| `historical` | `deployment`, `end`, `extendToDate`, `format`, `frameworks`, `gpus`, `metric`, `model`, `precisions`, `priceSource`, `sequence`, `start`, `target`, `tcoBasis`, `vendors` | +| `inference` | `allPoints`, `best`, `date`, `dates`, `deployment`, `end`, `format`, `frameworks`, `gpus`, `metric`, `model`, `optimal`, `percentile`, `power`, `precisions`, `priceSource`, `runId`, `sequence`, `spec`, `start`, `tcoBasis`, `topologies`, `unofficialrun`, `userCosts`, `userPowers`, `vendors`, `xmetric`, `xmode`, `xstat`, `serviceCompare`, `serviceBaseline`, `serviceComparator`, `serviceTarget`, `roleShare`, `powerFit` | +| `options` | `format` | +| `overview` | `compare`, `engine`, `format`, `hwrows`, `models`, `ref`, `rows`, `tier` | +| `profit-estimator` | `cacheHitMode`, `cachedInputPrice`, `costProvider`, `customCosts`, `date`, `dates`, `end`, `gpus`, `inputPrice`, `labCut`, `model`, `outputPrice`, `percentile`, `powerBasis`, `precisions`, `priceSource`, `runId`, `sequence`, `start`, `target`, `tcoBasis`, `unofficialrun`, `utilization` | +| `profit-estimator-per-gigawatt` | `cacheHitMode`, `cachedInputPrice`, `costProvider`, `customCosts`, `date`, `dates`, `end`, `gpus`, `inputPrice`, `labCut`, `model`, `outputPrice`, `percentile`, `powerBasis`, `precisions`, `priceSource`, `runId`, `sequence`, `start`, `target`, `tcoBasis`, `unofficialrun`, `utilization` | +| `rankings` | `format`, `kind`, `model`, `scenario` | +| `reliability` | `asOf`, `format`, `gpus`, `range` | +| `submissions` | `direction`, `limit`, `lines`, `mode`, `offset`, `onChangeOnly`, `search`, `sort` | +| `video` | `artifact`, `cell`, `compare`, `costs`, `gpuBasis`, `page`, `phase`, `run`, `selected`, `slot`, `source`, `view`, `workload`, `xAxis`, `yAxis` | Cache reuse returns `data.recipes` and the resolved `params.recipe`. Pass a returned key as `recipe` (the UI uses `c_recipe`) to select the same TP/EP/DP-attention, worker, GPU, speculation, offload and fingerprint combination. Omitted or stale keys use the shared dashboard default. Runs use that recipe if available, otherwise their own best-covered recipe; inspect each bar's source row for its identity. Overlay-only configurations expose their own recipe choices. Layout orientation and label placement are presentation-only controls. @@ -82,6 +82,11 @@ OperatorX is feature-gated in navigation and uses page-owned contract. Zoom, theme, axis scale, labels, media playback and report expansion are renderer state. GPU interactive downsampling does not alter returned raw data or statistics. +Power boundary labels are GPU Level Measured, GPU Level Provisioned (TDP), All in +Provisioned, and All in Measured. The last combines measured GPU power with modeled +unmeasured components and PUE; it is not a wall-meter measurement. These labels and +collapsed power-assumption/availability notes do not change metric IDs, API selectors, +or calculations. Profit comparison `powerLabel` display text follows the same names. The GPU statistics table includes startup and warmup for all chips in the selected series, regardless of chip visibility. It is separate from serving-window power, J/token and selected-time-window calculations. Run telemetry is DB-first with an @@ -95,6 +100,141 @@ displays `UMBP MoRI SGLang` through October 9, 2026 in America/New_York neither API selectors nor response data, framework/hardware keys, or raw CSV exports, so no API or OpenAPI contract change is required. +The `/inference` Power Timeline is a different surface from the raw `gpu-metrics` +explorer. Its existing read-only `/api/gpu-metrics?series=power` source returns +one-second per-device buckets, role assignments and validation-source identities; +benchmark `power_audit` supplies the exact serving-window bounds. The public +`/api/v1/views/gpu-metrics` projection alone does not return that combined evidence +and must not be described as a serving-window or role-pool projection. + +Timeline share fields `i_ptaxis`, `i_ptlines`, `i_ptwindow`, `i_ptfocus`, `i_ptutility` +and `i_ptconc` are renderer state, not parameters of the raw explorer API. `i_ptconc` +keeps the chart's rows at one concurrency, so platforms can be read at one load. Window-only +display selects retained bucket timestamps within the recorded inclusive bounds; +serving-relative display uses `(bucket UTC ms - window start UTC ms) / 1000`. +Missing, nonfinite or non-increasing bounds omit that trace in either mode. +No boundary samples are interpolated, and no stored window statistics or energy +values are recomputed. Per-GPU/mean/role-pool modes reuse the existing device identity +and pool-sum calculations. Focus only dims other traces and the utility toggle only +adds registry references. The validated-window summary below the chart lists each +trace's run, attempt, validation file and per-run telemetry source (database, or +GitHub artifact fallback when any requested series was read live), then per pool the +row's validated average as stored, the largest drawn pool bucket inside the window, +and GPUs × registry TDP. Raw telemetry, full-record statistics and API responses +are unchanged by all six display settings. The reusable client helpers are +`parsePowerTimelineParams`, `powerTimelineSampleX` and `tracePools` in +`components/inference/utils/powerTimeline.ts`, and `sumPowerAt` in +`components/gpu-power/power-series.ts`. + +Date comparisons are presentation-only as well. With compared dates, `/inference` keeps the +date-comparison chart on every x-axis mode, including concurrency, draws `?unofficialrun=` +rows beside the compared series, and the Timeline display follows the same per-date legend +toggles and colours. The inference view already returns `comparisons` (one projection per +`dates` or `start`/`end` entry) and `overlays` (one per unofficial run) for every `xmode`. +The per-date toggles are renderer state, not query keys, so no API or OpenAPI contract +change is required. + +Perf Ruler share state (`i_rulers`) also stays in the browser. Both the primary +scatter chart and date/run comparison chart restore it; run-qualified curve IDs +retain `~rRUN_ID`. Tooltip scrolling and viewport limits only keep existing actions +reachable. Neither changes returned values or adds a views API parameter. + +## Fixed-sequence service comparisons + +The inference view exposes the same mean/median selector as the dashboard through +`xstat=median|mean` (default median). Mean streaming speed is **1 / mean TPOT**; +it is not the arithmetic mean of per-request speeds. Mean TTFT and E2E use their +recorded mean values. Missing means remain missing; the view never substitutes a +median. AgentX keeps its selected `percentile`, and concurrency has no statistic: +`params.xstat` resolves to null in both cases, while `xAxis.statistic` records the +effective percentile or null. + +Equal-service comparisons and the same-concurrency diagnostic are API-only analysis. +The dashboard has no source-pair selector, target input, comparison curve or matched-concurrency +table. Its role and power-fit panels remain available. + +`serviceCompare=true` adds `serviceSources`, `equalServiceCurve`, and (when +`serviceTarget` is present) `equalServiceComparison`. Select exact opaque +`serviceBaseline` and `serviceComparator` keys returned by `serviceSources`, +encoded with `URLSearchParams`; omitted selections use its first two entries. +Each source's `label` is display text only: hardware and date, plus precision, topology, +run or other details only where two sources would otherwise look the same. +Stale explicit selections remain unavailable. Missing target returns null, not an +invented operating point. Streaming-speed targets are tok/s/user; TTFT/E2E targets +are seconds. Concurrency remains a separate observed-load diagnostic, not an +equal-service comparison axis. + +That diagnostic is `matchedConcurrency`, also returned by `serviceCompare=true`: the +two selected sources paired at every concurrency either one observed. Each side is +`observed` (J/output token, mean W/GPU and streaming speed at the selected statistic, +named by `interactivityField`, plus its point identity), `missing`, or `ambiguous` +when one source's observations at that load disagree; all are listed and none is +chosen. `changePercent` is `100 × (comparator / baseline − 1)` only when both sides +were observed. Same-load pairs usually serve different speeds, so interpret this +diagnostic separately from an equal-service comparison. + +The service-comparison API uses `equal-service-comparison.ts` and consumes scoped +observed points after chart coverage/limits, before frontier and best-per-SKU +pruning, with power-comparison clones excluded. `allPoints=true` restores clipped +observations. Source keys retain hardware, precision, source run, source date, +recipe, topology and workload identity; changing display dates does not create a +new measured source. The source run is the logical curve snapshot +(`curve_workflow_run_id` / `curve_date`, see +[Append-Only Curve Extensions](./data-pipeline.md#append-only-curve-extensions)), so +points an append-only run stitched onto an older curve stay one source even when their +telemetry producer or exporter hashes differ. Recipe, image and topology remain separate +configuration identities. Rows without a snapshot id, such as unofficial overlays, retain +their own run URL, measured date and telemetry producer/exporter hashes, and +an unknown run never joins distinct rows. Comparisons never join different sources into one +interpolation bracket. All three metrics use bounded numerical linear +interpolation of the underlying quantities, then compute +`100 × (comparator / baseline − 1)`. No log-axis interpolation, extrapolation, +missing-endpoint bridging or averaging of conflicting duplicate x-values occurs. + +Metrics are mean measured GPU board W/GPU, whole-deployment output tokens/s, and +validated measured GPU J/output token. Each estimate returns its bracket point +IDs, source/run/topology/recipe identities, endpoint x/value, and whether it was +interpolated. Unavailable comparisons return explicit reasons and null metrics, +not zero. This is an operating-point comparison, not proof of a hardware-only +causal effect. Official, historical and unofficial source identities remain +visible. + +`roleShare=true` returns `roleEnergyShares` and `rolePoints` through the same shared +role helper. Each role point carries prefill and decode mean W/GPU, role-local +prefill J/input and decode J/output, and the output-token reconstruction below; +any missing figure is null. Role points follow the response's `xAxis.field`, including +the derived `p75_e2e_norm_intvty` / `p90_e2e_norm_intvty` axes; `equalServiceComparison` +does not interpolate on those axes. +Validated disaggregated prefill J/input is multiplied by same-window aggregate +J/output ÷ J/input, then compared with decode J/output. The percentage denominator +is reconstructed prefill + decode energy on one output-token basis. Missing or +invalid role measurements are omitted. This view does not mix pool-local token +denominators or claim the reconstructed sum was independently measured. + +`powerFit=true` returns `powerFits`: per source, an ordinary least-squares line of +measured mean W/GPU (over all allocated GPUs) against whole-deployment output tok/s +per allocated GPU. Each fit reports intercept `P₀`, slope `m` (J/output token), R², +n and the fitted x-range, plus registry `tdpWatts` and every observation with its +point identity. Fewer than three distinct output rates return `fit: null` with +`reason: "too-few-points"`. `P₀` is an extrapolated intercept, not measured idle +power, and R² is null when power did not vary. + +These analytical results are JSON-only: `format=csv` with any analysis enabled returns +400, rather than silently exporting only the primary chart. Ordinary CSV retains +its existing plotted-point contract. + +| Surface | Dashboard control / share parameter | Read-only API coverage | +| ---------------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------- | +| Fixed-sequence statistic | `i_mstat` | `xstat` | +| Equal-service and matched-concurrency analysis | API-only; no dashboard control or share parameter | `serviceCompare`, `serviceBaseline`, `serviceComparator`, `serviceTarget` | +| Prefill / decode roles | `i_roleshare` | `roleShare` | +| Power vs output-rate fit | `i_powerfit` | `powerFit` | + +The scatter chart's Frontier points table (shown with `i_frontier` on a measured +power metric) lists the drawn cross-platform frontier with each point's run and +attempt, and exports it as CSV. The views API has no global-frontier parameter, so +that table remains dashboard-only. + ## Verification scope Route tests cover baseline views and selected extension behavior, including @@ -118,7 +258,53 @@ OperatorX 的入口受功能开关控制,页面使用专属的 `/api/v1/operat GPU 视图优先读取已存遥测,缺少存储数据时回退到产物。全记录统计使用所选文件、 主机序列的全部芯片摘要,包含启动与 warmup;已有摘要为空时不补算,缺失读数不补零。 芯片显隐和图表降采样不改变该统计,也不改变 serving-window 或 J/token 的计算口径。 + +`/inference` 的 Power Timeline 与原始 `gpu-metrics` 浏览器是两个不同视图。 +时间线通过现有只读 `/api/gpu-metrics?series=power` 获取逐秒采样、设备角色和验证来源, +再结合基准测试 `power_audit` 中的服务窗口边界;公开的 `/api/v1/views/gpu-metrics` +本身不返回这组完整证据,不能称为服务窗口或 GPU 池投影视图。 +六个 `i_pt*` 分享字段只控制显示,不是原始浏览器 API 的参数;`i_ptconc` 只保留同一并发数的 +行,便于在同一负载下对照多个平台。窗口模式仅显示已记录 +边界内的采样点(含边界),服务起点模式将 UTC 时间减去窗口起点后换算为秒。 +边界缺失、非有限值或结束不晚于开始时,不绘制对应曲线;不插值补点,也不重算已存储的 +窗口统计或能耗。图下的有效测量窗口汇总列出每条曲线的运行、尝试次数、验证文件和 +遥测来源(数据库,或有序列需实时读取时的 GitHub 产物回退),并按 GPU 池列出基准测试行 +存储的有效平均值、窗口内所绘曲线的最大值,以及 GPU 数 × 注册表 TDP。GPU 池模式复用已有设备角色和求和函数;聚焦仅调暗其他曲线,参考线 +使用硬件注册表。原始遥测、全记录统计和 API 响应均不因这些显示设置而改变。 响应使用 private, no-store,上游 503 保留为错误响应。 +日期对比同样只影响显示。选择对比日期后,`/inference` 在所有 X 轴模式(包括并发数)下都使用日期对比图, +并与对比序列一同绘制 `?unofficialrun=` 数据;时间线显示沿用同一套按日期切换的图例和配色。 +只读 inference 视图已对每种 `xmode` 返回 `comparisons`(每个 `dates` 或 `start`/`end` +条目一份投影)和 `overlays`(每个非官方运行一份)。按日期显隐属于渲染状态,不是查询参数, +因此无需修改 API 或 OpenAPI 契约。 + +同等服务条件下的对比和相同并发下的诊断仅通过 API 提供。仪表板不提供这两项分析的来源选择、 +目标值输入、对比曲线或同并发表格;prefill/decode 角色分析和功耗拟合面板继续保留。 +`serviceCompare`、`serviceBaseline`、`serviceComparator`、`serviceTarget` 仍为 API 查询参数, +不对应仪表板控件或分享参数。固定长度工作负载的统计量、角色分析和功耗拟合的分享参数与 API +参数仍一一对应:`i_mstat` → `xstat`、`i_roleshare` → `roleShare`、`i_powerfit` → `powerFit`。 + +`serviceSources` 中各数据源的 `label` 仅供显示,由硬件和日期组成;只有两个数据源无法区分时, +才补充精度、拓扑、运行等信息。选择数据源时应使用其 `key`。 +数据源的运行标识取自逻辑曲线快照(`curve_workflow_run_id` / `curve_date`):append-only 运行 +拼接到旧曲线上的数据点仍视为同一数据源,即使其 telemetry producer 或 exporter hash 不同; +测试配置指纹、镜像和拓扑仍用于区分不同配置。没有快照标识的行(例如非官方叠加数据)继续按各自的 +运行 URL、实测日期和 telemetry producer/exporter hash 区分数据源;运行未知的行不会与其他行合并为同一数据源。 + +`serviceCompare=true` 还返回 `matchedConcurrency`:按并发数逐行配对两个所选数据源,任一方在 +该并发数下有观测即列出一行;每侧为 `observed`、`missing`,或同一负载下观测值不一致时的 +`ambiguous`(全部列出、不选其一);仅当两侧都有观测时才计算变化百分比。同一并发下两者 +的服务速度通常不同,因此该表只作诊断,不替代同等服务对比。`roleShare=true` 另返回 +`rolePoints`(各角色 W/GPU、按本池 token 计的能耗及按输出 token 重建的能耗)。角色数据点沿用 +响应的 `xAxis.field`,包括派生的 `p75_e2e_norm_intvty` / `p90_e2e_norm_intvty` 横轴; +`equalServiceComparison` 不在这些横轴上插值。 +`powerFit=true` 返回 `powerFits`:每个数据源以最小二乘法拟合平均 W/GPU 与每个已分配 +GPU 的输出 tok/s,给出 `P₀`、`m`(J/输出 token)、R²、点数、拟合范围和注册表 TDP; +不同输出速率少于 3 个时不拟合。`P₀` 是外推截距,不是实测空载功耗。这些分析结果只支持 +JSON。散点图的前沿点表(`i_frontier` 开启且为实测功耗 +指标时显示)列出跨平台前沿各点的运行和尝试次数,并可导出 CSV;只读 API 没有全局前沿 +参数。 + 测试覆盖契约同步及代表性的筛选行为,并未穷举所有参数组合。生产数据库上的 完整 UI/API 对照仍需集成审查,不能仅凭单元测试宣称已完成。 diff --git a/docs/data-pipeline.md b/docs/data-pipeline.md index a7a365103..a567c323c 100644 --- a/docs/data-pipeline.md +++ b/docs/data-pipeline.md @@ -725,8 +725,3 @@ empty receipt says `no_8k1k_points`, never that power coverage was validated. The workflow retains both the input manifest and verification receipt. Cache invalidation errors fail the workflow instead of being swallowed. Imported P75/P90 ledger edits trigger the existing reviewed override workflow. - -The dashboard availability panel uses scoped points before Y-metric filtering, -including visible unofficial overlays. It distinguishes schema-2 validation, -other validated data, missing verdicts, withheld measurements, unavailable metrics, -and non-applicable separate-pool metrics without filling missing values. diff --git a/docs/data-transforms.md b/docs/data-transforms.md index 636be8966..16e9bc9e4 100644 --- a/docs/data-transforms.md +++ b/docs/data-transforms.md @@ -64,7 +64,26 @@ Returns `{ chartData: InferenceData[][], hardwareConfig: HardwareConfig }`. - `tpPerMw` — `(tputPerGpu * 1000) / hardwarePower` (GPU power in kW, result in tok/s/MW). - Cost-per-million fields — GPU hourly cost divided by tokens-per-hour (in millions): `costh` / `costr` for owning-at-large-hyperscaler-volume / 3-year-rental pricing respectively. Three token variants exist: combined (`costh`/`costr`), output-only (`costhOutput`/`costrOutput`), and input-only (`costhi`/`costri`). The former Neocloud tier (`costn`/`costni`/`costnOutput`) was removed; its share-link keys alias to the hyperscaler-volume metric. - Infrastructure purchasing-power fields divide tokens per hour by the same hourly costs. Total (`tokensPerDollarH`/`R`), output-only (`outputTokensPerDollarH`/`R`), and input-only (`inputTokensPerDollarH`/`R`) are separate USD Y-axis metrics. The former ¥-priced `tokensPerRmb*` axes and the Neocloud `*PerDollarN` axes were removed; their share-link keys alias to the matching hyperscaler-volume $ metric. The dashboard defaults to the hyperscaler-volume total-tokens-per-dollar axis (`y_tokensPerDollarH`). Links created for the removed API-price `i_metric=y_tokensPerDollar` axis resolve to the same hyperscaler-volume variant. -- Energy fields — `jTotal` / `jOutput` / `jInput`: `(hardwarePower * 1000) / tputPerGpu` (Joules per token, where power in kW is converted to W). +- Energy fields — `jTotal` / `jOutput` / `jInput`: `(hardwarePower * 1000) / tputPerGpu` (Joules per token, where power in kW is converted to W). These divide one GPU's all-in power by that GPU's reported throughput; for disaggregated rows `output_tput_per_gpu` is per decode GPU, so `jOutput` ignores the prefill pool. Its semantics are frozen; the whole-deployment view lives in the power boundaries below. +- Power-boundary fields — see [Power boundaries](#power-boundaries). + +#### Power boundaries + +`lib/power-basis.ts` derives the three boundaries beyond GPU-measured telemetry as W per allocated GPU plus J per successful output token. `buildDerivedChartFields` calls `buildPowerBasisChartFields(entry, specs)` right after the measured fields, so the official chart, the `?unofficialrun=` overlay (`transformBenchmarkRows`), and Historical Trends (`rowToLightweightPoint` with `requestedMetrics`) share one formula set; each key is gated by `wants(key)` so selective output stays exact. + +| Basis | W / GPU | J / output token | Fields | +| --------------------------------------------- | -------------------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------------------------ | +| B1 GPU measured | `avg_power_w` | `joules_per_output_token` | existing `measuredAvgPower`, `measuredJPerOutputToken` (unchanged) | +| B2 GPU provisioned | `HW_REGISTRY.tdp` | `W × N_alloc ÷ total output tok/s` | `gpuProvisionedWatts`, `gpuProvisionedJPerOutputToken` | +| B3 utility provisioned (all-in) | `HW_REGISTRY.power × 1000` | `W × N_alloc ÷ total output tok/s` | `utilityProvisionedWatts`, `utilityProvisionedJPerOutputToken` | +| B4 utility modeled (measured → chassis → PUE) | `modeledSystemPower.deploymentFacilityWatts ÷ modeledSystemPower.gpuCount` | `B1 J/out × (B4 W ÷ B1 W)` | `utilityModeledWatts`, `utilityModeledJPerOutputToken` | + +- **N_alloc and total throughput** (`powerBasisNormalization`). Aggregate rows report output per allocated GPU, so `N_alloc` cancels and the builder uses `W ÷ output_tput_per_gpu` without trusting display counts (legacy ingest can encode TP × EP twice). Fixed-sequence disaggregated rows (`disagg && benchmark_type === 'single_turn'`) report output per decode GPU while the deployment also powers the prefill pool, so `total output tok/s = output_tput_per_gpu × num_decode_gpu` and `N_alloc = num_prefill_gpu + num_decode_gpu`; the article's 4P+4D counts eight GPUs, and B3 energy is `(P + D) / D` times `jOutput` on such rows. Other disaggregated benchmark types emit no provisioned energy because it is not verifiable in-app whether AgentX throughput already divides by all GPUs. +- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. `deploymentFacilityWatts` is chassis AC × PUE with PUE applied exactly once inside `estimateChassisPower`, divided by the physical measured GPU count (not `modeledGpuCount`, which over-counts partially allocated chassis: a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8, and the unit test pins that divisor). This is a different quantity from the existing `modeledChassisPowerPerGpu` (chassis AC ÷ modeled GPU count, no PUE). Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator). +- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware such as GB200/GB300 NVL72 (`reason: 'hardware'`), and telemetry/topology failures withhold it. The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here). +- **Ordering invariant** (unit-tested on real B200 telemetry): B3 ≥ B4 ≥ B1 and B2 ≥ B1 for both W/GPU and J/out on the same point. +- **Reconstructed prefill energy** (`reconstructedPrefillJPerOutputToken`, `utils/role-energy.ts`). For validated (`power_valid === 1`, schema 2) disaggregated rows, `prefill_joules_per_input_token × (joules_per_output_token ÷ joules_per_input_token)` carries the prefill pool's energy onto the output-token axis: the ratio is the served input:output token count because schema-2 aggregate energy has one numerator. With `decode_joules_per_output_token` it sums back to the deployment's J/out. It feeds only the `i_pcompare=roles` comparison on the energy axis (PowerX Figure 7) and is never a y-axis of its own; aggregate rows and rows missing any of the four inputs omit it. +- Historical Trends substitutes `output_tput_per_gpu := tput_per_gpu` for legacy rows lacking output throughput; provisioned energies in trends inherit that fallback. **GPU specs lookup** happens inside `buildDerivedChartFields`. `getGpuSpecs(hwKey)` (`lib/constants.ts`) splits on `[-_]` to extract the base GPU token (for example, `"b200_trt_mtp"` becomes `"b200"`) and looks it up in `HW_REGISTRY`. Missing keys return zeroed specs, producing `0` cost/energy values rather than crashing. @@ -140,6 +159,10 @@ Both charts share the same Y-axis options. The `y` field is the default `AggData - `title` / `titleZh`: bilingual dropdown and chart titles. - `polarity`: whether higher or lower values are preferable. The derived chart definitions combine this with x-axis polarity to produce the concrete Pareto corner. -**X-axis selection** takes precedence over the Y metric: selecting an input-energy metric keeps the chosen Interactivity, TTFT, or E2E axis. Fixed-sequence TTFT uses the measured median; agentic TTFT uses the selected percentile. Official data, unofficial overlays, and Replay share the resolved axis. Missing or zero-filled TTFT values are omitted instead of becoming plotted latency measurements. The legacy input-metric override (`inputTputPerGpu` declares `p90_ttft`) applies only to callers without a global x-axis mode. +**X-axis selection** takes precedence over the Y metric: selecting an input-energy metric keeps the chosen Interactivity, TTFT, or E2E axis. Fixed-sequence service axes use `i_mstat=median|mean` (default `median`); Agentic axes keep the selected percentile. Mean interactivity resolves the separate `mean_tpot_intvty` field, derived as `1 / mean_tpot` with TPOT in seconds, not the arithmetic mean of per-request speeds (`mean_intvty`, retained unchanged). Mean TTFT and E2E use the explicit `mean_ttft` and `mean_e2el` measurements. Missing, non-finite or nonpositive mean measurements are omitted, never replaced by median or raw interactivity. Axes and headings identify Mean/Median; changing the effective field clears Perf Rulers. Official data, unofficial overlays, Replay, tables and CSV share the resolved axis. Concurrency ignores this statistic. Missing or zero-filled TTFT values are omitted instead of becoming plotted latency measurements. The legacy input-metric override (`inputTputPerGpu` declares `p90_ttft`) applies only to callers without a global x-axis mode. + +**Observed concurrency** — `i_xmode=concurrency` / API `xmode=concurrency` resolves `conc` through the same `resolveXAxisField` helper for official data, unofficial overlays and the read-only inference view. It has no percentile prefix or preferred X direction: `useChartData` removes roofline directions, and the API returns no frontier/best flags and resolves optimal/best filtering off. Missing, non-finite or nonpositive concurrency is not replaced with latency. Metric-bearing observed loads bypass cost/latency display clipping; tables and CSV retain the same X values. Replay is not offered for this mode. + +`pointTopologyKey` supplies exact topology filter keys without concurrency, hardware, precision, date, run or recipe identity. The API includes `topologyKey` on each point so a caller can reuse it as a selector. Rendering adds the omitted identity dimensions before joining observed loads; see [Observed concurrency sweeps](d3-charts.md#observed-concurrency-sweeps). A topology filter does not imply that two platforms' measured concurrency sets, software or other experimental conditions match. **Limits** — both charts include `y_cost_limit: 5` (clip cost-per-million metrics above $5/M tokens) and `y_latency_limit: 60` (clip x-axis outliers beyond 60s when TTFT is on x). Tokens-per-dollar metrics are not cost-clipped; clipped cost points remain available to the dashed continuation layer. diff --git a/docs/index.md b/docs/index.md index be35d7e2e..5d7230ff3 100644 --- a/docs/index.md +++ b/docs/index.md @@ -11,6 +11,7 @@ Design rationale and non-obvious conventions. See [CLAUDE.md](../CLAUDE.md) for - [API Skill Examples](./inferencex-api-examples.md) — Install the public skill, query benchmarks, export measured PowerX data, and explain empty results - [PowerX System Power](./powerx-system-power.md) — Pinned chassis model, measured-input guards, assumptions, and reproducible article exports +- [PowerX Permanent View](./powerx-permanent-view.md) — Power boundaries as gated Measured Energy metrics, `i_metric`/`i_rulers` share links, missing-value states - [PowerX Persistence and Recovery](./powerx-persistence-recovery.md) — Telemetry receipts, migration prerequisites, and targeted repair - [API Skill Releases](./inferencex-skills-release.md) — Prepare an immutable package, verify clean installations and agent exports, and publish through the package-specific workflow - [API Skill Discovery](./inferencex-skills-discovery.md) — Accept or reject implicit skill discovery in fresh Codex and Claude Code projects diff --git a/docs/powerx-permanent-view.md b/docs/powerx-permanent-view.md new file mode 100644 index 000000000..ba51b2e70 --- /dev/null +++ b/docs/powerx-permanent-view.md @@ -0,0 +1,362 @@ +# PowerX Permanent View + +How the PowerX article's power-boundary figures live inside `/inference` instead of a +separate page. The measured-power UI, the four power boundaries, and their share links all +sit in the existing ↑↑↓↓-gated **Measured Energy** metric group. + +## Purpose + +The article compares one deployment's power at four boundaries, from the GPU board to the +utility meter, both as W per GPU and as J per output token. The dashboard already had +GPU-measured telemetry (`measured*` metrics) and an ungated all-in provisioned energy +(`jOutput`). This view adds the other boundaries as ordinary y-axis metrics so every chart +feature (zoom, tooltip, Perf Ruler, Optimal Only, Table, CSV, PNG, `?unofficialrun=` overlay, +share URL) applies unchanged. One chart engine, no second chart component. + +## Gate + +The Measured Energy group is `gated: true` in `metric-registry.ts`. `ChartControls` hides a +gated group while the gate is locked unless the selected `i_metric` belongs to it, so a +shared link to any boundary renders for a reader who never unlocked the gate, while the +group stays out of the selector otherwise. The boundary metrics are members of that group +(`POWER_BASIS_METRIC_CONFIG_KEYS` is spread into it) so they inherit exactly this behaviour. + +## Boundaries + +| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source | +| --------------------- | ---------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------- | +| `gpu-measured` | GPU measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged | +| `gpu-provisioned` | GPU provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` | +| `utility-provisioned` | Utility provisioned (all-in) | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) | +| `utility-modeled` | Utility modeled (PUE) | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` chassis AC × PUE 1.3 (applied once), ÷ measured GPUs | + +Formulas, the all-GPU normalization (`N_alloc` = prefill + decode GPUs for disaggregated +rows) and the null rules are specified in +[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis model +itself in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its +per-decode-GPU normalization; its labels and the boundary metric's `all GPUs` label keep the +two distinguishable in the selector, the availability list and CSV headers. + +Every boundary metric has `polarity: 'lower'`. Energy metrics therefore get the usual +lower-is-better Pareto frontier. The three watt keys are listed in `POWER_CURVE_METRICS` +(`utils/powerCurves.ts`), so `ScatterGraph`/`GPUGraph` draw the upper power envelope for them +exactly as for `y_measuredAvgPower`, and Optimal Only keeps every load point; the declared +`lower_*` roofline only drives the ascending Table sort. They stay out of +`isMeasuredPowerCurveMetric`, like `y_modeledChassisPowerPerGpu`, so telemetry-only decorations +do not apply. `measured-power-direction.test.ts` pins both halves. A flat provisioned series +(TDP or all-in W is one constant per hardware) has a single envelope vertex per hardware, so no +envelope line is drawn for it — only the points. + +## Why the metric key carries the boundary + +There is no `i_pbasis` URL parameter. The Boundary select in `MeasuredMetricControls` is +presentation over `selectedYAxisMetric`, exactly like Per / Scope / Statistic / Display / +Unit: `measured-metric-config.ts` gives every Measured Energy key a `basis` and resolves each +control change to the nearest registered key. Consequences: + +- One share link per figure, and `i_metric` already round-trips through every existing + surface (share button, Historical Trends hand-off, Table, CSV header). +- Existing links are byte-identical: all `measured*` keys carry `basis: 'gpu-measured'`, and + `MEASURED_METRIC_DEFAULTS` did not change. +- Choosing a derived boundary snaps to its canonical combination (whole deployment, average + watts, joules per output token). Changing any telemetry-only setting while on a derived + boundary returns to GPU measured, so no control ever points at a key that does not exist. +- A gated metric selected through a shared URL keeps its Measured controls, as the + `measured*` siblings already do. +- The Y-axis dropdown's collapsed "Measured Power" / "Measured Energy" option for the other + family resolves to `MEASURED_METRIC_DEFAULTS[family]` (GPU measured), so switching family + from the dropdown resets the boundary; the resolver itself keeps `basis` on a family change + (`changeMeasuredMetricConfig(key, { family })`), which is what a future family control in + the Measured row should call. + +The Boundary select tracks `inference_power_basis_changed { basis, family }` in addition to the +existing `inference_y_axis_metric_selected` fired by `ChartControls`. + +## Missing values + +A point without a value for the selected boundary is omitted from that series only (the +builders never emit `{ y: 0 }`, and both the official and overlay paths filter by +`metricKey in point`). + +The chart caption (`data-testid="power-basis-assumptions"`) names the boundary, the formula in +words, the PUE constant and the chassis-model revision so a screenshot records its method. +The pinned tooltip's "Modeled system power" block (`tooltipUtils.ts` `modeledSystemPowerHTML`) +renders only for `measured*` keys and `y_modeledChassisPowerPerGpu`, so on the boundary keys the +caption is the only per-chart provenance; the caption does not promise more. +When no point in the selection reports the selected telemetry axis, the chart's empty state says +the selection has no measured GPU power (`noMeasuredDataHint`) instead of the generic hint. + +## Article figure → share link + +Base: `/inference?g_model=&i_seq=8k/1k&i_prec=` plus the hardware legend +selection. Append: + +| Figure | Content | Parameters | +| ------ | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| 2 | H200 W/GPU, all four boundaries on one chart | `&i_metric=y_measuredAvgPower&i_pcompare=boundaries` (one boundary at a time: `&i_metric=y_gpuProvisionedWatts` / `y_utilityProvisionedWatts` / `y_utilityModeledWatts`) | +| 3 | H200 J/output token, all four boundaries | `&i_metric=y_measuredJPerOutputToken&i_pcompare=boundaries` (single boundary: the matching `…JPerOutputToken` key) | +| 4 / 5 | B200 vs B300, GB200 vs GB300 at fixed X | `&i_metric=&i_rulers=` (Perf Ruler, T1) | +| 6 | Prefill vs decode W/GPU with the whole deploy | `&i_metric=y_measuredAvgPower&i_pcompare=roles` | +| 7 | Reconstructed request energy, prefill share | `&i_metric=y_measuredJPerOutputToken&i_pcompare=roles` (prefill share in the clone's tooltip) | +| 1 | Measured power over the benchmark job | `&i_metric=y_measuredPowerTimeline` (Display → Timeline, see below) | + +Pinned official dashboard points now offer **View PowerX** when the feature gate is unlocked +or a measured-metric share link is active. The lazy in-page dialog reads the existing +`GET /api/v1/gpu-metrics-point?id=N` API, without entering a run ID or changing chart filters. +This also works in the hardware/date comparison view. The per-chip chart and statistics span +the recorded job (including startup/warmup), not just the audited serving window. Fixed-sequence +points do not request AgentX server-metric overlays. Missing telemetry is shown as unavailable, +not zero power. Unofficial points have no database ID and retain **View power trace**, which +resolves their run/audit provenance automatically. No API contract or ingestion changes are +needed for this presentation-only entry point. + +The `/gpu-metrics` page keeps the raw per-run explorer (every metric, one artifact at a time); +the timeline below is the chart-scoped view of the same artifacts. + +## Comparison series (`i_pcompare`) + +`i_pcompare=boundaries|roles` (default `''`, "Off") overlays sibling series on the selected +metric without changing it. `utils/power-compare.ts` is the single source: `useChartData` and +the `?unofficialrun=` processor (`processOverlayChartDataWithClipping`) both call +`expandPowerCompareSeries`, which appends one clone per sibling to every base point — same +`x`, `y` taken from the sibling's field, `powerVariant` set — and leaves the base points +untouched, so a chart without a comparison is byte-identical to before. A point lacking a +sibling's value contributes nothing to that series (never a 0), so each boundary's and +role's availability rules carry through unchanged. + +| Mode | Base series | Siblings | Fields | +| ------------ | ------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `boundaries` | the selected boundary (`basis` of the metric key) | the other three boundaries | `measuredAvgPower` / `POWER_BASIS_FIELDS[*].watts`, or `measuredJPerOutputToken` / `POWER_BASIS_FIELDS[*].energy` | +| `roles` | the selected scope (`all`, `prefill`, `decode`) | the other two worker pools | W: `measuredAvgPower`, `measuredPrefillAvgPower`, `measuredDecodeAvgPower`; J/out: `measuredJPerOutputToken`, `reconstructedPrefillJPerOutputToken`, `measuredDecodeJPerOutputToken` | + +Siblings exist only where the metric names the whole-deployment **average W/chip** or **J per +output token** (the only quantities every boundary and role publishes on one axis); role +energy additionally excludes the prefill J per input token scope. Elsewhere +`powerCompareVariants` returns nothing and the Compare select disables the option. A link that +arrives with an inapplicable mode keeps the parameter, so switching back to an applicable +setting resumes it. + +Rendering (`ScatterGraph`): the series key is `scatterSeriesKey(point)` = +`_[-v-]` (`utils/point-identity.ts`), so rooflines, frontiers, +Optimal Only, line labels, overflow continuations and the perf ruler treat each sibling as its +own series; `parseScatterSeriesKey` recovers hardware, precision and variant wherever the key +was previously split on `_`. Siblings keep the hardware colour (overlay runs keep the run +colour) and take a per-variant `stroke-dasharray` (`powerVariantDash`); clone points render +at 0.6 opacity behind their base and carry `data-power-variant`. Line labels are placed per +hardware _and_ sibling: the base series keeps its plain hardware label, while a sibling appends +` · ` (`powerLineLabelSuffix`: TDP, All-in, PUE modeled, Prefill GPUs, …) and, when a +boundary is flat on a watts axis, its shared value (`B300 (SGLang) · TDP 1.2 kW`), so an exported +PNG explains its dashed lines without the legend; each pill carries `data-series-id` +(`::` for a sibling) and `data-power-variant`. The legend appends one +line-swatch row per series present (base first); rows toggle chart-local visibility +(`hiddenPowerVariants`, not in the URL) and hover-highlight that series across every hardware. +`scatterPointConfigId` includes the variant so a clone never replaces its base in a D3 join. +Tooltips add a "Series" line; on the energy axis a role clone also reports its share of the +reconstructed request energy. Table adds a "Series" column and CSV a trailing "Power Series" +column only while clones are present. Comparison clones are excluded from `bestSeriesPerSku`, +the legend points table and the date-comparison `GPUGraph`. Unofficial-run pills read `✕ ` (`getOverlayLineLabel`); the branch stays in the +legend and a short run tag (` · main`, ` · …-`) is appended only when several overlay +runs draw the same hardware. + +Figure 7's reconstruction lives in `utils/role-energy.ts`: schema-2 aggregate energy has one +numerator, so `J/out ÷ J/in` is the served input:output token ratio and +`prefill_joules_per_input_token × (J/out ÷ J/in)` is the prefill pool's energy per output +token; with `decode_joules_per_output_token` it sums back to the deployment's J/out when the +pool energies partition the total. `buildDerivedChartFields` emits it as +`reconstructedPrefillJPerOutputToken` for validated (`power_valid === 1`, schema 2) +disaggregated rows only; it is never a y-axis of its own. + +The Compare select tracks `inference_power_compare_changed { mode, family }`; legend rows +track `inference_power_compare_series_toggled { series, visible }`. + +## Power timeline (Display → Timeline) + +`y_measuredPowerTimeline` is the third value of the Measured Power **Display** control +(`watts` / `tdp` / `timeline`, `MeasuredPowerDisplay` in `measured-metric-config.ts`). Its +registry field aliases `measuredAvgPower`, so the point set, the Table view and the share link +are those of the measured average; only the chart body changes: +`ChartDisplay` renders `ui/PowerTimeline.tsx` instead of `ScatterGraph` when the resolved +config is `display: 'timeline'`. + +- **Join.** Each validated row's `power_audit.source` is `power_validation_.json`. + Two collectors publish the telemetry behind it (`utils/powerTimeline.ts`): + - single-node runners upload one `gpu_metrics_` CSV artifact per config + (`benchmark-tmpl.yml`), so the source names the artifact exactly + (`telemetryArtifactForPoint`); + - Slurm / Dynamo disaggregated runners upload one `power_audit_` bundle per + concurrency sweep whose `LOGS/power/samples.csv` (DCGM, `/` per device) + covers every concurrency. The API cuts it into one series per `power_validation_*.json` the + bundle contains (`components/gpu-power/power-audit-bundle.ts`: the validation file's + `selected_window` ± 60 s, roles from its `per_gpu_role`, manifest `expected_devices` as the + fallback) and labels each with that file name in `series.source`, which `joinPowerTimeline` + matches before falling back to the artifact name. + + Nothing is matched by hardware or concurrency. Rows whose telemetry is missing (expired + artifact, bundle over the download cap, another collector) are not drawn, never + estimated. Trace keys are + `:` (`traceKeyForPoint`), unique per point in both collectors. + +- **Fetch.** One request per workflow run in the visible points + (`planPowerTimelineRequests`, at most `POWER_TIMELINE_MAX_RUNS`; a deep-linked trace's run + goes first, then runs holding a point the legend shows (`?unofficialrun=` overlay runs before + official ones), then runs whose points are all hidden — `prioritizeRun` / `prioritizeRuns` — so + an overlay the user asked for, or a pair left visible for comparison, is never the run that gets + dropped), + narrowed with + `prefix=` to the common RESULT_FILENAME prefix so a nightly sweep's other models are not + downloaded. `/api/gpu-metrics?series=power` first uses persisted telemetry and returns one-second per-GPU buckets + (`components/gpu-power/power-series.ts`, ~1 MB for a 25-config run instead of ~27 MB of + raw rows). Missing stored telemetry falls back to artifacts; database failures are explicit errors. + See [persistence and repair](./powerx-persistence-recovery.md) for cache freshness and coverage receipts. + With `series=power` the artifact fallback also downloads `power_audit_*` bundles whose name + shares the prefix (a bundle names the sweep, so the match runs both ways), reads only + `LOGS/power/samples.csv`, `LOGS/power/manifest.json` and the top-level + `power_validation_*.json` entries, and skips bundles above 256 MiB (the GB200 nw8 sweep is + 215 MB). A bundle-cut series carries `source` and `devices[] { id, role? }`; CSV series carry + neither. Runner timestamps are UTC wall clock and are parsed as such + (`parseTelemetryTimestampUtc`); `new Date('2026/09/12 20:19:57')` would read browser local + time and misplace the audit window. +- **Legend with overlays.** Once `?unofficialrun=` data is in, the chart reads + `localOfficialOverride`, so the timeline legend writes the unified selection + (`setUnifiedOverlaySelection`, `computeToggle` solo semantics) exactly as `ScatterGraph` + does; the context's `toggleHwType` alone would change nothing visible there. + +- **Date comparison.** With chip configs and comparison dates selected, `ChartDisplay` passes + `comparison`, and official traces become the compared (date, chip config) series of the + date-comparison `GPUGraph`. `useComparisonSeries` gives both displays the same series, run + numbers and colours. The legend lists one row per series with a trace candidate, grouped + under its hardware; a click calls `toggleActiveDate` with the same solo semantics, and a + soloed date hides the other dates' traces. Overlay runs stay in one _Unofficial run_ group + in their run colour and still follow `activeOverlayHwTypes`. The tooltip header, the focus + chip and the summary's hardware column add the date or run + (`B200 (SGLang) · 2026-09-23`); end labels lead with it only while more than one entry is + visible (`2026-09-23 c1`). + +- **Drawing.** One trace per config, mean of its GPUs (legend switch: one line per GPU), + coloured by hardware for official rows and by `overlayRunColor(runIndex)` for + `?unofficialrun=` rows; legend toggles follow `activeHwTypes` / `activeOverlayHwTypes` + exactly as in `ScatterGraph`. The whole job is drawn faint and the `power_audit` window + emphasized (`data-segment="full" | "window"`). Rated TDP is a dashed reference per hardware; + the all-in provisioned line is an opt-in legend switch because it halves the traces' + vertical resolution. X axis: wall clock (UTC) when the visible traces come from one run, + otherwise seconds since each trace's start; both are a toolbar toggle. `c` labels sit + at the end of the emphasized segment; labels that would overprint stack one row apart + (`stackTraceLabels`). +- **Pools (Figure 1).** The legend switch _Prefill / decode pools_ (shown when a visible + trace carries worker roles) sums the board power of each role's GPUs + (`tracePools` / `sumPowerAt`) and draws one line per pool — prefill dashed `7 3`, decode + `2 3`, the same dashes as the `i_pcompare=roles` series — labelled `c · prefill` / + `· decode`, on a _GPU pool power (W)_ axis. Traces without roles draw their deployment total + as one `all` line so single-node configs stay comparable. Reference lines become pool size × + rated TDP per (hardware, pool size): roles of one hardware that hold the same GPU count share + one line labelled `prefill / decode ×16` (`data-pool` on `.power-reference` lists the roles, + `groupPoolsBySize`), and labels of lines at equal watts stack upward (`referenceLabelSlots`) + instead of overprinting; and the + all-in switch scales the same way. The tooltip names the pool, its summed watts against the + pool TDP, and the mean / min / max per GPU inside it. Pool mode and per-GPU lines are + mutually exclusive. +- **Deep link.** A pinned tooltip on any metric of the measured family, in the scatter chart or + the date-comparison `GPUGraph`, offers _View power trace_ (`data-action="view-power-trace"`, + official and `?unofficialrun=` points alike). It switches the Display to Timeline in place and records the point's trace key in a + one-shot module store (`requestPowerTraceFocus`); the timeline consumes it on mount, dims + every other trace, switches to pool mode when the trace has roles, and shows a _Focused on …_ + chip (`data-testid="power-timeline-focus"`) with _Show all_ to clear. The link's `href` is + the current page with `i_metric=y_measuredPowerTimeline` only, so open-in-new-tab lands on + the unfocused timeline; once the focus is applied it is written to `i_ptfocus` like the other + timeline settings. +- **Same load across platforms.** The toolbar _Concurrency_ select (`i_ptconc`, default all) + keeps the chart's rows at one load, so, for example, GB200 and GB300 prefill/decode pools draw + side by side; `i_ptaxis=serving` aligns them at each validated window's start. With two or + more hardware types visible, every line label leads with the hardware label. A deep link + resets the filter so the focused trace is never filtered out. +- **Validated-window summary.** `ui/PowerTimelineSummary.tsx` lists each drawn trace's + hardware and config, run link, attempt, validation file, per-run telemetry source + (`database`, or `GitHub artifact fallback` when `/api/gpu-metrics` read any requested series + live) and window length. Per pool (all GPUs, prefill, decode) it shows the GPU count, the + row's validated average W/GPU as stored, the largest drawn 1-s pool sum inside the window + (`summarizeTraceWindow`) and pool TDP = GPUs × `HW_REGISTRY.tdp`. No average is recomputed. +- **State.** Axis mode (`i_ptaxis`), line mode (`i_ptlines`: mean / `gpu` / `pool`), window-only + display (`i_ptwindow`), the focused trace (`i_ptfocus`), the all-in switch (`i_ptutility`) and + the concurrency filter (`i_ptconc`) are `PowerTimeline` component state mirrored into the URL by the component itself + (`parsePowerTimelineParams` on mount, `setUrlParams` on change); defaults serialize as `''`. + Hover highlight is never shared. See + [Dashboard read-only views](./dashboard-readonly-views.md) for the renderer-only status of + these fields against the raw `gpu-metrics` API. +- **Analytics.** `inference_power_timeline_loaded { traces, missing, runs }`, + `inference_power_timeline_axis_changed { mode }`, + `inference_power_timeline_lines_changed { lines: 'mean' | 'gpu' | 'pool' }`, + `inference_power_timeline_utility_toggled { enabled }`, + `inference_power_timeline_concurrency_changed { concurrency }`, + `inference_power_trace_opened { hwKey, conc, overlay }` (scatter tooltip action), + `inference_power_timeline_focus_cleared`. + +## Analysis panels + +Below the measured chart, `ui/PowerAnalysisPanels.tsx` offers the opt-in role and power-fit panels, and the scatter chart +offers the frontier-points table. They read the chart's scoped observed points +(`observedPoints`: official and `?unofficialrun=` rows, comparison clones excluded), keep one +source per curve snapshot and recipe (`equalServiceSourceKey`: a stitched append-only curve is +one source keyed by `curve_workflow_run_id`; rows without a snapshot id key by their own run), +and colour overlay sources with `overlayRunColor`. + +| Figures | Panel | Switch (share param) | Helper | +| ------------ | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | ------------------------------------------------------ | +| 12 / 13 / 14 | Role group: W/GPU by role, role-local J/input and J/output (not added), J/output token by role with the total, prefill share | _Prefill / decode roles_ (`i_roleshare=1`) | `getRolePoints` in `utils/equal-service-comparison.ts` | +| 15 | Least-squares fit of mean W/GPU against output tok/s per allocated GPU: points, line, dashed extension to zero, P₀, P₀ ÷ TDP, m, R², n, range | _Power vs output-rate fit_ (`i_powerfit=1`) | `utils/power-fit.ts` | +| 16 | Frontier points: the drawn cross-platform frontier and each point's run and attempt | Legend _Pareto frontier_ (`i_frontier`) on a measured metric | `utils/frontier-points.ts` | + +- **Role group** follows the chart's X axis, including the trace-derived P75/P90 E2E-normalized + interactivity axes (their values live on `point.x`); equal-service interpolation stays limited + to observed service fields and reports `unsupported-axis` there. +- **Roles** use validated disaggregated rows. Each panel names its denominator. Missing role + telemetry is omitted, never drawn as zero, and share points stay unconnected. +- **Fit** needs three distinct output rates per source. Output is whole-deployment tok/s over + all allocated GPUs, on the same basis as the mean W/GPU. P₀ is an extrapolated intercept, + not measured idle power, and R² describes only that line. +- **Frontier** lists `globalParetoFrontier`'s own output, so the table is exactly what is + drawn. Ties keep the first point, official before overlay. +- **Source labels** (`getEqualServiceSources`) read hardware and snapshot date, adding + precision, topology, snapshot run, attempt, recipe, image or point only where two sources + would otherwise look the same. The opaque key stays the exact identity. +- Role and fit plots export PNG and CSV; the frontier table exports CSV only. Export subtitles + name the workload and sources. The views API returns `rolePoints` and `powerFits` + ([Dashboard read-only views](./dashboard-readonly-views.md#fixed-sequence-service-comparisons)); + it has no global-frontier parameter. +- Analytics: `inference_power_roles_toggled` and `inference_power_fit_toggled` (`{ enabled }`), + and chart-button events under `power_roles`, `power_fit` and `frontier_points`. + +Equal-service comparisons and the matched-concurrency diagnostic remain available through the +read-only inference API with `serviceCompare`, `serviceBaseline`, `serviceComparator` and +`serviceTarget`. They have no dashboard control, panel, table or share parameter. The API retains +source identity, bounded interpolation, missing-data reasons and same-load pairing semantics; +see [Dashboard read-only views](./dashboard-readonly-views.md#fixed-sequence-service-comparisons). + +同等服务条件下的对比和相同并发下的诊断继续通过只读 inference API 提供,查询参数为 +`serviceCompare`、`serviceBaseline`、`serviceComparator` 和 `serviceTarget`。仪表板不提供 +对应的控件、面板、表格或分享参数;角色分析、功耗拟合和前沿点表继续保留。API 的来源标识、 +有界插值、缺失原因和同并发配对规则保持不变,详见[仪表板只读视图](./dashboard-readonly-views.md#fixed-sequence-service-comparisons)。 + +## Tests + +- `lib/power-basis.test.ts` and `lib/chart-utils.test.ts` — the six boundary values through + the real builder, and the same fields on `?unofficialrun=` overlay rows. +- `gpu-power/power-series.test.ts`, `power-audit-bundle.test.ts` — one-second buckets with + `null` gaps, a partial pool is a gap rather than a lower sum, and bundle rows stay on their + host/device and role inside the padded window. +- `utils/powerTimeline.test.ts` — overlay runs are fetched ahead of official runs. +- `api/gpu-metrics/route*.test.ts` — the stored digest shape, DB-first serving, database + failure as 503, known missing hosts, and the retained-inventory recount before a CSV + fallback. +- `cypress/component/power-timeline.cy.tsx`, `power-compare.cy.tsx` — overlay-run colour and + the overlay hardware filter; in date comparison, per-date trace colours, legend solo toggles + and date-prefixed end labels. +- `cypress/component/gpu-graph.cy.tsx`, `cypress/e2e/inference-chart.cy.ts` and + `lib/d3-chart/layers/rooflines.test.ts` — `?unofficialrun=` runs stay on the date-comparison + `GPUGraph` in their run colour and dash on the interactivity and concurrency axes, per-curve + dashes survive display updates, and concurrency sweeps split per date, run and topology. +- `utils/matched-concurrency.test.ts`, `utils/power-fit.test.ts`, `utils/powerTimeline.test.ts` + — signed same-concurrency deltas, the least-squares fit and R², disaggregated fits on output + per allocated GPU, and the peak pool power inside the validated window. +- `cypress/component/power-analysis-panels.cy.tsx`, `frontier-points-panel.cy.tsx`, + `power-timeline.cy.tsx` and `cypress/e2e/powerx-compare.cy.ts` — the article panels on + `?unofficialrun=` overlay rows. diff --git a/docs/powerx-persistence-recovery.md b/docs/powerx-persistence-recovery.md index 1ebf2edd2..922e4ecc7 100644 --- a/docs/powerx-persistence-recovery.md +++ b/docs/powerx-persistence-recovery.md @@ -192,9 +192,13 @@ receipt recovery error also blocks the complete count, even when older data rema ## Full-record statistics The point detail, run explorer and public `/api/v1/views/gpu-metrics` projection use the -stored per-GPU digest for the existing full-record statistics, with name mappings only. +current-version per-GPU digest for the existing full-record statistics, with name mappings +only. Unversioned or outdated digests (`stats_version` ≠ `GPU_STATS_VERSION`) are recomputed +read-only from retained DB samples with the shared ingest algorithm; incomplete retained +samples leave statistics empty for the source-gap recovery path (see +[statistics upgrades](./data-pipeline.md#full-record-statistics-upgrades-migration-017)). Units and percentile/stddev definitions are unchanged. -Zero is a value; missing metrics or an empty digest remain missing. Live, un-ingested +Zero is a value; missing metrics or an empty current-version digest remain missing. Live, un-ingested artifact data still computes statistics in the browser. These tables include startup and warmup. They are not serving-window power, energy, or user-selected-window statistics. diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md index e1e1e5426..4775f66de 100644 --- a/docs/powerx-system-power.md +++ b/docs/powerx-system-power.md @@ -16,6 +16,28 @@ The pinned source currently identifies itself as **DRAFT / pending human verification**. Numerical parity establishes implementation equivalence, not empirical chassis calibration. +## Updating the model for historical results + +Modeled power is derived from retained measurements when the browser or a shared +views API transforms a benchmark row. Changing the model does not rewrite the +original GPU measurements or require a per-run database backfill. + +1. Update `REVISION` in `packages/app/scripts/generate-system-power-reference.py` + to the intended clean Python model commit, and update the recorded assumptions + when required. +2. Run that script with the path to the pinned model checkout to regenerate + `system-power-model.profiles.json` and `system-power-model.reference.json`. + If equations or load-dependent components changed, update the TypeScript + implementation too; regenerating constants alone is insufficient. +3. Run the system-power model parity and admission tests, then deploy the app. + Existing browser sessions need the updated bundle. Derived API responses need + the normal authenticated cache invalidation or cache expiry; deployment alone + does not establish that every cached response uses the new revision. +4. Regenerate frozen CSV/JSON exports separately. If the revised model needs + inputs that were never recorded, those rows stay unavailable until the input + gap is resolved. A new benchmark's power must not be attached to an older + benchmark's throughput. + ## Boundary and assumptions The input is measured mean GPU power during a validated serving window. The @@ -95,15 +117,20 @@ facility kW/GPU used to calculate capacity per GW. Consequently, revenue, compute expense, license fee, and profit scale together; profit margin does not change. Electricity expense is not recomputed separately. -This opt-in AgentX estimate requires validated schema-v2 telemetry and a -single-node chassis supported by the pinned model. Validated 1/2/4-GPU allocations -use the existing full-chassis extrapolation: fill an eight-GPU server with whole -replicas at the measured per-GPU power and throughput, then divide modeled facility -power by eight. This assumes replica co-location does not change performance or -power; it is not a measurement of a partly idle server. The chart, tooltip, and CSV -label every extrapolated estimate, including interpolation with one partial knot. -Unsupported GB200/GB300 chassis, multi-node layouts, allocations that cannot tile -eight GPUs, and missing/invalid measurements stay unavailable with distinct reasons. +This opt-in AgentX estimate requires validated schema-v2 telemetry and chassis +supported by the pinned model. Fully measured eight-GPU chassis are supported +on a single node, per measured worker host, or across an aggregate multinode +deployment without per-worker telemetry at the deployment-mean GPU power +(`topologyBasis: 'uniform-hosts'`; symmetric TP/PP/DP shards load each host alike). +Validated single-node 1/2/4-GPU allocations use full-chassis extrapolation: fill +an eight-GPU server with whole replicas at the measured per-GPU power and +throughput, then divide modeled facility power by eight. This assumes replica +co-location does not change performance or power; it is not a measurement of a +partly idle server. The chart, tooltip, and CSV label every extrapolated estimate, +including interpolation with one partial knot. Unsupported GB200/GB300 chassis, +partial multi-host allocations, disaggregated deployments without per-worker +telemetry, allocations that cannot tile eight GPUs, and missing/invalid +measurements stay unavailable with distinct reasons. The ordinary 8K/1K transformation keeps its existing admission policy. At an exact frontier point, use that point's modeled power. Between points, @@ -202,6 +229,13 @@ dropping that replicate. ## 中文说明 +模型结果在浏览器或共享 views API 转换 benchmark 数据时计算,不写回原始 GPU +测量值。更新模型时,先修改生成脚本中的固定版本及相关假设,再生成 profiles 和 +reference JSON;如果公式或随负载变化的组件有改动,还需同步 TypeScript 实现。 +通过一致性及准入测试后部署,刷新浏览器,并使派生 API 缓存失效或等待其过期。 +冻结的 CSV/JSON 需另行导出。通常无需逐 run 回填数据库;若新模型需要历史记录中 +没有的输入,应保留不可用状态,也不能把新一轮测得的功耗配到旧吞吐结果上。 + PowerX 的系统功耗结果以实测 GPU 功率为输入,使用固定版本的功耗模型估算 8-GPU 机箱的 AC 输入功率,再单独应用 PUE 得到设施功率估计。CPU 和 DRAM 利用率 均假设为 20%;这些是模型参数,不是实测利用率。完整平台配置、源码版本和校验和 diff --git a/docs/state-ownership.md b/docs/state-ownership.md index e90df44ed..0c32bcf50 100644 --- a/docs/state-ownership.md +++ b/docs/state-ownership.md @@ -113,12 +113,66 @@ them. Axis and presentation state: - selected x-axis and y-axis metrics, percentile, and effective x-axis mode +- the Measured controls (Boundary / Per / Scope / Statistic / Display / Unit) own no state: `measured-metric-config.ts` resolves each selection to the nearest registered metric key and writes it back to `selectedYAxisMetric`, so the power boundary rides on `i_metric` (there is deliberately no `i_pbasis`; see [PowerX Permanent View](./powerx-permanent-view.md)) +- the Measured Power Display value `timeline` (`y_measuredPowerTimeline`) swaps the chart body for `PowerTimeline`; its axis mode, line mode, window-only display, focused trace, all-in reference switch and concurrency filter are component state that `PowerTimeline` itself reads and writes through `useUrlState` as `i_ptaxis` / `i_ptlines` / `i_ptwindow` / `i_ptfocus` / `i_ptutility` / `i_ptconc`, so they never pass through the display domain - token-revenue price source (`i_revenue`): normalized uncached/cached/output pricing or the selected model's live OpenRouter catalog prices - scale, optimal-point, label, contrast, legend, and overlay controls Display changes stay in this domain. A contrast or label toggle therefore does not notify filter-only or data-only consumers. +**`usePerfRulerStore`** (perf rulers, `i_rulers`) + +Completed Perf Rulers on the primary inference chart (`chart-0`) are provider state, +exposed through a dedicated `PerfRulerStoreContext` +(`packages/app/src/components/inference/perf-ruler-store.ts`) rather than the display domain: a +ruler commit would otherwise rerender every display consumer, and harnesses that mount +`InferenceContextsProvider` with static mock values would silently lose ruler +interactivity. The store is scoped to one chart id on purpose. The replay chart +(`replay-chart-0`) draws the same curve classes under the same provider, so a shared +store would render every ruler twice and let the replay's prune pass delete rulers the +main chart still shows. `ScatterGraph` binds when its `chartId` is `chart-0` (ChartDisplay +mounts it as the primary chart) and falls back to component-local state when no matching +store is present; the date-comparison `GPUGraph` draws no rulers. + +`ScatterGraph` reads `state`/`setState` from the store, so the existing reducers, refs, +and draw passes are unchanged. The one thing that moved is the axis reset: +`usePerfRulerAxisReset` adjusts state during the chart's render, which is only legal for +the chart's own state, so for persisted rulers the same reset (either axis metric changes +→ rulers, draft, and pending share-link rulers are cleared) runs inside the provider. Its +key, `persistedPerfRulerAxisKey`, describes the graph `ChartDisplay` renders as `chart-0`, +not `graphs[0]`: `graphs` is always `[interactivity, e2e]` and ChartDisplay shows the e2e +graph for every non-interactivity x mode, so the key picks the graph by `selectedXAxisMode` +(as `bestHwTypes` does), uses its `x_scale_field` (which carries the percentile), and folds +the mode itself in because the derived agentic modes override the e2e graph's x field inside +ChartDisplay only. The first chart definition (null → key) and a reload of the same axes +(key → null → key) are not axis changes, so share-link rulers survive the load. + +Restoring from a link is a two-phase commit because of a load race. Rulers parsed from +`i_rulers` start as `pending`; data, `i_gpus`, and `?unofficialrun=` +overlays all arrive after the chart's first draw, and the chart prunes any committed +ruler whose curve path is absent from the DOM. The chart therefore commits a pending +ruler only once BOTH of its curve paths exist (hidden-at-opacity-0 counts as present, as +for prune), clamping the iso-x to the pair's overlap through the drawn paths. The commit +runs inside the perf-ruler D3 layer's draw pass (`drawPerfRuler`), not in a React effect: +the chart's first draw happens in a `D3Chart`-local re-render after its container is +measured, which re-renders nothing in `ScatterGraph`, so an effect keyed on the chart's +props could miss the first draw and leave resolvable rulers pending for the session. No +commit happens while the ruler mode is off (the power envelope forces it off on +non-measured axes; the mode-off effect discards pending rulers instead). Rulers whose +curves never appear stay pending — invisible and unserialized — until an axis change, +toggle-off, or clear discards them; this is a deliberate choice over a "data settled" +readiness predicate (fragile across the async order of availability, benchmark, overlay, +and derived-metric fetches). The visible cost is that a link whose curves the recipient's +view lacks lands with the ruler switch on and nothing drawn, and a later legend change on +the same axes can surface the ruler. A non-empty restore switches the ruler mode on for +that chart instance (rulers only render in mode) and fires +`interactivity_perf_ruler_shared_load` once per opened link, with `count` = the number of +rulers the link carried, the first time any of them renders — never per commit batch, so +the event counts links, not curve arrivals. Because the URL is read in a `useState` +initialiser, a retained-provider tab switch onto `/inference?i_rulers=…` does not +re-import the param — the same limitation as the other `i_*` toggles. + **`useInferenceActions`** All inference commands and setters. The context value and every exposed action keep @@ -271,8 +325,8 @@ How the GPU-across-time comparison works in the inference tab: 2. `useChartData` (in `InferenceProvider`) calls `buildComparisonDates()` to deduplicate and exclude the main `effectiveRunDate`. 3. `useQueries` fires one `useBenchmarks(model, date)` request per comparison date in parallel, alongside the main date query. 4. **Date stamping**: Each row from a comparison query is overwritten with `{ date: comparisonDates[i], actualDate: r.date }`. The `actualDate` field preserves the real DB date. Without this stamp, `activeDates` (keyed by user-selected date strings like `2025-01-15_h100-sxm`) would never match the rows' `date` field, so the toggle set would have no effect. -5. `activeDates` is a `Set` of `${date}_${gpuKey}` composite keys. It is initialised to all IDs whenever `allDateIds` changes (effect at line 473). Users toggle individual overlays on/off. -6. Rows from all dates are merged into a single `rows` array and passed through `transformBenchmarkRows` — the chart renders all of them on the same axes, coloured by GPU + date. +5. `activeDates` is a `Set` of `${date}_${gpuKey}` composite keys. It is initialised to all IDs whenever `allDateIds` changes (effect at line 473). Users toggle individual overlays on/off: the `GPUGraph` legend and, on the Measured Power Timeline display, the `PowerTimeline` legend both call `toggleActiveDate` (`computeToggle` solo semantics). `ChartDisplay`'s `visibleDateComparisonRows` applies the same set to the table, CSV export and PowerX analysis panels. `?unofficialrun=` rows are not date series, so `activeDates` never hides them; they follow `activeOverlayHwTypes`. +6. Rows from all dates are merged into a single `rows` array and passed through `transformBenchmarkRows` — the chart renders all of them on the same axes, coloured by GPU + date. `useComparisonSeries` owns the series order, run numbers and colours, so `GPUGraph` and `PowerTimeline` draw one (date, GPU) pair in the same colour. **When the latest date is selected as the main run date**: `useChartData` maps the selected date to `''` if it equals `latestAvailableDate`, reusing the no-date query key from the materialized view rather than firing a duplicate request. @@ -365,3 +419,24 @@ Dashboard scope membership is declared by `shareParamScopes` in `packages/app/src/lib/dashboard-routes.ts`. Tests enforce completeness and route-specific share behavior, so this document deliberately does not duplicate a manually maintained parameter table. + +One entry needs a note on its encoding: `i_rulers` (inference scope, default `''`) is +`serializePerfRulers` output — `isoX|curveA|curveB` per ruler joined by `;`, where the +curve ids are the rendered roofline path identity classes (`roofline-_`, +`overlay-roofline-__run`, optionally `__`) and the +iso-x is in DATA space rounded to four significant digits. `parsePerfRulers` only accepts +curve ids of that identity-class shape — the shared `roofline-path` marker class or any +other zoom-group node would match many paths and draw a ruler between arbitrary curves. +Only committed rulers are serialized (never the draft or still-pending link rulers), and +the chart prunes them against the rendered curves, so a link written from one data set +degrades to fewer rulers, never to an error. Overlay ids depend on the run order in `unofficialruns`, which the share +link already carries. + +`i_pcompare` (inference scope, default `''`) is the power comparison mode, `boundaries` or +`roles`, owned by `InferenceProvider` as `powerCompare` (display context) with +`setPowerCompare` (actions). It is written as `''` for `none`. It is deliberately NOT cleared +when the metric changes to one without a common axis for the siblings: the chart then draws +the metric alone and the Measured controls show a hint, so the link's intent survives a +detour through another setting. Which comparison rows a reader hid from the legend is +`ScatterGraph` component state (`hiddenPowerVariants`), never shared — see +[PowerX Permanent View](./powerx-permanent-view.md#comparison-series-i_pcompare). diff --git a/packages/app/cypress/component/comparison-series-theme.cy.tsx b/packages/app/cypress/component/comparison-series-theme.cy.tsx new file mode 100644 index 000000000..da37217c0 --- /dev/null +++ b/packages/app/cypress/component/comparison-series-theme.cy.tsx @@ -0,0 +1,50 @@ +import { useTheme } from 'next-themes'; + +import { useComparisonSeries } from '@/components/inference/hooks/useComparisonSeries'; +import { ThemeProvider } from '@/components/ui/theme-provider'; +import { APP_THEMES } from '@/lib/themes'; + +import { mountWithProviders } from '../support/test-utils'; + +function ComparisonColor() { + const { setTheme } = useTheme(); + const { allGraphs } = useComparisonSeries(); + + return ( + <> + + + + {allGraphs[0]?.color} + + ); +} + +describe('comparison chart themes', () => { + for (const highContrast of [false, true]) { + it(`keeps decorative themes on the dark comparison palette (${highContrast ? 'high contrast' : 'standard'})`, () => { + mountWithProviders( + + + , + { inference: { selectedGPUs: ['h100'], highContrast } }, + ); + + cy.contains('button', 'dark').click(); + cy.get('[data-testid="comparison-color"]') + .invoke('text') + .then((darkColor) => { + expect(darkColor.length).to.be.greaterThan(0); + for (const theme of ['csgo', 'gta']) { + cy.contains('button', theme).click(); + cy.get('[data-testid="comparison-color"]').should('have.text', darkColor); + } + }); + }); + } +}); diff --git a/packages/app/cypress/component/frontier-points-panel.cy.tsx b/packages/app/cypress/component/frontier-points-panel.cy.tsx new file mode 100644 index 000000000..4d1af5c0a --- /dev/null +++ b/packages/app/cypress/component/frontier-points-panel.cy.tsx @@ -0,0 +1,98 @@ +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; + +import type { InferenceData } from '@/components/inference/types'; +import FrontierPointsPanel from '@/components/inference/ui/FrontierPointsPanel'; +import { globalParetoFrontier } from '@/components/inference/utils/global-pareto'; +import { createMockInferenceData } from '../support/mock-data'; +import { mountWithProviders } from '../support/test-utils'; + +const RUN = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs'; +const OVERLAY_RUN_URL = `${RUN}/33348766792/attempts/2`; +const point = (overrides: Partial) => + createMockInferenceData({ + hwKey: 'b200_sglang', + framework: 'sglang', + precision: 'fp8', + physicalChips: 4, + tp: 4, + decode_tp: 4, + date: '2026-09-18', + run_url: `${RUN}/35317697106/attempts/1`, + image: 'lmsysorg/sglang:nightly-dev-cu13-20260918-20518d85', + ...overrides, + }); +const gb300 = (overrides: Partial) => + point({ + hwKey: 'gb300_dynamo-sglang', + framework: 'dynamo-sglang', + disagg: true, + physicalChips: 8, + num_prefill_gpu: 4, + num_decode_gpu: 4, + run_url: `${RUN}/35319969159/attempts/1`, + ...overrides, + }); +// Interactivity (higher is better) against J/output token (lower is better). +const official = [ + point({ id: 1, conc: 128, x: 31.6, y: 0.824 }), + point({ id: 2, conc: 1, x: 186.9, y: 9.065 }), + gb300({ id: 3, conc: 1, x: 206.3, y: 12.275 }), + gb300({ id: 4, conc: 4, x: 172.2, y: 4.634 }), + gb300({ id: 5, conc: 64, x: 72.2, y: 1.311 }), +]; +// Faster and cheaper than GB300 c1, so the overlay point displaces it. +const overlay = [ + point({ + id: 0, + hwKey: 'mi355x_sglang', + conc: 1, + x: 215, + y: 11, + run_url: OVERLAY_RUN_URL, + image: 'lmsysorg/sglang-rocm:v0.5.18-rocm720-mi35x-20260828', + }), +]; +const eligible = [...official, ...overlay]; +const frontier = globalParetoFrontier(eligible, true, false); + +function mountPanel() { + cy.viewport(1280, 1000); + mountWithProviders( + +
+ entry.hwKey.split('_')[0].toUpperCase()} + hardwareColor={() => 'rgb(200, 0, 0)'} + /> +
+
, + { unofficial: { runIndexByUrl: { [OVERLAY_RUN_URL]: 0 } } }, + ); +} + +describe('FrontierPointsPanel', () => { + it('lists each frontier point with its run and marks the ?unofficialrun= point', () => { + mountPanel(); + cy.get('[data-testid="frontier-points-row"]').should('have.length', 5); + cy.get('[data-testid="frontier-points-row"]') + .filter(':contains("MI355X")') + .should('contain.text', 'unofficial') + .and('contain.text', 'attempt 2') + .find('a') + .should('have.attr', 'href', OVERLAY_RUN_URL) + .and('have.text', 'run 33348766792'); + cy.get('[data-testid="frontier-points-row"]') + .filter(':contains("MI355X")') + .find('th > span') + .first() + .should('have.attr', 'style') + .and('contain', 'var(--overlay-run-0)'); + cy.get('[data-testid="frontier-points-table"]').should('not.contain.text', '206.3'); + }); +}); diff --git a/packages/app/cypress/component/gpu-graph.cy.tsx b/packages/app/cypress/component/gpu-graph.cy.tsx index bc19342c8..698a51042 100644 --- a/packages/app/cypress/component/gpu-graph.cy.tsx +++ b/packages/app/cypress/component/gpu-graph.cy.tsx @@ -1,19 +1,34 @@ import GPUGraph from '@/components/inference/ui/GPUGraph'; import { InferenceContextsProvider } from '@/components/inference/InferenceContext'; -import { useState } from 'react'; +import type { ChartDefinition, InferenceData } from '@/components/inference/types'; +import { + UnofficialRunContext, + type UnofficialRunContextType, +} from '@/components/unofficial-run-provider'; +import { useState, type ReactElement } from 'react'; import { mountWithProviders } from '../support/test-utils'; import { createMockInferenceData, createMockChartDefinition, createMockHardwareConfig, createMockInferenceContextValues, + createMockUnofficialRunContext, } from '../support/mock-data'; import { Precision, Sequence } from '@/lib/data-mappings'; +import { overlayRooflineDasharray, overlayRunColor } from '@/lib/overlay-run-style'; +import { computeToggle } from '@/lib/toggle-set'; +import { POWER_TIMELINE_METRIC_KEY } from '@/components/inference/utils/powerTimeline'; import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; const defaultChartDef = createMockChartDefinition(); const hwConfig = createMockHardwareConfig(); +// GPUGraph reads the unofficial-run context; no run is loaded unless a test says so. +const mountGpuGraph = ( + tree: ReactElement, + overrides: Parameters[1] = {}, +) => mountWithProviders(tree, { unofficial: {}, ...overrides }); + describe('GPUGraph', () => { it('renders SVG within chart container', () => { const data = [ @@ -26,7 +41,7 @@ describe('GPUGraph', () => { }), ]; - mountWithProviders( + mountGpuGraph(
{ }); it('shows empty state when data is empty', () => { - mountWithProviders( + mountGpuGraph(
{ }); it('explains missing role-local energy in GPU comparison mode', () => { - mountWithProviders( + mountGpuGraph(
{ }); it('localizes the Chinese comparison empty state', () => { - mountWithProviders( + mountGpuGraph(
{ }), ]; - mountWithProviders( + mountGpuGraph(
{ cy.get('[data-testid="gpu-graph"] svg .visible-shape').should('have.length.greaterThan', 0); }); - it('draws the historical-power ring and reports measured-point coverage', () => { - const data = [ - createMockInferenceData({ - hwKey: 'h100', - x: 32, - y: 2.1, - date: '2025-03-01', - precision: Precision.FP4, - power_tier: 'legacy', - }), - createMockInferenceData({ - hwKey: 'h100', - x: 64, - y: 1.8, - date: '2025-03-01', - precision: Precision.FP4, - power_tier: 'certified', - }), - ]; - - mountWithProviders( -
- -
, - { - inference: { - hardwareConfig: hwConfig, - selectedGPUs: ['h100'], - selectedDates: ['2025-03-01'], - selectedDateRange: { startDate: '', endDate: '' }, - activeDates: new Set(['2025-03-01_h100']), - selectedPrecisions: [Precision.FP4], - selectedYAxisMetric: 'y_measuredJPerOutputToken', - hideNonOptimal: false, - }, - }, - ); - - cy.get('#test-gpu-measured-power svg .legacy-power-ring').should('have.length', 1); - cy.get('[data-testid="measured-power-summary"]') - .should('contain.text', 'Showing 2 of 2 measured points') - .and('contain.text', '1/1 validated') - .and('contain.text', '1/1 historical'); - }); - it('shows spec decoding only on hover while retaining the offload halo', () => { const data = [ createMockInferenceData({ @@ -271,7 +234,7 @@ describe('GPUGraph', () => { }), ]; - mountWithProviders( + mountGpuGraph(
{ y_tpPerGpu_roofline: 'upper_left', }); - mountWithProviders( + mountGpuGraph(
{ }), ]; - mountWithProviders( + mountGpuGraph(
{ }), ]; - mountWithProviders( + mountGpuGraph(
{ ); } - it('reveals off-boundary measurements and historical rings without changing power envelopes or axes', () => { - mountWithProviders(); + it('reveals off-boundary measurements without changing power envelopes or axes', () => { + mountGpuGraph(); cy.get('#gpu-show-all-measurements').should('not.exist'); cy.get('#gpu-power-curves .dot-group').should('have.length', 6); - cy.get('#gpu-power-curves .legacy-power-ring').should('have.length', 2); - cy.get('[data-testid="measured-power-summary"]') - .should('contain.text', 'Showing 6 of 12 measured points') - .and('contain.text', '4/6 validated') - .and('contain.text', '2/6 historical'); cy.get('#gpu-power-curves .roofline-path') .should('have.length', 2) .each(($path) => { @@ -582,11 +540,6 @@ describe('GPU comparison power envelopes', () => { ); cy.get('#gpu-hide-non-optimal').click({ force: true }); cy.get('#gpu-power-curves .dot-group').should('have.length', 12); - cy.get('#gpu-power-curves .legacy-power-ring').should('have.length', 6); - cy.get('[data-testid="measured-power-summary"]').should( - 'contain.text', - 'Showing 12 of 12 measured points', - ); cy.get('#gpu-hide-non-optimal').should('have.attr', 'data-state', 'unchecked'); cy.get('#gpu-power-curves svg').should(($current) => { expect( @@ -599,7 +552,6 @@ describe('GPU comparison power envelopes', () => { cy.get('#gpu-show-all-measurements').should('not.exist'); cy.get('#gpu-hide-non-optimal').click({ force: true }); cy.get('#gpu-power-curves .dot-group').should('have.length', 6); - cy.get('#gpu-power-curves .legacy-power-ring').should('have.length', 2); }); cy.get('#gpu-power-curves .line-label').should('have.length', 2); cy.contains('button', 'Hide older date').click(); @@ -609,18 +561,15 @@ describe('GPU comparison power envelopes', () => { }); it('keeps boundary measurements by default toward lower latency', () => { - mountWithProviders(); + mountGpuGraph(); cy.get('#gpu-power-curves .dot-group').should('have.length', 6); cy.get('#gpu-power-curves .roofline-path') .should('have.length', 2) .each(($path) => expect($path.attr('d')).to.contain('C')); - cy.get('#gpu-power-curves [data-testid="power-curve-description"]') - .should('contain.text', 'upper power boundary') - .and('contain.text', 'not efficiency frontiers'); }); it('uses the same boundary toggle for percent TDP and fleet percentiles while preserving energy Pareto', () => { - mountWithProviders( + mountGpuGraph( , @@ -630,7 +579,6 @@ describe('GPU comparison power envelopes', () => { cy.get('#gpu-hide-non-optimal').should('have.attr', 'data-state', 'checked'); cy.get('#gpu-power-curves .dot-group').should('have.length', 6); cy.get('#gpu-power-curves .roofline-path').should('have.length', 2); - cy.get('[data-testid="power-curve-description"]').should('contain', '不代表能效 Pareto 前沿'); cy.get('#gpu-show-all-measurements').should('not.exist'); cy.get('#gpu-hide-non-optimal').click({ force: true }); cy.get('#gpu-power-curves .dot-group').should('have.length', 12); @@ -645,6 +593,209 @@ describe('GPU comparison power envelopes', () => { cy.get('#gpu-power-curves .roofline-path') .should('have.length', 2) .each(($path) => expect($path.attr('d')).to.contain('C')); - cy.get('[data-testid="power-curve-description"]').should('not.exist'); + }); +}); + +const runUrl = (id: number) => `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${id}`; + +describe('GPU comparison with unofficial runs and load sweeps', () => { + const OVERLAY_RUN_URL = runUrl(31415926535); + const DATES = ['2025-03-01', '2025-03-15']; + const ALL_SERIES = new Set(DATES.map((date) => `${date}_h100`)); + + const official = (date: string, conc: number, y: number, tp = 8) => + createMockInferenceData({ + hwKey: 'h100', + date, + conc, + tp, + x: conc, + y, + precision: Precision.FP4, + run_url: runUrl(DATES.indexOf(date) + 1000), + }); + const overlay = (conc: number, y: number) => + createMockInferenceData({ + hwKey: 'b200', + date: '2025-03-20', + conc, + tp: 8, + x: conc, + y, + precision: Precision.FP4, + run_url: OVERLAY_RUN_URL, + }); + + function Comparison({ + chartDefinition, + data, + overlayPoints, + unofficial, + context, + }: { + chartDefinition: ChartDefinition; + data: InferenceData[]; + overlayPoints: InferenceData[]; + unofficial: UnofficialRunContextType; + context?: Parameters[0]; + }) { + const [activeDates, setActiveDates] = useState(new Set(ALL_SERIES)); + const [overlayHw, setOverlayHw] = useState(new Set(['b200'])); + const value = createMockInferenceContextValues({ + hardwareConfig: hwConfig, + selectedGPUs: ['h100'], + selectedDates: DATES, + selectedDateRange: { startDate: '', endDate: '' }, + activeDates, + toggleActiveDate: (id: string) => + setActiveDates((prev) => computeToggle(prev, id, ALL_SERIES)), + selectedPrecisions: [Precision.FP4], + showLineLabels: true, + ...context, + }); + return ( + + + +
+ +
+
+
+ ); + } + + const mountComparison = ( + chartDefinition: ChartDefinition, + data: InferenceData[], + overlayPoints: InferenceData[], + context?: Parameters[0], + ) => { + const unofficial = createMockUnofficialRunContext({ + isUnofficialRun: true, + unofficialRunInfos: [ + { + id: 31415926535, + name: 'Run Sweep', + branch: 'feat/power-sweep', + sha: 'abc123', + createdAt: '2025-03-20T00:00:00Z', + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + runIndexByUrl: { [OVERLAY_RUN_URL]: 0 }, + }); + mountWithProviders( + , + ); + }; + + it('keeps unofficial runs on the date comparison in the run color and dash', () => { + mountComparison( + createMockChartDefinition({ chartType: 'interactivity', y_tpPerGpu_roofline: 'upper_left' }), + DATES.flatMap((date, d) => [8, 16, 32].map((x, i) => official(date, x, 300 - i * 60 + d))), + [8, 16, 32].map((x, i) => overlay(x, 400 - i * 60)), + ); + + cy.get('#gpu-overlay .unofficial-overlay-pt').should('have.length', 3); + cy.get('#gpu-overlay .unofficial-overlay-pt .overlay-x').each(($marker) => { + expect($marker.attr('stroke')).to.equal(overlayRunColor(0)); + }); + cy.get('#gpu-overlay .roofline-overlay-run0_b200_fp4') + .should('have.attr', 'stroke', overlayRunColor(0)) + .and('have.attr', 'stroke-dasharray', overlayRooflineDasharray(0)); + cy.get('.sidebar-legend') + .should('contain.text', 'UNOFFICIAL: feat/power-sweep') + .and('contain.text', '✕ B200'); + cy.get('#gpu-overlay .line-label').should('have.length', 3).and('contain.text', '✕ B200'); + + // A date toggle hides that official series; the unofficial run stays. + cy.get('.sidebar-legend label').contains('2025-03-15').click(); + cy.get('#gpu-overlay .dot-group').should('have.length', 3); + cy.get('#gpu-overlay .unofficial-overlay-pt').should('have.length', 3); + + cy.contains('button', 'Hide overlay hardware').click(); + cy.get('#gpu-overlay .unofficial-overlay-pt').should('not.exist'); + cy.get('#gpu-overlay .roofline-overlay-run0_b200_fp4').should('not.exist'); + cy.get('.sidebar-legend').should('not.contain.text', 'UNOFFICIAL'); + }); + + it('draws concurrency load sweeps per date, run and topology without frontier tools', () => { + mountComparison( + createMockChartDefinition({ + chartType: 'interactivity', + x_scale_field: 'conc', + y_tpPerGpu_roofline: 'upper_left', + }), + [ + // A load sweep is not a frontier: the dip at c16 stays on the line. + ...DATES.flatMap((date, d) => [ + official(date, 8, 300 + d), + official(date, 16, 250 + d), + official(date, 32, 320 + d), + ]), + // Another topology on the later date is a separate sweep. + official(DATES[1], 8, 150, 4), + official(DATES[1], 16, 180, 4), + ], + [overlay(8, 400), overlay(16, 380), overlay(32, 450)], + ); + + cy.get('#gpu-overlay .roofline-path').should('have.length', 4); + cy.get('#gpu-overlay .roofline-path').each(($path) => { + expect($path.attr('d'), 'straight segments between observations').not.to.contain('C'); + }); + cy.get('#gpu-overlay .roofline-path[class*="2025-03-01_h100_fp4"]') + .invoke('attr', 'd') + .should('match', /^M[^L]+L[^L]+L[^L]+$/u); + cy.get('#gpu-overlay .roofline-path[class*="overlay-run0_b200_fp4"]').should( + 'have.attr', + 'stroke-dasharray', + overlayRooflineDasharray(0), + ); + cy.get('#gpu-overlay .dot-group').should('have.length', 8); + cy.get('#gpu-hide-non-optimal').should('not.exist'); + }); + + it('opens the power trace from an official point on the concurrency comparison', () => { + const setSelectedYAxisMetric = cy.stub().as('setMetric'); + const audited = (date: string, conc: number, y: number): InferenceData => ({ + ...official(date, conc, y), + power_audit: { + source: `power_validation_h100_conc${conc}.json`, + } as InferenceData['power_audit'], + }); + mountComparison( + createMockChartDefinition({ chartType: 'interactivity', x_scale_field: 'conc' }), + DATES.flatMap((date, d) => [audited(date, 8, 300 + d), audited(date, 16, 320 + d)]), + [], + { selectedYAxisMetric: 'y_measuredAvgPower', setSelectedYAxisMetric }, + ); + + cy.get('#gpu-overlay .dot-group .visible-shape').first().click({ force: true }); + cy.get('[data-chart-tooltip]:visible [data-action="view-power-trace"]').click(); + cy.get('@setMetric').should('have.been.calledWith', POWER_TIMELINE_METRIC_KEY); }); }); diff --git a/packages/app/cypress/component/inference-chart-controls.cy.tsx b/packages/app/cypress/component/inference-chart-controls.cy.tsx index ded6bd9ba..3ddb08726 100644 --- a/packages/app/cypress/component/inference-chart-controls.cy.tsx +++ b/packages/app/cypress/component/inference-chart-controls.cy.tsx @@ -469,15 +469,20 @@ describe('Inference ChartControls grouped measured metrics', () => { cy.get(`input[aria-label="${searchLabel}"]`).type(group); cy.get('[data-select-option][data-value^="y_measured"]').should(($options) => { const values = [...$options].map((option) => option.dataset.value); - expect(values).to.have.length(13); - expect(new Set(values).size).to.equal(13); + // Thirteen measured axes plus the Timeline display of measured power. + expect(values).to.have.length(14); + expect(new Set(values).size).to.equal(14); + expect(values).to.include('y_measuredPowerTimeline'); }); cy.get(`input[aria-label="${searchLabel}"]`).clear().type(power); - cy.get('[data-select-option]') - .should('have.length', 1) - .and('have.text', power) - .and('have.attr', 'data-value', 'y_measuredP75Power') - .and('have.attr', 'aria-pressed', 'true'); + // 'All in Measured Power per Chip' also contains the family name; only + // the family option may display exactly that name. + cy.get('[data-select-option]').should(($options) => { + const exact = [...$options].filter((option) => option.textContent?.trim() === power); + expect(exact).to.have.length(1); + expect(exact[0].dataset.value).to.equal('y_measuredP75Power'); + expect(exact[0].getAttribute('aria-pressed')).to.equal('true'); + }); cy.get('[data-testid="option-help-y_measuredP75Power"]').should('exist'); cy.get(`input[aria-label="${searchLabel}"]`).clear().type(energy); cy.get('[data-select-option]') @@ -861,7 +866,8 @@ describe('Inference axis selector — Chinese Agentic controls', () => { ); cy.get('[data-testid="inference-secondary-controls"] > button').click(); cy.get('[data-testid="x-axis-mode-selector"]').should('contain.text', '交互性').click(); - cy.get('[role="grid"] [data-select-option]').should('have.length', 4); + cy.get('[role="grid"] [data-select-option]').should('have.length', 5); + cy.get('[data-testid="x-axis-mode-concurrency"]').should('have.text', '并发数'); cy.get('[data-testid="x-axis-mode-e2e-normalized-interactivity"]').should( 'have.text', '端到端归一化交互性', diff --git a/packages/app/cypress/component/measured-metric-controls.cy.tsx b/packages/app/cypress/component/measured-metric-controls.cy.tsx new file mode 100644 index 000000000..c1d163ced --- /dev/null +++ b/packages/app/cypress/component/measured-metric-controls.cy.tsx @@ -0,0 +1,72 @@ +import { useState } from 'react'; +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; + +import { MeasuredMetricControls } from '@/components/inference/ui/MeasuredMetricControls'; + +function Controls() { + const [metric, setMetric] = useState('y_measuredAvgPower'); + return ( +
+ + {metric} +
+ ); +} + +describe('Power boundary labels', () => { + for (const locale of ['en', 'zh'] as const) { + it(`keeps metric selection and the modeled-component note clear (${locale})`, () => { + const width = locale === 'en' ? 1280 : 390; + cy.viewport(width, 720); + cy.mount( + + + , + ); + const labels = + locale === 'en' + ? [ + 'GPU Level Measured', + 'GPU Level Provisioned (TDP)', + 'All in Provisioned', + 'All in Measured', + ] + : ['GPU 实测功耗', 'GPU 额定功耗(TDP)', '整体预配功耗', '整体实测功耗']; + cy.get('[data-testid="all-in-measured-note"]').should('not.exist'); + cy.get('[data-testid="measured-power-basis"]').click(); + cy.get('[role="option"]').should(($options) => { + expect([...$options].map((option) => option.textContent?.trim())).to.deep.equal(labels); + }); + cy.get('[role="option"][data-value="gpu-provisioned"]').click(); + cy.get('[data-testid="selected-metric"]').should('have.text', 'y_gpuProvisionedWatts'); + cy.get('[data-testid="measured-power-basis"]').click(); + cy.get('[role="option"][data-value="utility-provisioned"]').click(); + cy.get('[data-testid="selected-metric"]').should('have.text', 'y_utilityProvisionedWatts'); + cy.get('[data-testid="measured-power-basis"]').click(); + cy.get('[role="option"][data-value="utility-modeled"]').click(); + cy.get('[role="option"]').should('not.exist'); + cy.get('[data-testid="selected-metric"]').should('have.text', 'y_utilityModeledWatts'); + cy.get('[data-testid="all-in-measured-note"]') + .should('be.visible') + .and( + 'contain', + locale === 'en' ? 'unmeasured components are modeled' : '未实测的组件功耗由模型估算', + ); + cy.get('[data-testid="measured-metric-controls"]').should(($controls) => { + const box = $controls[0].getBoundingClientRect(); + expect(box.left).to.be.at.least(0); + expect(box.right).to.be.at.most(width); + expect($controls[0].scrollWidth).to.be.at.most($controls[0].clientWidth); + }); + cy.screenshot(`power-boundary-note-${locale}`, { overwrite: true }); + cy.get('[data-testid="measured-power-basis"]').click(); + cy.get('[role="option"][data-value="gpu-measured"]').click(); + cy.get('[data-testid="selected-metric"]').should('have.text', 'y_measuredAvgPower'); + cy.get('[data-testid="all-in-measured-note"]').should('not.exist'); + // Finish selection assertions before a screenshot can dismiss the open Radix menu. + cy.get('[data-testid="measured-power-basis"]').click(); + cy.get('[role="option"]').should('be.visible'); + cy.screenshot(`power-boundary-options-${locale}`, { overwrite: true }); + }); + } +}); diff --git a/packages/app/cypress/component/power-analysis-panels.cy.tsx b/packages/app/cypress/component/power-analysis-panels.cy.tsx new file mode 100644 index 000000000..2993c0e71 --- /dev/null +++ b/packages/app/cypress/component/power-analysis-panels.cy.tsx @@ -0,0 +1,163 @@ +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; +import { useState } from 'react'; + +import type { AggDataEntry, InferenceData } from '@/components/inference/types'; +import PowerAnalysisPanels from '@/components/inference/ui/PowerAnalysisPanels'; +import { writeUrlParams } from '@/lib/url-state'; + +import { createMockInferenceData } from '../support/mock-data'; +import { mountWithProviders } from '../support/test-utils'; + +const metric = (y: number) => ({ y, roof: false }); +const point = (overrides: Partial = {}) => + createMockInferenceData({ + hwKey: 'b200_sglang', + hw: 'B200', + framework: 'sglang', + precision: 'fp8', + physicalChips: 4, + tp: 4, + decode_tp: 4, + conc: 8, + date: '2026-09-23', + run_url: 'https://example.invalid/runs/900000001', + benchmark_type: 'single_turn', + mean_tpot_intvty: 20, + output_tput_per_gpu: 50, + measuredAvgPower: metric(400), + measuredJPerOutputToken: metric(10), + ...overrides, + }); +const OVERLAY_RUN_URL = 'https://example.invalid/runs/900000002'; +const overlayPoints = [100, 200].map((rate, index) => + point({ + id: 3 + index, + hwKey: 'b300_sglang', + run_url: OVERLAY_RUN_URL, + conc: 2 ** index, + output_tput_per_gpu: rate, + measuredAvgPower: metric(600 + 2 * rate), + measuredJPerOutputToken: metric(8 + 8 * index), + }), +); +const fitLadder = [20, 60, 120, 260].map((rate, index) => + point({ + id: 20 + index, + conc: 2 ** index, + mean_tpot_intvty: 100 - 20 * index, + output_tput_per_gpu: rate, + measuredAvgPower: metric(250 + 0.5 * rate), + measuredJPerOutputToken: metric((250 + 0.5 * rate) / rate), + }), +); + +function PanelsHarness({ + initial, + added = [], + overlay = overlayPoints, + xField = 'mean_tpot_intvty', +}: { + initial: InferenceData[]; + added?: InferenceData[]; + overlay?: InferenceData[]; + xField?: keyof AggDataEntry; +}) { + const [extra, setExtra] = useState([]); + return ( + +
+ {added.length > 0 && ( + + )} + +
+
+ ); +} + +const providers = { + inference: {}, + unofficial: { runIndexByUrl: { [OVERLAY_RUN_URL]: 0 } }, +}; + +describe('PowerAnalysisPanels', () => { + beforeEach(() => { + cy.on('uncaught:exception', (error) => { + if (error.message.includes('ResizeObserver loop')) return false; + }); + // URL state survives the previous component mount. + writeUrlParams({ i_roleshare: '0', i_powerfit: '0' }); + }); + + it('keeps official and overlay power fits without the retired comparison controls', () => { + mountWithProviders(, providers); + cy.get('[data-testid="power-analysis-panels"]').should('be.visible'); + cy.get('[data-testid="power-fit-toggle"]').check(); + cy.get( + '[data-testid="power-analysis-test-power-fit-plot"] circle.point[fill="var(--overlay-run-0)"]', + ).should('have.length', 2); + cy.get('[data-testid="power-analysis-test-power-fit-plot"] circle.point') + .not('[fill="var(--overlay-run-0)"]') + .should('have.length', 4); + cy.get('[data-testid="power-fit-row"]').should('have.length', 2); + cy.get('[data-testid^="equal-service-"]').should('not.exist'); + cy.get('[data-testid^="matched-concurrency"]').should('not.exist'); + }); + + it('adds late official observations to an enabled fit without losing the overlay', () => { + mountWithProviders(, providers); + cy.get('[data-testid="power-fit-toggle"]').check(); + cy.get('[data-testid="power-fit-row"]').should('have.length', 1); + cy.get('[data-testid="add-source"]').click(); + cy.get('[data-testid="power-fit-row"]').should('have.length', 2); + cy.get('[data-testid="power-analysis-test-power-fit-plot"] circle.point').should( + 'have.length', + 6, + ); + cy.get( + '[data-testid="power-analysis-test-power-fit-plot"] circle.point[fill="var(--overlay-run-0)"]', + ).should('have.length', 2); + }); + + for (const percentile of ['p75', 'p90']) { + it(`renders role observations on the derived ${percentile.toUpperCase()} axis`, () => { + const rolePoints = [...fitLadder.slice(0, 2), overlayPoints[0]].map( + (entry, index) => ({ + ...entry, + x: [31.2, 24.8, 28][index], + benchmark_type: 'agentic_traces', + disagg: true, + physicalChips: 8, + num_prefill_gpu: 4, + num_decode_gpu: 4, + measuredPrefillAvgPower: metric(300), + measuredDecodeAvgPower: metric(500), + }), + ); + mountWithProviders( + , + providers, + ); + cy.get('[data-testid="role-share-toggle"]').check(); + cy.get('[data-testid="power-analysis-test-role-power-plot"] circle.point').should( + 'have.length', + 6, + ); + cy.get( + '[data-testid="power-analysis-test-role-power-plot"] circle.point[fill="var(--overlay-run-0)"]', + ).should('have.length', 2); + }); + } +}); diff --git a/packages/app/cypress/component/power-compare.cy.tsx b/packages/app/cypress/component/power-compare.cy.tsx new file mode 100644 index 000000000..f2a51b485 --- /dev/null +++ b/packages/app/cypress/component/power-compare.cy.tsx @@ -0,0 +1,183 @@ +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; + +import ScatterGraph from '@/components/inference/ui/ScatterGraph'; +import type { InferenceData } from '@/components/inference/types'; +import { expandPowerCompareSeries } from '@/components/inference/utils/power-compare'; +import { Precision } from '@/lib/data-mappings'; +import { overlayRunColor } from '@/lib/overlay-run-style'; + +import { + createMockChartDefinition, + createMockHardwareConfig, + createMockInferenceData, +} from '../support/mock-data'; +import { mountWithProviders } from '../support/test-utils'; + +// Power comparison series (`i_pcompare`) on `?unofficialrun=` overlays: role +// clones draw in the overlay run colour with the role dash, and their legend +// rows toggle them. + +const OVERLAY_RUN_ID = 31415926535; +const OVERLAY_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${OVERLAY_RUN_ID}`; +const hwConfig = createMockHardwareConfig(); +const chartDefinition = createMockChartDefinition({ + chartType: 'interactivity', + y_measuredAvgPower: 'measuredAvgPower.y', + y_measuredAvgPower_roofline: 'lower_left', +}); + +const metric = (y: number) => ({ y, roof: false }); + +/** + * Three measured points with both roles available. Power falls as x rises so + * every point sits on the upper power envelope the chart draws for watt axes. + */ +function measuredCurve(hwKey: string, run_url?: string): InferenceData[] { + return [ + [8, 700, 840, 500], + [16, 600, 820, 450], + [32, 500, 800, 400], + ].map(([x, measured, prefill, decode]) => + createMockInferenceData({ + hwKey, + x, + conc: x, + y: measured, + precision: Precision.FP4, + run_url, + disagg: true, + measuredAvgPower: metric(measured), + measuredPrefillAvgPower: metric(prefill), + measuredDecodeAvgPower: metric(decode), + }), + ); +} + +function mountCompare(data: InferenceData[], overlay: InferenceData[]) { + mountWithProviders( + +
+ +
+
, + { + inference: { + selectedYAxisMetric: 'y_measuredAvgPower', + hardwareConfig: hwConfig, + activeHwTypes: new Set(['b200', 'h100']), + hwTypesWithData: new Set(['b200', 'h100']), + selectedPrecisions: [Precision.FP4], + hideNonOptimal: false, + showLineLabels: false, + }, + unofficial: { + isUnofficialRun: true, + activeOverlayHwTypes: new Set(['h100']), + allOverlayHwTypes: new Set(['h100']), + runIndexByUrl: { [OVERLAY_RUN_URL]: 0, [String(OVERLAY_RUN_ID)]: 0 }, + unofficialRunInfos: [ + { + id: OVERLAY_RUN_ID, + name: 'powerx-compare', + branch: 'powerx-compare', + sha: 'abc000', + createdAt: '2026-09-01T00:00:00Z', + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + }, + }, + ); +} + +const svg = '#power-compare-test svg'; +const legend = '#power-compare-test [data-testid="chart-legend"]'; + +describe('ScatterGraph power comparison series', () => { + beforeEach(() => { + cy.on('uncaught:exception', (error) => { + if (error.message.includes('ResizeObserver loop')) return false; + }); + }); + + it('draws role siblings for ?unofficialrun= overlays in the run colour with the role dash', () => { + const official = expandPowerCompareSeries(measuredCurve('b200'), 'y_measuredAvgPower', 'roles'); + const overlay = expandPowerCompareSeries( + measuredCurve('h100', OVERLAY_RUN_URL), + 'y_measuredAvgPower', + 'roles', + ); + mountCompare(official, overlay); + + cy.get(`${svg} .roofline-path[data-power-variant="prefill"]`).should( + 'have.attr', + 'stroke-dasharray', + '7 3', + ); + cy.get(`${svg} .unofficial-overlay-pt`).should('have.length', 9); + cy.get(`${svg} .overlay-roofline-path[data-power-variant="decode"]`) + .should('have.length', 1) + .and('have.attr', 'stroke', overlayRunColor(0)) + .and('have.attr', 'stroke-dasharray', '2 3'); + cy.get(`${svg} .overlay-roofline-path:not([data-power-variant])`).should('have.length', 1); + cy.get(legend).within(() => { + cy.contains('All GPUs').should('exist'); + cy.contains('Prefill GPUs').should('exist'); + cy.contains('Decode GPUs').click(); + }); + cy.get(`${svg} .overlay-roofline-path[data-power-variant="decode"]`).should( + 'have.css', + 'opacity', + '0', + ); + cy.get(`${svg} .unofficial-overlay-pt`).then(($points) => { + const hidden = [...$points].filter((point) => getComputedStyle(point).opacity === '0'); + expect(hidden).to.have.length(3); + }); + }); + + it('keeps the base series lit when its legend row is hovered', () => { + const official = expandPowerCompareSeries(measuredCurve('b200'), 'y_measuredAvgPower', 'roles'); + const overlay = expandPowerCompareSeries( + measuredCurve('h100', OVERLAY_RUN_URL), + 'y_measuredAvgPower', + 'roles', + ); + mountCompare(official, overlay); + + cy.get(`${svg} .roofline-path[data-power-variant="prefill"]`).should('exist'); + // Base points and rooflines carry no variant id (only role siblings are + // cloned), so the base row has to map back onto them. + cy.get(legend).contains('label', 'All GPUs').trigger('mouseover'); + // Wait for the siblings to settle first: opacity transitions over 150 ms, + // so a base check taken at once would still read the pre-hover value. + cy.get(`${svg} .roofline-path[data-power-variant="prefill"]`).should( + 'have.css', + 'opacity', + '0.15', + ); + cy.get(`${svg} .roofline-path:not([data-power-variant])`).should('have.css', 'opacity', '1'); + cy.get(legend).contains('label', 'All GPUs').trigger('mouseout'); + cy.get(`${svg} .roofline-path[data-power-variant="prefill"]`).should( + 'have.css', + 'opacity', + '1', + ); + }); +}); diff --git a/packages/app/cypress/component/power-metric-availability.cy.tsx b/packages/app/cypress/component/power-metric-availability.cy.tsx deleted file mode 100644 index 6a6c7b1a6..000000000 --- a/packages/app/cypress/component/power-metric-availability.cy.tsx +++ /dev/null @@ -1,88 +0,0 @@ -import { useState } from 'react'; -import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; -import { - PowerMetricAvailability, - PowerMetricAvailabilityPanel, -} from '@/components/inference/ui/PowerMetricAvailability'; -import { createMockInferenceData } from '../support/mock-data'; -import { mountWithProviders } from '../support/test-utils'; -import { Precision, Model, Sequence } from '@/lib/data-mappings'; - -const source = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/123/attempts/2'; -const base = { model: Model.Qwen3_5, precision: Precision.FP8, run_url: source }; -const missing = createMockInferenceData({ ...base, hwKey: 'gb200' }); -const invalid = createMockInferenceData({ - ...base, - hwKey: 'mi355x', - power_valid: 0, - power_invalid_reasons: ['sampling_gap_exceeded'], -}); -const measured = createMockInferenceData({ - ...base, - hwKey: 'b200', - power_valid: 1, - power_metric_schema_version: 2, - measuredAvgPower: { y: 640, roof: false }, -}); -function Panel() { - const [metric, select] = useState('y_measuredAvgPower'); - return ( - - ); -} - -describe('PowerX metric availability', () => { - it('explains missing GB and withheld AMD data, exposes source links, and switches metrics', () => { - cy.mount(); - cy.contains('1 of 3 points have this metric'); - cy.contains('Validation failed: 1'); - cy.contains('Metric not reported: 1'); - cy.contains('summary', 'source details').click(); - cy.contains('sampling_gap_exceeded'); - cy.contains('a', 'Source run').should('have.attr', 'href', source); - cy.contains('summary', 'all measured metrics').click(); - cy.contains('button', 'Measured Prefill Power per Chip').click(); - cy.contains('No separate worker pools: 3'); - cy.contains('0 of 3 points have this metric'); - }); - it('uses Chinese copy for the same coverage states', () => { - cy.mount( - - - , - ); - cy.contains('3 个数据点中有 1 个提供此指标'); - cy.contains('验证失败: 1'); - cy.contains('summary', '所有实测指标的可用性').click(); - cy.contains('缺失值不会被替换为零或 TDP 估算值'); - }); - it('counts filtered unofficial rows before the metric filter and respects hidden overlay hardware', () => { - mountWithProviders( - , - { - inference: { - selectionPoints: [missing], - activeHwTypes: new Set(['gb200']), - selectedModel: Model.Qwen3_5, - selectedSequence: Sequence.EightK_OneK, - selectedPrecisions: [Precision.FP8], - }, - unofficial: { - isUnofficialRun: true, - activeOverlayHwTypes: new Set(['b200']), - getOverlayData: () => ({ data: [measured, invalid], hardwareConfig: {} }), - }, - }, - ); - cy.contains('1 of 2 points have this metric'); - cy.contains('Metric not reported: 1'); - cy.get('[data-testid="power-metric-availability"]').should( - 'not.contain', - 'Validation failed: 1', - ); - }); -}); diff --git a/packages/app/cypress/component/power-telemetry.cy.tsx b/packages/app/cypress/component/power-telemetry.cy.tsx new file mode 100644 index 000000000..332a5fde4 --- /dev/null +++ b/packages/app/cypress/component/power-telemetry.cy.tsx @@ -0,0 +1,205 @@ +import { QueryClient, QueryClientProvider, useQuery } from '@tanstack/react-query'; +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; +import { useState } from 'react'; + +import { TelemetryDisplayControls } from '@/components/gpu-power/TelemetryDisplayControls'; +import { + DEFAULT_TELEMETRY_DISPLAY, + type TelemetryDisplayState, +} from '@/components/gpu-power/telemetry-smoothing'; +import { PowerTelemetryView } from '@/components/inference/agentic-point/power-telemetry-view'; +import type { GpuMetricsPointPayload, GpuMetricSeries } from '@/hooks/api/use-gpu-metrics-point'; +import { registerAnalyticsClient } from '@/lib/analytics'; + +const ID = 206887; +const endpoint = `/api/v1/gpu-metrics-point?id=${ID}`; +const queryKey = ['gpu-metrics-point', ID]; +const chart = '[data-testid="power-telemetry-chart"]'; + +function series(id: number, host: string, base: number): GpuMetricSeries { + const data = [0, 1, 2].flatMap((second) => + [0, 1].map((index) => ({ + timestamp: `2026-09-21T00:00:0${second}Z`, + index, + power: base + index * 100 + second * 10, + temperature: 40 + index + second, + })), + ); + return { + id, + artifactName: 'gpu_metrics_qwen35_b200', + configKey: 'qwen35_b200_tp8', + fileName: `${host}/gpu_metrics.csv`, + vendor: 'nvidia', + sampleIntervalS: 1, + sampleCount: data.length, + gpuCount: 2, + startedAt: data[0].timestamp, + endedAt: data.at(-1)!.timestamp, + sidecars: {}, + benchmarkResultIds: [ID], + stats: [0, 1].map((gpuIndex) => ({ + gpuIndex, + metric: 'power_w', + count: 3, + min: base + gpuIndex * 100, + max: base + gpuIndex * 100 + 20, + mean: base + gpuIndex * 100 + 10, + median: base + gpuIndex * 100 + 10, + p95: base + gpuIndex * 100 + 19, + p99: base + gpuIndex * 100 + 19.8, + stddev: Math.sqrt(200 / 3), + })), + data, + }; +} + +const payload: GpuMetricsPointPayload = { + benchmarkResultId: ID, + series: [series(1, 'host-a', 500), series(2, 'host-b', 700)], +}; + +function QueryStatus() { + const { status } = useQuery({ queryKey, enabled: false }); + return {status}; +} + +function mountPoint(path = '/inference/agentic/206887') { + const client = new QueryClient({ defaultOptions: { queries: { retry: false, gcTime: 0 } } }); + cy.mount( + + +
+ + +
+
+
, + ); + return client; +} + +function Controls() { + const [display, setDisplay] = useState(DEFAULT_TELEMETRY_DISPLAY); + return ( + <> + + {JSON.stringify(display)} + + ); +} + +describe('PowerX telemetry interactions', () => { + beforeEach(() => { + registerAnalyticsClient({ capture: cy.stub().as('capture') }); + }); + + it('changes display, window and chip aggregation independently and tracks each action', () => { + cy.mount(); + cy.get('[data-testid="display-mode-points"]').click(); + cy.get('#display-window').should('not.exist'); + cy.get('[data-testid="display-series-mean"]').click(); + cy.get('[data-testid="display-mode-rolling"]').click(); + cy.get('#display-window').click(); + cy.get('[role="option"]').contains('60 s').click(); + cy.get('output').should('have.text', '{"mode":"rolling","windowS":60,"series":"mean"}'); + cy.get('@capture').should('have.been.calledWith', 'power_test_display_mode_changed', { + mode: 'points', + }); + cy.get('@capture').should('have.been.calledWith', 'power_test_series_mode_changed', { + series: 'mean', + }); + cy.get('@capture').should('have.been.calledWith', 'power_test_smoothing_window_changed', { + windowS: 60, + }); + cy.get('@capture').its('callCount').should('eq', 4); + }); + + it('shows loading, distinguishes an initial failure from missing data, and retries', () => { + let failed = true; + cy.intercept('GET', endpoint, (request) => { + request.reply(failed ? { statusCode: 503, delay: 150, body: {} } : { body: payload }); + }).as('telemetry'); + mountPoint(); + cy.get('[data-testid="power-telemetry-loading"]').should('contain.text', 'Loading PowerX'); + cy.wait('@telemetry'); + cy.get('[data-testid="power-telemetry-query-error"]').should('contain.text', 'Failed to load'); + cy.get('[data-testid="power-telemetry-missing"]').should('not.exist'); + cy.then(() => { + failed = false; + }); + cy.contains('button', 'Retry').click(); + cy.wait('@telemetry'); + cy.get(chart).find('svg .point').should('have.length', 6); + cy.get('@capture').should( + 'have.been.calledWith', + 'inference_agentic_power_telemetry_retry_clicked', + ); + }); + + it('shows a genuine missing point without offering an error retry', () => { + cy.intercept('GET', endpoint, { statusCode: 404, body: {} }); + mountPoint(); + cy.get('[data-testid="power-telemetry-missing"]').should('contain.text', `#${ID}`); + cy.get('[data-testid="power-telemetry-query-error"]').should('not.exist'); + cy.get(chart).should('not.exist'); + }); + + it('preserves a rendered chart when a background refetch fails', () => { + let failed = false; + cy.intercept('GET', endpoint, (request) => { + request.reply(failed ? { statusCode: 503, body: {} } : { body: payload }); + }).as('telemetry'); + const client = mountPoint(); + cy.wait('@telemetry'); + cy.get(chart).find('svg .point').should('have.length', 6); + cy.then(() => { + failed = true; + return client.refetchQueries({ queryKey }); + }); + cy.wait('@telemetry'); + cy.get('[data-testid="query-status"]').should('have.text', 'error'); + cy.get(chart).find('svg .point').should('have.length', 6); + cy.get('[data-testid="power-telemetry-query-error"]').should('not.exist'); + }); + + for (const [locale, width] of [ + ['en', 1280], + ['zh', 390], + ] as const) { + it(`keeps chip filters scoped to a host and renders ${locale} at ${width}px`, () => { + cy.viewport(width, 900); + cy.intercept('GET', endpoint, { body: payload }); + mountPoint(`${locale === 'zh' ? '/zh' : ''}/inference/agentic/${ID}`); + cy.get(chart).find('svg .point').should('have.length', 6); + cy.get('[data-testid="power-telemetry-sample-count"]').should('have.text', '6'); + cy.get('table tbody tr').first().find('td').eq(4).should('have.text', '510.0'); + cy.get('[data-testid="chart-legend"]') + .contains(locale === 'zh' ? '芯片 0' : 'Chip 0') + .click(); + cy.get(chart).find('svg .point').should('have.length', 3); + cy.get('#power-telemetry-series-select').click(); + cy.get('[role="option"]').contains('host-b').click(); + cy.get(chart).find('svg .point').should('have.length', 6); + cy.get('table tbody tr').first().find('td').eq(4).should('have.text', '710.0'); + cy.get('[data-testid="power-telemetry-display-mode-points"]').click(); + cy.get('[data-testid="power-telemetry-display-series-mean"]').click(); + cy.get(chart).find('svg .point').should('have.length', 3); + cy.get('[data-testid="power-telemetry-view"]').should(($view) => { + expect($view[0].scrollWidth).to.be.at.most($view[0].clientWidth + 1); + }); + cy.get('[data-testid="power-telemetry-metric-select"]').should( + 'contain.text', + locale === 'zh' ? '功耗' : 'Power', + ); + cy.get('[data-testid="power-telemetry-view"]').screenshot( + `power-telemetry-${locale}-${width}`, + ); + }); + } +}); diff --git a/packages/app/cypress/component/power-timeline.cy.tsx b/packages/app/cypress/component/power-timeline.cy.tsx new file mode 100644 index 000000000..bfd8ae8ad --- /dev/null +++ b/packages/app/cypress/component/power-timeline.cy.tsx @@ -0,0 +1,458 @@ +import { PathnameContext } from 'next/dist/shared/lib/hooks-client-context.shared-runtime'; +import { useState } from 'react'; + +import type { GpuPowerSeries, GpuPowerSeriesResponse } from '@/components/gpu-power/power-series'; +import { InferenceContextsProvider } from '@/components/inference/InferenceContext'; +import PowerTimeline from '@/components/inference/ui/PowerTimeline'; +import type { InferenceData } from '@/components/inference/types'; +import { traceKeyForPoint } from '@/components/inference/utils/powerTimeline'; +import { Model, Precision, Sequence } from '@/lib/data-mappings'; +import { overlayRunColor } from '@/lib/overlay-run-style'; +import { computeToggle } from '@/lib/toggle-set'; + +import { + createMockHardwareConfig, + createMockInferenceContextValues, + createMockInferenceData, + createMockUnofficialRunContext, +} from '../support/mock-data'; +import { mountWithProviders } from '../support/test-utils'; + +// PowerTimeline joins chart points to `gpu_metrics_` artifacts +// by the `power_audit.source` file name and draws one trace per config. Overlay +// runs keep their run colour and follow the overlay hardware filter. + +const RUN_ID = '34716669498'; +const RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${RUN_ID}`; +const OVERLAY_RUN_ID = '31415926535'; +const OVERLAY_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${OVERLAY_RUN_ID}`; +const SECOND_RUN_ID = '34716669499'; +const SECOND_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${SECOND_RUN_ID}`; +const START_MS = Date.UTC(2026, 8, 12, 20, 20, 0); +const hwConfig = createMockHardwareConfig(); +const HW_TYPES = new Set(['b200', 'h100']); + +const resultName = (hardware: string, conc: number) => + `dsv4_8k1k_fp4_sglang_tp8-pp1-dcp1-pcp1-ep1-dpafalse_disagg-false_spec-none_conc${conc}_${hardware}-host-0123456789abcdef0123`; +const WINDOW = { + // Window covers the last 20 s of a 60 s job. + window_start_unix: (START_MS + 40_000) / 1000, + window_end_unix: (START_MS + 60_000) / 1000, +}; + +function measuredPoint( + hwKey: string, + conc: number, + watts: number, + overrides: Partial = {}, +): InferenceData { + return createMockInferenceData({ + hwKey, + conc, + tp: 8, + x: conc, + y: watts, + precision: Precision.FP4, + run_url: RUN_URL, + measuredAvgPower: { y: watts, roof: true }, + measuredPowerTimeline: { y: watts, roof: true }, + power_audit: { + source: `power_validation_${resultName(hwKey.split('_')[0], conc)}.json`, + ...WINDOW, + }, + ...overrides, + }); +} + +/** 61 one-second buckets: idle 200 W, ramps to `peak` inside the window. */ +function series(hwKey: string, conc: number, peak: number, gpus = [0, 1]): GpuPowerSeries { + const t = Array.from({ length: 61 }, (_, i) => i); + return { + artifact: `gpu_metrics_${resultName(hwKey.split('_')[0], conc)}`, + startMs: START_MS, + bucketSeconds: 1, + gpus, + t, + power: gpus.map((gpu) => t.map((second) => (second >= 40 ? peak + gpu * 10 : 200 + gpu))), + }; +} + +const response: GpuPowerSeriesResponse = { + runInfo: { + id: Number(RUN_ID), + name: 'Run Sweep', + branch: 'main', + sha: 'abc123', + createdAt: '2026-09-12T20:00:00Z', + url: RUN_URL, + conclusion: 'success', + status: 'completed', + }, + series: [series('b200', 16, 700), series('b200', 64, 900)], +}; + +const Y_LABEL = 'Measured Average Power per Chip over Time (W)'; + +function providerOverrides( + unofficial: Parameters[0] = {}, + activeHwTypes?: readonly string[], +): Parameters[1] { + return { + inference: { + selectedModel: Model.DeepSeek_V4_Pro, + selectedSequence: Sequence.EightK_OneK, + selectedYAxisMetric: 'y_measuredPowerTimeline', + hardwareConfig: hwConfig, + activeHwTypes: new Set(activeHwTypes ?? HW_TYPES), + hwTypesWithData: new Set(HW_TYPES), + }, + unofficial, + }; +} + +function mountTimeline( + data: InferenceData[], + options: { + overlay?: Parameters[0]['overlayData']; + unofficial?: Parameters[0]; + activeHwTypes?: readonly string[]; + } = {}, +) { + mountWithProviders( + +
+ +
+
, + providerOverrides(options.unofficial, options.activeHwTypes), + ); +} + +/** One measured run at mount; a button adds a second run to the same plot. */ +function GrowingTimeline() { + const [data, setData] = useState(() => [measuredPoint('b200', 16, 700)]); + return ( + + +
+ +
+
+ ); +} + +const overlayResponse = (traces: GpuPowerSeries[]): GpuPowerSeriesResponse => ({ + runInfo: { + ...response.runInfo, + id: Number(OVERLAY_RUN_ID), + url: OVERLAY_RUN_URL, + }, + series: traces, +}); + +const overlayData = (data: InferenceData[]) => ({ + data, + hardwareConfig: hwConfig, + label: 'powerx-timeline', + runUrl: OVERLAY_RUN_URL, +}); + +/** Context overrides for one loaded `?unofficialrun=` whose overlay legend shows `hw`. */ +function overlayRun(hw: string): Parameters[0] { + return { + isUnofficialRun: true, + unofficialRunInfos: [ + { + id: Number(OVERLAY_RUN_ID), + name: 'powerx-timeline', + branch: 'powerx-timeline', + sha: 'abc000', + createdAt: '2026-09-12T00:00:00Z', + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + runIndexByUrl: { [OVERLAY_RUN_URL]: 0, [OVERLAY_RUN_ID]: 0 }, + activeOverlayHwTypes: new Set([hw]), + }; +} + +const svg = () => cy.get('[data-testid="power-timeline-chart-svg"]'); + +describe('PowerTimeline', () => { + beforeEach(() => { + cy.on('uncaught:exception', (error) => { + if (error.message.includes('ResizeObserver loop')) return false; + }); + }); + + it('colours overlay-run traces by run and honours the overlay hardware filter', () => { + const overlayPoint = measuredPoint('h200', 16, 500, { + run_url: OVERLAY_RUN_URL, + }); + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { + body: response, + }).as('official'); + cy.intercept('POST', `/api/gpu-metrics?runId=${OVERLAY_RUN_ID}*`, { + body: overlayResponse([series('h200', 16, 500)]), + }).as('overlay'); + mountTimeline([measuredPoint('b200', 16, 700)], { + overlay: overlayData([overlayPoint]), + unofficial: createMockUnofficialRunContext(overlayRun('h200')), + }); + cy.wait(['@official', '@overlay']); + + svg().within(() => { + cy.get('path.power-trace[data-run-index="0"][data-segment="window"]') + .should('have.length', 1) + .and('have.attr', 'stroke', overlayRunColor(0)); + cy.get('path.power-trace[data-hw="b200"][data-segment="window"]').should('have.length', 1); + // Two runs: the axis defaults to elapsed time so traces overlap by phase. + cy.get('text').contains('Time since telemetry start').should('exist'); + // Reference lines cover both hardware SKUs. + cy.get('.power-reference[data-reference="tdp"]').should('have.length', 2); + }); + cy.get('[data-testid="chart-legend"]').should('contain.text', '✕ powerx-timeline'); + }); + + it('joins a run that arrives after mount without a hook-shape warning', () => { + const secondResponse: GpuPowerSeriesResponse = { + runInfo: { + ...response.runInfo, + id: Number(SECOND_RUN_ID), + url: SECOND_RUN_URL, + }, + series: [series('h100', 16, 500)], + }; + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { + body: response, + }).as('first'); + cy.intercept('POST', `/api/gpu-metrics?runId=${SECOND_RUN_ID}*`, { + body: secondResponse, + }).as('second'); + cy.stub(console, 'error').as('consoleError'); + mountWithProviders(, providerOverrides()); + cy.wait('@first'); + svg().find('path.power-trace[data-segment="window"]').should('have.length', 1); + + cy.get('[data-testid="add-run"]').click(); + cy.wait('@second'); + svg().find('path.power-trace[data-segment="window"]').should('have.length', 2); + // React logs this when a memo's dependency array changes length between + // renders; one query per run used to be spread into that array. + cy.get('@consoleError').then((stub) => { + const calls = (stub as unknown as { args: unknown[][] }).args; + const shapeWarnings = calls.filter((args) => + args.some((a) => typeof a === 'string' && a.includes('changed size between renders')), + ); + expect(shapeWarnings, JSON.stringify(shapeWarnings)).to.have.length(0); + }); + }); + + it('marks overlay traces in the same-load summary with their run colour', () => { + const overlayPoint = measuredPoint('h200', 16, 500, { + run_url: OVERLAY_RUN_URL, + }); + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { + body: response, + }).as('official'); + cy.intercept('POST', `/api/gpu-metrics?runId=${OVERLAY_RUN_ID}*`, { + body: overlayResponse([series('h200', 16, 500)]), + }).as('overlay'); + mountTimeline([measuredPoint('b200', 16, 700)], { + overlay: overlayData([overlayPoint]), + unofficial: createMockUnofficialRunContext(overlayRun('h200')), + }); + cy.wait(['@official', '@overlay']); + cy.get( + `[data-testid="power-timeline-summary-trace"][data-trace="${traceKeyForPoint(overlayPoint)}"]`, + ) + .should('contain.text', 'unofficial') + .find('th > span') + .first() + .should('have.attr', 'style') + .and('contain', overlayRunColor(0)); + }); + + it('follows the date comparison series: colours, legend toggles and labels', () => { + const EARLIER_RUN_ID = '34600000001'; + const EARLIER_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${EARLIER_RUN_ID}`; + const DATES = ['2026-09-11', '2026-09-12']; + const ALL_SERIES = new Set(DATES.map((date) => `${date}_b200`)); + cy.intercept('POST', `/api/gpu-metrics?runId=${EARLIER_RUN_ID}*`, { + body: { + runInfo: { + ...response.runInfo, + id: Number(EARLIER_RUN_ID), + url: EARLIER_RUN_URL, + }, + series: [series('b200', 16, 600)], + }, + }).as('earlier'); + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { + body: response, + }).as('later'); + + function DateComparison() { + const [activeDates, setActiveDates] = useState(new Set(ALL_SERIES)); + const value = createMockInferenceContextValues({ + selectedModel: Model.DeepSeek_V4_Pro, + selectedSequence: Sequence.EightK_OneK, + selectedYAxisMetric: 'y_measuredPowerTimeline', + hardwareConfig: hwConfig, + activeHwTypes: new Set(HW_TYPES), + hwTypesWithData: new Set(HW_TYPES), + selectedGPUs: ['b200'], + selectedDates: DATES, + selectedDateRange: { startDate: '', endDate: '' }, + activeDates, + toggleActiveDate: (id: string) => + setActiveDates((prev) => computeToggle(prev, id, ALL_SERIES)), + }); + return ( + +
+ +
+
+ ); + } + mountWithProviders( + + + , + { unofficial: {} }, + ); + cy.wait(['@earlier', '@later']); + + // One legend row per compared date, under the hardware, as in GPUGraph. + cy.get('[data-testid="chart-legend"] .gpu-legend-title').should('have.text', 'B200'); + cy.get('[data-testid="chart-legend"] label').then(($labels) => { + const rows = $labels.toArray().map((label) => label.textContent?.trim()); + expect(rows).to.include.members(DATES); + }); + svg().within(() => { + cy.get('path.power-trace[data-segment="window"]').then(($paths) => { + expect($paths).to.have.length(2); + const strokes = new Set($paths.toArray().map((path) => path.getAttribute('stroke'))); + expect(strokes.size, 'each date has its own colour').to.equal(2); + }); + cy.get('text.power-trace-label').then(($labels) => { + expect($labels.toArray().map((label) => label.textContent)).to.have.members( + DATES.map((date) => `${date} c16`), + ); + }); + }); + + // Soloing the later date removes the earlier date's trace. + cy.get('[data-testid="chart-legend"] label').contains(DATES[1]).click(); + svg().within(() => { + cy.get('path.power-trace[data-segment="window"]') + .should('have.length', 1) + .and('have.attr', 'data-trace-key', `${RUN_ID}:${resultName('b200', 16)}`); + cy.get('text.power-trace-label').should('have.length', 1).and('have.text', 'c16'); + }); + }); + + it('follows dates-only range comparison (i_dstart/i_dend, no ~r runs)', () => { + // Range endpoints alone (empty selectedDates) must still drive per-date + // colours, legend toggles and end labels — the dates-only share/reload path. + const EARLIER_RUN_ID = '34600000002'; + const EARLIER_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${EARLIER_RUN_ID}`; + const DATES = ['2026-09-01', '2026-09-09'] as const; + const ALL_SERIES = new Set(DATES.map((date) => `${date}_b200`)); + cy.intercept('POST', `/api/gpu-metrics?runId=${EARLIER_RUN_ID}*`, { + body: { + runInfo: { ...response.runInfo, id: Number(EARLIER_RUN_ID), url: EARLIER_RUN_URL }, + series: [series('b200', 16, 600)], + }, + }).as('range-earlier'); + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { body: response }).as('range-later'); + + function DatesOnlyRange() { + const [activeDates, setActiveDates] = useState(new Set(ALL_SERIES)); + const value = createMockInferenceContextValues({ + selectedModel: Model.DeepSeek_V4_Pro, + selectedSequence: Sequence.EightK_OneK, + selectedYAxisMetric: 'y_measuredPowerTimeline', + hardwareConfig: hwConfig, + activeHwTypes: new Set(HW_TYPES), + hwTypesWithData: new Set(HW_TYPES), + selectedGPUs: ['b200'], + selectedDates: [], + selectedDateRange: { startDate: DATES[0], endDate: DATES[1] }, + activeDates, + toggleActiveDate: (id: string) => + setActiveDates((prev) => computeToggle(prev, id, ALL_SERIES)), + }); + return ( + +
+ +
+
+ ); + } + mountWithProviders( + + + , + { unofficial: {} }, + ); + cy.wait(['@range-earlier', '@range-later']); + + cy.get('[data-testid="chart-legend"] label').then(($labels) => { + const rows = $labels.toArray().map((label) => label.textContent?.trim()); + expect(rows).to.include.members([...DATES]); + expect(rows.join(' ')).not.to.match(/#\d/u); + }); + svg().within(() => { + cy.get('path.power-trace[data-segment="window"]').should('have.length', 2); + cy.get('text.power-trace-label').then(($labels) => { + expect($labels.toArray().map((label) => label.textContent)).to.have.members([ + `${DATES[0]} c16`, + `${DATES[1]} c16`, + ]); + }); + }); + cy.get('[data-testid="chart-legend"] label').contains(DATES[0]).click(); + svg().within(() => { + cy.get('path.power-trace[data-segment="window"]').should('have.length', 1); + cy.get('text.power-trace-label').should('have.text', 'c16'); + }); + }); +}); diff --git a/packages/app/cypress/component/scatter-graph.cy.tsx b/packages/app/cypress/component/scatter-graph.cy.tsx index 0f157fea1..7ab5903c7 100644 --- a/packages/app/cypress/component/scatter-graph.cy.tsx +++ b/packages/app/cypress/component/scatter-graph.cy.tsx @@ -465,6 +465,64 @@ describe('ScatterGraph', () => { ); }); + it('explains why All in Measured has no points', () => { + mountWithProviders( +
+ +
, + { + inference: { + hardwareConfig: hwConfig, + activeHwTypes: new Set(['b200_trt']), + hwTypesWithData: new Set(), + selectedYAxisMetric: 'y_utilityModeledWatts', + }, + unofficial: {}, + }, + ); + + cy.contains('No values are available for All in Measured in this selection.').should( + 'be.visible', + ); + cy.contains('No measurements to plot for this selection.').should('not.exist'); + }); + + it('localizes the All in Measured explanation', () => { + mountWithProviders( + +
+ +
+
, + { + inference: { + hardwareConfig: hwConfig, + activeHwTypes: new Set(['b200_trt']), + hwTypesWithData: new Set(), + selectedYAxisMetric: 'y_utilityModeledJPerOutputToken', + }, + unofficial: {}, + }, + ); + + cy.contains('当前选择没有可用的整体实测功耗数值。').should('be.visible'); + cy.contains('当前选择没有可绘制的测量数据。').should('not.exist'); + }); + for (const selectedYAxisMetric of ['y_tpPerGpu', 'y_measuredPrefillJPerInputToken'] as const) { it(`offers targeted quick-filter recovery on ${selectedYAxisMetric} without changing model, precision or date`, () => { mountWithProviders( @@ -925,14 +983,19 @@ describe('ScatterGraph', () => { cy.get('#test-scatter-overlay-labels svg .line-label') .filter('[data-line-key]:not([data-line-key^="overlay-"])') .should('have.length.greaterThan', 0); - // The exact branch that crashed the production page remains visible in the - // overlay line label and legend after ScatterGraph's render-time updates. - cy.get('#test-scatter-overlay-labels svg .line-label[data-line-key^="overlay-"]') - .find('text') - .should('contain.text', runBranch); + // The pill names the hardware behind the ✕ marker, parsed like an official + // pill; the long branch that crashed the production page stays in the legend. + cy.get('#test-scatter-overlay-labels svg .line-label[data-line-key^="overlay-"] .ll-text') + .should('have.text', '✕ B200 (TRTLLM)') + .and('not.contain.text', runBranch); cy.get( '#test-scatter-overlay-labels svg .line-label[data-line-key^="overlay-"] .ll-gpu', - ).should('not.exist'); + ).should('have.text', 'B200'); + // b200_trt is active only in the overlay legend (official rows: h100), so + // the overlay pill must stay visible after the filter-sync effect. + cy.get('#test-scatter-overlay-labels svg .line-label[data-line-key^="overlay-"]') + .should('have.attr', 'data-visible', '1') + .and('have.css', 'opacity', '1'); cy.get('#test-scatter-overlay-labels [data-testid="chart-legend"]').should( 'contain.text', runBranch, @@ -1094,7 +1157,7 @@ describe('ScatterGraph', () => { cy.get('#test-scatter-singleton-overlay-label svg .line-label[data-line-key^="overlay-"]') .should('have.length', 1) .find('text') - .should('contain.text', 'tileRT'); + .should('have.text', '✕ B200 (TRTLLM)'); cy.get('#test-scatter-singleton-overlay-label svg').then(($svg) => { const svg = $svg[0]; @@ -2875,13 +2938,6 @@ describe('Power envelopes', () => { .should('have.length', 3) .each(($point) => cy.wrap($point).should('have.css', 'opacity', '1')); cy.get('#scatter-show-all-measurements').should('not.exist'); - cy.get('[data-testid="measured-power-summary"]') - .should('contain.text', 'Showing 3 of 4 measured points') - .and('contain.text', '1/2 historical'); - cy.get('#power-sweep .dot-group') - .filter((_, element) => element.style.opacity !== '0') - .find('.legacy-power-ring') - .should('have.length', 1); cy.get('#power-sweep .dot-group') .filter((_, element) => element.style.opacity === '0') .should('have.css', 'pointer-events', 'none'); @@ -2893,11 +2949,6 @@ describe('Power envelopes', () => { .should('have.length', 4) .each(($point) => cy.wrap($point).should('have.css', 'opacity', '1')); cy.get('#power-sweep .roofline-path').should('have.attr', 'd', boundary); - cy.get('[data-testid="measured-power-summary"]').should( - 'contain.text', - 'Showing 4 of 4 measured points', - ); - cy.get('#power-sweep .legacy-power-ring').should('have.length', 2); cy.get('#scatter-show-all-measurements').should('not.exist'); cy.get('#scatter-hide-non-optimal').click({ force: true }); cy.get('#power-sweep .dot-group') @@ -2920,7 +2971,6 @@ describe('Power envelopes', () => { cy.contains('button', 'Energy').click(); cy.get('#scatter-hide-non-optimal').should('have.attr', 'data-state', 'checked'); cy.get('#power-sweep .roofline-path[data-curve-kind="pareto"]').should('have.length', 1); - cy.get('[data-testid="power-curve-description"]').should('not.exist'); cy.get('#scatter-show-all-measurements').should('not.exist'); }); diff --git a/packages/app/cypress/e2e/certified-power-filter.cy.ts b/packages/app/cypress/e2e/certified-power-filter.cy.ts index 30098dbff..d5cd54b52 100644 --- a/packages/app/cypress/e2e/certified-power-filter.cy.ts +++ b/packages/app/cypress/e2e/certified-power-filter.cy.ts @@ -114,12 +114,9 @@ describe('Validated vs historical measured power', () => { }); }); - it('rings legacy points on a measured axis and filters them via Quick Filters', () => { + it('filters validated and historical measured points via Quick Filters', () => { visitCertifiedPowerChart(); - cy.get('.legacy-power-ring').should('not.exist'); - cy.get('[data-testid="legacy-power-key"]').should('not.exist'); - cy.get('[data-testid="yaxis-metric-selector"]').click('right'); cy.contains('[data-slot="select-item"]', 'Measured Power') .scrollIntoView() @@ -127,16 +124,8 @@ describe('Validated vs historical measured power', () => { .click(); cy.get('[data-slot="select-content"]').should('not.exist'); - cy.get('[data-testid="measured-power-summary"]') - .should('contain.text', 'Showing 2 of 6 measured points') - .and('contain.text', '1/3 validated') - .and('contain.text', '1/3 historical') - .and('contain.text', 'Best per SKU and Optimal Only are enabled'); - - cy.get('.dot-group[data-hw-key^="b200"] .legacy-power-ring').should('exist'); - cy.get('.dot-group[data-hw-key^="mi300x"] .legacy-power-ring').should('not.exist'); - cy.get('[data-testid="legacy-power-key"]').should('be.visible'); - cy.screenshot('legacy-power-rings', { capture: 'viewport' }); + cy.get('.dot-group[data-hw-key^="b200"]').should('exist'); + cy.get('.dot-group[data-hw-key^="mi300x"]').should('exist'); cy.get('[data-testid="scatter-quick-filters"]').click(); cy.get('[data-testid="quick-filters-dialog"]').should('be.visible'); @@ -152,8 +141,6 @@ describe('Validated vs historical measured power', () => { cy.get('[data-testid="quick-filters-selected-count"]').should('contain.text', '1 selected'); cy.get('.dot-group[data-hw-key^="b200"]').should('not.exist'); cy.get('.dot-group[data-hw-key^="mi300x"]').should('exist'); - cy.get('.legacy-power-ring').should('not.exist'); - cy.get('[data-testid="legacy-power-key"]').should('not.exist'); cy.get('[data-testid="inference-chart-display"] svg').should('exist'); cy.screenshot('certified-only-filter', { capture: 'viewport' }); @@ -165,8 +152,7 @@ describe('Validated vs historical measured power', () => { 'false', ); cy.get('[data-testid="quick-filters-dialog"]').contains('button', 'Done').click(); - cy.get('.dot-group[data-hw-key^="b200"] .legacy-power-ring').should('exist'); - cy.get('[data-testid="legacy-power-key"]').should('be.visible'); + cy.get('.dot-group[data-hw-key^="b200"]').should('exist'); }); it('restores a shared i_power=certified link with the toggle pre-selected', () => { @@ -177,7 +163,6 @@ describe('Validated vs historical measured power', () => { cy.get('.dot-group[data-hw-key^="mi300x"]').should('exist'); cy.get('.dot-group[data-hw-key^="b200"]').should('not.exist'); - cy.get('[data-testid="legacy-power-key"]').should('not.exist'); cy.get('[data-testid="scatter-quick-filters"]').click(); cy.get('[data-testid="quick-filter-power-certified"]').should( @@ -202,10 +187,6 @@ describe('Validated vs historical measured power', () => { cy.get('#scatter-hide-non-optimal').should('have.attr', 'data-state', 'checked'); cy.get('#scatter-show-all-measurements').should('not.exist'); visiblePowerPoints().should('have.length', 4); - cy.get('[data-testid="measured-power-summary"]').should( - 'contain.text', - 'Showing 4 of 6 measured points', - ); // Let the initial ResizeObserver update reach the SVG before saving geometry. cy.get('[data-testid="d3-chart-svg"]').should(($svg) => { @@ -239,7 +220,6 @@ describe('Validated vs historical measured power', () => { cy.get('#scatter-hide-non-optimal').click(); visiblePowerPoints().should('have.length', 4); - visiblePowerPoints().find('.legacy-power-ring').should('have.length', 2); cy.get('.roofline-path').should(($current) => { expect(Array.from($current, (curve) => curve.getAttribute('d'))).to.deep.equal(geometry); }); diff --git a/packages/app/cypress/e2e/inference-chart.cy.ts b/packages/app/cypress/e2e/inference-chart.cy.ts index bd8946401..3f00279c4 100644 --- a/packages/app/cypress/e2e/inference-chart.cy.ts +++ b/packages/app/cypress/e2e/inference-chart.cy.ts @@ -1,3 +1,4 @@ +import type { InferenceData } from '@/components/inference/types'; import { interceptVrPublicationData, VR_FIXTURE_DATE, @@ -54,8 +55,8 @@ const boundaryRows = (runUrl: string | null) => })); function interceptMeasuredComparison( - official = measuredRows(null), - overlay = measuredRows(OVERLAY_RUN_URL), + official: object[] = measuredRows(null), + overlay: object[] = measuredRows(OVERLAY_RUN_URL), ) { cy.intercept('GET', '/api/v1/availability', { body: official.slice(0, 1) }); cy.intercept('GET', '/api/v1/benchmarks*', { body: official }); @@ -107,11 +108,14 @@ function assertVisibleMeasuredValues(selector: string, expected: number[]) { describe('Inference Chart', () => { before(() => { + cy.intercept('GET', '/api/v1/availability').as('chartAvailability'); + cy.intercept('GET', '/api/v1/benchmarks*').as('chartBenchmarks'); cy.viewport(1440, 900); cy.window().then((win) => { win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); }); cy.visit('/inference'); + cy.wait(['@chartAvailability', '@chartBenchmarks']); }); it('renders the inference chart display wrapper', () => { @@ -580,7 +584,9 @@ describe('AgentX replaces a complete curve while preserving an unofficial compar }); }); -it('hydrates a direct PowerX metric link and shows availability for the selected workload', () => { +it('hydrates a direct PowerX metric link', () => { + cy.intercept('GET', '/api/v1/availability').as('powerLinkAvailability'); + cy.intercept('GET', '/api/v1/benchmarks*').as('powerLinkBenchmarks'); cy.viewport(1440, 900); cy.visit('/inference/qwen-3-5?i_seq=8k%2F1k&i_prec=fp8&i_metric=y_measuredPowerPercentTdp', { onBeforeLoad(win) { @@ -589,6 +595,7 @@ it('hydrates a direct PowerX metric link and shows availability for the selected cy.spy(win.console, 'error').as('powerLinkConsoleErrors'); }, }); + cy.wait(['@powerLinkAvailability', '@powerLinkBenchmarks']); cy.get('[data-testid="yaxis-metric-selector"]').should('contain', 'Measured Power'); cy.get('[data-testid="measured-power-display"]').should('contain', 'TDP'); cy.get('[data-testid="measured-power-statistic-average"]').should( @@ -596,17 +603,6 @@ it('hydrates a direct PowerX metric link and shows availability for the selected 'aria-pressed', 'true', ); - cy.get('[data-testid="power-metric-availability"]').should( - 'contain', - 'Current workload and hardware selection', - ); - cy.contains('summary', 'Availability of all measured metrics').click(); - cy.get('[data-testid="power-metric-availability"]').within(() => { - cy.contains('button', 'Measured P75 Fleet Power per Chip').should('contain', '/'); - cy.contains('button', 'Measured Joules per Output Token').click(); - }); - cy.get('[data-testid="yaxis-metric-selector"]').should('contain', 'Measured Energy'); - cy.get('[data-testid="measured-energy-denominator"]').should('contain', 'Output'); cy.get('@powerLinkConsoleErrors').should('not.be.calledWithMatch', /hydrat/i); }); @@ -637,6 +633,29 @@ it('replots measured settings for official and unofficial data and preserves ove assertMeasuredValues('.dot-group', [2, 3, 4, 5]); }); +it('says when chosen chip configs report no measured power, and plots them on other metrics', () => { + // The config has benchmarks but no power telemetry, like GB200/GB300 NVL72 on DSR1 8K/1K. + interceptMeasuredComparison(singleTurnRows(null), []); + cy.viewport(1440, 900); + cy.visit( + '/inference?g_model=DeepSeek-V4-Pro&i_seq=1k%2F1k&i_prec=fp4&i_metric=y_measuredAvgPower&i_gpus=b300_sglang', + { onBeforeLoad: unlockAgenticGate }, + ); + cy.get('[data-testid="gpu-multiselect"]').should('contain', 'B300 (SGLang)'); + cy.get('[data-testid="scatter-empty-state"]') + .should('have.attr', 'data-reason', 'selection') + .and('contain', 'No measured GPU power is reported for this selection.'); + cy.get('[data-testid="yaxis-metric-selector"]').click('right'); + cy.contains('[data-slot="select-item"]', /^Token Throughput per Chip/u) + .scrollIntoView() + .click(); + cy.get('[data-testid="scatter-empty-state"]').should('not.exist'); + cy.get('[data-testid="inference-chart-display"] svg .dot-group').should( + 'have.length.at.least', + 1, + ); +}); + it('uses Optimal Only to filter power boundary dots without replacing official or overlay curves', () => { interceptMeasuredComparison(boundaryRows(null), boundaryRows(OVERLAY_RUN_URL)); cy.viewport(1440, 900); @@ -693,7 +712,6 @@ it('uses Optimal Only to filter power boundary dots without replacing official o cy.get('[data-testid="chart-figure"] h2').should('contain', 'Measured Joules per Output Token'); cy.get('#scatter-hide-non-optimal').should('have.attr', 'data-state', 'checked'); cy.get('#scatter-show-all-measurements').should('not.exist'); - cy.get('[data-testid="power-curve-description"]').should('not.exist'); assertVisibleMeasuredValues('.dot-group', [2, 4, 5]); assertVisibleMeasuredValues('.unofficial-overlay-pt', [3, 5, 6]); cy.get(curves) @@ -826,3 +844,75 @@ describe('VR default date preference', () => { assertVrDate(VR_LATEST_FIXTURE_DATE); }); }); + +const withTp4 = (rows: ReturnType) => [ + ...rows, + ...rows.map((row) => ({ + ...row, + id: row.id + 10000, + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + })), +]; + +const assertObservedLoads = (expected: number[]) => { + for (const selector of ['.dot-group', '.unofficial-overlay-pt']) { + cy.get( + `[data-testid="inference-chart-display"] svg ${selector}`, + ).should(($points) => { + expect($points).to.have.length(expected.length); + expect([...$points].map((element) => element.__data__.x).sort((a, b) => a - b)).to.deep.equal( + expected, + ); + for (const element of $points) expect(getComputedStyle(element).opacity).to.equal('1'); + }); + } +}; + +describe('Date comparison with unofficial runs', () => { + for (const [axis, xMode] of [ + ['interactivity', ''], + ['concurrency', '&i_xmode=concurrency'], + ]) { + // Every measurement is shown (i_optimal=0), so both runs plot all four loads. + it(`keeps ?unofficialrun= overlays on the ${axis} date comparison`, () => { + interceptMeasuredComparison(); + cy.viewport(1440, 900); + cy.visit( + `/inference?g_model=DeepSeek-V4-Pro&unofficialrun=${OVERLAY_RUN_ID}&i_seq=1k%2F1k&i_prec=fp4&i_metric=y_measuredAvgPower&i_gpus=b300_sglang&i_dstart=${SINGLE_TURN_DATE}&i_dend=${SINGLE_TURN_DATE}&i_optimal=0${xMode}`, + { onBeforeLoad: unlockAgenticGate }, + ); + cy.wait('@measuredOverlay'); + cy.get('[data-testid="gpu-graph"]').should('exist'); + cy.get('[data-testid="scatter-graph"]').should('not.exist'); + assertMeasuredValues('.dot-group', [450, 460, 470, 480]); + assertMeasuredValues('.unofficial-overlay-pt', [450, 460, 470, 480]); + cy.get('[aria-label="Dismiss measured-comparison"]').click(); + cy.get('[data-testid="gpu-graph"] .unofficial-overlay-pt').should('not.exist'); + assertMeasuredValues('.dot-group', [450, 460, 470, 480]); + }); + } +}); + +describe('Observed concurrency and exact topology', () => { + it('applies the exact-topology filter to official and ?unofficialrun= overlay loads', () => { + const officialRun = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/800001'; + const officialBase = boundaryRows(null).map((row) => ({ ...row, run_url: officialRun })); + const overlayBase = boundaryRows(OVERLAY_RUN_URL); + interceptMeasuredComparison(withTp4(officialBase), withTp4(overlayBase)); + cy.viewport(1440, 900); + cy.visit( + `/inference?g_model=DeepSeek-V4-Pro&unofficialrun=${OVERLAY_RUN_ID}&i_seq=1k%2F1k&i_prec=fp4&i_metric=y_measuredAvgPower&i_xmode=concurrency&i_optimal=1&i_best=1`, + { onBeforeLoad: unlockAgenticGate }, + ); + cy.wait('@measuredOverlay'); + assertObservedLoads([1, 1, 2, 2, 8, 8, 48, 48]); + cy.get('[data-testid="scatter-quick-filters"]').click(); + cy.contains('[data-testid="quick-filter-topology-options"] button', /GPU=?4.*TP=?4/u).click(); + cy.get('body').type('{esc}'); + cy.get('[data-testid="quick-filters-dialog"]').should('not.exist'); + assertObservedLoads([1, 2, 8, 48]); + }); +}); diff --git a/packages/app/cypress/e2e/powerx-compare.cy.ts b/packages/app/cypress/e2e/powerx-compare.cy.ts new file mode 100644 index 000000000..159b01683 --- /dev/null +++ b/packages/app/cypress/e2e/powerx-compare.cy.ts @@ -0,0 +1,218 @@ +import { assertShareLinkParams } from '../support/share-link'; + +// The PowerX article panels inside the gated Measured Energy group. +// Deterministic disaggregated rows carry validated whole-deployment and +// per-role telemetry; the overlay run adds H200 rows. + +const MODEL = 'dsv4'; +const DATE = '2026-09-01'; +const OVERLAY_RUN_ID = '31415926535'; +const OVERLAY_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${OVERLAY_RUN_ID}`; + +/** + * conc, interactivity (tok/s/user), output tok/s per GPU, measured W per GPU, + * prefill W per GPU, decode W per GPU, J per output token, decode J per output token. + * Input tokens outnumber output tokens 8:1, so J/in = J/out ÷ 8 and the prefill + * pool's J/in reconstructs to J/out − decode J/out on the output-token axis. + */ +const CONFIGS: [number, number, number, number, number, number, number, number][] = [ + [16, 90, 200, 500, 800, 400, 2.4, 1.6], + [64, 60, 400, 600, 820, 450, 1.6, 1], + [256, 30, 800, 700, 840, 500, 1, 0.6], +]; +const TOKEN_RATIO = 8; + +let rowId = 990000; +const rows = (runUrl: string | null, hardware: 'b200' | 'h200') => + CONFIGS.map(([conc, intvty, outputTput, watts, prefillWatts, decodeWatts, jOut, decodeJOut]) => ({ + id: runUrl ? 0 : rowId++, + hardware, + framework: 'dynamo-sglang', + model: MODEL, + precision: 'fp4', + spec_method: 'none', + disagg: true, + is_multinode: true, + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + isl: 8192, + osl: 1024, + conc, + offload_mode: 'off', + benchmark_type: 'single_turn', + image: 'dynamo:test', + metrics: { + median_intvty: intvty, + median_itl: 1 / intvty, + median_e2el: 20, + median_ttft: 0.5, + tput_per_gpu: outputTput * 9, + output_tput_per_gpu: outputTput, + input_tput_per_gpu: outputTput * 8, + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: watts, + avg_total_gpu_power_w: watts * 8, + prefill_avg_power_w: prefillWatts, + decode_avg_power_w: decodeWatts, + joules_per_output_token: jOut, + joules_per_input_token: jOut / TOKEN_RATIO, + prefill_joules_per_input_token: (jOut - decodeJOut) / TOKEN_RATIO, + decode_joules_per_output_token: decodeJOut, + }, + workers: null, + date: DATE, + run_url: runUrl, + })); + +const availability = [ + { + model: MODEL, + isl: 8192, + osl: 1024, + precision: 'fp4', + hardware: 'b200', + framework: 'dynamo-sglang', + spec_method: 'none', + disagg: true, + benchmark_type: 'single_turn', + date: DATE, + }, +]; + +function interceptRows(officialRunUrl: string) { + const official = rows(null, 'b200').map((row, index) => ({ + ...row, + curve_workflow_run_id: 27182818284, + curve_date: DATE, + run_url: index === 0 ? officialRunUrl : `${officialRunUrl}0`, + power_audit: { + producer_sha: index === 0 ? 'producer-a' : 'producer-b', + exporter_image_sha256: index === 0 ? 'exporter-a' : 'exporter-b', + }, + })); + cy.intercept('GET', '/api/v1/availability', { body: availability }).as('availability'); + cy.intercept('GET', '/api/v1/benchmarks*', { body: official }).as('benchmarks'); + cy.intercept('GET', '/api/v1/workflow-info*', { + body: { runs: [], changelogs: [], configs: [] }, + }); +} + +/** Overlay rows 10% cheaper per token than the official rows, so they own the frontier. */ +function interceptCheaperOverlayRows() { + const cheaper = rows(OVERLAY_RUN_URL, 'h200').map((row) => ({ + ...row, + metrics: { + ...row.metrics, + joules_per_output_token: row.metrics.joules_per_output_token * 0.9, + joules_per_input_token: row.metrics.joules_per_input_token * 0.9, + }, + })); + cy.intercept('GET', '/api/unofficial-run*', { + body: { + runInfos: [ + { + id: OVERLAY_RUN_ID, + name: 'powerx-panels', + branch: 'powerx-panels', + sha: 'abc000', + createdAt: `${DATE}T00:00:00Z`, + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + benchmarks: cheaper, + evaluations: [], + }, + }).as('unofficialRun'); +} + +function visitChart(extraParams: string, officialRunUrl: string) { + interceptRows(officialRunUrl); + cy.visit(`/inference?g_model=DeepSeek-V4-Pro&i_seq=8k/1k&i_prec=fp4${extraParams}`, { + onBeforeLoad(win) { + win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); + win.localStorage.setItem('inferencex-feature-gate', '1'); + }, + }); + cy.wait(['@availability', '@benchmarks']); + cy.get('[data-testid="inference-chart-display"]').should('exist'); + cy.get('[data-testid="chart-figure"]').should('have.length.at.least', 1); +} + +describe('PowerX article panels', () => { + const PANELS = '&i_roleshare=1&i_powerfit=1&i_frontier=1'; + const RETIRED_COMPARE = + '&i_servicecompare=1&i_servicebase=old-baseline&i_servicepeer=old-comparator&i_servicetarget=40'; + const OFFICIAL_RUN_URL = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/27182818284'; + const OVERLAY_COLOR = 'var(--overlay-run-0)'; + + beforeEach(() => { + cy.on('uncaught:exception', (error) => { + if (error.message === 'ResizeObserver loop completed with undelivered notifications.') { + return false; + } + }); + }); + + it('keeps role, fit and frontier panels while ignoring retired comparison share state', () => { + interceptCheaperOverlayRows(); + visitChart( + `&unofficialrun=${OVERLAY_RUN_ID}&i_metric=y_measuredJPerOutputToken${PANELS}${RETIRED_COMPARE}`, + OFFICIAL_RUN_URL, + ); + cy.wait('@unofficialRun'); + cy.get('[data-testid="power-analysis-panels"]').should('be.visible'); + cy.get('[data-testid^="equal-service-"]').should('not.exist'); + cy.get('[data-testid^="matched-concurrency"]').should('not.exist'); + + // Role power: prefill and decode W/GPU for both sources; the overlay in its run colour. + cy.get('[data-testid="chart-0-role-power-plot"] circle.point').should( + 'have.length', + CONFIGS.length * 4, + ); + cy.get(`[data-testid="chart-0-role-power-plot"] circle.point[fill="${OVERLAY_COLOR}"]`).should( + 'have.length', + CONFIGS.length * 2, + ); + cy.get('[data-testid="chart-0-role-share-plot"] circle.point').should( + 'have.length', + CONFIGS.length * 2, + ); + + // Output per allocated GPU is output per decode GPU × 4 ÷ 8: 100, 200 and 400 tok/s. + // W/GPU 500, 600, 700 → P0 450 W, m 0.643 J/token, R² 0.964. + cy.get('[data-testid="power-fit-row"]').should('have.length', 2); + cy.get('[data-testid="power-fit-row"]') + .filter(':contains("B200")') + .should('contain.text', '450') + .and('contain.text', '45% · 1,000 W') + .and('contain.text', '0.643') + .and('contain.text', '0.964') + .and('contain.text', '100.0–400.0'); + cy.get('[data-testid="power-fit-row"]') + .filter(':contains("H200")') + .should('contain.text', '64% · 700 W'); + + // The cheaper overlay rows own the frontier; each lists its run. + cy.get('[data-testid="frontier-points-row"]') + .should('have.length', CONFIGS.length) + .each(($row) => { + expect($row.text()).to.include('unofficial'); + expect($row.find('a').attr('href')).to.eq(OVERLAY_RUN_URL); + }); + assertShareLinkParams({ + i_servicecompare: null, + i_servicebase: null, + i_servicepeer: null, + i_servicetarget: null, + i_roleshare: '1', + i_powerfit: '1', + i_frontier: '1', + }); + }); +}); diff --git a/packages/app/cypress/e2e/powerx-timeline.cy.ts b/packages/app/cypress/e2e/powerx-timeline.cy.ts new file mode 100644 index 000000000..1a624390f --- /dev/null +++ b/packages/app/cypress/e2e/powerx-timeline.cy.ts @@ -0,0 +1,198 @@ +import { assertShareLinkParams } from '../support/share-link'; + +// The Measured Power "Timeline" display (`y_measuredPowerTimeline`) restores its +// shared view on a `?unofficialrun=` trace. Intercepted rows carry the +// `power_audit.source` / `run_url` provenance that names each point's +// `gpu_metrics_*` artifact; `/api/gpu-metrics?series=power` returns matching +// one-second series. + +const MODEL = 'dsv4'; +const DATE = '2026-09-01'; +const RUN_ID = '34716669498'; +const RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${RUN_ID}`; +const OVERLAY_RUN_ID = '31415926535'; +const OVERLAY_RUN_URL = `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${OVERLAY_RUN_ID}`; +const START_MS = Date.UTC(2026, 8, 1, 20, 0, 0); + +/** conc, interactivity (tok/s/user), output tok/s per GPU, measured W per GPU */ +const CONFIGS: [number, number, number, number][] = [ + [16, 90, 200, 500], + [64, 60, 400, 600], + [256, 30, 800, 700], +]; + +const resultName = (hardware: string, conc: number) => + `dsv4_8k1k_fp4_sglang_tp8-pp1-dcp1-pcp1-ep1-dpafalse_disagg-false_spec-none_conc${conc}_${hardware}-host-0123456789abcdef0123`; + +let rowId = 980000; +const rows = (runUrl: string, hardware: 'b200' | 'h200') => + CONFIGS.map(([conc, intvty, outputTput, watts]) => ({ + id: hardware === 'b200' ? rowId++ : 0, + hardware, + framework: 'sglang', + model: MODEL, + precision: 'fp4', + spec_method: 'none', + disagg: false, + is_multinode: false, + prefill_tp: 8, + decode_tp: 8, + num_prefill_gpu: 8, + num_decode_gpu: 8, + isl: 8192, + osl: 1024, + conc, + offload_mode: 'off', + benchmark_type: 'single_turn', + image: 'sglang:test', + metrics: { + median_intvty: intvty, + median_itl: 1 / intvty, + median_e2el: 20, + median_ttft: 0.5, + tput_per_gpu: outputTput * 9, + output_tput_per_gpu: outputTput, + input_tput_per_gpu: outputTput * 8, + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: watts, + avg_total_gpu_power_w: watts * 8, + joules_per_output_token: watts / outputTput, + }, + workers: null, + power_audit: { + source: `power_validation_${resultName(hardware, conc)}.json`, + sample_count: 160, + window_start_unix: (START_MS + 40_000) / 1000, + window_end_unix: (START_MS + 60_000) / 1000, + observed_gpu_ids: ['0', '1', '2', '3', '4', '5', '6', '7'], + }, + date: DATE, + run_url: runUrl, + })); + +const seriesFor = (hardware: string, runUrl: string, runId: string) => ({ + runInfo: { + id: Number(runId), + name: 'Run Sweep', + branch: hardware === 'b200' ? 'main' : 'powerx-timeline', + sha: 'abc123', + createdAt: `${DATE}T20:00:00Z`, + url: runUrl, + conclusion: 'success', + status: 'completed', + }, + series: CONFIGS.map((config) => { + const [conc] = config; + const watts = config[3]; + const t = Array.from({ length: 61 }, (_, i) => i); + return { + artifact: `gpu_metrics_${resultName(hardware, conc)}`, + startMs: START_MS, + bucketSeconds: 1, + gpus: [0, 1], + t, + power: [0, 1].map((gpu) => t.map((second) => (second >= 40 ? watts + gpu : 150 + gpu))), + }; + }), +}); + +const availability = [ + { + model: MODEL, + isl: 8192, + osl: 1024, + precision: 'fp4', + hardware: 'b200', + framework: 'sglang', + spec_method: 'none', + disagg: false, + benchmark_type: 'single_turn', + date: DATE, + }, +]; + +function interceptRows() { + cy.intercept('GET', '/api/v1/availability', { body: availability }).as('availability'); + cy.intercept('GET', '/api/v1/benchmarks*', { body: rows(RUN_URL, 'b200') }).as('benchmarks'); + cy.intercept('GET', '/api/v1/workflow-info*', { + body: { runs: [], changelogs: [], configs: [] }, + }); + cy.intercept('POST', `/api/gpu-metrics?runId=${RUN_ID}*`, { + body: seriesFor('b200', RUN_URL, RUN_ID), + }).as('series'); +} + +function interceptOverlay() { + cy.intercept('GET', '/api/unofficial-run*', { + body: { + runInfos: [ + { + id: OVERLAY_RUN_ID, + name: 'powerx-timeline', + branch: 'powerx-timeline', + sha: 'abc000', + createdAt: `${DATE}T00:00:00Z`, + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + benchmarks: rows(OVERLAY_RUN_URL, 'h200'), + evaluations: [], + }, + }).as('unofficialRun'); + cy.intercept('POST', `/api/gpu-metrics?runId=${OVERLAY_RUN_ID}*`, { + body: seriesFor('h200', OVERLAY_RUN_URL, OVERLAY_RUN_ID), + }).as('overlaySeries'); +} + +function visitChart(extraParams: string) { + interceptRows(); + cy.visit(`/inference?g_model=DeepSeek-V4-Pro&i_seq=8k/1k&i_prec=fp4${extraParams}`, { + onBeforeLoad(win) { + win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); + win.localStorage.setItem('inferencex-feature-gate', '1'); + }, + }); + cy.wait(['@availability', '@benchmarks']); + cy.get('[data-testid="inference-chart-display"]').should('exist'); +} + +describe('PowerX measured power timeline', () => { + beforeEach(() => { + cy.on('uncaught:exception', (error) => { + if (error.message === 'ResizeObserver loop completed with undelivered notifications.') { + return false; + } + }); + }); + + it('restores the shared serving-window view and focus on an unofficial trace', () => { + interceptOverlay(); + const focus = `${OVERLAY_RUN_ID}:${resultName('h200', 64)}`; + visitChart( + `&i_metric=y_measuredPowerTimeline&unofficialrun=${OVERLAY_RUN_ID}&i_ptaxis=serving&i_ptlines=gpu&i_ptwindow=window&i_ptfocus=${encodeURIComponent(focus)}`, + ); + cy.wait(['@series', '@unofficialRun', '@overlaySeries']); + cy.get('[data-testid="power-timeline-focus"]').should('contain.text', 'c64'); + cy.get('[data-testid="power-timeline-per-gpu"]').should('have.attr', 'data-state', 'checked'); + cy.get('[data-testid="power-timeline-window-only"]').should( + 'have.attr', + 'data-state', + 'checked', + ); + cy.get('path.power-trace[data-segment="full"]').should('not.exist'); + cy.get('path.power-trace[data-run-index="0"][data-segment="window"]').should('exist'); + assertShareLinkParams({ + i_metric: 'y_measuredPowerTimeline', + i_ptaxis: 'serving', + i_ptlines: 'gpu', + i_ptwindow: 'window', + i_ptfocus: focus, + }); + cy.get('[data-testid="power-timeline-focus-clear"]').click(); + assertShareLinkParams({ i_ptfocus: null, i_ptlines: 'gpu', i_ptwindow: 'window' }); + }); +}); diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts index b584c8a7e..80dad7963 100644 --- a/packages/app/cypress/e2e/profit-estimator.cy.ts +++ b/packages/app/cypress/e2e/profit-estimator.cy.ts @@ -94,10 +94,22 @@ const chart = () => cy.get('[data-testid="profit-estimator-chart"]'); const chartSvg = () => chart().find('svg').filter(':has(.chart-root)').first(); const bars = () => chart().find('rect.bar'); +function assertDisclosureOpen(testId: string, open: boolean) { + cy.get(`[data-testid="${testId}"]`).should(($details) => { + expect($details[0].open, `${testId} native disclosure state`).to.equal(open); + const content = $details[0].querySelector('p'); + expect(content, `${testId} content`).not.to.equal(null); + // Cypress visibility omits native closed-details rendering in some browsers. + if (content && typeof content.checkVisibility === 'function') { + expect(content.checkVisibility(), `${testId} browser visibility`).to.equal(open); + } + }); +} + // Clear the preceding chart before each case changes the viewport. describe('Profit estimator power option', { testIsolation: true }, () => { for (const locale of ['en', 'zh'] as const) { - it(`prices DeepSeek Flash partial chassis with visible assumptions and CSV labels (${locale})`, () => { + it(`prices DeepSeek Flash partial chassis with a one-line power note and CSV labels (${locale})`, () => { stubOpenRouter(); cy.viewport(locale === 'en' ? 1280 : 393, 900); cy.intercept('GET', '/api/v1/benchmarks*', { @@ -148,10 +160,16 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.get('[data-testid="profit-power-unavailable"]') .should('contain', 'GB300') .and('contain', hardwareReason); - cy.get('[data-testid="profit-power-note"]').should( - 'contain', - locale === 'en' ? 'partly idle server' : '部分 GPU 闲置', - ); + assertDisclosureOpen('profit-power-unavailable', false); + cy.get('[data-testid="profit-power-unavailable"] > summary').click(); + assertDisclosureOpen('profit-power-unavailable', true); + cy.get('[data-testid="profit-power-unavailable"] > p').should('be.visible'); + cy.get('[data-testid="profit-power-note"]') + .should('contain', locale === 'en' ? 'All in Measured' : '整体实测功耗') + .and('not.contain', locale === 'en' ? 'unmeasured components' : '未实测的组件'); + cy.get('[data-testid="profit-power-assumptions"]').should('not.exist'); + cy.get('[data-testid="profit-power-unavailable"] > summary').click(); + assertDisclosureOpen('profit-power-unavailable', false); cy.get('[data-testid="profit-power-note"]').then(($note) => { const box = $note[0].getBoundingClientRect(); expect(box.left).to.be.at.least(0); @@ -285,7 +303,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.get('#profit-target').should('have.value', '45'); chart().find('text.revenue-label').should('have.length', 6); chart().should('contain', 'B200').and('contain', 'B300').and('contain', 'MI355X'); - chart().should('contain', 'Measured + modeled').and('contain', 'Provisioned'); + chart().should('contain', 'All in Measured').and('contain', 'All in Provisioned'); cy.get('[data-testid="profit-power-unavailable"]').should('contain', 'GB300'); }); @@ -297,7 +315,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => { .invoke('text') .then((original) => { cy.get('#profit-power').click(); - cy.get('[role="option"]').contains('Measured + modeled power').click(); + cy.get('[role="option"]').contains('All in Measured').click(); // These existing fixtures intentionally have throughput but no validated power. cy.get('[data-testid="profit-power-unavailable"]').should( 'contain', @@ -308,7 +326,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.get('[data-testid="profit-price-source-selector"]').should('contain', 'Moonshot'); cy.get('[data-testid="profit-estimator-chart"]').should('not.exist'); cy.get('#profit-power').click(); - cy.get('[role="option"]').contains('Provisioned power').click(); + cy.get('[role="option"]').contains('All in Provisioned').click(); bars().its('length').should('be.greaterThan', 0); chart().should('have.text', original); }); diff --git a/packages/app/cypress/e2e/ttft-x-axis-toggle.cy.ts b/packages/app/cypress/e2e/ttft-x-axis-toggle.cy.ts index 114b1e2c0..316316d69 100644 --- a/packages/app/cypress/e2e/ttft-x-axis-toggle.cy.ts +++ b/packages/app/cypress/e2e/ttft-x-axis-toggle.cy.ts @@ -140,12 +140,12 @@ describe('X-Axis Mode Toggle (inference chart)', () => { interceptDerivedAgenticMetrics(); }); - it('defaults to Interactivity and offers all four full names in one axis dropdown', () => { + it('defaults to Interactivity and offers every full name in one axis dropdown', () => { cy.get('[data-testid="scenario-selector"]').should('contain.text', 'Agentic'); cy.get('[data-testid="x-axis-mode-selector"]').should('contain.text', 'Interactivity'); cy.get('[data-testid="x-axis-mode-buttons"]').should('not.exist'); openXAxisMenu(); - cy.get('[role="grid"] [data-select-option]').should('have.length', 4); + cy.get('[role="grid"] [data-select-option]').should('have.length', 5); cy.get('[data-testid="x-axis-mode-e2e-normalized-interactivity"]') .should('have.text', 'E2E Normalized Interactivity') .and('have.attr', 'aria-pressed', 'false'); @@ -212,7 +212,7 @@ describe('X-Axis Mode Toggle (inference chart)', () => { cy.get('#chart-0 [data-testid="offload-halo-key"]').should('not.exist'); }); - it('shows the selected percentile in the Interactivity axis label', () => { + it('shows the fixed p90 percentile in the Interactivity axis label', () => { // Explicitly select the mode — do not rely on the agentic default mode. selectXAxisMode('interactivity', 'Interactivity'); // Agentic plots percentile fields (p90_intvty), so the axis label carries it. @@ -274,22 +274,6 @@ describe('X-Axis Mode Toggle (inference chart)', () => { 'contain.text', 'P90 E2E Normalized Interactivity (tok/s/user)', ); - - cy.get('[data-testid="percentile-selector"]').click(); - cy.contains('[role="option"]', 'p75').click(); - cy.get('[data-testid="chart-figure"] h2').should( - 'contain.text', - 'P75 E2E Normalized Interactivity', - ); - - // The percentile selector is shared page state for the whole suite — - // restore the p90 default so later tests assert against a known value. - cy.get('[data-testid="percentile-selector"]').click(); - cy.contains('[role="option"]', 'p90').click(); - cy.get('[data-testid="chart-figure"] h2').should( - 'contain.text', - 'P90 E2E Normalized Interactivity', - ); }); it('does not change the chart while browsing options, and Escape cancels', () => { @@ -312,15 +296,12 @@ describe('X-Axis Mode Toggle (inference chart)', () => { ); }); - it('follows the percentile selector in the Interactivity axis label', () => { - // Select p75 here rather than inheriting it from another test — the axis - // label must track the selector on its own. + it('keeps the agentic Latency Percentile control removed and labels the fixed p90 basis', () => { selectXAxisMode('interactivity', 'Interactivity'); - cy.get('[data-testid="percentile-selector"]').click(); - cy.contains('[role="option"]', 'p75').click(); + cy.get('[data-testid="percentile-selector"]').should('not.exist'); cy.get('[data-testid="chart-figure"] svg').should( 'contain.text', - 'P75 Interactivity (tok/s/user)', + 'P90 Interactivity (tok/s/user)', ); }); }); @@ -352,15 +333,15 @@ describe('X-axis mode URL param', () => { cy.get('[data-testid="chart-figure"] h2').should('contain.text', 'Time To First Token'); }); - // AgentX publishes on P90, so the percentile control is insider-only. With - // the gate locked it must not render, and the chart must still plot P90. - it('hides the percentile selector behind the feature gate and defaults to P90', () => { + // Agentic inference charts fix the x-axis latency basis at P90. The old + // Latency Percentile dropdown is gone even when the feature gate is unlocked. + it('never shows the latency percentile selector and labels the fixed P90 basis', () => { interceptAgenticData(); interceptDerivedAgenticMetrics(); cy.visit('/inference?i_seq=agentic-traces', { onBeforeLoad(win) { win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); - win.localStorage.removeItem('inferencex-feature-gate'); + unlockAgenticGate(win); }, }); @@ -479,11 +460,12 @@ describe('Label defaults for fixed-sequence scenarios', () => { cy.get('[data-testid="chart-figure"] h2').should('contain.text', 'Time To First Token'); }); - it('offers only the three supported axes for fixed sequences', () => { + it('offers only the four supported axes for fixed sequences', () => { interceptFixedSequenceData(); cy.visit('/inference?i_seq=8k%2F1k'); openXAxisMenu(); - cy.get('[role="grid"] [data-select-option]').should('have.length', 3); + cy.get('[role="grid"] [data-select-option]').should('have.length', 4); + cy.get('[data-testid="x-axis-mode-concurrency"]').should('be.visible'); cy.get('[data-testid="x-axis-mode-interactivity"]').should('be.visible'); cy.get('[data-testid="x-axis-mode-e2e"]').should('be.visible'); cy.get('[data-testid="x-axis-mode-ttft"]').should('be.visible'); @@ -586,8 +568,8 @@ const expectCoordinates = (selector: string, expected: number[]) => { }); }; const expectInteractivity = () => { - cy.get('#chart-0 .x-axis-label').should('have.text', 'Interactivity (tok/s/user)'); - cy.get('[data-testid="chart-figure"] h2').should('contain.text', 'vs. Interactivity'); + cy.get('#chart-0 .x-axis-label').should('have.text', 'Median Interactivity (tok/s/user)'); + cy.get('[data-testid="chart-figure"] h2').should('contain.text', 'vs. Median Interactivity'); expectCoordinates('.dot-group', [80, 40, 20]); expectCoordinates('.unofficial-overlay-pt', [60, 30]); }; diff --git a/packages/app/cypress/support/mock-data.ts b/packages/app/cypress/support/mock-data.ts index 413f38fab..6d5823812 100644 --- a/packages/app/cypress/support/mock-data.ts +++ b/packages/app/cypress/support/mock-data.ts @@ -91,11 +91,13 @@ export function createMockHardwareConfig(): HardwareConfig { // --------------------------------------------------------------------------- export function createMockChartDefinition(overrides?: Partial): ChartDefinition { + const chartType = overrides?.chartType ?? 'e2e'; + const x = overrides?.x ?? (chartType === 'interactivity' ? 'median_intvty' : 'median_e2el'); return { - chartType: 'e2e', + chartType, heading: 'End-to-End Latency vs Throughput', - x: 'conc' as keyof AggDataEntry, - x_label: 'Concurrency', + x, + x_label: chartType === 'interactivity' ? 'Interactivity' : 'End-to-end Latency (s)', y: 'mean_e2el' as keyof AggDataEntry, y_label: 'Mean E2E Latency (ms)', y_tpPerGpu: 'tput_per_gpu', @@ -103,8 +105,9 @@ export function createMockChartDefinition(overrides?: Partial): y_tpPerGpu_title: 'Throughput per Chip', y_tpPerGpu_roofline: 'upper_right', ...overrides, - x_scale_field: overrides?.x_scale_field ?? String(overrides?.x ?? 'conc'), - x_labelZh: overrides?.x_labelZh ?? '并发数', + x_scale_field: overrides?.x_scale_field ?? String(x), + x_labelZh: + overrides?.x_labelZh ?? (chartType === 'interactivity' ? '交互性' : '端到端延迟(s)'), }; } @@ -224,7 +227,8 @@ export function createMockInferenceContextValues( openRouterPricingError: null, setTokenRevenuePriceSource: namedStub('setTokenRevenuePriceSource'), selectedPercentile: 'p90', - setSelectedPercentile: namedStub('setSelectedPercentile'), + fixedSequenceStatistic: 'median', + setFixedSequenceStatistic: namedStub('setFixedSequenceStatistic'), selectedXAxisMetric: null, setSelectedXAxisMetric: namedStub('setSelectedXAxisMetric'), selectedE2eXAxisMetric: null, @@ -232,6 +236,8 @@ export function createMockInferenceContextValues( setSelectedXAxisMode: namedStub('setSelectedXAxisMode'), scaleType: 'auto', setScaleType: namedStub('setScaleType'), + powerCompare: 'none' as const, + setPowerCompare: namedStub('setPowerCompare'), quickFilters: { vendors: [], frameworks: [], deployment: [], spec: [], power: [] }, availableQuickFilters: { vendors: [], frameworks: [], deployment: [], spec: [], power: [] }, setQuickFilterVendors: namedStub('setQuickFilterVendors'), @@ -239,6 +245,7 @@ export function createMockInferenceContextValues( setQuickFilterDeployment: namedStub('setQuickFilterDeployment'), setQuickFilterSpec: namedStub('setQuickFilterSpec'), setQuickFilterPower: namedStub('setQuickFilterPower'), + setQuickFilterTopologies: namedStub('setQuickFilterTopologies'), isLegendExpanded: true, setIsLegendExpanded: namedStub('setIsLegendExpanded'), hideNonOptimal: false, diff --git a/packages/app/cypress/support/share-link.ts b/packages/app/cypress/support/share-link.ts new file mode 100644 index 000000000..29081a970 --- /dev/null +++ b/packages/app/cypress/support/share-link.ts @@ -0,0 +1,20 @@ +/** + * The address bar is stripped clean after load (share-link state lives in the + * in-memory store), so the Share popover is where share-link state must land. + * Asserts several parameters in one popover round trip; a `null` expectation + * means the parameter must be absent (stripped as a default). + */ +export function assertShareLinkParams(expected: Record): void { + cy.get('[data-testid="share-button"]').first().click(); + cy.get('[data-testid="share-url-input"]') + .invoke('val') + .should((value) => { + const params = new URL(String(value)).searchParams; + for (const [key, param] of Object.entries(expected)) { + expect(params.get(key), `${key} in the share link`).to.eq(param); + } + }); + // The link field is read-only, so dismiss the popover from the document. + cy.get('body').type('{esc}'); + cy.get('[data-testid="share-url-input"]').should('not.exist'); +} diff --git a/packages/app/src/app/api/v1/views/extensions.test.ts b/packages/app/src/app/api/v1/views/extensions.test.ts index 16f2ea239..0e38c701a 100644 --- a/packages/app/src/app/api/v1/views/extensions.test.ts +++ b/packages/app/src/app/api/v1/views/extensions.test.ts @@ -416,7 +416,7 @@ describe('new dashboard projections', () => { expect(output.skipped).toEqual([]); expect(output.rows.map((row: { powerLabel: string }) => row.powerLabel)).toEqual( powerBasis === 'compare' - ? ['Provisioned', 'Measured + modeled · Full-chassis extrapolation'] + ? ['All in Provisioned', 'All in Measured · Full-chassis extrapolation'] : ['Full-chassis extrapolation'], ); } diff --git a/packages/app/src/app/api/v1/views/inference/route.test.ts b/packages/app/src/app/api/v1/views/inference/route.test.ts index 612162960..6c14fb6be 100644 --- a/packages/app/src/app/api/v1/views/inference/route.test.ts +++ b/packages/app/src/app/api/v1/views/inference/route.test.ts @@ -1,6 +1,12 @@ import { NextRequest } from 'next/server'; import { beforeEach, describe, expect, it, vi } from 'vitest'; +import { + buildEqualServiceComparison, + getEqualServiceSources, +} from '@/components/inference/utils/equal-service-comparison'; +import { buildInferenceSeries } from '@/lib/views-api/series'; +import { Sequence } from '@/lib/data-mappings'; import type { BenchmarkRow } from '@/lib/api'; const { mockGetLatestBenchmarks, mockGetBenchmarksForRun, mockUnofficialRun, mockGetDb } = @@ -96,6 +102,172 @@ beforeEach(() => { }); describe('GET /api/v1/views/inference', () => { + it('compares stitched observations before frontier pruning and preserves each producer endpoint', async () => { + const rows = ['h200', 'mi300x'].flatMap((hardware, index) => + [20, 60].map((x, position) => + makeRow({ + hardware, + conc: position + 1, + curve_workflow_run_id: 900, + curve_date: '2026-03-02', + date: position === 0 ? '2026-03-01' : '2026-03-02', + run_url: `https://github.com/org/repo/actions/runs/${777 + position}`, + power_audit: { + producer_sha: `producer-${position}`, + exporter_image_sha256: `exporter-${position}`, + }, + metrics: { + ...makeRow().metrics, + median_intvty: x, + avg_power_w: 200 + 200 * position + 50 * index, + output_tput_per_gpu: 100 + 100 * position + 50 * index, + joules_per_output_token: 4 - 2 * position - index, + power_valid: 1, + power_metric_schema_version: 2, + }, + }), + ), + ); + mockGetLatestBenchmarks.mockResolvedValue(rows); + const projected = buildInferenceSeries(rows, { + sequence: Sequence.EightK_OneK, + percentile: 'p90', + precisions: ['fp8'], + metricConfigKey: 'y_measuredAvgPower', + xmode: 'interactivity', + xmetric: 'p90_ttft', + gpus: [], + quickFilters: { vendors: [], frameworks: [], deployment: [], spec: [], power: [] }, + optimal: true, + best: true, + }); + const sources = getEqualServiceSources(projected.observedPoints); + const expected = buildEqualServiceComparison(projected.observedPoints, { + baseline: sources[0].key, + comparator: sources[1].key, + target: 40, + xField: 'median_intvty', + }); + const response = await GET( + request( + '/api/v1/views/inference?model=DeepSeek-R1-0528&metric=measuredAvgPower&serviceCompare=true&serviceTarget=40', + ), + ); + const body = await response.json(); + expect(response.status).toBe(200); + expect(body.serviceSources).toHaveLength(2); + expect(body.serviceSources).toEqual(sources); + expect(body.equalServiceComparison.metrics.meanWattsPerGpu.changePercent).toBeCloseTo(100 / 6); + expect(body.equalServiceComparison.metrics.outputTokensPerSecond.changePercent).toBeCloseTo( + 100 / 3, + ); + expect(body.equalServiceComparison.metrics.joulesPerOutputToken.changePercent).toBeCloseTo( + -100 / 3, + ); + for (const [metric, values] of Object.entries(expected.metrics)) { + expect(body.equalServiceComparison.metrics[metric].changePercent).toBe(values.changePercent); + expect(body.equalServiceComparison.metrics[metric].baseline.value).toBe( + values.baseline?.value, + ); + } + expect( + body.equalServiceComparison.metrics.meanWattsPerGpu.baseline.endpoints.map( + (endpoint: { point: { id: number } }) => endpoint.point.id, + ), + ).toEqual(rows.slice(0, 2).map((row) => row.id)); + expect( + body.equalServiceComparison.metrics.meanWattsPerGpu.baseline.endpoints.map( + (endpoint: { point: { runUrl: string } }) => endpoint.point.runUrl, + ), + ).toEqual(rows.slice(0, 2).map((row) => row.run_url)); + expect(body.matchedConcurrency.rows).toMatchObject([ + { concurrency: 1, baseline: { status: 'observed' }, comparator: { status: 'observed' } }, + { concurrency: 2, baseline: { status: 'observed' }, comparator: { status: 'observed' } }, + ]); + expect(body).not.toHaveProperty('observedPoints'); + expect(body.equalServiceCurve).toHaveLength(2); + }); + + it('includes unofficial sources and validated role shares using one output-token denominator', async () => { + const row = makeRow({ + hardware: 'gb200', + framework: 'trt', + disagg: true, + num_prefill_gpu: 4, + num_decode_gpu: 4, + prefill_num_workers: 1, + decode_num_workers: 1, + metrics: { + ...makeRow().metrics, + avg_power_w: 400, + power_valid: 1, + power_metric_schema_version: 2, + joules_per_input_token: 1, + joules_per_output_token: 8, + prefill_joules_per_input_token: 0.4, + decode_joules_per_output_token: 4.8, + }, + }); + mockGetLatestBenchmarks.mockResolvedValue([row]); + mockUnofficialRun.mockImplementation(() => + Response.json({ + benchmarks: [{ ...row, id: 999, run_url: 'https://github.com/org/repo/actions/runs/999' }], + evaluations: [], + }), + ); + const response = await GET( + request( + '/api/v1/views/inference?model=DeepSeek-R1-0528&metric=measuredAvgPower&serviceCompare=true&roleShare=true&unofficialrun=999', + ), + ); + const body = await response.json(); + expect(response.status).toBe(200); + expect(body.serviceSources).toHaveLength(2); + expect(body.roleEnergyShares).toHaveLength(2); + expect( + body.roleEnergyShares + .map((item: { point: { id: number } }) => item.point.id) + .sort((a: number, b: number) => a - b), + ).toEqual([row.id, 999].sort((a, b) => a - b)); + for (const share of body.roleEnergyShares) { + expect(share).toMatchObject({ + prefill: 3.2, + decode: 4.8, + total: 8, + prefillShare: 40, + decodeShare: 60, + }); + } + expect(body.overlays[0]).not.toHaveProperty('observedPoints'); + }); + + it('resolves the fixed-sequence statistic and rejects unknown values', async () => { + mockGetLatestBenchmarks.mockResolvedValue([ + makeRow({ metrics: { ...makeRow().metrics, mean_tpot: 0.025, mean_intvty: 99 } }), + ]); + const response = await GET( + request('/api/v1/views/inference?model=DeepSeek-R1-0528&metric=tpPerGpu&xstat=mean'), + ); + const body = await response.json(); + expect(body.params.xstat).toBe('mean'); + expect(body.xAxis).toMatchObject({ field: 'mean_tpot_intvty', statistic: 'mean' }); + expect(body.series[0].points[0].x).toBe(40); + const diagnostic = await GET( + request( + '/api/v1/views/inference?model=DeepSeek-R1-0528&metric=tpPerGpu&xstat=mean&xmode=concurrency', + ), + ); + const diagnosticBody = await diagnostic.json(); + expect(diagnosticBody.params.xstat).toBeNull(); + expect(diagnosticBody.xAxis.statistic).toBeNull(); + const invalid = await GET( + request('/api/v1/views/inference?model=DeepSeek-R1-0528&xstat=average'), + ); + expect(invalid.status).toBe(400); + const invalidBody = await invalid.json(); + expect(invalidBody.param).toBe('xstat'); + }); + it('defaults Qwen to the cross-framework winner while best=false keeps both engines', async () => { const rows = ['vllm', 'sglang'].flatMap((framework) => [16, 64].map((conc) => diff --git a/packages/app/src/app/api/v1/views/inference/route.ts b/packages/app/src/app/api/v1/views/inference/route.ts index 14c3f61a1..8ae6f0d7d 100644 --- a/packages/app/src/app/api/v1/views/inference/route.ts +++ b/packages/app/src/app/api/v1/views/inference/route.ts @@ -1,7 +1,29 @@ import { GET as derived } from '@/app/api/v1/derived-agentic-metrics/route'; import { preferVrDefaultRun, VR_DEFAULT_RUN } from '@/components/inference/default-run-preference'; import { NORMALIZED_TOKEN_REVENUE_PRICING } from '@/components/inference/token-revenue'; -import type { TokenRevenuePricing } from '@/components/inference/types'; +import { + buildEqualServiceComparison, + equalServiceSourceKey, + getEqualServiceComparisonCurve, + getEqualServiceSources, + getPrefillSharePoints, + getRolePoints, + type EqualServiceComparison, + type EqualServiceEstimate, +} from '@/components/inference/utils/equal-service-comparison'; +import { + buildMatchedConcurrencyTable, + type MatchedConcurrencySide, + type MatchedConcurrencyTable, +} from '@/components/inference/utils/matched-concurrency'; +import { buildPowerFits } from '@/components/inference/utils/power-fit'; +import { resolveServiceField } from '@/components/inference/utils/resolveXAxisField'; +import { pointTopologyKey } from '@/components/inference/utils/topology-filter'; +import type { + InferenceData, + AggDataEntry, + TokenRevenuePricing, +} from '@/components/inference/types'; import type { DerivedAgenticMetricMap } from '@/hooks/api/use-derived-agentic-metrics'; import { fetchOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing'; import type { TcoBasis } from '@/lib/constants'; @@ -16,6 +38,7 @@ import { import { ViewsApiParamError, runViewsRoute } from '@/lib/views-api/errors'; import { parseNumberMap, + parseNumberParam, validateParams as validateViewParams, parseBoolParam, parseDateParam, @@ -91,6 +114,7 @@ interface InferenceViewParams { readonly precisionsExplicit: boolean; readonly metric: string; readonly xmode: SeriesXMode; + readonly xstat: 'mean' | 'median'; readonly xmetric: string; readonly percentile: string; readonly date?: string; @@ -105,6 +129,7 @@ interface InferenceViewParams { readonly best: boolean; readonly tcoBasis: TcoBasis; readonly power: readonly ('certified' | 'legacy')[]; + readonly topologies: readonly string[]; readonly allPoints: boolean; readonly userCosts: Record; readonly userPowers: Record; @@ -146,6 +171,7 @@ function buildView( precisions: resolvedPrecisions, metricConfigKey: parseMetricParam(params.metric), xmode: params.xmode, + fixedSequenceStatistic: params.xstat, xmetric: params.xmetric, gpus: params.gpus, quickFilters: { @@ -156,6 +182,7 @@ function buildView( // Measured-power tier pills are a dashboard-only affordance; the API // returns every row regardless of power certification, like the default view. power: [...params.power], + topologies: [...params.topologies], }, optimal: params.optimal, best: params.best, @@ -170,7 +197,7 @@ function buildView( return { resolvedPrecisions, result }; } -function csvRows(data: InferenceViewData) { +function csvRows(data: { result: Pick }) { return data.result.series.flatMap((entry) => entry.points.map((point) => ({ hwKey: entry.hwKey, @@ -184,6 +211,7 @@ function csvRows(data: InferenceViewData) { x: point.x, y: point.y, concurrency: point.concurrency, + topologyKey: point.topologyKey, tp: point.tp, date: point.date, runId: point.runId ?? '', @@ -196,6 +224,63 @@ function csvRows(data: InferenceViewData) { ); } +/** Keep public endpoint provenance without serializing internal chart objects. */ +function observedPointIdentity(point: InferenceData) { + return { + id: point.id ?? null, + sourceKey: equalServiceSourceKey(point), + hwKey: point.hwKey, + precision: point.precision, + concurrency: point.conc, + topologyKey: pointTopologyKey(point), + date: point.actualDate ?? point.date, + runUrl: point.run_url ?? null, + recipeFingerprint: point.recipe_fingerprint ?? null, + image: point.image ?? null, + }; +} + +function publicServiceComparison(comparison: EqualServiceComparison) { + const estimate = (value: EqualServiceEstimate | null) => + value === null + ? null + : { + ...value, + endpoints: value.endpoints.map(({ point, ...endpoint }) => ({ + ...endpoint, + point: observedPointIdentity(point), + })), + }; + return { + ...comparison, + metrics: Object.fromEntries( + Object.entries(comparison.metrics).map(([key, metric]) => [ + key, + { ...metric, baseline: estimate(metric.baseline), comparator: estimate(metric.comparator) }, + ]), + ), + }; +} + +function publicMatchedSide(side: MatchedConcurrencySide) { + if (side.status === 'observed') + return { status: side.status, values: side.values, point: observedPointIdentity(side.point) }; + if (side.status === 'ambiguous') + return { status: side.status, points: side.points.map(observedPointIdentity) }; + return side; +} + +function publicMatchedConcurrency(table: MatchedConcurrencyTable) { + return { + ...table, + rows: table.rows.map((row) => ({ + ...row, + baseline: publicMatchedSide(row.baseline), + comparator: publicMatchedSide(row.comparator), + })), + }; +} + export function GET(request: NextRequest) { return runViewsRoute('inference', async () => { validateViewParams(request.nextUrl.searchParams, VIEW_QUERY_PARAMS['inference']); @@ -217,6 +302,7 @@ export function GET(request: NextRequest) { requestedXMode === 'e2e-normalized-interactivity' && sequence !== Sequence.AgenticTraces ? 'interactivity' : requestedXMode; + const xstat = parseEnumParam(search.get('xstat'), 'xstat', ['mean', 'median'], 'median'); const xmetric = parseEnumParam(search.get('xmetric'), 'xmetric', XMETRIC_VALUES, 'p90_ttft'); const percentile = parseEnumParam( search.get('percentile'), @@ -231,15 +317,41 @@ export function GET(request: NextRequest) { const frameworks = parseFrameworkFamiliesParam(search.get('frameworks')); const deployment = parseDeploymentParam(search.get('deployment')); const spec = parseSpecModesParam(search.get('spec')); - const optimal = parseBoolParam(search.get('optimal'), 'optimal', true); - const best = parseBoolParam( + const requestedOptimal = parseBoolParam(search.get('optimal'), 'optimal', true); + const requestedBest = parseBoolParam( search.get('best'), 'best', !isBestPerSkuDefaultOff(displayName as Model, sequence), ); const format = parseFormatParam(search.get('format')); + const serviceCompare = parseBoolParam(search.get('serviceCompare'), 'serviceCompare', false); + const roleShare = parseBoolParam(search.get('roleShare'), 'roleShare', false); + const powerFit = parseBoolParam(search.get('powerFit'), 'powerFit', false); + const requestedServiceBaseline = search.get('serviceBaseline'); + const requestedServiceComparator = search.get('serviceComparator'); + const serviceTarget = search.has('serviceTarget') + ? parseNumberParam(search.get('serviceTarget'), 'serviceTarget', 0, { min: Number.MIN_VALUE }) + : null; + if (serviceTarget !== null && serviceTarget <= 0) + throw new ViewsApiParamError('serviceTarget', 'serviceTarget must be positive'); + if (format === 'csv' && (serviceCompare || roleShare || powerFit)) + throw new ViewsApiParamError( + 'format', + 'Equal-service, role-share and power-fit panels require format=json', + ); const tcoBasis = parseTcoBasisParam(search.get('tcoBasis')); const power = parseListParam(search.get('power'), 'power', POWER_TIER_ORDER); + // Topology keys are opaque values returned by this view, not case-folded names. + const topologies = [ + ...new Set( + (search.get('topologies') ?? '') + .split(',') + .map((key) => key.trim()) + .filter(Boolean), + ), + ].toSorted(); + const optimal = xmode === 'concurrency' ? false : requestedOptimal; + const best = xmode === 'concurrency' ? false : requestedBest; const allPoints = parseBoolParam(search.get('allPoints'), 'allPoints', false); const userCosts = parseNumberMap(search.get('userCosts'), 'userCosts'); const userPowers = parseNumberMap(search.get('userPowers'), 'userPowers'); @@ -265,6 +377,7 @@ export function GET(request: NextRequest) { precisionsExplicit: precisions.length > 0, metric, xmode, + xstat, xmetric, percentile, ...(date ? { date } : {}), @@ -278,6 +391,7 @@ export function GET(request: NextRequest) { best, tcoBasis, power, + topologies, allPoints, userCosts, userPowers, @@ -326,10 +440,16 @@ export function GET(request: NextRequest) { await getCachedBenchmarks([...dbModelKeys], VR_DEFAULT_RUN.date, true), ); const data = await project(rows, params); + const observedPoints = [...data.result.observedPoints]; + // Service/role panels pool every scope's observed points; the payload omits them. + const collectObserved = ({ observedPoints: points, ...result }: InferenceSeriesResult) => { + observedPoints.push(...points); + return result; + }; const comparisons = await Promise.all( scopes.map(async (scope) => { const comparison = await project(await fetchRows(scope.params), scope.params); - return { entry: scope.entry, ...comparison.result }; + return { entry: scope.entry, ...collectObserved(comparison.result) }; }), ); const overlayRows = await unofficialRows(request); @@ -339,16 +459,80 @@ export function GET(request: NextRequest) { overlayRows.filter((row) => row.run_url === url), params, ); - return { runUrl: url, ...overlay.result }; + return { runUrl: url, ...collectObserved(overlay.result) }; }), ); + const serviceSources = serviceCompare ? getEqualServiceSources(observedPoints) : []; + const serviceBaseline = requestedServiceBaseline ?? serviceSources[0]?.key ?? ''; + const serviceComparator = requestedServiceComparator ?? serviceSources[1]?.key ?? ''; + const serviceOptions = { + baseline: serviceBaseline, + comparator: serviceComparator, + xField: data.result.xAxis.field as keyof AggDataEntry, + }; + const servicePanels = serviceCompare + ? { + serviceSources, + equalServiceComparison: + serviceTarget === null + ? null + : publicServiceComparison( + buildEqualServiceComparison(observedPoints, { + ...serviceOptions, + target: serviceTarget, + }), + ), + equalServiceCurve: getEqualServiceComparisonCurve(observedPoints, serviceOptions).map( + publicServiceComparison, + ), + matchedConcurrency: publicMatchedConcurrency( + buildMatchedConcurrencyTable(observedPoints, { + baseline: serviceBaseline, + comparator: serviceComparator, + // The dashboard reads streaming speed at the selected statistic. + interactivityField: resolveServiceField('median_intvty', { + isAgentic: sequence === Sequence.AgenticTraces, + percentile, + fixedSequenceStatistic: xstat, + }), + }), + ), + } + : {}; + const rolePanel = roleShare + ? { + roleEnergyShares: getPrefillSharePoints(observedPoints, serviceOptions.xField).map( + ({ point, ...energy }) => ({ + ...energy, + point: observedPointIdentity(point), + decodeShare: 100 - energy.prefillShare, + }), + ), + rolePoints: getRolePoints(observedPoints, serviceOptions.xField).map( + ({ point, ...role }) => ({ ...role, point: observedPointIdentity(point) }), + ), + } + : {}; + const fitPanel = powerFit + ? { + powerFits: buildPowerFits(observedPoints).map(({ observations, ...fit }) => ({ + ...fit, + observations: observations.map(({ point, ...observation }) => ({ + ...observation, + point: observedPointIdentity(point), + })), + })), + } + : {}; + const resolvedParams = { model: displayName, sequence: sequence as string, precisions: data.resolvedPrecisions, metric, xmode, + xstat: sequence === Sequence.AgenticTraces || xmode === 'concurrency' ? null : xstat, xmetric, percentile, date: date ?? null, @@ -358,6 +542,7 @@ export function GET(request: NextRequest) { frameworks, deployment, spec, + topologies, optimal, best, format, @@ -369,6 +554,12 @@ export function GET(request: NextRequest) { priceSource, dates, unofficialrun: search.get('unofficialrun') ?? null, + serviceCompare, + serviceBaseline: serviceCompare ? serviceBaseline : null, + serviceComparator: serviceCompare ? serviceComparator : null, + serviceTarget: serviceCompare ? serviceTarget : null, + roleShare, + powerFit, }; if (format === 'csv') { @@ -396,6 +587,9 @@ export function GET(request: NextRequest) { comparisons, overlays, pricing, + ...servicePanels, + ...rolePanel, + ...fitPanel, }); }); } diff --git a/packages/app/src/app/api/v1/views/options/route.ts b/packages/app/src/app/api/v1/views/options/route.ts index 90bde1be0..a1e766c22 100644 --- a/packages/app/src/app/api/v1/views/options/route.ts +++ b/packages/app/src/app/api/v1/views/options/route.ts @@ -152,6 +152,7 @@ function buildOptionsPayload() { specMethods: [...SPEC_METHOD_KEYS].toSorted(), percentiles: PERCENTILE_OPTIONS, xAxisModes: X_AXIS_MODES, + fixedSequenceStatistics: ['median', 'mean'], scaleModes: ['auto', 'linear', 'log'], metrics, quickFilters: { @@ -190,6 +191,10 @@ function buildOptionsPayload() { metric: DEFAULT_METRIC_CONFIG_KEY, percentile: 'p90', xmode: 'interactivity', + xstat: 'median', + serviceCompare: false, + roleShare: false, + powerFit: false, xmetric: 'p90_ttft', scale: 'auto', precisions: 'auto', diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx index 71b6a961e..9dcb489ca 100644 --- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx +++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx @@ -61,6 +61,7 @@ import { captionControlTriggerClassName, ResultContext } from '@/components/ui/r import { SearchableSelect } from '@/components/ui/searchable-select'; import { Skeleton } from '@/components/ui/skeleton'; import { TooltipProvider } from '@/components/ui/tooltip'; +import { ALL_IN_MEASURED_NOTE, POWER_BASIS_LABELS } from '@/lib/power-basis'; import { useComparisonChangelogs } from '@/hooks/api/use-comparison-changelogs'; import { useOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing'; import { useOpenDropdown } from '@/hooks/useOpenDropdown'; @@ -210,17 +211,19 @@ const STRINGS = { powerTooltip: 'Change only the power budget used to scale the same benchmark result to one GW. Pricing, throughput, utilization and unit costs stay the same.', powerOptions: { - provisioned: 'Provisioned power', - modeled: 'Measured + modeled power', + provisioned: POWER_BASIS_LABELS['utility-provisioned'].en, + modeled: POWER_BASIS_LABELS['utility-modeled'].en, compare: 'Compare both', }, powerBarLabels: { - provisioned: 'Provisioned', - modeled: 'Measured + modeled', + provisioned: POWER_BASIS_LABELS['utility-provisioned'].en, + modeled: POWER_BASIS_LABELS['utility-modeled'].en, extrapolated: 'Full-chassis extrapolation', }, - powerPreview: - 'PowerX estimate · Same target, throughput, pricing and unit costs. GPU power comes from the same serving-frontier points; power between them is estimated linearly. Server overhead is modeled, with PUE 1.3 and 10% headroom. Full-chassis extrapolation fills an eight-GPU server with replicas of the measured 1/2/4-GPU workload at the same per-GPU power and throughput; it does not measure a partly idle server. AgentX system power is not yet qualified.', + powerPreview: `${ALL_IN_MEASURED_NOTE.en} AgentX system power is not yet qualified.`, + powerDetails: + 'GPU power is interpolated between the same throughput points. Includes PUE 1.3 and 10% headroom. Full-chassis extrapolation fills an eight-GPU server with replicas of the measured 1/2/4-GPU workload at the same per-GPU power and throughput; it does not measure a partly idle server.', + unavailableEstimates: (count: number) => `Unavailable estimates (${count})`, pricingGroup: 'Pricing Config', costProviderLabel: 'Cost Provider', costProviderTooltip: @@ -339,13 +342,19 @@ const STRINGS = { powerTooltip: '仅更改将同一基准测试结果换算为每 GW 收益时采用的功耗预算。价格、吞吐量、利用率和单位成本保持不变。', powerOptions: { - provisioned: '预配功耗', - modeled: '实测 GPU + 系统功耗估算', + provisioned: POWER_BASIS_LABELS['utility-provisioned'].zh, + modeled: POWER_BASIS_LABELS['utility-modeled'].zh, compare: '对比两种估算方式', }, - powerBarLabels: { provisioned: '预配功耗', modeled: '实测 + 估算', extrapolated: '整机外推' }, - powerPreview: - 'PowerX 估算 · 两种方式采用相同的目标交互性、吞吐量、价格和单位成本。GPU 功耗取自同一组性能前沿数据点,点间功耗采用线性估算。服务器开销由模型估算,PUE 为 1.3,功耗余量为 10%。整机外推假设在八卡服务器上部署多个相同的实测单卡、双卡或四卡实例,每卡功耗和吞吐量保持不变;它不代表部分 GPU 闲置时的整机实测功耗。AgentX 系统功耗模型尚未完成验证。', + powerBarLabels: { + provisioned: POWER_BASIS_LABELS['utility-provisioned'].zh, + modeled: POWER_BASIS_LABELS['utility-modeled'].zh, + extrapolated: '整机外推', + }, + powerPreview: `${ALL_IN_MEASURED_NOTE.zh} AgentX 系统功耗模型尚未完成验证。`, + powerDetails: + 'GPU 功耗在相同的吞吐量数据点间插值,计入 PUE 1.3 和 10% 功耗余量。整机外推假设八卡服务器部署多个相同的实测单卡、双卡或四卡实例,每卡功耗和吞吐量保持不变;它不代表部分 GPU 闲置时的整机实测功耗。', + unavailableEstimates: (count: number) => `无法估算(${count} 项)`, pricingGroup: '定价配置', costProviderLabel: '成本供应商', costProviderTooltip: @@ -1409,13 +1418,21 @@ function ProfitEstimatorInner({ {powerControlsEnabled && (

{t.powerLabel}: {t.powerOptions[powerBasis]} - {powerBasis !== 'provisioned' && <>. {t.powerPreview}}

)} {basis === 'gw-year' && powerBasis !== 'provisioned' && fullEstimate.skipped.length > 0 && ( -

- {powerUnavailable} -

+
+ track('profit_estimator_power_unavailable_toggled')} + > + {t.unavailableEstimates(fullEstimate.skipped.length)} + +

{powerUnavailable}

+
)}
{ + return value !== null && typeof value === 'object' && !Array.isArray(value); +} + function parseJsonObject(text: string | undefined): Record | null { if (text === undefined) return null; try { const parsed: unknown = JSON.parse(text); - return parsed !== null && typeof parsed === 'object' && !Array.isArray(parsed) - ? (parsed as Record) - : null; + return isRecord(parsed) ? parsed : null; } catch { return null; } } -function isRecord(value: unknown): value is Record { - return value !== null && typeof value === 'object' && !Array.isArray(value); -} - function manifestSlot(hostname: string, gpuIndex: number): string { return `${hostname}#${gpuIndex}`; } @@ -266,7 +264,7 @@ export function cutPowerAuditBundle( const manifest = parseJsonObject(files.get(BUNDLE_MANIFEST_ENTRY)); const contextFiles = [...files] .filter(([name]) => isContextEntry(name)) - .sort(([a], [b]) => (a < b ? -1 : a > b ? 1 : 0)); + .sort(([a], [b]) => compareText(a, b)); const smiFiles = [...files] .filter(([name]) => isSmiCsv(name)) .map(([name, text]) => { @@ -287,13 +285,7 @@ export function cutPowerAuditBundle( if (samplesText === undefined) return []; const { samples, devices } = parseSamples(samplesText); if (samples.length === 0) return []; - return cutPowerAuditSamples( - artifact, - samples, - devices, - validations, - parseJsonObject(files.get(BUNDLE_MANIFEST_ENTRY)), - ); + return cutPowerAuditSamples(artifact, samples, devices, validations, manifest); } /** Apply bundle window cuts to the richer SMI CSVs without requiring DCGM UUID sidecars. */ diff --git a/packages/app/src/components/gpu-power/power-series.ts b/packages/app/src/components/gpu-power/power-series.ts index 8ff75fabb..0f98d1e80 100644 --- a/packages/app/src/components/gpu-power/power-series.ts +++ b/packages/app/src/components/gpu-power/power-series.ts @@ -51,6 +51,8 @@ export interface GpuPowerSeries { } export interface GpuPowerSeriesResponse { + /** Where the route read the series: the ingested digest or the run's GitHub artifacts. */ + source?: 'database' | 'github'; runInfo: GpuPowerRunInfo; series: GpuPowerSeries[]; /** Coverage of requested validation identities only, never the full run/sweep. */ diff --git a/packages/app/src/components/gpu-power/stored-power-series.ts b/packages/app/src/components/gpu-power/stored-power-series.ts index 964d44812..80dfb6b96 100644 --- a/packages/app/src/components/gpu-power/stored-power-series.ts +++ b/packages/app/src/components/gpu-power/stored-power-series.ts @@ -3,15 +3,12 @@ import type { GpuMetricSeries } from '@semianalysisai/inferencex-db/queries/gpu- import { cutPowerAuditSamples, cutPowerAuditCsvs, + isRecord, type PowerAuditDevice, type PowerAuditSample, } from './power-audit-bundle'; import { bucketPowerFiles, type GpuPowerSeries } from './power-series'; -function isRecord(value: unknown): value is Record { - return value !== null && typeof value === 'object' && !Array.isArray(value); -} - export class StoredTelemetryIncompleteError extends Error { readonly artifact: string; diff --git a/packages/app/src/components/gpu-power/types.ts b/packages/app/src/components/gpu-power/types.ts index 6c82d5616..d4bdbf005 100644 --- a/packages/app/src/components/gpu-power/types.ts +++ b/packages/app/src/components/gpu-power/types.ts @@ -200,10 +200,6 @@ export function getAvailableMetrics(data: GpuMetricRow[]): GpuMetricConfig[] { return ALL_METRIC_OPTIONS.filter((m) => data.some((row) => Number.isFinite(row[m.key]))); } -/** - * Detect GPU SKU from an artifact name and return its TDP in watts. - * Artifact names look like: gpu_metrics_dsr1_1k8k_fp8_sglang_tp8_..._h200-nb_0 - */ /** TDP for a known hardware key, e.g. the benchmark point's own `hardware`. */ export function tdpForHardware(hardware: string | undefined): { sku: string; tdp: number } | null { const key = hardware?.toLowerCase(); @@ -211,6 +207,10 @@ export function tdpForHardware(hardware: string | undefined): { sku: string; tdp return entry ? { sku: key!.toUpperCase(), tdp: entry.tdp } : null; } +/** + * Detect GPU SKU from an artifact name and return its TDP in watts. + * Artifact names look like: gpu_metrics_dsr1_1k8k_fp8_sglang_tp8_..._h200-nb_0 + */ export function detectTdpFromArtifactName( artifactName: string, ): { sku: string; tdp: number } | null { diff --git a/packages/app/src/components/inference/InferenceContext.tsx b/packages/app/src/components/inference/InferenceContext.tsx index 4b874817a..1bbbaf41d 100644 --- a/packages/app/src/components/inference/InferenceContext.tsx +++ b/packages/app/src/components/inference/InferenceContext.tsx @@ -43,6 +43,7 @@ import type { InferenceDataContextType, InferenceDisplayContextType, InferenceFiltersContextType, + PowerCompare, TokenRevenuePriceSource, } from '@/components/inference/types'; import { resolveMetricConfigKey } from '@/components/inference/metric-registry'; @@ -62,6 +63,15 @@ import { useUrlStateSync, } from '@/hooks/useChartContext'; import { useUrlState } from '@/hooks/useUrlState'; +import { serializePerfRulers } from '@/lib/d3-chart/layers/perf-ruler'; +import type { FixedSequenceStatistic } from '@/components/inference/utils/resolveXAxisField'; +import { parsePowerCompare } from '@/components/inference/utils/power-compare'; +import { + PERSISTED_PERF_RULER_CHART_ID, + PerfRulerStoreContext, + persistedPerfRulerAxisKey, + usePerfRulerStoreValue, +} from '@/components/inference/perf-ruler-store'; import { useParetoHighlightToggle } from './hooks/useParetoHighlightToggle'; import { useOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing'; import { DEFAULT_Y_AXIS_METRIC } from '@/lib/url-state'; @@ -200,9 +210,12 @@ export function resolveE2eXAxisMetric( mode: XAxisMode, sequence: Parameters[0], percentile: string, + fixedSequenceStatistic: FixedSequenceStatistic = 'median', ): string | null { if (mode === 'ttft') { - return sequenceKind(sequence) === 'agentic' ? `${percentile}_ttft` : 'median_ttft'; + return sequenceKind(sequence) === 'agentic' + ? `${percentile}_ttft` + : `${fixedSequenceStatistic}_ttft`; } if (mode === 'e2e') return null; return requestedMetric; @@ -489,20 +502,31 @@ export function InferenceProvider({ xAxisModeFromUrlRef.current = true; setRequestedXAxisMode(mode); }, []); - // Latency percentile applied to the chart x-axis for agentic scenarios. - // Values: 'p90' | 'p99'. Non-agentic charts ignore. - const [selectedPercentile, setSelectedPercentile] = useState( - () => getUrlParam('i_pctl') || 'p90', - ); + const [fixedSequenceStatistic, setFixedSequenceStatistic] = + useState('median'); + useEffect(() => { + setFixedSequenceStatistic(getUrlParam('i_mstat') === 'mean' ? 'mean' : 'median'); + // eslint-disable-next-line react-hooks/exhaustive-deps + }, []); + + // Agentic x-axis latency basis is fixed at p90 (no Latency Percentile control). + const selectedPercentile = 'p90'; const selectedE2eXAxisMetric = resolveE2eXAxisMetric( requestedE2eXAxisMetric, selectedXAxisMode, effectiveSequence, selectedPercentile, + fixedSequenceStatistic, ); const [scaleType, setScaleType] = useState<'auto' | 'linear' | 'log'>( () => (getUrlParam('i_scale') as 'auto' | 'linear' | 'log') || 'auto', ); + // Comparison series on a gated power metric (`i_pcompare`). Kept while the + // metric changes: a key without a common axis simply yields no siblings, and + // the Measured controls say so, so a link's intent survives a detour. + const [powerCompare, setPowerCompare] = useState(() => + parsePowerCompare(getUrlParam('i_pcompare')), + ); // ── Quick filters (vendor / framework / deployment / mtp-stp / power tier) ── // Coarse pre-filters applied to the point set. Empty = no constraint. @@ -531,8 +555,9 @@ export function InferenceProvider({ const [quickFilterDeployment, setQuickFilterDeployment] = useState([]); const [quickFilterSpec, setQuickFilterSpec] = useState([]); const [quickFilterPower, setQuickFilterPower] = useState([]); + const [quickFilterTopologies, setQuickFilterTopologies] = useState([]); useEffect(() => { - const parse = (key: 'i_vendor' | 'i_fw' | 'i_disagg' | 'i_spec' | 'i_power') => { + const parse = (key: 'i_vendor' | 'i_fw' | 'i_disagg' | 'i_spec' | 'i_power' | 'i_topology') => { const v = getUrlParam(key); return v ? v.split(',').filter(Boolean) : []; }; @@ -543,11 +568,13 @@ export function InferenceProvider({ const deployment = parseDeploymentModes(parse('i_disagg')); const spec = parse('i_spec') as SpecMode[]; const power = parsePowerTiers(parse('i_power')); + const topologies = parse('i_topology'); if (vendors.length > 0) setQuickFilterVendors(vendors); if (frameworks.length > 0) setQuickFilterFrameworks(frameworks); if (deployment.length > 0) setQuickFilterDeployment(deployment); if (spec.length > 0) setQuickFilterSpec(spec); if (power.length > 0) setQuickFilterPower(power); + if (topologies.length > 0) setQuickFilterTopologies(topologies); }, [getUrlParam, setQuickFilterFrameworks]); const quickFilters = useMemo( () => ({ @@ -556,6 +583,7 @@ export function InferenceProvider({ deployment: quickFilterDeployment, spec: quickFilterSpec, power: quickFilterPower, + topologies: quickFilterTopologies, }), [ quickFilterVendors, @@ -563,6 +591,7 @@ export function InferenceProvider({ quickFilterDeployment, quickFilterSpec, quickFilterPower, + quickFilterTopologies, ], ); // Historical Trends hides Quick Filters, so never apply invisible selections there. @@ -790,6 +819,8 @@ export function InferenceProvider({ !isUnofficialRun && !hasExplicitRunSelection && selectedRunDateRev === 0, + powerCompare, + fixedSequenceStatistic, ); // For GPU comparison date picker — use shared availability data from global filters @@ -1063,6 +1094,21 @@ export function InferenceProvider({ const refreshing = !availabilityError && chartDataRefreshing; const error = availabilityError || workflowError || chartDataError; + // ── Perf rulers (persisted chart) ──────────────────────────────────────── + // The axis identity follows the graph ChartDisplay renders as `chart-0` + // (picked by x mode, like `bestHwTypes` below), so an x-mode switch that + // swaps the rendered chart or its x units clears the rulers the same way + // the chart's own `usePerfRulerAxisReset` does for local state. + const perfRulerStore = usePerfRulerStoreValue( + PERSISTED_PERF_RULER_CHART_ID, + getUrlParam('i_rulers'), + persistedPerfRulerAxisKey(graphs, selectedXAxisMode, selectedYAxisMetric), + ); + const iRulersStr = useMemo( + () => serializePerfRulers(perfRulerStore.state), + [perfRulerStore.state], + ); + // ── Toggle sets ─────────────────────────────────────────────────────────── const { @@ -1636,7 +1682,7 @@ export function InferenceProvider({ { i_metric: selectedYAxisMetric, i_revenue: usesTokenSalePricing(selectedYAxisMetric) ? tokenRevenuePriceSource : 'normalized', - i_pctl: selectedPercentile, + i_mstat: fixedSequenceStatistic, i_gpus: selectedGPUs.join(','), i_dates: selectedDates.join(','), i_dstart: selectedDateRange.startDate, @@ -1668,6 +1714,9 @@ export function InferenceProvider({ i_disagg: quickFilterDeployment.join(','), i_spec: quickFilterSpec.join(','), i_power: quickFilterPower.join(','), + i_topology: quickFilterTopologies.join(','), + i_rulers: iRulersStr, + i_pcompare: powerCompare === 'none' ? '' : powerCompare, }, [ selectedYAxisMetric, @@ -1675,6 +1724,7 @@ export function InferenceProvider({ selectedXAxisMetric, selectedE2eXAxisMetric, selectedXAxisMode, + fixedSequenceStatistic, scaleType, selectedGPUs, selectedDates, @@ -1699,6 +1749,9 @@ export function InferenceProvider({ quickFilterDeployment, quickFilterSpec, quickFilterPower, + quickFilterTopologies, + iRulersStr, + powerCompare, ], ); @@ -1913,7 +1966,9 @@ export function InferenceProvider({ selectedXAxisMetric, selectedE2eXAxisMetric, selectedXAxisMode, + fixedSequenceStatistic, scaleType, + powerCompare, isLegendExpanded, hideNonOptimal, showAllMeasurements, @@ -1938,7 +1993,9 @@ export function InferenceProvider({ selectedXAxisMetric, selectedE2eXAxisMetric, selectedXAxisMode, + fixedSequenceStatistic, scaleType, + powerCompare, isLegendExpanded, hideNonOptimal, showAllMeasurements, @@ -1969,15 +2026,17 @@ export function InferenceProvider({ setSelectedPrecisions: setSelectedPrecisionsAndClear, setSelectedYAxisMetric: setSelectedYAxisMetricAndClear, setTokenRevenuePriceSource, - setSelectedPercentile, + setFixedSequenceStatistic, setSelectedXAxisMetric, setSelectedXAxisMode: handleSetXAxisMode, setScaleType, + setPowerCompare, setQuickFilterVendors, setQuickFilterFrameworks, setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, setIsLegendExpanded, setHideNonOptimal, setShowAllMeasurements, @@ -2009,7 +2068,9 @@ export function InferenceProvider({ display={displayValue} actions={actionsValue} > - {children} + + {children} + [ { value: 'point', label: t.perPoint, testId: 'detail-view-point' }, { value: 'timeline', label: t.requestTimeline, testId: 'detail-view-timeline' }, + { value: 'power', label: t.powerX, testId: 'detail-view-power' }, { value: 'aggregates', label: t.aggregatesAcrossConfigs, testId: 'detail-view-aggregates' }, { value: 'logs', label: t.logs, testId: 'detail-view-logs' }, ], @@ -347,6 +351,8 @@ export function AgenticPointDetail({ id }: Props) { {view === 'logs' ? ( + ) : view === 'power' ? ( + ) : view === 'aggregates' ? ( aggregatesQuery.isError ? ( { t: number; value: number }[]; +} + +const percent = (points: readonly { t: number; value: number }[]) => + points.map((p) => ({ t: p.t, value: p.value * 100 })); + +/** + * Menu of overlay candidates, in display order. Only sources whose series is + * non-empty for the point are offered (see `availableOverlaySources`). + */ +export const OVERLAY_SOURCES: readonly OverlaySource[] = [ + { + key: 'decodeTps', + label: { en: 'Decode throughput', zh: 'Decode 吞吐量' }, + unit: 'tok/s', + color: '#8b5cf6', + points: (m) => m.decodeTps, + }, + { + key: 'prefillTps', + label: { en: 'Prefill throughput', zh: 'Prefill 吞吐量' }, + unit: 'tok/s', + color: '#06b6d4', + points: (m) => m.prefillTps, + }, + { + key: 'kvCacheUsage', + label: { en: 'KV cache utilization', zh: 'KV cache 利用率' }, + unit: '%', + color: '#f59e0b', + points: (m) => percent(m.kvCacheUsage), + }, + { + key: 'hostKvCacheUsage', + label: { en: 'Host KV cache utilization', zh: '主机 KV cache 利用率' }, + unit: '%', + color: '#d97706', + points: (m) => percent(m.hostKvCacheUsage), + }, + { + key: 'prefixCacheHitRate', + label: { en: 'Prefix cache hit rate', zh: 'Prefix cache 命中率' }, + unit: '%', + color: '#10b981', + points: (m) => percent(m.prefixCacheHitRate), + }, + { + key: 'prefixCacheHitsTps', + label: { en: 'Prefix cache hits', zh: 'Prefix cache 命中量' }, + unit: 'tok/s', + color: '#14b8a6', + points: (m) => m.prefixCacheHitsTps, + }, + { + key: 'queueDepth', + label: { en: 'Queue depth (running + waiting)', zh: '队列深度(运行中 + 等待中)' }, + unit: 'req', + color: '#ec4899', + points: (m) => m.queueDepth.map((p) => ({ t: p.t, value: p.total })), + }, +]; + +/** Sources that have at least one sample for this point, in menu order. */ +export function availableOverlaySources( + metrics: TraceServerMetrics | null | undefined, +): OverlaySource[] { + if (!metrics) return []; + return OVERLAY_SOURCES.filter((source) => source.points(metrics).length > 0); +} diff --git a/packages/app/src/components/inference/agentic-point/power-telemetry-view.tsx b/packages/app/src/components/inference/agentic-point/power-telemetry-view.tsx new file mode 100644 index 000000000..f850c431f --- /dev/null +++ b/packages/app/src/components/inference/agentic-point/power-telemetry-view.tsx @@ -0,0 +1,448 @@ +'use client'; + +import { useMemo, useState } from 'react'; + +import GpuMetricsChart, { + GPU_COLORS, + type TelemetryOverlaySeries, +} from '@/components/gpu-power/GpuPowerChart'; +import GpuStatsTable from '@/components/gpu-power/GpuStatsTable'; +import { TelemetryDisplayControls } from '@/components/gpu-power/TelemetryDisplayControls'; +import { + DEFAULT_TELEMETRY_DISPLAY, + toAbsoluteMs, + type TelemetryDisplayState, +} from '@/components/gpu-power/telemetry-smoothing'; +import { + type GpuMetricKey, + type GpuMetricRow, + ALL_METRIC_OPTIONS, + getAvailableMetrics, + getGpuMetricLabel, +} from '@/components/gpu-power/types'; +import { Card } from '@/components/ui/card'; +import ChartLegend from '@/components/ui/chart-legend'; +import { Label } from '@/components/ui/label'; +import { RetryableQueryError } from '@/components/ui/retryable-query-error'; +import { + Select, + SelectContent, + SelectItem, + SelectTrigger, + SelectValue, +} from '@/components/ui/select'; +import { useGpuMetricsPoint, type GpuMetricSeries } from '@/hooks/api/use-gpu-metrics-point'; +import { useTraceServerMetrics } from '@/hooks/api/use-trace-server-metrics'; + +import { availableOverlaySources } from './overlay-sources'; +import { track } from '@/lib/analytics'; +import { useLocale } from '@/lib/use-locale'; + +const STRINGS = { + en: { + loading: 'Loading PowerX telemetry…', + error: 'Failed to load PowerX telemetry.', + missing: + 'No PowerX telemetry is stored for benchmark point #{id}. Telemetry may not have been collected or ingested.', + series: 'Telemetry series', + metric: 'Metric', + vendor: 'Collector', + samples: 'Samples', + chips: 'Chips', + interval: 'Sample interval', + window: 'Recorded window', + sharedNote: + 'This series covers the whole benchmark job, including server start-up and warm-up, so summary rows below span more than the measured serving window.', + perGpuStats: 'Per-chip statistics', + chip: 'Chip', + secondsUnit: 's', + resetFilter: 'Show all chips', + overlayToggle: 'Overlay server metric', + overlayNone: 'None', + overlayLoading: 'Loading server metrics…', + overlayError: 'Server metrics failed to load; overlays are unavailable.', + overlayUnavailable: 'This point has no server-metric series to overlay.', + overlayAligned: 'The overlay is aligned to the telemetry by wall-clock timestamps.', + overlayRelative: + 'The trace has no wall-clock timestamps, so the overlay and the telemetry are both aligned at their own t=0.', + }, + zh: { + loading: '正在加载 PowerX 遥测数据……', + error: 'PowerX 遥测数据加载失败。', + missing: '基准测试数据点 #{id} 没有存储的 PowerX 遥测数据。遥测数据可能尚未采集或入库。', + series: '遥测序列', + metric: '指标', + vendor: '采集器', + samples: '样本数', + chips: '芯片数', + interval: '采样间隔', + window: '记录时间窗口', + sharedNote: + '该序列覆盖整个基准测试任务,包括服务启动与 warmup 阶段,因此下方统计范围大于实际测量的服务窗口。', + perGpuStats: '单芯片统计信息', + chip: '芯片', + secondsUnit: '秒', + resetFilter: '显示全部芯片', + overlayToggle: '叠加服务端指标', + overlayNone: '无', + overlayLoading: '正在加载服务端指标……', + overlayError: '服务端指标加载失败,无法叠加显示。', + overlayUnavailable: '该数据点没有可叠加的服务端指标序列。', + overlayAligned: '叠加曲线已按绝对时间戳与遥测数据对齐。', + overlayRelative: 'trace 缺少绝对时间戳,因此叠加曲线与遥测数据均从各自的 t=0 开始对齐。', + }, +} as const; + +const VENDOR_LABEL: Record = { nvidia: 'nvidia-smi', amd: 'amd-smi' }; + +/** + * Single-node CSVs come from the vendor CLI; multinode power bundles record + * their own producer (e.g. `srt-slurm.dcgm-power`) in the context sidecar. + */ +export function collectorLabel(series: Pick): string { + const context = series.sidecars?.context; + const producer = + context && typeof context === 'object' ? (context as { producer?: unknown }).producer : null; + if (typeof producer === 'string' && producer.trim() !== '') return producer; + return VENDOR_LABEL[series.vendor] ?? series.vendor; +} + +interface Props { + id: number; + enabled: boolean; + /** The point's hardware key, for the TDP reference line. */ + hardware?: string; + /** Fixed-sequence points do not have AgentX server-metric overlays. */ + serverMetricsEnabled?: boolean; +} + +function seriesLabel(series: GpuMetricSeries, total: number): string { + return total > 1 ? `${series.artifactName} · ${series.fileName}` : series.artifactName; +} + +/** + * PowerX tab of the per-point detail page: the full-resolution chip telemetry + * recorded while this benchmark point ran, read from the ingest-time digest + * (migration 016) rather than from GitHub artifacts. + */ +export function PowerTelemetryView({ id, enabled, hardware, serverMetricsEnabled = true }: Props) { + const locale = useLocale(); + const t = STRINGS[locale]; + const query = useGpuMetricsPoint(id, enabled); + const seriesList = query.data?.series ?? []; + + const [seriesSelection, setSeriesSelection] = useState<{ id: number; seriesId: number } | null>( + null, + ); + const selectedSeries = + (seriesSelection?.id === id + ? seriesList.find((series) => series.id === seriesSelection.seriesId) + : undefined) ?? seriesList[0]; + const data: GpuMetricRow[] = useMemo(() => selectedSeries?.data ?? [], [selectedSeries]); + const availableMetrics = useMemo(() => getAvailableMetrics(data), [data]); + + const [metricSelection, setMetricSelection] = useState('power'); + const metricKey: GpuMetricKey = availableMetrics.some((m) => m.key === metricSelection) + ? metricSelection + : 'power'; + const metricConfig = ALL_METRIC_OPTIONS.find((m) => m.key === metricKey)!; + const allGpuIndices = useMemo( + () => [...new Set(data.map((row) => row.index))].toSorted((a, b) => a - b), + [data], + ); + // Hidden chips are scoped to the series they were hidden on so switching + // series never carries over a stale filter. + const [hiddenSelection, setHiddenSelection] = useState<{ + seriesId: number; + hidden: number[]; + } | null>(null); + const hiddenGpus = useMemo( + () => + new Set( + hiddenSelection && hiddenSelection.seriesId === selectedSeries?.id + ? hiddenSelection.hidden + : [], + ), + [hiddenSelection, selectedSeries?.id], + ); + const visibleGpus = useMemo( + () => new Set(allGpuIndices.filter((gpuIndex) => !hiddenGpus.has(gpuIndex))), + [allGpuIndices, hiddenGpus], + ); + const toggleGpu = (gpuIndex: number) => { + if (!selectedSeries) return; + track('inference_agentic_power_gpu_toggled', { id, gpuIndex }); + const next = new Set(hiddenGpus); + if (next.has(gpuIndex)) next.delete(gpuIndex); + else next.add(gpuIndex); + setHiddenSelection({ seriesId: selectedSeries.id, hidden: [...next] }); + }; + const [isLegendExpanded, setIsLegendExpanded] = useState(true); + const [display, setDisplay] = useState(DEFAULT_TELEMETRY_DISPLAY); + + // Server-metric overlay. The series are fetched as soon as the tab opens so + // the menu can list exactly the metrics this point has; one source at a time. + const metricsQuery = useTraceServerMetrics(id, enabled && serverMetricsEnabled); + const serverMetrics = metricsQuery.data; + const overlaySources = useMemo(() => availableOverlaySources(serverMetrics), [serverMetrics]); + const [overlaySelection, setOverlaySelection] = useState<{ id: number; key: string } | null>( + null, + ); + const overlayKey = overlaySelection?.id === id ? overlaySelection.key : 'none'; + const overlaySource = overlaySources.find((source) => source.key === overlayKey) ?? null; + // Trace timeslices carry epoch-ns starts, so both series can share wall-clock + // time. A zero startNs means the trace only has relative time. + const overlayAbsolute = Boolean(serverMetrics && serverMetrics.startNs > 0); + const overlay = useMemo(() => { + if (!overlaySource || !serverMetrics || !selectedSeries) return null; + const originMs = overlayAbsolute + ? serverMetrics.startNs / 1e6 + : new Date(selectedSeries.startedAt).getTime(); + return { + label: overlaySource.label[locale], + unit: overlaySource.unit, + color: overlaySource.color, + points: toAbsoluteMs(overlaySource.points(serverMetrics), originMs), + }; + }, [overlaySource, serverMetrics, selectedSeries, overlayAbsolute, locale]); + const overlayNote = ((): string | null => { + if (metricsQuery.isLoading) return t.overlayLoading; + if (metricsQuery.isError) return t.overlayError; + if (overlaySources.length === 0) return t.overlayUnavailable; + if (!overlay) return null; + return overlayAbsolute ? t.overlayAligned : t.overlayRelative; + })(); + + if (!enabled) return null; + + if (query.isLoading) { + return ( +
+ {t.loading} +
+ ); + } + if (query.isError && !query.data) { + return ( + + ); + } + if (!selectedSeries) { + return ( +
+ {t.missing.replace('{id}', String(id))} +
+ ); + } + + const durationS = Math.max( + 0, + (new Date(selectedSeries.endedAt).getTime() - new Date(selectedSeries.startedAt).getTime()) / + 1000, + ); + const numberLocale = locale === 'zh' ? 'zh-CN' : undefined; + + return ( +
+ +
+
+
{t.vendor}
+
{collectorLabel(selectedSeries)}
+
+
+
{t.samples}
+
+ {selectedSeries.sampleCount.toLocaleString(numberLocale)} +
+
+
+
{t.chips}
+
{selectedSeries.gpuCount}
+
+
+
{t.interval}
+
+ {selectedSeries.sampleIntervalS === null + ? '—' + : `${selectedSeries.sampleIntervalS.toFixed(2)} ${t.secondsUnit}`} +
+
+
+
{t.window}
+
+ {new Date(selectedSeries.startedAt).toLocaleTimeString(numberLocale)} ·{' '} + {Math.round(durationS).toLocaleString(numberLocale)} {t.secondsUnit} +
+
+
+
+ {seriesList.length > 1 && ( +
+ + +
+ )} +
+ + +
+
+ + {serverMetricsEnabled && ( +
+
+ + +
+ {overlayNote && ( + + {overlayNote} + + )} +
+ )} +
+ + + ({ + name: `${t.chip} ${gpuIndex}`, + hw: String(gpuIndex), + label: `${t.chip} ${gpuIndex}`, + color: GPU_COLORS[gpuIndex % GPU_COLORS.length], + isActive: visibleGpus.has(gpuIndex), + onClick: () => toggleGpu(gpuIndex), + }))} + onItemRemove={(hw) => { + const gpuIndex = Number(hw); + if (visibleGpus.has(gpuIndex)) toggleGpu(gpuIndex); + }} + isLegendExpanded={isLegendExpanded} + onExpandedChange={(expanded) => { + setIsLegendExpanded(expanded); + track('inference_agentic_power_legend_expanded', { id, expanded }); + }} + actions={ + hiddenGpus.size === 0 + ? [] + : [ + { + id: 'power-telemetry-show-all-chips', + label: t.resetFilter, + onClick: () => { + track('inference_agentic_power_gpu_reset_filter', { id }); + setHiddenSelection(null); + }, + }, + ] + } + /> + } + caption={ + + {getGpuMetricLabel(metricConfig, locale)} · {t.sharedNote} + + } + /> + + + +

{t.perGpuStats}

+ +
+
+ ); +} diff --git a/packages/app/src/components/inference/agentic-point/use-detail-view.ts b/packages/app/src/components/inference/agentic-point/use-detail-view.ts index 1a97a85cc..cbdd7a43d 100644 --- a/packages/app/src/components/inference/agentic-point/use-detail-view.ts +++ b/packages/app/src/components/inference/agentic-point/use-detail-view.ts @@ -6,10 +6,11 @@ import { useClientSearchParams } from '@/hooks/useClientSearch'; import { track } from '@/lib/analytics'; import { replaceClientSearch } from '@/lib/client-navigation'; -export type DetailView = 'point' | 'timeline' | 'aggregates' | 'logs'; +const DETAIL_VIEWS = ['point', 'timeline', 'power', 'aggregates', 'logs'] as const; +export type DetailView = (typeof DETAIL_VIEWS)[number]; const isDetailView = (value: string | null): value is DetailView => - value === 'point' || value === 'timeline' || value === 'aggregates' || value === 'logs'; + (DETAIL_VIEWS as readonly string[]).includes(value ?? ''); /** URL-persisted detail view (`?view=`; per-point is the unadorned default). */ export function useDetailView(): [DetailView, (nextView: DetailView) => void] { diff --git a/packages/app/src/components/inference/axis-metric-explanations.ts b/packages/app/src/components/inference/axis-metric-explanations.ts index 220703207..0307fc094 100644 --- a/packages/app/src/components/inference/axis-metric-explanations.ts +++ b/packages/app/src/components/inference/axis-metric-explanations.ts @@ -187,11 +187,11 @@ function provisionedJoules(tokenType: TokenType): MetricExplanation { /** Validation-status note appended to every Measured Energy explanation. */ const MEASURED_TIER_NOTE_EN = ' Validated points passed the current PowerX telemetry checks. Historical points are real ' + - "older measurements but lack the information needed to confirm today's method; a dotted ring " + - 'marks them. Filter either status under Quick Filters → Measured Power.'; + "older measurements but lack the information needed to confirm today's method. Filter either " + + 'status under Quick Filters → Measured Power.'; const MEASURED_TIER_NOTE_ZH = - '已验证数据点通过了当前 PowerX 遥测检查。历史数据点来自真实的旧版测量,但缺少按当前方法完成验证所需的信息;' + - '图表以虚线圆环标记这类数据点。可在快捷筛选的“实测功耗”中按测量状态筛选。'; + '已验证数据点通过了当前 PowerX 遥测检查。历史数据点来自真实的旧版测量,但缺少按当前方法完成验证所需的信息。' + + '可在快捷筛选的“实测功耗”中按测量状态筛选。'; type MeasuredPhase = 'run' | 'prefill' | 'decode'; @@ -431,6 +431,116 @@ export const METRIC_EXPLANATIONS: Record = { zh: '% TDP = 每芯片实测平均功耗(W)÷ 额定 TDP(W)× 100', }, }, + measuredPowerTimeline: { + description: { + en: + `The per-second accelerator power samples behind each measured average, drawn over ` + + `the whole benchmark job (server start, warmup, and the validated measurement window, ` + + `which is emphasized). One trace per config, mean of its GPUs by default; the rated TDP ` + + `is a dashed reference per hardware. Configs whose telemetry artifact is missing are ` + + `listed under the chart rather than estimated.${MEASURED_TIER_NOTE_EN}`, + zh: + `每个实测平均值背后的逐秒加速器功耗采样,覆盖整个基准测试任务(服务启动、warmup ` + + `以及被突出显示的有效测量窗口)。每个配置一条曲线,默认取其 GPU 的平均值;` + + `每种硬件的额定 TDP 以虚线作为参考。缺少遥测产物的配置会列在图表下方,而不会用估算值代替。${ + MEASURED_TIER_NOTE_ZH + }`, + }, + formula: { + en: 'W(t) = mean over GPUs of the sampled power draw in each one-second bucket', + zh: 'W(t) = 每个一秒时间桶内各 GPU 功耗采样值的平均', + }, + }, + gpuProvisionedWatts: { + description: { + en: + 'Rated accelerator TDP from the hardware registry, shown as a flat per-chip value so ' + + 'measured power can be read against the GPU-only provisioning boundary. It does not ' + + 'depend on the run.', + zh: + '取硬件注册表中的加速器额定 TDP,以每芯片恒定值显示,用于对照 GPU 侧的额定供电边界与实测功耗。' + + '该值与具体运行无关。', + }, + formula: { + en: 'W/GPU = rated TDP (W)', + zh: 'W/GPU = 额定 TDP(W)', + }, + }, + gpuProvisionedJPerOutputToken: { + description: { + en: + 'Energy per output token if every allocated accelerator drew exactly its rated TDP for ' + + 'the whole run. Disaggregated deployments count prefill and decode GPUs together, so ' + + 'this is the GPU-only provisioning boundary the measured J/token can be compared against.', + zh: + '假设所有已分配加速器在整个运行中恒以额定 TDP 耗电时的每输出 token 能耗。' + + '分离式部署将 prefill 与 decode GPU 一并计入,因此它是可与实测 J/token 对照的 GPU 侧额定边界。', + }, + formula: { + en: 'J/tok = rated TDP (W) × allocated GPUs ÷ total output tokens per second', + zh: 'J/tok = 额定 TDP(W)× 已分配 GPU 数 ÷ 总输出 token 吞吐(tok/s)', + }, + }, + utilityProvisionedWatts: { + description: { + en: + 'All-in provisioned power per chip from the hardware registry: the utility-side capacity ' + + 'a data center reserves for one accelerator including host, networking, cooling and ' + + 'power-conversion overheads. It is a flat value independent of the run.', + zh: + '取硬件注册表中的每芯片整体预配功耗(all-in):数据中心为单张加速器预留的电源侧容量,' + + '包含主机、网络、散热与电源转换开销。该值为恒定值,与运行无关。', + }, + formula: { + en: 'W/GPU = all-in provisioned power per GPU (kW) × 1000', + zh: 'W/GPU = 每 GPU 整体预配功耗(kW)× 1000', + }, + }, + utilityProvisionedJPerOutputToken: { + description: { + en: + 'Energy per output token at the all-in provisioned power boundary, normalized by every ' + + 'allocated accelerator. It differs from the public All-in Provisioned J per Output Token ' + + 'metric only for disaggregated runs, where that metric normalizes by decode GPUs alone.', + zh: + '在整体预配功耗边界下的每输出 token 能耗,按全部已分配加速器归一。' + + '仅在分离式运行中与公开的 All-in Provisioned J per Output Token 指标不同,后者只按 decode GPU 归一。', + }, + formula: { + en: 'J/tok = all-in provisioned power per GPU (W) × allocated GPUs ÷ total output tokens per second', + zh: 'J/tok = 每 GPU 整体预配功耗(W)× 已分配 GPU 数 ÷ 总输出 token 吞吐(tok/s)', + }, + }, + utilityModeledWatts: { + description: { + en: + 'Modeled facility power per allocated accelerator: measured GPU power is scaled to chassis ' + + 'AC by the system power model and then multiplied once by PUE. Only hardware with a known ' + + 'eight-GPU chassis profile on 8k1k runs is supported; NVL72 systems show no value.', + zh: + '每已分配加速器的整体实测功耗:先由系统功耗模型将 GPU 实测功耗换算为机箱交流功耗,再乘以一次 PUE。' + + '仅支持在 8k1k 运行中具有已知八卡机箱模型的硬件;NVL72 系统不显示数值。', + }, + formula: { + en: 'W/GPU = modeled chassis AC power (W) × PUE ÷ allocated GPUs', + zh: 'W/GPU = 机箱交流建模功耗(W)× PUE ÷ 已分配 GPU 数', + }, + }, + utilityModeledJPerOutputToken: { + description: { + en: + 'Measured energy per output token scaled to the modeled facility boundary, so its ratio to ' + + 'measured GPU energy equals the ratio of modeled facility power to measured GPU power. ' + + 'Missing where the system power model or validated measured power is unavailable.', + zh: + '将实测每输出 token 能耗按整体实测功耗边界缩放,其与 GPU 实测能耗之比等于整体实测功耗与 GPU 实测功耗之比。' + + '系统功耗模型或通过验证的实测功耗缺失时不显示。', + }, + formula: { + en: 'J/tok = measured J per output token × modeled facility W per GPU ÷ measured W per GPU', + zh: 'J/tok = 实测每输出 token 能耗 × 每 GPU 整体实测功耗(W)÷ 每 GPU 实测功耗(W)', + }, + }, }; /** @@ -438,9 +548,15 @@ export const METRIC_EXPLANATIONS: Record = { * the percentile prefix. Mirrors the branch logic in `resolveXAxisField` plus * the derived agentic x-axis mode handled in `ChartDisplay`. */ -export type XAxisKind = 'interactivity' | 'e2eLatency' | 'ttft' | 'e2eNormalizedInteractivity'; +export type XAxisKind = + | 'concurrency' + | 'interactivity' + | 'e2eLatency' + | 'ttft' + | 'e2eNormalizedInteractivity'; export const X_AXIS_KINDS: readonly XAxisKind[] = [ + 'concurrency', 'interactivity', 'e2eLatency', 'ttft', @@ -457,11 +573,18 @@ export interface XAxisExplanation { } const zhPctl = (pctl: string | null): string => - pctl === null ? '' : pctl === 'Median' ? '中位' : `${pctl} `; + pctl === null ? '' : pctl === 'Median' ? '中位' : pctl === 'Mean' ? '平均' : `${pctl} `; const enPctl = (pctl: string | null): string => (pctl === null ? '' : `${pctl} `); export const X_AXIS_EXPLANATIONS: Record = { + concurrency: { + name: { en: () => 'Concurrency', zh: () => '并发数' }, + description: { + en: 'The configured number of concurrent requests in each observed benchmark. This is a load setting, not a higher-is-better score. All observed load points are retained; lines only connect the same serving topology and run.', + zh: '每个实测基准配置的并发请求数。这是负载设置,不是越高越好的性能分数。保留全部实测负载点,连线仅连接同一服务拓扑、同一次运行的数据。', + }, + }, interactivity: { name: { en: (pctl) => `${enPctl(pctl)}Interactivity (tok/s/user)`, @@ -471,7 +594,8 @@ export const X_AXIS_EXPLANATIONS: Record = { en: 'Interactivity is the rate at which a single user receives generated tokens while the ' + 'model streams its answer — how quickly new words appear on screen. Higher values feel ' + - 'snappier; operators trade it against batch throughput.', + 'snappier; operators trade it against batch throughput. For fixed-sequence Mean, the rate ' + + 'is 1 divided by mean TPOT in seconds, not the arithmetic mean of per-request rates.', zh: '交互性(interactivity)指模型流式输出回答时,单个用户接收生成 token 的速率——' + '即新内容出现在屏幕上的快慢。数值越高体验越流畅;运营方需要在交互性与批量吞吐量之间权衡。', @@ -545,6 +669,7 @@ export function resolveXAxisKind( isDerivedNormalizedInteractivity: boolean; }, ): XAxisKind { + if (opts.xAxisField === 'conc') return 'concurrency'; if (opts.isDerivedNormalizedInteractivity) return 'e2eNormalizedInteractivity'; if (opts.xAxisField.endsWith('ttft')) return 'ttft'; return chartType === 'e2e' ? 'e2eLatency' : 'interactivity'; @@ -553,9 +678,11 @@ export function resolveXAxisKind( /** * Extract the percentile word from a resolved x-axis label (e.g. * "P90 Time To First Token (s)" → "P90"). The chart pipelines always render - * the percentile prefix in this English form, including on /zh pages. + * percentiles in English; fixed-sequence statistics are localized on /zh. */ export function xAxisPercentileFromLabel(xAxisLabel: string): string | null { + if (xAxisLabel.startsWith('平均')) return 'Mean'; + if (xAxisLabel.startsWith('中位')) return 'Median'; const match = /^(?Median|Mean|P\d+(?:\.\d+)?)\s/iu.exec(xAxisLabel); if (!match?.groups?.pctl) return null; const pctl = match.groups.pctl; diff --git a/packages/app/src/components/inference/hooks/chart-data-core.ts b/packages/app/src/components/inference/hooks/chart-data-core.ts index 0851b8040..45d7499aa 100644 --- a/packages/app/src/components/inference/hooks/chart-data-core.ts +++ b/packages/app/src/components/inference/hooks/chart-data-core.ts @@ -12,9 +12,15 @@ import { supportsChartTokenMetric, type TokenMetricType } from '@/lib/supplement * Chart x-axis variant selected by the dropdown in the Chart panel. The * inference provider and ChartDisplay import this single definition. */ -export type XAxisMode = 'ttft' | 'e2e' | 'interactivity' | 'e2e-normalized-interactivity'; +export type XAxisMode = + | 'ttft' + | 'e2e' + | 'interactivity' + | 'e2e-normalized-interactivity' + | 'concurrency'; export const X_AXIS_MODES: readonly XAxisMode[] = [ + 'concurrency', 'ttft', 'e2e', 'interactivity', @@ -115,7 +121,7 @@ const X_LABEL_STAT_PREFIX_RE = /^(?:Median|Mean|P75|P90|P95|P99(?:\.9)?)\b\s*/iu * existing leading statistic word (e.g. the TTFT override's "P90 Time To * First Token (s)") or prefixes the percentile when the configured label has * none (e.g. "Interactivity (tok/s/user)" → "P90 Interactivity (tok/s/user)"). - * Only call for agentic sequences — fixed-seq labels must stay untouched. + * Also used for the explicit Mean/Median labels on fixed-sequence service axes. */ export function applyAgenticPercentileToXLabel(label: string, pctlWord: string): string { return X_LABEL_STAT_PREFIX_RE.test(label) diff --git a/packages/app/src/components/inference/hooks/useChartData.ts b/packages/app/src/components/inference/hooks/useChartData.ts index 9f6498b5a..3401a6002 100644 --- a/packages/app/src/components/inference/hooks/useChartData.ts +++ b/packages/app/src/components/inference/hooks/useChartData.ts @@ -27,19 +27,24 @@ import type { ChartDefinition, HardwareConfig, InferenceData, + PowerCompare, RenderableGraph, TokenRevenuePriceSource, TokenRevenuePricing, YAxisMetricKey, } from '@/components/inference/types'; import { partitionChartDataByLimits } from '@/components/inference/utils'; +import { expandPowerCompareSeries } from '@/components/inference/utils/power-compare'; import { parseComparisonEntry } from '@/components/inference/utils/comparisonEntry'; import { computeAvailableQuickFilters, EMPTY_QUICK_FILTERS, type QuickFilters, } from '@/components/inference/utils/quickFilters'; -import { resolveXAxisField } from '@/components/inference/utils/resolveXAxisField'; +import { + resolveXAxisField, + type FixedSequenceStatistic, +} from '@/components/inference/utils/resolveXAxisField'; import { benchmarkQueryOptions, useBenchmarks } from '@/hooks/api/use-benchmarks'; import type { BenchmarkRow } from '@/lib/api'; import { benchmarkCurveDate, dedupeAgenticHistoryRuns } from '@/lib/benchmark-run-selection'; @@ -116,6 +121,9 @@ export function useChartData( tcoBasis: TcoBasis = DEFAULT_TCO_BASIS, /** Opt-in from the inference page only; explicit date/run/history views opt out. */ allowDefaultRunPreference = false, + /** Sibling boundary / role series appended to a gated power metric (`i_pcompare`). */ + powerCompare: PowerCompare = 'none', + fixedSequenceStatistic: FixedSequenceStatistic = 'median', ) { // When the selected date is the latest available, use '' (empty string) to match // the initial no-date query key, reusing the eagerly-fetched benchmarks from the @@ -354,19 +362,24 @@ export function useChartData( isAgentic, percentile: selectedPercentile, xAxisMode: selectedXAxisMode, + fixedSequenceStatistic, }); const naturalX = resolved.naturalX as keyof AggDataEntry; const xAxisField = resolved.xAxisField as keyof AggDataEntry; const { isTtftOverride } = resolved; const ttftPctl = isTtftOverride ? xAxisField.replace(/_ttft$/u, '') : 'p90'; - const ttftPctlWord = ttftPctl === 'median' ? 'Median' : ttftPctl.toUpperCase(); + const ttftPctlWord = + ttftPctl === 'median' ? 'Median' : ttftPctl === 'mean' ? 'Mean' : ttftPctl.toUpperCase(); const ttftLabel = `${ttftPctlWord} Time To First Token (s)`; const ttftLabelZh = `${ttftPctlWord} 首 token 延迟 (s)`; let xAxisLabel = chartDef.x_label; let xAxisLabelZh = chartDef.x_labelZh; - if (resolved.branch === 'user-input-override') { + if (resolved.branch === 'concurrency') { + xAxisLabel = 'Concurrency'; + xAxisLabelZh = '并发数'; + } else if (resolved.branch === 'user-input-override') { const labelKey = `${selectedYAxisMetric}_x_label` as keyof ChartDefinition; const labelZhKey = `${selectedYAxisMetric}_x_labelZh` as keyof ChartDefinition; if (effectiveXMetric === chartDef[`${selectedYAxisMetric}_x` as keyof ChartDefinition]) { @@ -398,7 +411,9 @@ export function useChartData( : selectedXAxisMode === undefined ? (chartDef[headingKey] as string) || chartDef.heading : chartDef.heading; - if (isAgentic) { + if (resolved.branch === 'concurrency') { + chartHeading = 'vs. Concurrency'; + } else if (isAgentic) { const pctlWord = selectedPercentile.toUpperCase(); xAxisLabel = applyAgenticPercentileToXLabel(xAxisLabel, pctlWord); xAxisLabelZh = applyAgenticPercentileToXLabel(xAxisLabelZh, pctlWord); @@ -406,6 +421,22 @@ export function useChartData( /^(?vs\.\s+)(?:(?:Median|Mean|P75|P90|P95|P99(?:\.9)?)\s+)?/iu, `$1${pctlWord} `, ); + } else { + const word = xAxisField.startsWith('mean_') + ? 'Mean' + : xAxisField.startsWith('median_') + ? 'Median' + : xAxisField.split('_')[0].toUpperCase(); + const wordZh = word === 'Mean' ? '平均' : '中位'; + xAxisLabel = applyAgenticPercentileToXLabel(xAxisLabel, word); + xAxisLabelZh = applyAgenticPercentileToXLabel(xAxisLabelZh, word).replace( + /^(?:Mean|Median)\s+/u, + wordZh, + ); + chartHeading = chartHeading.replace( + /^(?vs\.\s+)(?:(?:Median|Mean|P75|P90|P95|P99(?:\.9)?)\s+)?/iu, + `$1${word} `, + ); } // The x-axis is "flipped" only when the good-direction reverses @@ -418,11 +449,13 @@ export function useChartData( xAxisField !== naturalX && !(chartDef.chartType === 'e2e' && isTtftOverride); const rooflineOverrides: Partial = {}; - if (xAxisFlipped) { + if (xAxisFlipped || resolved.branch === 'concurrency') { for (const key of Object.keys(chartDef) as (keyof ChartDefinition)[]) { if (typeof key === 'string' && key.endsWith('_roofline')) { const dir = chartDef[key] as string | undefined; - if (dir && dir in FLIP_MAP) { + if (resolved.branch === 'concurrency') { + (rooflineOverrides as any)[key] = undefined; + } else if (dir && dir in FLIP_MAP) { (rooflineOverrides as any)[key] = flipRooflineDirection(dir as RooflineDirection); } } @@ -485,6 +518,7 @@ export function useChartData( selectedXAxisMetric, selectedE2eXAxisMetric, selectedPercentile, + fixedSequenceStatistic, selectedSequence, tokenRevenuePriceSource, ], @@ -538,8 +572,14 @@ export function useChartData( ); const hasMetric = metricData.length > 0; const isTtftX = typeof xAxisField === 'string' && xAxisField.endsWith('_ttft'); + // Comparison clones are appended after the remap so they share the + // base point's x and differ only in y and `powerVariant`. const mappedData = hasMetric - ? metricData.map((d) => remapInferencePoint(d, metricKey, xAxisField)) + ? expandPowerCompareSeries( + metricData.map((d) => remapInferencePoint(d, metricKey, xAxisField)), + selectedYAxisMetric, + powerCompare, + ) : []; const isAgentic = selectedSequence === Sequence.AgenticTraces; @@ -576,6 +616,7 @@ export function useChartData( compareGpuPair, selectedPercentile, quickFilters, + powerCompare, ]); // Points that pass every scope filter but NOT the y-metric coverage filter. diff --git a/packages/app/src/components/inference/hooks/useChartData.x-axis-wiring.test.tsx b/packages/app/src/components/inference/hooks/useChartData.x-axis-wiring.test.tsx index 88412ce51..4318fa12c 100644 --- a/packages/app/src/components/inference/hooks/useChartData.x-axis-wiring.test.tsx +++ b/packages/app/src/components/inference/hooks/useChartData.x-axis-wiring.test.tsx @@ -7,6 +7,8 @@ import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; import type { BenchmarkRow } from '@/lib/api'; import { Model, Sequence } from '@/lib/data-mappings'; import { resolveScatterXAxisScale } from '@/components/inference/utils/x-axis-scale'; +import { processOverlayChartData } from '@/components/inference/utils'; +import { transformBenchmarkRows } from '@/lib/benchmark-transform'; import { buildReplayTimeline } from '@/components/inference/replay/buildReplayTimeline'; const mocks = vi.hoisted(() => ({ @@ -107,9 +109,11 @@ let result: ReturnType | undefined; function Probe({ mode, energy = false, + statistic = 'median', metric = energy ? 'y_measuredPrefillJPerInputToken' : 'y_inputTputPerGpu', }: { mode?: XAxisMode; + statistic?: 'mean' | 'median'; energy?: boolean; metric?: string; }) { @@ -135,6 +139,15 @@ function Probe({ undefined, undefined, mode, + undefined, + undefined, + undefined, + undefined, + undefined, + undefined, + undefined, + undefined, + statistic, ); return null; } @@ -154,6 +167,47 @@ afterEach(() => { }); describe('useChartData x-axis scale wiring', () => { + it.each([ + ['interactivity', 'interactivity', 'mean_tpot_intvty', 25, 'Mean Interactivity'], + ['ttft', 'e2e', 'mean_ttft', 8, 'Mean Time To First Token'], + ['e2e', 'e2e', 'mean_e2el', 30, 'Mean End-to-end Latency'], + ] as const)( + 'keeps official, overlay and replay mean %s numerically aligned', + (mode, chartType, field, expectedX, label) => { + mocks.rows[0].metrics = { + ...mocks.rows[0].metrics, + mean_tpot: 0.04, + mean_intvty: 777, + mean_ttft: 8, + mean_e2el: 30, + }; + mocks.rows[1].metrics = { ...mocks.rows[1].metrics, mean_intvty: 888 }; + act(() => root.render()); + const graph = result!.graphs.find((g) => g.chartDefinition.chartType === chartType)!; + expect(graph.chartDefinition.x_scale_field).toBe(field); + expect(graph.chartDefinition.x_label).toContain(label); + expect(graph.chartDefinition.heading).toContain(label); + expect(graph.data.map((point) => point.x)).toEqual([expectedX]); + expect(graph.data[0].mean_intvty).toBe(777); + const { chartData } = transformBenchmarkRows(mocks.rows); + const overlay = processOverlayChartData( + chartData[chartType === 'interactivity' ? 0 : 1], + chartType, + 'y_tpPerGpu', + null, + { + selectedXAxisMode: mode, + fixedSequenceStatistic: 'mean', + }, + ); + expect(overlay.map((point) => point.x)).toEqual([expectedX]); + const replay = buildReplayTimeline(mocks.rows, graph.chartDefinition, 'y_tpPerGpu', null, [ + 'fp8', + ]); + expect(replay.configs.map((series) => series.template.x)).toEqual([expectedX]); + }, + ); + it.each([ ['interactivity', 'interactivity', 'median_intvty'], ['ttft', 'e2e', 'median_ttft'], @@ -204,8 +258,8 @@ describe('useChartData x-axis scale wiring', () => { ]); expect(energyGraph?.chartDefinition).toMatchObject({ x_scale_field: 'median_intvty', - x_label: 'Interactivity (tok/s/user)', - heading: 'vs. Interactivity', + x_label: 'Median Interactivity (tok/s/user)', + heading: 'vs. Median Interactivity', y_measuredPrefillJPerInputToken_roofline: 'lower_right', }); if (!energyGraph) throw new Error('useChartData did not produce the interactivity graph'); @@ -252,8 +306,8 @@ describe('useChartData x-axis scale wiring', () => { expect(graph?.data.map((point) => point.x)).toEqual([12, 12]); expect(graph?.chartDefinition).toMatchObject({ x_scale_field: 'median_e2el', - x_label: 'End-to-end Latency (s)', - heading: 'vs. End-to-end Latency', + x_label: 'Median End-to-end Latency (s)', + heading: 'vs. Median End-to-end Latency', }); }); diff --git a/packages/app/src/components/inference/hooks/useComparisonSeries.ts b/packages/app/src/components/inference/hooks/useComparisonSeries.ts new file mode 100644 index 000000000..eae3c1653 --- /dev/null +++ b/packages/app/src/components/inference/hooks/useComparisonSeries.ts @@ -0,0 +1,130 @@ +'use client'; + +import { useMemo } from 'react'; +import { useTheme } from 'next-themes'; + +import { useInferenceDisplay, useInferenceFilters } from '@/components/inference/InferenceContext'; +import { + buildRunNumbering, + comparisonEntrySortValue, + resolveComparisonEntries, +} from '@/components/inference/utils/comparisonEntry'; +import { useThemeColors } from '@/hooks/useThemeColors'; +import { getModelSortIndex } from '@/lib/constants'; +import { generateGpuDateColors, generateHighContrastGpuDateColors } from '@/lib/dynamic-colors'; +import { isDarkTheme } from '@/lib/themes'; + +/** One (comparison entry, chip config) series of the date comparison view. */ +export interface ComparisonSeries { + date: string; + hwKey: string; + /** `${date}_${hwKey}`: the key of the `activeDates` toggle set. */ + id: string; + color: string; +} + +/** + * Series, run numbers and colours of the date comparison view. GPUGraph and + * the Power Timeline share them so one (date, chip config) pair reads the + * same in both displays. + */ +export function useComparisonSeries(providedRunNumbering?: Map) { + const { selectedGPUs, selectedDateRange, selectedDates } = useInferenceFilters(); + const { highContrast } = useInferenceDisplay(); + const { resolvedTheme } = useTheme(); + + // Shared date+GPU pairs. `dates` holds comparison-series entries (plain dates + // and/or specific-run entries); a same-day range endpoint is dropped when that + // date also has run entries (resolveComparisonEntries), then sorted earliest → + // latest so a day's runs read #1 → #N. + const gpuDatePairs = useMemo(() => { + const deduplicated = resolveComparisonEntries(selectedDates, selectedDateRange); + deduplicated.sort((a, b) => { + const [ta, ia] = comparisonEntrySortValue(a); + const [tb, ib] = comparisonEntrySortValue(b); + return ta - tb || ia - ib; + }); + const sortedGPUs = [...selectedGPUs].toSorted( + (a, b) => getModelSortIndex(a) - getModelSortIndex(b) || a.localeCompare(b), + ); + return { dates: deduplicated, sortedGPUs }; + }, [selectedDateRange, selectedDates, selectedGPUs]); + + // Run numbers for legend/line labels. Prefer the stable numbering passed by + // the parent (shared with the changelog, so labels match it and removed runs + // leave a gap); fall back to gap-free numbering of the on-chart series. + const runNumbering = useMemo( + () => providedRunNumbering ?? buildRunNumbering(gpuDatePairs.dates), + [providedRunNumbering, gpuDatePairs.dates], + ); + + const graphIdentifiers = useMemo(() => { + const ids: string[] = []; + gpuDatePairs.sortedGPUs.forEach((gpu) => + gpuDatePairs.dates.forEach((date) => ids.push(`${date}_${gpu}`)), + ); + return ids; + }, [gpuDatePairs]); + + // High contrast keys off the GPU (not `date_gpu`) so each hardware config + // gets exactly one hue; the dates within a config are separated by the + // lightness ramp built below rather than by unrelated hues. + const { resolveColor, getCssColor } = useThemeColors({ + highContrast, + identifiers: graphIdentifiers, + hcKeys: gpuDatePairs.sortedGPUs, + }); + + // Dynamic GPU×date color map + const gpuDateColorMap = useMemo(() => { + const { dates, sortedGPUs } = gpuDatePairs; + if (sortedGPUs.length === 0 || dates.length === 0) return {}; + const theme = isDarkTheme(resolvedTheme) ? 'dark' : 'light'; + return generateGpuDateColors(sortedGPUs, dates.length, theme); + }, [gpuDatePairs, resolvedTheme]); + + // High-contrast GPU×date color map: one iwanthue hue per GPU, ramped across + // the compared dates so a config's runs stay recognisably the same color + // while still reading oldest → newest. + const hcGpuDateColorMap = useMemo(() => { + const { dates, sortedGPUs } = gpuDatePairs; + if (!highContrast || sortedGPUs.length === 0 || dates.length === 0) return {}; + const theme = isDarkTheme(resolvedTheme) ? 'dark' : 'light'; + const baseColors: Record = {}; + for (const gpu of sortedGPUs) baseColors[gpu] = getCssColor(resolveColor(gpu)); + return generateHighContrastGpuDateColors(baseColors, dates.length, theme); + }, [gpuDatePairs, highContrast, resolvedTheme, resolveColor, getCssColor]); + + const allGraphs = useMemo(() => { + const { dates, sortedGPUs } = gpuDatePairs; + const result: ComparisonSeries[] = []; + sortedGPUs.forEach((gpu) => { + dates.forEach((date, dateIndex) => { + const id = `${date}_${gpu}`; + const compositeKey = `${dateIndex}_${gpu}`; + const dynamicColor = gpuDateColorMap[compositeKey]; + result.push({ + date, + hwKey: gpu, + id, + color: highContrast + ? hcGpuDateColorMap[compositeKey] || getCssColor(resolveColor(gpu)) + : dynamicColor || 'var(--foreground)', + }); + }); + }); + return result; + }, [gpuDatePairs, gpuDateColorMap, hcGpuDateColorMap, highContrast, resolveColor, getCssColor]); + + const paletteIdentity = useMemo( + () => + [ + resolvedTheme ?? 'system', + highContrast ? 'high-contrast' : 'standard', + ...allGraphs.map(({ id, color }) => `${id}:${color}`), + ].join('|'), + [resolvedTheme, highContrast, allGraphs], + ); + + return { gpuDatePairs, runNumbering, allGraphs, paletteIdentity, resolveColor, getCssColor }; +} diff --git a/packages/app/src/components/inference/hooks/usePowerTraceAction.ts b/packages/app/src/components/inference/hooks/usePowerTraceAction.ts new file mode 100644 index 000000000..09b559126 --- /dev/null +++ b/packages/app/src/components/inference/hooks/usePowerTraceAction.ts @@ -0,0 +1,56 @@ +'use client'; + +import { useCallback, type RefObject } from 'react'; + +import { useInferenceActions } from '@/components/inference/InferenceContext'; +import type { InferenceData } from '@/components/inference/types'; +import { + POWER_TIMELINE_METRIC_KEY, + requestPowerTraceFocus, + traceKeyForPoint, +} from '@/components/inference/utils/powerTimeline'; +import { track } from '@/lib/analytics'; +import type { D3ChartHandle } from '@/lib/d3-chart/D3Chart/types'; + +/** + * "View power trace" on a pinned tooltip (official or overlay point): the + * same-tab click stays in-page — remember which trace to emphasise, switch + * the metric to the Timeline display, and let the anchor's href keep + * serving open-in-new-tab. Listeners are attached per pin because the + * tooltip HTML is replaced on every pin. + */ +export function usePowerTraceAction(chartRef: RefObject) { + const { setSelectedYAxisMetric } = useInferenceActions(); + return useCallback( + (tooltipEl: HTMLElement, d: InferenceData, overlay: boolean) => { + const action = tooltipEl.querySelector('[data-action="view-power-trace"]'); + const traceKey = traceKeyForPoint(d); + if (!action || !traceKey) return; + action.addEventListener('click', (actionEvent) => { + actionEvent.stopPropagation(); + // Modifier / auxiliary clicks keep the anchor's own behaviour: the + // href opens this chart's timeline in a new tab or window. + const mouse = actionEvent as MouseEvent; + if ( + mouse.button !== 0 || + mouse.metaKey || + mouse.ctrlKey || + mouse.shiftKey || + mouse.altKey + ) { + return; + } + actionEvent.preventDefault(); + requestPowerTraceFocus(traceKey); + chartRef.current?.dismissTooltip(); + setSelectedYAxisMetric(POWER_TIMELINE_METRIC_KEY); + track('inference_power_trace_opened', { + hwKey: String(d.hwKey), + conc: d.conc, + overlay, + }); + }); + }, + [chartRef, setSelectedYAxisMetric], + ); +} diff --git a/packages/app/src/components/inference/measured-metric-config.test.ts b/packages/app/src/components/inference/measured-metric-config.test.ts index 47c356f22..41276d400 100644 --- a/packages/app/src/components/inference/measured-metric-config.test.ts +++ b/packages/app/src/components/inference/measured-metric-config.test.ts @@ -1,6 +1,10 @@ import { describe, expect, it } from 'vitest'; -import { MEASURED_ENERGY_METRIC_CONFIG_KEYS, METRIC_CONFIG_KEYS } from './metric-registry'; +import { + MEASURED_ENERGY_METRIC_CONFIG_KEYS, + METRIC_CONFIG_KEYS, + POWER_BASIS_METRIC_CONFIG_KEYS, +} from './metric-registry'; import { changeMeasuredMetricConfig, getMeasuredMetricConfig, @@ -8,20 +12,27 @@ import { } from './measured-metric-config'; describe('measured metric configuration', () => { - it.each(MEASURED_ENERGY_METRIC_CONFIG_KEYS)( - 'round-trips the existing share-link metric %s', - (key) => { - const config = getMeasuredMetricConfig(key); - expect(config).toBeDefined(); - expect(changeMeasuredMetricConfig(key, {})).toBe(key); - expect(changeMeasuredMetricConfig('y_tpPerGpu', config!)).toBe(key); - }, - ); + it.each([ + 'y_measuredAvgPower', + 'y_measuredPowerTimeline', + 'y_measuredJPerOutputToken', + 'y_measuredWhPerSuccessfulQuery', + ] as const)('round-trips the existing share-link metric %s', (key) => { + const config = getMeasuredMetricConfig(key); + expect(config).toBeDefined(); + expect(changeMeasuredMetricConfig(key, {})).toBe(key); + expect(changeMeasuredMetricConfig('y_tpPerGpu', config!)).toBe(key); + }); it('does not group unrelated metrics or unknown persisted values', () => { const grouped = METRIC_CONFIG_KEYS.filter((key) => getMeasuredMetricConfig(key)); - expect(grouped).toHaveLength(13); - expect(new Set(grouped)).toEqual(new Set(MEASURED_ENERGY_METRIC_CONFIG_KEYS)); + expect(grouped).toHaveLength(20); + expect(new Set(grouped)).toEqual( + new Set([...MEASURED_ENERGY_METRIC_CONFIG_KEYS, ...POWER_BASIS_METRIC_CONFIG_KEYS]), + ); + for (const key of MEASURED_ENERGY_METRIC_CONFIG_KEYS) { + expect(getMeasuredMetricConfig(key)?.basis, key).toBe('gpu-measured'); + } expect(getMeasuredMetricConfig('y_modeledChassisPowerPerGpu')).toBeUndefined(); expect(getMeasuredMetricConfig('y_removedMetric')).toBeUndefined(); expect(getMeasuredMetricConfig('')).toBeUndefined(); @@ -42,6 +53,7 @@ describe('measured metric configuration', () => { it('keeps fleet percentiles, role averages and TDP normalization distinct', () => { expect(getMeasuredMetricConfig('y_measuredP90Power')).toEqual({ family: 'power', + basis: 'gpu-measured', scope: 'all', statistic: 'p90', display: 'watts', @@ -78,6 +90,7 @@ describe('measured metric configuration', () => { ); expect(getMeasuredMetricConfig('y_measuredPrefillJPerInputToken')).toEqual({ family: 'energy', + basis: 'gpu-measured', scope: 'prefill', denominator: 'input', unit: 'joules', diff --git a/packages/app/src/components/inference/measured-metric-config.ts b/packages/app/src/components/inference/measured-metric-config.ts index 555fbef2e..ff28546be 100644 --- a/packages/app/src/components/inference/measured-metric-config.ts +++ b/packages/app/src/components/inference/measured-metric-config.ts @@ -1,17 +1,27 @@ +import type { PowerBasis } from '@/lib/power-basis'; import type { MetricConfigKey } from './metric-registry'; export type MeasuredMetricFamily = 'power' | 'energy'; type MeasuredScope = 'all' | 'prefill' | 'decode'; +/** + * How whole-deployment average power is shown: per-chip watts, percent of + * TDP, or the per-second telemetry trace behind the average (`timeline`, which + * ChartDisplay renders with `PowerTimeline` instead of the scatter chart). + */ +export type MeasuredPowerDisplay = 'watts' | 'tdp' | 'timeline'; export type MeasuredMetricConfig = | { family: 'power'; + /** Power boundary the key plots; only `gpu-measured` publishes the other dimensions. */ + basis: PowerBasis; scope: MeasuredScope; statistic: 'average' | 'p75' | 'p90'; - display: 'watts' | 'tdp'; + display: MeasuredPowerDisplay; } | { family: 'energy'; + basis: PowerBasis; scope: MeasuredScope; denominator: 'input' | 'output' | 'total' | 'query'; unit: 'joules' | 'wattHours'; @@ -19,9 +29,10 @@ export type MeasuredMetricConfig = export type MeasuredMetricConfigChange = Partial<{ family: MeasuredMetricFamily; + basis: PowerBasis; scope: MeasuredScope; statistic: 'average' | 'p75' | 'p90'; - display: 'watts' | 'tdp'; + display: MeasuredPowerDisplay; denominator: 'input' | 'output' | 'total' | 'query'; unit: 'joules' | 'wattHours'; }>; @@ -31,51 +42,78 @@ export const MEASURED_METRIC_DEFAULTS = { energy: 'y_measuredJPerOutputToken', } as const satisfies Record; +const measured = { basis: 'gpu-measured' } as const; + // Presentation settings resolve to existing metrics; they do not own chart state. const MEASURED_METRIC_CONFIGS: readonly (readonly [MetricConfigKey, MeasuredMetricConfig])[] = [ - ['y_measuredAvgPower', { family: 'power', scope: 'all', statistic: 'average', display: 'watts' }], - ['y_measuredP75Power', { family: 'power', scope: 'all', statistic: 'p75', display: 'watts' }], - ['y_measuredP90Power', { family: 'power', scope: 'all', statistic: 'p90', display: 'watts' }], + [ + 'y_measuredAvgPower', + { family: 'power', ...measured, scope: 'all', statistic: 'average', display: 'watts' }, + ], + [ + 'y_measuredP75Power', + { family: 'power', ...measured, scope: 'all', statistic: 'p75', display: 'watts' }, + ], + [ + 'y_measuredP90Power', + { family: 'power', ...measured, scope: 'all', statistic: 'p90', display: 'watts' }, + ], [ 'y_measuredPrefillAvgPower', - { family: 'power', scope: 'prefill', statistic: 'average', display: 'watts' }, + { family: 'power', ...measured, scope: 'prefill', statistic: 'average', display: 'watts' }, ], [ 'y_measuredDecodeAvgPower', - { family: 'power', scope: 'decode', statistic: 'average', display: 'watts' }, + { family: 'power', ...measured, scope: 'decode', statistic: 'average', display: 'watts' }, ], [ 'y_measuredPowerPercentTdp', - { family: 'power', scope: 'all', statistic: 'average', display: 'tdp' }, + { family: 'power', ...measured, scope: 'all', statistic: 'average', display: 'tdp' }, + ], + [ + 'y_measuredPowerTimeline', + { family: 'power', ...measured, scope: 'all', statistic: 'average', display: 'timeline' }, ], [ 'y_measuredJPerInputToken', - { family: 'energy', scope: 'all', denominator: 'input', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'all', denominator: 'input', unit: 'joules' }, ], [ 'y_measuredJPerOutputToken', - { family: 'energy', scope: 'all', denominator: 'output', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'all', denominator: 'output', unit: 'joules' }, ], [ 'y_measuredJPerTotalToken', - { family: 'energy', scope: 'all', denominator: 'total', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'all', denominator: 'total', unit: 'joules' }, ], [ 'y_measuredPrefillJPerInputToken', - { family: 'energy', scope: 'prefill', denominator: 'input', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'prefill', denominator: 'input', unit: 'joules' }, ], [ 'y_measuredDecodeJPerOutputToken', - { family: 'energy', scope: 'decode', denominator: 'output', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'decode', denominator: 'output', unit: 'joules' }, ], [ 'y_measuredJPerSuccessfulQuery', - { family: 'energy', scope: 'all', denominator: 'query', unit: 'joules' }, + { family: 'energy', ...measured, scope: 'all', denominator: 'query', unit: 'joules' }, ], [ 'y_measuredWhPerSuccessfulQuery', - { family: 'energy', scope: 'all', denominator: 'query', unit: 'wattHours' }, + { family: 'energy', ...measured, scope: 'all', denominator: 'query', unit: 'wattHours' }, ], + // Derived boundaries publish one canonical combination per family: whole + // deployment, average watts, joules per output token (lib/power-basis.ts). + ...( + [ + ['gpu-provisioned', 'y_gpuProvisionedWatts', 'y_gpuProvisionedJPerOutputToken'], + ['utility-provisioned', 'y_utilityProvisionedWatts', 'y_utilityProvisionedJPerOutputToken'], + ['utility-modeled', 'y_utilityModeledWatts', 'y_utilityModeledJPerOutputToken'], + ] as const satisfies readonly (readonly [PowerBasis, MetricConfigKey, MetricConfigKey])[] + ).flatMap(([basis, watts, energy]): (readonly [MetricConfigKey, MeasuredMetricConfig])[] => [ + [watts, { family: 'power', basis, scope: 'all', statistic: 'average', display: 'watts' }], + [energy, { family: 'energy', basis, scope: 'all', denominator: 'output', unit: 'joules' }], + ]), ]; export function getMeasuredMetricConfig(metric: string): MeasuredMetricConfig | undefined { @@ -83,6 +121,8 @@ export function getMeasuredMetricConfig(metric: string): MeasuredMetricConfig | return config ? { ...config } : undefined; } +const OTHER_DIMENSIONS = ['scope', 'statistic', 'display', 'denominator', 'unit'] as const; + export function changeMeasuredMetricConfig( metric: string, change: MeasuredMetricConfigChange, @@ -93,6 +133,21 @@ export function changeMeasuredMetricConfig( current?.family === family ? current : getMeasuredMetricConfig(MEASURED_METRIC_DEFAULTS[family])!; + // Choosing a derived boundary snaps the other dimensions to its canonical + // combination. Changing any of those dimensions while on a derived boundary + // returns to GPU-measured telemetry, the only basis that publishes variants, + // so every control change lands on a real key. A family switch keeps the + // boundary: the metric key is what carries it. + const changesOtherDimension = OTHER_DIMENSIONS.some((key) => change[key] !== undefined); + const basis = + change.basis ?? (changesOtherDimension ? 'gpu-measured' : (current ?? config).basis); + if (basis !== 'gpu-measured') { + return ( + MEASURED_METRIC_CONFIGS.find( + ([, candidate]) => candidate.family === family && candidate.basis === basis, + )?.[0] ?? MEASURED_METRIC_DEFAULTS[family] + ); + } let scope = change.scope ?? config.scope; if (config.family === 'power') { @@ -103,6 +158,7 @@ export function changeMeasuredMetricConfig( MEASURED_METRIC_CONFIGS.find( ([, candidate]) => candidate.family === 'power' && + candidate.basis === 'gpu-measured' && candidate.scope === scope && candidate.statistic === statistic && candidate.display === display, @@ -123,6 +179,7 @@ export function changeMeasuredMetricConfig( MEASURED_METRIC_CONFIGS.find( ([, candidate]) => candidate.family === 'energy' && + candidate.basis === 'gpu-measured' && candidate.scope === scope && candidate.denominator === denominator && candidate.unit === unit, diff --git a/packages/app/src/components/inference/metric-registry.test.ts b/packages/app/src/components/inference/metric-registry.test.ts index 28045728e..db4011826 100644 --- a/packages/app/src/components/inference/metric-registry.test.ts +++ b/packages/app/src/components/inference/metric-registry.test.ts @@ -19,6 +19,7 @@ import { metricCostTier, metricForCostTier, metricOptionTitle, + POWER_BASIS_METRIC_CONFIG_KEYS, resolveMetricConfigKey, tokenMetricTypeForConfigKey, } from './metric-registry'; @@ -39,6 +40,10 @@ describe('metric registry', () => { expect(e2e.y_costh_roofline).toBe('lower_left'); expect(interactivity.y_measuredPowerPercentTdp_roofline).toBe('lower_right'); expect(e2e.y_measuredPowerPercentTdp_roofline).toBe('lower_left'); + for (const key of POWER_BASIS_METRIC_CONFIG_KEYS) { + expect(interactivity[`${key}_roofline`], key).toBe('lower_right'); + expect(e2e[`${key}_roofline`], key).toBe('lower_left'); + } }); it('preserves metric-specific x overrides and bilingual labels', () => { @@ -210,7 +215,11 @@ describe('metric registry', () => { ); const measuredGroup = METRIC_CONTROL_GROUPS.find((group) => group.label === 'Measured Energy'); - expect(measuredGroup?.metrics).toBe(MEASURED_ENERGY_METRIC_CONFIG_KEYS); + expect(measuredGroup?.gated).toBe(true); + expect(measuredGroup?.metrics).toEqual([ + ...MEASURED_ENERGY_METRIC_CONFIG_KEYS, + ...POWER_BASIS_METRIC_CONFIG_KEYS, + ]); }); it('classifies measured-energy config keys', () => { diff --git a/packages/app/src/components/inference/metric-registry.ts b/packages/app/src/components/inference/metric-registry.ts index ccba918b1..05c19a7ce 100644 --- a/packages/app/src/components/inference/metric-registry.ts +++ b/packages/app/src/components/inference/metric-registry.ts @@ -386,6 +386,79 @@ export const METRIC_REGISTRY = { titleZh: '实测平均功耗占 TDP 百分比', polarity: 'lower', }, + // The per-second telemetry behind `measuredAvgPower`. The field aliases the + // same average so the table view, availability panel, and share links keep + // working; ChartDisplay swaps the scatter chart for `PowerTimeline`, which + // fetches each point's `gpu_metrics_*` artifact and draws the trace. The + // label leads with "Measured Average Power", like the %TDP display, so a + // "Measured Power" search still finds only the family option. + measuredPowerTimeline: { + field: 'measuredPowerTimeline.y', + label: 'Measured Average Power per Chip over Time (W)', + labelZh: '每芯片实测平均功耗时间线(W)', + title: 'Measured Average Power per Chip over Time', + titleZh: '每芯片实测平均功耗时间线', + polarity: 'lower', + }, + // Power boundaries beyond GPU-measured telemetry (`lib/power-basis.ts`). + // Each boundary publishes W per allocated GPU and J per output token; the + // Boundary select in the Measured controls resolves to these keys, so the + // metric key alone carries the boundary in share links. Keys deliberately + // lack the `measured` prefix: they are spec constants or model output, not + // telemetry, so the telemetry-only decorations must not treat them as such. + gpuProvisionedWatts: { + field: 'gpuProvisionedWatts.y', + label: 'GPU Level Provisioned Power per Chip (TDP, W)', + labelZh: '每芯片 GPU 额定功耗(TDP,W)', + title: 'GPU Level Provisioned Power per Chip (TDP)', + titleZh: '每芯片 GPU 额定功耗(TDP)', + polarity: 'lower', + }, + gpuProvisionedJPerOutputToken: { + field: 'gpuProvisionedJPerOutputToken.y', + label: 'GPU Level Provisioned J per Output Token (TDP, J/tok)', + labelZh: '每输出 token GPU 额定能耗(TDP,J/tok)', + title: 'GPU Level Provisioned Joules per Output Token (TDP)', + titleZh: '每输出 token GPU 额定焦耳能耗(TDP)', + polarity: 'lower', + }, + utilityProvisionedWatts: { + field: 'utilityProvisionedWatts.y', + label: 'All in Provisioned Power per Chip (W)', + labelZh: '每芯片整体预配功耗(W)', + title: 'All in Provisioned Power per Chip', + titleZh: '每芯片整体预配功耗', + polarity: 'lower', + }, + // Unlike the ungated `jOutput`, which divides by output tokens per decode + // GPU, this normalizes by every allocated GPU (prefill + decode). + utilityProvisionedJPerOutputToken: { + field: 'utilityProvisionedJPerOutputToken.y', + label: 'All in Provisioned J per Output Token, all GPUs (J/tok)', + labelZh: '每输出 token 整体预配能耗,按全部 GPU 归一(J/tok)', + title: 'All in Provisioned Joules per Output Token, all GPUs', + titleZh: '每输出 token 整体预配焦耳能耗,按全部 GPU 归一', + polarity: 'lower', + }, + // Names follow POWER_BASIS_LABELS (lib/power-basis.ts) so the Table, export + // titles and API docs read the same as the Boundary select: All in Provisioned + // / 整体预配功耗 and All in Measured / 整体实测功耗. + utilityModeledWatts: { + field: 'utilityModeledWatts.y', + label: 'All in Measured Power per Chip (W)', + labelZh: '每芯片整体实测功耗(W)', + title: 'All in Measured Power per Chip', + titleZh: '每芯片整体实测功耗', + polarity: 'lower', + }, + utilityModeledJPerOutputToken: { + field: 'utilityModeledJPerOutputToken.y', + label: 'All in Measured J per Output Token (J/tok)', + labelZh: '每输出 token 整体实测能耗(J/tok)', + title: 'All in Measured Joules per Output Token', + titleZh: '每输出 token 整体实测焦耳能耗', + polarity: 'lower', + }, } as const satisfies Record; export type MetricKey = keyof typeof METRIC_REGISTRY; @@ -580,8 +653,7 @@ export interface MetricControlGroup { /** * The runner-telemetry y-axes in the "Measured Energy" control group. * Exported (and referenced by the group below, so the two cannot drift) for - * consumers that treat measured axes specially — the legacy-power point ring, - * tooltip tier line, and footer legend key. + * consumers that treat measured axes specially, such as the tooltip tier line. */ export const MEASURED_ENERGY_METRIC_CONFIG_KEYS = [ 'y_measuredPrefillAvgPower', @@ -597,6 +669,7 @@ export const MEASURED_ENERGY_METRIC_CONFIG_KEYS = [ 'y_measuredJPerSuccessfulQuery', 'y_measuredWhPerSuccessfulQuery', 'y_measuredPowerPercentTdp', + 'y_measuredPowerTimeline', ] as const satisfies readonly MetricConfigKey[]; const MEASURED_ENERGY_METRIC_CONFIG_KEY_SET: ReadonlySet = new Set( @@ -618,6 +691,36 @@ export function isRoleLocalMeasuredEnergyConfigKey(configKey: string): boolean { return ROLE_LOCAL_MEASURED_ENERGY_METRIC_CONFIG_KEY_SET.has(configKey); } +/** + * The derived power-boundary y-axes (GPU provisioned, utility provisioned, + * utility modeled) that share the gated Measured Energy group and its + * Boundary select. They are kept out of `MEASURED_ENERGY_METRIC_CONFIG_KEYS` + * on purpose: spec constants and model output carry no telemetry tier, so the + * tier tooltip line does not apply to them. + */ +export const POWER_BASIS_METRIC_CONFIG_KEYS = [ + 'y_gpuProvisionedWatts', + 'y_gpuProvisionedJPerOutputToken', + 'y_utilityProvisionedWatts', + 'y_utilityProvisionedJPerOutputToken', + 'y_utilityModeledWatts', + 'y_utilityModeledJPerOutputToken', +] as const satisfies readonly MetricConfigKey[]; + +const POWER_BASIS_METRIC_CONFIG_KEY_SET: ReadonlySet = new Set( + POWER_BASIS_METRIC_CONFIG_KEYS, +); + +/** Whether a y-axis config key plots a derived power boundary (B2–B4). */ +export function isPowerBasisConfigKey(configKey: string): boolean { + return POWER_BASIS_METRIC_CONFIG_KEY_SET.has(configKey); +} + +/** The two All in Measured (utility-modeled) axes; they alone need the chassis power model. */ +export function isAllInMeasuredConfigKey(configKey: string): boolean { + return configKey === 'y_utilityModeledWatts' || configKey === 'y_utilityModeledJPerOutputToken'; +} + export const MODELED_SYSTEM_POWER_METRIC_CONFIG_KEY = 'y_modeledChassisPowerPerGpu'; /** Whether a y-axis config key plots the modeled chassis AC power metric. */ @@ -671,10 +774,13 @@ export const METRIC_CONTROL_GROUPS: readonly MetricControlGroup[] = [ // Runner power telemetry and the chassis model built on it are still being // validated, so both groups stay behind the ↑↑↓↓ feature gate until the // measurements are stable enough to publish. + // The derived boundaries ride along so the same gate and the same + // shared-URL exception (a gated metric selected by `i_metric` still renders + // while locked) apply to them. { label: 'Measured Energy', labelZh: '实测能耗', - metrics: MEASURED_ENERGY_METRIC_CONFIG_KEYS, + metrics: [...MEASURED_ENERGY_METRIC_CONFIG_KEYS, ...POWER_BASIS_METRIC_CONFIG_KEYS], gated: true, }, { diff --git a/packages/app/src/components/inference/perf-ruler-store.ts b/packages/app/src/components/inference/perf-ruler-store.ts new file mode 100644 index 000000000..d2aa927b2 --- /dev/null +++ b/packages/app/src/components/inference/perf-ruler-store.ts @@ -0,0 +1,181 @@ +'use client'; + +import { + type Dispatch, + type SetStateAction, + createContext, + useCallback, + useContext, + useMemo, + useRef, + useState, +} from 'react'; + +import { track } from '@/lib/analytics'; +import { perfRulerAxisMetricKey } from '@/hooks/usePerfRulerAxisReset'; +import { + EMPTY_PERF_RULER_STATE, + MAX_PERF_RULERS, + type PerfRulerMeasurement, + type PerfRulerState, + clearPerfRulers, + parsePerfRulers, +} from '@/lib/d3-chart/layers/perf-ruler'; + +/** + * @file perf-ruler-store.ts + * @description Provider-owned Perf Ruler state for the primary inference + * chart, so completed rulers persist in share links (`i_rulers`). Lives + * beside `InferenceContext` rather than inside it so `ScatterGraph` can + * consume the store without importing the (heavily mocked) provider module. + * ChartDisplay mounts it as `chart-0`; the date-comparison `GPUGraph` draws + * no rulers, and the replay chart keeps local state on purpose. + */ + +/** + * The chart instance whose Perf Rulers persist in share links. `ChartDisplay` + * mounts the primary chart as `chart-${graphIndex}` and only graph 0 is ever + * visible; the replay chart (`replay-chart-0`) draws interpolated frames of + * the same curves and must NOT bind, or every ruler would render twice and + * the replay's prune pass could delete rulers the main chart still shows. + */ +export const PERSISTED_PERF_RULER_CHART_ID = 'chart-0'; + +/** + * Perf-ruler store for the persisted chart. Lives in its own context rather + * than the Display domain so a ruler commit does not rerender every display + * consumer, and so harnesses that mount `InferenceContextsProvider` with + * static mock values (no store) keep the chart's component-local fallback. + * + * `state` holds COMMITTED rulers: the D3 layer renders them and `i_rulers` + * serializes them. `pending` holds rulers parsed from the share link whose + * curves may not have been drawn yet — data, `i_gpus`, comparison dates, + * and `?unofficialrun=` overlays all arrive after the chart's first draw, + * and the chart prunes any committed ruler whose curve path is absent from + * the DOM. The chart therefore commits a pending ruler only once BOTH of + * its curve paths exist (see the perf-ruler decoration effect in + * ScatterGraph); rulers whose curves never appear stay pending, invisible + * and unserialized, until an axis change or an explicit clear discards them. + */ +export interface PerfRulerStore { + chartId: string; + state: PerfRulerState; + setState: Dispatch>; + pending: readonly PerfRulerMeasurement[] | null; + /** + * Commit share-link rulers whose curves now exist (`resolved`, iso-x + * already clamped to the pair's overlap) and keep `remaining` pending. + */ + commitPending: ( + resolved: readonly PerfRulerMeasurement[], + remaining: readonly PerfRulerMeasurement[] | null, + ) => void; + /** Drop share-link rulers that were never committed (toggle-off, clear). */ + discardPending: () => void; +} + +/** + * Axis identity of the chart `ChartDisplay` renders as `chart-0`, for the + * store's axis reset. `graphs` is always `[interactivity, e2e]`, but + * ChartDisplay shows the e2e graph for every non-interactivity x mode + * (`visibleGraphs`), so the rendered chart — not `graphs[0]` — is what the + * rulers were placed on. The x mode itself is part of the identity as well: + * the derived agentic modes (e2e-normalized interactivity, …) override the + * e2e graph's `x_scale_field` inside ChartDisplay only, so here the same + * definition still reads `_e2el` for those modes. Percentile changes + * are already encoded in `x_scale_field`. Null while no graph exists. + */ +export function persistedPerfRulerAxisKey( + graphs: readonly { chartDefinition: { chartType: string; x_scale_field: string } }[], + xAxisMode: string, + yAxisMetric: string, +): string | null { + const wantedType = xAxisMode === 'interactivity' ? 'interactivity' : 'e2e'; + const graph = + graphs.find((candidate) => candidate.chartDefinition.chartType === wantedType) ?? graphs[0]; + if (!graph) return null; + return perfRulerAxisMetricKey(`${xAxisMode}:${graph.chartDefinition.x_scale_field}`, yAxisMetric); +} + +/** Provided by `InferenceProvider`; exported for chart component tests. */ +export const PerfRulerStoreContext = createContext(undefined); + +/** The persisted-ruler store, or undefined outside `InferenceProvider`. */ +export function usePerfRulerStore(): PerfRulerStore | undefined { + return useContext(PerfRulerStoreContext); +} + +/** + * Owns the persisted perf-ruler state. Exported so component tests can host a + * real store around a chart without the full provider. + * + * `axisMetricKey` is the persisted chart's axis identity + * ({@link persistedPerfRulerAxisKey}), or null while no chart definition exists. + * The axis reset runs HERE, not through `usePerfRulerAxisReset` in the chart: + * that hook adjusts state during the chart's render, which is only legal for + * the chart's own state — updating a provider's state from a child's render + * is a React error. Same semantics: a change of either axis metric clears + * committed rulers (redrawn curves would give a ratio nobody placed) and + * discards pending ones (they were placed on the old axes). The null → key + * transition on first data is not a change, so share-link rulers survive + * the load; the x-mode fallback for fixed sequences also settles before any + * chart definition exists. + */ +export function usePerfRulerStoreValue( + chartId: string, + initialSerialized: string | undefined, + axisMetricKey: string | null, +): PerfRulerStore { + const [state, setState] = useState(EMPTY_PERF_RULER_STATE); + const [pending, setPending] = useState(() => { + const parsed = parsePerfRulers(initialSerialized).rulers; + return parsed.length > 0 ? parsed : null; + }); + // `interactivity_perf_ruler_shared_load` fires once per store — once per + // opened link — with the number of rulers the link carried, the first time + // any of them renders. Rulers commit per curve arrival (below), so a + // per-commit event would count one link several times with partial counts. + const linkRulerCountRef = useRef(pending?.length ?? 0); + const sharedLoadReportedRef = useRef(false); + + const [appliedAxisMetricKey, setAppliedAxisMetricKey] = useState(axisMetricKey); + if (axisMetricKey !== null && axisMetricKey !== appliedAxisMetricKey) { + setAppliedAxisMetricKey(axisMetricKey); + if (appliedAxisMetricKey !== null) { + setState(clearPerfRulers); + setPending(null); + } + } + + const commitPending = useCallback( + ( + resolved: readonly PerfRulerMeasurement[], + remaining: readonly PerfRulerMeasurement[] | null, + ) => { + if (resolved.length > 0) { + setState((prev) => { + // Fresh ids from the live counter: a ruler placed by hand before the + // share-link rulers resolved must keep its own join key. + const rulers = [ + ...prev.rulers, + ...resolved.map((ruler, index) => ({ ...ruler, id: prev.nextId + index })), + ]; + while (rulers.length > MAX_PERF_RULERS) rulers.shift(); + return { rulers, draft: prev.draft, nextId: prev.nextId + resolved.length }; + }); + if (!sharedLoadReportedRef.current) { + sharedLoadReportedRef.current = true; + track('interactivity_perf_ruler_shared_load', { count: linkRulerCountRef.current }); + } + } + setPending(remaining); + }, + [], + ); + const discardPending = useCallback(() => setPending(null), []); + + return useMemo( + () => ({ chartId, state, setState, pending, commitPending, discardPending }), + [chartId, state, pending, commitPending, discardPending], + ); +} diff --git a/packages/app/src/components/inference/power-telemetry-dialog.tsx b/packages/app/src/components/inference/power-telemetry-dialog.tsx new file mode 100644 index 000000000..e20718a55 --- /dev/null +++ b/packages/app/src/components/inference/power-telemetry-dialog.tsx @@ -0,0 +1,51 @@ +'use client'; + +import type { InferenceData } from '@/components/inference/types'; +import { PowerTelemetryView } from '@/components/inference/agentic-point/power-telemetry-view'; +import { + Dialog, + DialogContent, + DialogDescription, + DialogHeader, + DialogTitle, +} from '@/components/ui/dialog'; +import { isPersistedBenchmarkId } from '@/lib/benchmark-id'; +import { useLocale } from '@/lib/use-locale'; + +const STRINGS = { + en: { point: 'Benchmark point', concurrency: 'Concurrency' }, + zh: { point: '基准测试数据点', concurrency: '并发数' }, +} as const; + +interface Props { + point: InferenceData; + onOpenChange: (open: boolean) => void; +} + +export function PowerTelemetryDialog({ point, onOpenChange }: Props) { + const t = STRINGS[useLocale()]; + if (!isPersistedBenchmarkId(point.id)) return null; + + return ( + + + + PowerX + + {t.point} #{point.id} · {point.hwKey} · {point.precision.toUpperCase()} ·{' '} + {t.concurrency} {point.conc} + + + + + + ); +} diff --git a/packages/app/src/components/inference/types.ts b/packages/app/src/components/inference/types.ts index c9f9e6c96..33a491b16 100644 --- a/packages/app/src/components/inference/types.ts +++ b/packages/app/src/components/inference/types.ts @@ -6,6 +6,9 @@ import type { Model, Sequence } from '@/lib/data-mappings'; import type { PowerTier } from '@/lib/power-tier'; import type { SystemPowerEstimate } from '@/lib/modeled-system-power'; import type { MetricKey } from './metric-registry'; +import type { FixedSequenceStatistic } from './utils/resolveXAxisField'; +import type { XAxisMode } from './hooks/chart-data-core'; +import type { PowerBasis } from '@/lib/power-basis'; export type { WorkerPower }; @@ -87,6 +90,8 @@ export interface AggDataEntry { 'p99.9_ttft': number; mean_tpot: number; mean_intvty: number; + /** Reciprocal of finite positive mean TPOT (seconds), not raw mean_intvty. */ + mean_tpot_intvty?: number; median_tpot: number; median_intvty: number; std_tpot: number; @@ -204,6 +209,13 @@ export interface AggDataEntry { actualDate?: string; /** URL to the GitHub Actions workflow run that produced this data point. */ run_url?: string; + /** + * Logical curve snapshot this row belongs to. An append-only run stitches new + * points onto an older run's curve; both keep their own `run_url` but share + * this snapshot identity. Undefined on legacy and unofficial rows. + */ + curve_date?: string; + curve_workflow_run_id?: number; /** Benchmark scenario: `single_turn` (fixed-seq isl/osl) or `agentic_traces`. */ benchmark_type?: string; /** ISL in tokens — null for agentic_traces. */ @@ -347,8 +359,67 @@ export interface InferenceData extends Partial void; setSelectedYAxisMetric: (metric: string) => void; setTokenRevenuePriceSource: (source: TokenRevenuePriceSource) => void; - setSelectedPercentile: (percentile: string) => void; + setFixedSequenceStatistic: (statistic: FixedSequenceStatistic) => void; setSelectedXAxisMetric: (metric: string | null) => void; - setSelectedXAxisMode: ( - mode: 'ttft' | 'e2e' | 'interactivity' | 'e2e-normalized-interactivity', - ) => void; + setSelectedXAxisMode: (mode: XAxisMode) => void; setScaleType: (type: 'auto' | 'linear' | 'log') => void; + setPowerCompare: (mode: PowerCompare) => void; setQuickFilterVendors: (vendors: string[]) => void; setQuickFilterFrameworks: (frameworks: string[]) => void; setQuickFilterDeployment: (modes: DeploymentMode[]) => void; setQuickFilterSpec: (modes: SpecMode[]) => void; setQuickFilterPower: (tiers: PowerTier[]) => void; + setQuickFilterTopologies: (topologies: string[]) => void; setIsLegendExpanded: (expanded: boolean) => void; setHideNonOptimal: (hide: boolean) => void; setShowAllMeasurements: (show: boolean) => void; diff --git a/packages/app/src/components/inference/ui/ActiveQuickFilters.tsx b/packages/app/src/components/inference/ui/ActiveQuickFilters.tsx index 2160444ee..a8f1eda3b 100644 --- a/packages/app/src/components/inference/ui/ActiveQuickFilters.tsx +++ b/packages/app/src/components/inference/ui/ActiveQuickFilters.tsx @@ -34,6 +34,7 @@ export function ActiveQuickFilters() { else if (category === 'deployment') actions.setQuickFilterDeployment(values as DeploymentMode[]); else if (category === 'spec') actions.setQuickFilterSpec(values as SpecMode[]); + else if (category === 'topologies') actions.setQuickFilterTopologies(values); else actions.setQuickFilterPower(values as PowerTier[]); }; @@ -56,7 +57,7 @@ export function ActiveQuickFilters() { onClick={() => { setCategory( category, - quickFilters[category].filter((item) => item !== value), + (quickFilters[category] ?? []).filter((item) => item !== value), ); track('inference_quick_filter_removed', { category, value, source: 'result_summary' }); }} diff --git a/packages/app/src/components/inference/ui/ChartControls.tsx b/packages/app/src/components/inference/ui/ChartControls.tsx index 203275706..38ff3d99f 100644 --- a/packages/app/src/components/inference/ui/ChartControls.tsx +++ b/packages/app/src/components/inference/ui/ChartControls.tsx @@ -20,7 +20,6 @@ import { import { ModelSelector, ScenarioSelector, - PercentileSelector, PrecisionSelector, } from '@/components/ui/chart-selectors'; import { DateRangePicker } from '@/components/ui/date-range-picker'; @@ -58,15 +57,15 @@ import { import { useOpenDropdown } from '@/hooks/useOpenDropdown'; import { ModelArchitectureInfoLink } from './ModelArchitectureInfoLink'; import { MetricExplanation } from './MetricExplanation'; -import { PowerMetricAvailability } from './PowerMetricAvailability'; import { MeasuredMetricControls } from './MeasuredMetricControls'; import { + changeMeasuredMetricConfig, getMeasuredMetricConfig, MEASURED_METRIC_DEFAULTS, type MeasuredMetricFamily, } from '../measured-metric-config'; import { XAxisModeSelector } from './XAxisModeSelector'; -import { showsTcoBasisSelector, Sequence, type Model, type Percentile } from '@/lib/data-mappings'; +import { showsTcoBasisSelector, Sequence, type Model } from '@/lib/data-mappings'; import { useLocale } from '@/lib/use-locale'; import { DEFAULT_Y_AXIS_METRIC } from '@/lib/url-state'; @@ -226,10 +225,10 @@ export default function ChartControls({ openRouterModelId, openRouterPricingLoading, openRouterPricingError, - selectedPercentile, selectedXAxisMetric, selectedXAxisMode, scaleType, + powerCompare, } = useInferenceDisplay(); const { setSelectedModel, @@ -237,11 +236,11 @@ export default function ChartControls({ setSelectedPrecisions, setSelectedYAxisMetric, setTokenRevenuePriceSource, - setSelectedPercentile, setSelectedGPUs, setSelectedDateRange, setSelectedXAxisMetric, setScaleType, + setPowerCompare, } = useInferenceActions(); // Y-axis options come from the canonical registry and need no API data. @@ -335,10 +334,14 @@ export default function ChartControls({ if (!config) return [option]; if (seen.has(config.family)) return []; seen.add(config.family); + // Keep the selected boundary (and other dimensions) when hopping between + // the power and energy families; fall back to the family default otherwise. const value = selectedConfig?.family === config.family ? selectedYAxisMetric - : MEASURED_METRIC_DEFAULTS[config.family]; + : selectedConfig + ? changeMeasuredMetricConfig(selectedYAxisMetric, { family: config.family }) + : MEASURED_METRIC_DEFAULTS[config.family]; return [ { value, @@ -457,8 +460,6 @@ export default function ChartControls({ showTcoBasis && isCostMetric(selectedYAxisMetric) && showsTcoBasisSelector(selectedModel, selectedSequence); - const showPercentile = - mounted && selectedSequence === Sequence.AgenticTraces && featureGateUnlocked; return ( @@ -467,9 +468,7 @@ export default function ChartControls({ legend={t.benchmarkControls} className={hideGpuComparison ? 'lg:col-span-2' : 'lg:col-span-3'} > -
+
- {/* AgentX publishes on P90, so the percentile control is an insider - affordance rather than a normal chart filter: it stays behind the - ↑↑↓↓ feature gate and the chart defaults to P90 without it. */} - {showPercentile && ( - setSelectedPercentile(p)} - data-testid="percentile-selector" - /> - )}
@@ -558,27 +547,15 @@ export default function ChartControls({ noResultsLabel={locale === 'zh' ? '无结果' : undefined} clearSearchLabel={locale === 'zh' ? '清除搜索' : undefined} /> - {mounted && !getMeasuredMetricConfig(selectedYAxisMetric) && ( - - )}
{mounted && getMeasuredMetricConfig(selectedYAxisMetric) && ( - <> - -
- -
- + )} {tcoVisible && ( diff --git a/packages/app/src/components/inference/ui/ChartDisplay.tsx b/packages/app/src/components/inference/ui/ChartDisplay.tsx index f23dacd94..768d87c48 100644 --- a/packages/app/src/components/inference/ui/ChartDisplay.tsx +++ b/packages/app/src/components/inference/ui/ChartDisplay.tsx @@ -7,13 +7,16 @@ import { BarChart3, Table2 } from 'lucide-react'; import chartDefinitions, { costTierLabel, costTierOptionLabel, - isMeasuredEnergyConfigKey, isModeledSystemPowerConfigKey, metricCostTier, tokenMetricTypeForConfigKey, type MetricKey, } from '@/components/inference/metric-registry'; import { metricRowLabel } from '@/components/inference/axis-metric-explanations'; +import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config'; +import { AIR_COOLED_SYSTEM_PUE } from '@/lib/modeled-system-power'; +import { SYSTEM_POWER_MODEL_REVISION } from '@/lib/system-power-model'; +import { ALL_IN_MEASURED_EMPTY, ALL_IN_MEASURED_NOTE } from '@/lib/power-basis'; import { applyTokenRevenuePricing, cachedInputPricePerMillion, @@ -28,6 +31,7 @@ import { } from '@/components/inference/InferenceContext'; import { useGlobalFilterSelection } from '@/components/GlobalFilterContext'; import type { + AggDataEntry, ChartDefinition, HardwareConfig, InferenceData, @@ -46,6 +50,8 @@ import { matchesQuickFilters } from '@/components/inference/utils/quickFilters'; import { bestSeriesPerSku } from '@/components/inference/utils/best-series-per-sku'; import InferenceTable from '@/components/inference/ui/InferenceTable'; import ScatterGraph from '@/components/inference/ui/ScatterGraph'; +import PowerTimeline from '@/components/inference/ui/PowerTimeline'; +import PowerAnalysisPanels from '@/components/inference/ui/PowerAnalysisPanels'; import { Card } from '@/components/ui/card'; import { ChartButtons } from '@/components/ui/chart-buttons'; import { ShareButton } from '@/components/ui/share-button'; @@ -94,7 +100,6 @@ import { ATOM_FOOTNOTE_MARKER, AtomEngineFootnote } from '@/components/ui/atom-e import { AgenticOptimizationNote } from '@/components/inference/ui/AgenticOptimizationNote'; import { CacheReuseLink } from '@/components/inference/ui/CacheReuseLink'; import { OffloadHaloLegendKey } from '@/components/inference/ui/OffloadHaloLegendKey'; -import { LegacyPowerLegendKey } from '@/components/inference/ui/LegacyPowerLegendKey'; import { ActiveQuickFilters } from '@/components/inference/ui/ActiveQuickFilters'; import { ResultContext } from '@/components/ui/result-context'; import { ModelLogo } from '@/components/ui/model-logo'; @@ -140,6 +145,15 @@ const STRINGS = { 'No benchmark data matches the current model, scenario, and filter selection. Adjust the filters above to see results.', noSystemPowerData: 'No system-power estimates are available for this selection. Choose 8K / 1K with validated GPU telemetry, supported hardware, and known eight-GPU chassis placement. Measured GPU power remains available separately where telemetry exists.', + // Boundary disclosures for the derived power axes (lib/power-basis.ts). + // Formulas in words; constants named so a screenshot records its method. + powerBasisAssumptions: { + 'gpu-provisioned': + 'GPU Level Provisioned (TDP) · Watts are the rated TDP per GPU from the hardware registry, so the power curve is flat per hardware. Joules per output token = TDP × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together. Hardware without a published TDP is omitted.', + 'utility-provisioned': + 'All in Provisioned · Watts are the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model), so the power curve is flat per hardware. Joules per output token = all-in W × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together, unlike the ungated All-in Provisioned J per Output Token, which divides per decode GPU.', + 'utility-modeled': `All in Measured · Measured GPU power carried through the modeled chassis (CPU, DRAM, platform, PSU losses) to the utility meter: modeled chassis AC × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled, applied once), divided by the measured GPUs; joules per output token scale measured joules by the same ratio. Chassis power model revision ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}. Available for 8K / 1K with validated telemetry on supported hardware only; NVL72 systems (GB200, GB300) and points without values are omitted.`, + }, vsTtft: (word: string) => `vs. ${word} Time To First Token`, vsE2eLatency: (pctl?: string) => pctl ? `vs. ${pctl} End-to-end Latency` : 'vs. End-to-end Latency', @@ -166,6 +180,13 @@ const STRINGS = { noChartData: '当前模型、场景与筛选条件下没有匹配的基准测试数据。请调整上方筛选条件查看结果。', noSystemPowerData: '当前选择没有可用的系统功耗估算。请选择 8K / 1K 场景;估算仅覆盖 GPU 遥测已验证、硬件受支持、八卡机箱位置已知的运行。存在遥测数据时,仍可单独查看 GPU 实测功耗。', + powerBasisAssumptions: { + 'gpu-provisioned': + 'GPU 额定功耗(TDP)· 功率取硬件注册表中每 GPU 的额定 TDP,因此每种硬件的功率曲线为水平线。每输出 token 能耗 = TDP × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入。未公布 TDP 的硬件不绘制。', + 'utility-provisioned': + '整体预配功耗 · 功率取硬件注册表中每 GPU 的全电源配置(all-in)市电功率(来源:SemiAnalysis Datacenter Industry Model),因此每种硬件的功率曲线为水平线。每输出 token 能耗 = all-in 功率 × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入,这与未加门控的“每输出 token 全电源配置能耗”按 decode GPU 计算不同。', + 'utility-modeled': `整体实测功耗 · 将 GPU 实测功耗经机箱功耗模型(CPU、DRAM、平台开销、PSU 损耗)推算至市电侧:机箱交流功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷,仅应用一次),再除以实测 GPU 数;每输出 token 能耗按同一比例放大实测能耗。机箱功耗模型版本 ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}。仅适用于 8K / 1K、遥测已验证且硬件受支持的运行;NVL72 系统(GB200、GB300)及缺少数值的数据点不绘制。`, + }, vsTtft: (word: string) => `vs. ${word === 'Median' ? '中位' : word} 首 token 延迟(TTFT)`, vsE2eLatency: (pctl?: string) => (pctl ? `vs. ${pctl} 端到端延迟` : 'vs. 端到端延迟'), }, @@ -189,7 +210,9 @@ function zhHeading(configured: string): string { const subjectZh = match?.groups && HEADING_SUBJECT_ZH[match.groups.subject]; if (!subjectZh) return configured; const pctl = match.groups?.pctl; - return `vs. ${pctl ? `${pctl} ` : ''}${subjectZh}`; + const statisticZh = + pctl === 'Mean' ? '平均' : pctl === 'Median' ? '中位' : pctl ? `${pctl} ` : ''; + return `vs. ${statisticZh}${subjectZh}`; } /** Presentation and data plumbing for trace-derived agentic x-axis modes. */ @@ -288,12 +311,23 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean selectedXAxisMetric, selectedE2eXAxisMetric, selectedPercentile, + fixedSequenceStatistic, selectedXAxisMode, tokenRevenuePricing, showLineLabels, + powerCompare, } = useInferenceDisplay(); const { setSelectedDates, setSelectedDatesFromRunExpansion, setIsLegendExpanded } = useInferenceActions(); + const selectedMeasuredConfig = getMeasuredMetricConfig(selectedYAxisMetric); + // The metric key carries the power boundary; the caption discloses it for + // the derived boundaries (there is no separate URL param). + const selectedPowerBasis = selectedMeasuredConfig?.basis; + // The Measured Power "Timeline" display swaps the scatter body for the + // per-second telemetry traces (PowerTimeline); table view and captions are + // unchanged because the metric key aliases the measured average. + const isPowerTimeline = + selectedMeasuredConfig?.family === 'power' && selectedMeasuredConfig.display === 'timeline'; const selectedBenchmarkType: 'single_turn' | 'agentic_traces' = selectedSequence === Sequence.AgenticTraces ? 'agentic_traces' : 'single_turn'; const workflowInfoBenchmarkType = @@ -454,8 +488,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean { isAgentic, selectedPercentile, + fixedSequenceStatistic, tcoBasis, selectedXAxisMode, + powerCompare, }, ); @@ -505,6 +541,8 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean selectedXAxisMetric, selectedE2eXAxisMetric, selectedPercentile, + fixedSequenceStatistic, + powerCompare, selectedXAxisMode, tokenRevenuePricing, tcoBasis, @@ -651,6 +689,25 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean }, [selectedPrecisions, quickFilters, selectedOfficialHwTypes, scopedActiveOverlayHwTypes], ); + // Date comparison (GPUGraph, Timeline): official rows follow the per-date + // legend toggles instead of the scatter hardware selection, and unofficial + // runs keep the overlay hardware selection. Boundary / role siblings are a + // same-run comparison, so neither side draws them here. + const visibleDateComparisonRows = useCallback( + (officialRows: InferenceData[], overlay: OverlayData | null | undefined) => ({ + officialRows: officialRows.filter( + (point) => + !point.powerVariant && + selectedPrecisions.includes(point.precision) && + matchesQuickFilters(point, quickFilters) && + activeDates.has(`${point.date}_${point.hwKey}`), + ), + overlayRows: visibleComparisonRows([], overlay).overlayRows.filter( + (point) => !point.powerVariant, + ), + }), + [selectedPrecisions, quickFilters, activeDates, visibleComparisonRows], + ); if (!loading && error) { console.error(error); @@ -747,7 +804,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const derivedSpec = useDerivedXAxis ? DERIVED_X_MODE_SPECS[selectedXAxisMode] : undefined; const renderableGraphs = useMemo(() => { - if (!isAgenticSequence) return visibleGraphs; + if (!isAgenticSequence || selectedXAxisMode === 'concurrency') return visibleGraphs; if (!derivedMetrics) { // Legacy AgentX axes can still render transient/non-persisted rows, which // have no ids to request. @@ -797,6 +854,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean isAgenticSequence, derivedSpec, derivedTargetIds.length, + selectedXAxisMode, visibleGraphs, derivedMetrics, selectedYAxisMetric, @@ -825,7 +883,9 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean

{isModeledSystemPowerConfigKey(selectedYAxisMetric) ? t.noSystemPowerData - : t.noChartData} + : selectedPowerBasis === 'utility-modeled' + ? ALL_IN_MEASURED_EMPTY[locale] + : t.noChartData}

, ] @@ -834,38 +894,33 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const isTimelineMode = Boolean( selectedDateRange.startDate && selectedDateRange.endDate && selectedGPUs.length > 0, ); - const replayAvailable = getViewMode(graphIndex) === 'chart' && !isTimelineMode; + const replayAvailable = + getViewMode(graphIndex) === 'chart' && + !isTimelineMode && + selectedXAxisMode !== 'concurrency'; // Chart-level notices: the KV-offload halo // key, the agentic optimization note, and the ATOM engine // footnote. Detected from the same data the chart plots — // official points plus any loaded unofficial-run overlay for // this chart type — so they moved out of the legend without // changing when they appear. - // GPU/date comparison renders GPUGraph, which plots official - // points only — skip the unofficial overlay there so the footer - // can't advertise a halo or ATOM series that isn't on the chart. + // GPU/date comparison renders GPUGraph, which plots the loaded + // unofficial runs next to the compared dates on every x-axis. const isGpuComparison = selectedGPUs.length > 0 && ((selectedDateRange.startDate && selectedDateRange.endDate) || selectedDates.length > 0); - const footerOverlay = isGpuComparison - ? undefined - : selectUnofficialOverlayForMode( - selectedXAxisMode, - graph.chartDefinition.chartType, - overlayDataByChartType, - ); + const footerOverlay = selectUnofficialOverlayForMode( + selectedXAxisMode, + graph.chartDefinition.chartType, + overlayDataByChartType, + ); const footerPoints = [ ...graph.data, ...(footerOverlay?.data ?? []), ...(footerOverlay?.clippedData ?? []).map((entry) => entry.point), ]; const hasOffloadHalo = footerPoints.some((point) => point.offload_mode === 'on'); - // Legacy-power rings render only on Measured Energy axes, so the - // key follows the same gate to never advertise an absent ring. - const hasLegacyPowerPoints = - isMeasuredEnergyConfigKey(selectedYAxisMetric) && - footerPoints.some((point) => point.power_tier === 'legacy'); const hasAtomSeries = footerPoints.some( (point) => point.framework !== undefined && @@ -875,14 +930,13 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean // here as the footer's last block rather than in the chart subtitle, // keeping the result-context header compact. const footerNotices = - hasOffloadHalo || hasLegacyPowerPoints || isAgenticSequence || hasAtomSeries ? ( + hasOffloadHalo || isAgenticSequence || hasAtomSeries ? ( <>
{hasOffloadHalo && } - {hasLegacyPowerPoints && } {isAgenticSequence && } {isAgenticSequence && !minimalChrome && } {hasAtomSeries && ( @@ -948,9 +1002,6 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean : undefined } onExportCsv={() => { - const candidateVisibleData = isTimelineMode - ? graph.data.filter((d) => activeDates.has(`${d.date}_${d.hwKey}`)) - : graph.data; const overlay = selectUnofficialOverlayForMode( selectedXAxisMode, graph.chartDefinition.chartType, @@ -959,9 +1010,9 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const { officialRows: visibleData, overlayRows: visibleOverlayRowsForExport, - } = isTimelineMode - ? { officialRows: candidateVisibleData, overlayRows: [] } - : visibleComparisonRows(candidateVisibleData, overlay); + } = isGpuComparison + ? visibleDateComparisonRows(graph.data, overlay) + : visibleComparisonRows(graph.data, overlay); const { headers, rows } = inferenceChartToCsv( visibleData, graph.model, @@ -1013,6 +1064,14 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean {getSequenceLabel(graph.sequence as Sequence, locale)}{' '} {metricChartTitle(graph.chartDefinition, selectedYAxisMetric, locale)}{' '} {(() => { + // The timeline's x axis is time, not the scatter x metric. + if (isPowerTimeline) return null; + if (selectedXAxisMode === 'concurrency') + return locale === 'zh' ? '与并发数的关系' : 'vs. Concurrency'; + if (!isAgenticSequence) { + const heading = String(graph.chartDefinition.heading); + return locale === 'zh' ? zhHeading(heading) : heading; + } const xField = graph.chartDefinition.x_scale_field; if (xField?.endsWith('_ttft')) { const percentile = xField.replace(/_ttft$/u, ''); @@ -1140,6 +1199,25 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean {t.systemPowerAssumptions}

)} + {selectedPowerBasis && selectedPowerBasis !== 'gpu-measured' && ( +

+ {t.powerBasisAssumptions[selectedPowerBasis]} +

+ )} + {selectedPowerBasis && + powerCompare === 'boundaries' && + selectedPowerBasis !== 'utility-modeled' && ( +

+ {ALL_IN_MEASURED_NOTE[locale]} +

+ )} {isUnofficialRun && selectedXAxisMode === 'e2e-normalized-interactivity' && (

@@ -1173,10 +1251,9 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean ], } : overlay; - const { officialRows, overlayRows } = visibleComparisonRows( - tableOfficialData, - tableOverlay, - ); + const { officialRows, overlayRows } = isGpuComparison + ? visibleDateComparisonRows(tableOfficialData, tableOverlay) + : visibleComparisonRows(tableOfficialData, tableOverlay); return ( <> {chartCaption} @@ -1189,15 +1266,49 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean ); } + if (isPowerTimeline) { + return ( +

+ entry.point), + ]} + overlayData={ + selectUnofficialOverlayForMode( + selectedXAxisMode, + graph.chartDefinition.chartType, + overlayDataByChartType, + ) ?? undefined + } + yLabel={metricLabel( + graph.chartDefinition, + selectedYAxisMetric, + locale, + )} + caption={chartCaption} + comparison={Boolean(isGpuComparison)} + runNumbering={runNumbering} + /> +
+ ); + } + return isGpuComparison ? ( !point.powerVariant)} xLabel={resolvedXLabel} yLabel={metricLabel(graph.chartDefinition, selectedYAxisMetric, locale)} chartDefinition={graph.chartDefinition} caption={chartCaption} + overlayData={footerOverlay ?? undefined} runNumbering={runNumbering} /> ) : ( @@ -1223,6 +1334,29 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
); })()} + {selectedMeasuredConfig && + !isPowerTimeline && + getViewMode(graphIndex) !== 'table' && + (() => { + const overlay = selectUnofficialOverlayForMode( + selectedXAxisMode, + graph.chartDefinition.chartType, + overlayDataByChartType, + ); + const { officialRows, overlayRows } = isGpuComparison + ? visibleDateComparisonRows(graph.data, overlay) + : visibleComparisonRows(graph.data, overlay); + return ( + + ); + })()} {replayAvailable && !minimalChrome && ( `attempt ${attempt}`, + unofficial: 'unofficial', + noRun: 'No run recorded', + }, + zh: { + title: '前沿点', + hardware: '硬件', + config: '配置', + concurrency: '并发', + run: '运行', + date: '日期', + attempt: (attempt: number) => `第 ${attempt} 次尝试`, + unofficial: '非官方', + noRun: '未记录运行', + }, +}; + +const number = d3.format(',.4~g'); + +export default function FrontierPointsPanel({ + chartId, + eligible, + frontier, + xLabel, + yLabel, + overlayPoints, + hardwareLabel, + hardwareColor, +}: { + chartId: string; + /** Observations that competed: the chart's visible, frontier-eligible points. */ + eligible: readonly InferenceData[]; + frontier: readonly InferenceData[]; + xLabel: string; + yLabel: string; + overlayPoints: readonly InferenceData[]; + hardwareLabel: (point: InferenceData) => string; + hardwareColor: (point: InferenceData) => string; +}) { + const locale = useLocale(); + const t = STRINGS[locale]; + const { runIndexByUrl } = useUnofficialRun(); + const overlaySet = useMemo(() => new Set(overlayPoints), [overlayPoints]); + const topologies = useMemo(() => [...new Set(eligible.map(pointTopologyKey))], [eligible]); + const sectionId = `${chartId}-frontier-points`; + const colorOf = (point: InferenceData) => + overlaySet.has(point) + ? overlayRunColor(overlayRunIndex(point.run_url, runIndexByUrl)) + : hardwareColor(point); + + const exportCsv = () => { + exportToCsv( + 'InferenceX_frontier_points', + [...FRONTIER_EXPORT_HEADERS, 'unofficial'], + frontier.map((point) => [...frontierExportRow(point), overlaySet.has(point)]), + [`x: ${xLabel}`, `y: ${yLabel}`], + ); + }; + + return ( +
+ + {t.title} + +
+ + + + + + + + + + + + + + {frontier.map((point, index) => { + const runId = runIdFromUrl(point.run_url); + const attempt = runAttemptFromUrl(point.run_url); + return ( + + + + + + + + + + ); + })} + +
{t.hardware}{t.config}{t.concurrency}{xLabel}{yLabel}{t.run}{t.date}
+ ● + {hardwareLabel(point)} + {overlaySet.has(point) && ( + {t.unofficial} + )} + + {[ + point.framework, + point.precision.toUpperCase(), + topologyLabel(pointTopologyKey(point), locale, topologies), + ] + .filter(Boolean) + .join(' · ')} + {point.conc}{number(point.x)}{number(point.y)} + {point.run_url && runId ? ( + + {`run ${runId}`} + + ) : ( + {t.noRun} + )} + {attempt !== null && ( + + {t.attempt(attempt)} + + )} + {point.actualDate ?? point.date}
+
+ +
+ ); +} diff --git a/packages/app/src/components/inference/ui/GPUGraph.tsx b/packages/app/src/components/inference/ui/GPUGraph.tsx index cc9a47640..6e3822566 100644 --- a/packages/app/src/components/inference/ui/GPUGraph.tsx +++ b/packages/app/src/components/inference/ui/GPUGraph.tsx @@ -1,5 +1,7 @@ 'use client'; +import { useFeatureGate } from '@/lib/use-feature-gate'; +import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config'; import { track } from '@/lib/analytics'; import { isPersistedBenchmarkId } from '@/lib/benchmark-id'; import { useEphemeralUrlState } from '@/hooks/useUrlState'; @@ -7,8 +9,6 @@ import { rememberChartStateInUrl } from '@/lib/url-state'; import * as d3 from 'd3'; import dynamic from 'next/dynamic'; import React, { useCallback, useEffect, useMemo, useRef, useState } from 'react'; -import { useTheme } from 'next-themes'; -import { isDarkTheme } from '@/lib/themes'; import { useInferenceActions, @@ -19,16 +19,20 @@ import { import ChartLegend from '@/components/ui/chart-legend'; import { Button } from '@/components/ui/button'; import { OFFICIAL_PREVIEW_SERIES } from '@/components/official-preview-notice'; -import { getHardwareConfig, getModelSortIndex, hardwareKeyMatchesAnyBase } from '@/lib/constants'; -import { getInferenceHardwareConfig } from '@/lib/inference-labels'; +import { useUnofficialRun } from '@/components/unofficial-run-provider'; +import { getHardwareConfig, hardwareKeyMatchesAnyBase } from '@/lib/constants'; +import { + getInferenceHardwareConfig, + getInferenceRunLabel, + getOverlayLineLabel, +} from '@/lib/inference-labels'; import { getChartWatermark, Sequence } from '@/lib/data-mappings'; -import { generateGpuDateColors, generateHighContrastGpuDateColors } from '@/lib/dynamic-colors'; import { useLocale } from '@/lib/use-locale'; import { formatNumber, getDisplayLabel, updateRepoUrl } from '@/lib/utils'; -import { useThemeColors } from '@/hooks/useThemeColors'; import { useTraceAvailability } from '@/hooks/api/use-trace-availability'; import { useLogAvailability } from '@/hooks/api/use-log-availability'; import { D3Chart } from '@/lib/d3-chart/D3Chart'; +import { CHART_TYPE, px } from '@/lib/d3-chart/typography'; import type { CustomLayerConfig, D3ChartHandle, @@ -41,13 +45,26 @@ import { applyNormalState, formatLargeNumber, getShapeKeyForPrecision, + HIT_AREA_RADIUS, logTickFormat, } from '@/lib/chart-rendering'; +import { computeTooltipPosition } from '@/lib/d3-chart/layers/scatter-points'; +import { + attachOverlayXMarkerHandlers, + overlayMarkerPosition, + xMarkerPath, +} from '@/lib/d3-chart/overlay-x-marker'; +import { + overlayRooflineDasharray, + overlayRunColor, + overlayRunIndex, +} from '@/lib/overlay-run-style'; import type { ParetoDirection } from '@/lib/chart-utils'; import { chartFrontier, upperPowerEnvelope, isPowerCurveMetric, + isPowerGaugeSeries, isMeasuredPowerCurveMetric, } from '@/components/inference/utils/powerCurves'; import type { @@ -55,13 +72,15 @@ import type { InferenceData, ScatterGraphProps, } from '@/components/inference/types'; +import { comparisonEntryLabel } from '@/components/inference/utils/comparisonEntry'; +import { groupConcurrencySeries } from '@/components/inference/utils/concurrency-series'; +import { matchesQuickFilters } from '@/components/inference/utils/quickFilters'; import { - buildRunNumbering, - comparisonEntryLabel, - comparisonEntrySortValue, - resolveComparisonEntries, -} from '@/components/inference/utils/comparisonEntry'; -import { generateGPUGraphTooltipContent } from '@/components/inference/utils/tooltipUtils'; + generateGPUGraphTooltipContent, + generateOverlayTooltipContent, +} from '@/components/inference/utils/tooltipUtils'; +import { useComparisonSeries } from '@/components/inference/hooks/useComparisonSeries'; +import { usePowerTraceAction } from '@/components/inference/hooks/usePowerTraceAction'; import { pointLabelText } from '@/components/inference/ui/point-label'; import { scatterPointConfigId } from '@/components/inference/utils/point-identity'; import { @@ -70,15 +89,10 @@ import { } from '@/components/inference/utils/knownIssueAnnotations'; import { matchKnownConfigIssues, pointMatchesIssue } from '@/lib/known-issues'; import { renderOffloadHalo } from '@/components/inference/utils/offload-halo'; -import { renderLegacyPowerRing } from '@/components/inference/utils/legacy-power-marker'; import { isMeasuredEnergyConfigKey, isRoleLocalMeasuredEnergyConfigKey, } from '@/components/inference/metric-registry'; -import { - countPowerTiers, - MeasuredPowerSummary, -} from '@/components/inference/ui/MeasuredPowerSummary'; import { keepPointLabelsInPlot, parallelismLabelBoxes, @@ -89,6 +103,14 @@ import { } from '@/components/inference/ui/line-label-layer'; import { QuickFiltersDialog } from '@/components/inference/ui/QuickFiltersDialog'; +const PowerTelemetryDialog = dynamic( + () => + import('@/components/inference/power-telemetry-dialog').then( + (module) => module.PowerTelemetryDialog, + ), + { ssr: false }, +); + const FixedSequenceLogDialog = dynamic(() => import('@/components/inference/log-viewer/fixed-sequence-log-dialog').then( (module) => module.FixedSequenceLogDialog, @@ -97,15 +119,22 @@ const FixedSequenceLogDialog = dynamic(() => const CHART_MARGIN = { top: 24, right: 10, bottom: 60, left: 60 }; +// Series id of an unofficial run: `overlay-run_`. Official series +// ids are `${date}_${hwKey}`, the `activeDates` key. +const OVERLAY_SERIES_PREFIX = 'overlay-run'; + // Label text combines the hw config (display label) and the date so // both dimensions of the GPU comparison view are legible on the chart, // not only the legend. Falls back to the raw hwKey if the config // lookup misses (legacy data). -function labelTextFor(pts: InferenceData[], numbering: Map): string { +function hardwareLabelFor(pts: InferenceData[]): string { const hwKey = String(pts[0].hwKey); const cfg = getInferenceHardwareConfig(hwKey, pts[0].model, pts); - const hwLabel = cfg ? getDisplayLabel(cfg) : hwKey; - return `${hwLabel} • ${comparisonEntryLabel(String(pts[0].date), numbering)}`; + return cfg ? getDisplayLabel(cfg) : hwKey; +} + +function labelTextFor(pts: InferenceData[], numbering: Map): string { + return `${hardwareLabelFor(pts)} • ${comparisonEntryLabel(String(pts[0].date), numbering)}`; } const GPU_STRINGS = { @@ -116,10 +145,6 @@ const GPU_STRINGS = { showAllMeasurements: 'Show all measurements', powerBoundaryInfo: 'Show only points on the upper measured power boundary. Turn off to show all measurements; the boundary stays the same. This is a power-load boundary, not an energy-efficiency frontier.', - powerCurves: - 'Smooth lines trace the upper power boundary across tested configurations. Dots are measured; lines are interpolated, not efficiency frontiers.', - powerOptimal: - 'A power Pareto frontier can contain a single point. Turn off Optimal Only to show the upper power boundary.', labels: 'Labels', parallelismLabels: 'Parallelism Labels', concurrencyLabels: '# Concurrent Sessions', @@ -130,6 +155,9 @@ const GPU_STRINGS = { noDataHint: 'Please change the model, sequence, precision, date range or chip selection.', noRoleEnergyDataHint: 'This dataset does not report role-level prefill/decode energy. Choose a different model, scenario, precision, date, or measured-energy metric.', + noMeasuredDataHint: + 'No measured GPU power is reported for this selection. Choose other chip configs, or a different model, scenario, precision or date.', + unofficialTitle: (branch: string) => `UNOFFICIAL: ${branch}`, }, zh: { logScale: '对数缩放', @@ -138,9 +166,6 @@ const GPU_STRINGS = { showAllMeasurements: '显示全部测量点', powerBoundaryInfo: '仅显示实测功率上边界上的点。关闭后显示全部测量点,边界曲线保持不变。这是功率负载边界,不是能效前沿。', - powerCurves: - '平滑曲线勾勒各测试配置的功耗上边界。数据点来自实测,曲线通过插值得到,不代表能效 Pareto 前沿。', - powerOptimal: '功耗的 Pareto 前沿可能只有一个点。关闭“仅最优”即可查看功耗上边界。', labels: '标签', parallelismLabels: '并行配置标签', concurrencyLabels: '并发会话数', @@ -151,6 +176,9 @@ const GPU_STRINGS = { noDataHint: '请调整模型、序列长度、精度、日期范围或芯片选项。', noRoleEnergyDataHint: '当前数据集未提供 Prefill/Decode 各角色的能耗数据。请选择其他模型、场景、精度、日期或实测能耗指标。', + noMeasuredDataHint: + '当前选择没有实测 GPU 功耗数据。请选择其他芯片配置,或更换模型、场景、精度或日期。', + unofficialTitle: (branch: string) => `非官方:${branch}`, }, } as const; @@ -163,6 +191,7 @@ const GPUGraph = React.memo( yLabel, chartDefinition, caption, + overlayData, runNumbering: providedRunNumbering, }: ScatterGraphProps) => { const { hardwareConfig } = useInferenceData(); @@ -208,23 +237,35 @@ const GPUGraph = React.memo( setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, } = useInferenceActions(); const locale = useLocale(); + const featureGateUnlocked = useFeatureGate(); + const showPowerTelemetry = + featureGateUnlocked || getMeasuredMetricConfig(selectedYAxisMetric) !== undefined; + const showPowerTelemetryRef = useRef(showPowerTelemetry); + showPowerTelemetryRef.current = showPowerTelemetry; const legendT = GPU_STRINGS[locale]; - const frontierDirection = chartDefinition[ - `${selectedYAxisMetric}_roofline` as keyof ChartDefinition - ] as ParetoDirection | undefined; + // The Concurrency axis plots observed load sweeps: no frontier or power + // envelope, same as ScatterGraph. + const isConcurrencyAxis = chartDefinition.x_scale_field === 'conc'; + const frontierDirection = isConcurrencyAxis + ? undefined + : (chartDefinition[`${selectedYAxisMetric}_roofline` as keyof ChartDefinition] as + | ParetoDirection + | undefined); const hideNonOptimal = Boolean(frontierDirection) && savedHideNonOptimal; const powerCurveMetric = isPowerCurveMetric(selectedYAxisMetric); const isMeasuredPowerAxis = isMeasuredPowerCurveMetric(selectedYAxisMetric); - const powerEnvelopeMode = powerCurveMetric && (isMeasuredPowerAxis || !hideNonOptimal); + const powerEnvelopeMode = + !isConcurrencyAxis && powerCurveMetric && (isMeasuredPowerAxis || !hideNonOptimal); const showAllMeasurements = isMeasuredPowerAxis ? !hideNonOptimal : savedShowAllMeasurements; - const isMeasuredEnergyAxis = isMeasuredEnergyConfigKey(selectedYAxisMetric); const noDataHint = isRoleLocalMeasuredEnergyConfigKey(selectedYAxisMetric) ? legendT.noRoleEnergyDataHint - : legendT.noDataHint; + : isMeasuredEnergyConfigKey(selectedYAxisMetric) + ? legendT.noMeasuredDataHint + : legendT.noDataHint; const ephemeralUrlState = useEphemeralUrlState(); - const { resolvedTheme } = useTheme(); const chartRef = useRef(null); const [quickFiltersOpen, setQuickFiltersOpen] = useState(false); // A framework lock (embed routes) is not a user filter, so it is not counted. @@ -233,6 +274,7 @@ const GPUGraph = React.memo( (lockedFrameworks ? 0 : quickFilters.frameworks.length) + quickFilters.deployment.length + quickFilters.power.length + + (quickFilters.topologies?.length ?? 0) + (selectedSequence === Sequence.AgenticTraces ? 0 : quickFilters.spec.length); const clearQuickFilters = useCallback(() => { setQuickFilterVendors([]); @@ -240,37 +282,49 @@ const GPUGraph = React.memo( setQuickFilterDeployment([]); setQuickFilterSpec([]); setQuickFilterPower([]); + setQuickFilterTopologies([]); }, [ setQuickFilterVendors, setQuickFilterFrameworks, setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, ]); - // Shared date+GPU pairs. `dates` holds comparison-series entries (plain dates - // and/or specific-run entries); a same-day range endpoint is dropped when that - // date also has run entries (resolveComparisonEntries), then sorted earliest → - // latest so a day's runs read #1 → #N. - const gpuDatePairs = useMemo(() => { - const deduplicated = resolveComparisonEntries(selectedDates, selectedDateRange); - deduplicated.sort((a, b) => { - const [ta, ia] = comparisonEntrySortValue(a); - const [tb, ib] = comparisonEntrySortValue(b); - return ta - tb || ia - ib; - }); - const sortedGPUs = [...selectedGPUs].toSorted( - (a, b) => getModelSortIndex(a) - getModelSortIndex(b) || a.localeCompare(b), - ); - return { dates: deduplicated, sortedGPUs }; - }, [selectedDateRange, selectedDates, selectedGPUs]); + const { runNumbering, allGraphs, paletteIdentity, resolveColor, getCssColor } = + useComparisonSeries(providedRunNumbering); - // Run numbers for legend/line labels. Prefer the stable numbering passed by - // the parent (shared with the changelog, so labels match it and removed runs - // leave a gap); fall back to gap-free numbering of the on-chart series. - const runNumbering = useMemo( - () => providedRunNumbering ?? buildRunNumbering(gpuDatePairs.dates), - [providedRunNumbering, gpuDatePairs.dates], + // Unofficial runs stay on this chart as their own (run, chip config) + // series in the run's colour, next to the compared dates. Same gates as + // ScatterGraph: precision, quick filters and overlay hardware selection; + // dismissing a run removes its rows from `overlayData`. + const { runIndexByUrl, unofficialRunInfos, activeOverlayHwTypes } = useUnofficialRun(); + const overlayPoints = useMemo( + () => + (overlayData?.data ?? []).filter( + (point) => + // Boundary / role siblings are a same-run comparison (see ChartDisplay). + !point.powerVariant && + selectedPrecisions.includes(point.precision) && + matchesQuickFilters(point, quickFilters) && + activeOverlayHwTypes.has(String(point.hwKey)), + ), + [overlayData, selectedPrecisions, quickFilters, activeOverlayHwTypes], + ); + const overlayPointSet = useMemo(() => new Set(overlayPoints), [overlayPoints]); + const overlayRunOf = useCallback( + (point: InferenceData) => overlayRunIndex(point.run_url ?? null, runIndexByUrl), + [runIndexByUrl], + ); + const overlaySeriesId = useCallback( + (point: InferenceData) => `${OVERLAY_SERIES_PREFIX}${overlayRunOf(point)}_${point.hwKey}`, + [overlayRunOf], + ); + const seriesIdOf = useCallback( + (point: InferenceData) => + overlayPointSet.has(point) ? overlaySeriesId(point) : `${point.date}_${point.hwKey}`, + [overlayPointSet, overlaySeriesId], ); // Removing a series from the legend should also drop it from the comparison @@ -290,94 +344,24 @@ const GPUGraph = React.memo( [selectedGPUs, selectedDates, setSelectedDates, removeActiveDate], ); - const graphIdentifiers = useMemo(() => { - const ids: string[] = []; - gpuDatePairs.sortedGPUs.forEach((gpu) => - gpuDatePairs.dates.forEach((date) => ids.push(`${date}_${gpu}`)), - ); - return ids; - }, [gpuDatePairs]); - - // High contrast keys off the GPU (not `date_gpu`) so each hardware config - // gets exactly one hue; the dates within a config are separated by the - // lightness ramp built below rather than by unrelated hues. - const { resolveColor, getCssColor } = useThemeColors({ - highContrast, - identifiers: graphIdentifiers, - hcKeys: gpuDatePairs.sortedGPUs, - }); - - // Dynamic GPU×date color map - const gpuDateColorMap = useMemo(() => { - const { dates, sortedGPUs } = gpuDatePairs; - if (sortedGPUs.length === 0 || dates.length === 0) return {}; - const theme = isDarkTheme(resolvedTheme) ? 'dark' : 'light'; - return generateGpuDateColors(sortedGPUs, dates.length, theme); - }, [gpuDatePairs, resolvedTheme]); - - // High-contrast GPU×date color map: one iwanthue hue per GPU, ramped across - // the compared dates so a config's runs stay recognisably the same color - // while still reading oldest → newest. - const hcGpuDateColorMap = useMemo(() => { - const { dates, sortedGPUs } = gpuDatePairs; - if (!highContrast || sortedGPUs.length === 0 || dates.length === 0) return {}; - const theme = isDarkTheme(resolvedTheme) ? 'dark' : 'light'; - const baseColors: Record = {}; - for (const gpu of sortedGPUs) baseColors[gpu] = getCssColor(resolveColor(gpu)); - return generateHighContrastGpuDateColors(baseColors, dates.length, theme); - }, [gpuDatePairs, highContrast, resolvedTheme, resolveColor, getCssColor]); - - const allGraphs = useMemo(() => { - const { dates, sortedGPUs } = gpuDatePairs; - const result: { date: string; color: string; hwKey: string; id: string }[] = []; - sortedGPUs.forEach((gpu) => { - dates.forEach((date, dateIndex) => { - const id = `${date}_${gpu}`; - const compositeKey = `${dateIndex}_${gpu}`; - const dynamicColor = gpuDateColorMap[compositeKey]; - result.push({ - date, - hwKey: gpu, - id, - color: highContrast - ? hcGpuDateColorMap[compositeKey] || getCssColor(resolveColor(gpu)) - : dynamicColor || 'var(--foreground)', - }); - }); + const groupedData = useMemo(() => { + const groups: Record = {}; + const add = (point: InferenceData) => { + const key = `${seriesIdOf(point)}_${point.precision}`; + (groups[key] ??= []).push(point); + }; + data.forEach((point) => { + if (selectedPrecisions.includes(point.precision)) add(point); }); - return result; - }, [gpuDatePairs, gpuDateColorMap, hcGpuDateColorMap, highContrast, resolveColor, getCssColor]); - - const paletteIdentity = useMemo( - () => - [ - resolvedTheme ?? 'system', - highContrast ? 'high-contrast' : 'standard', - ...allGraphs.map(({ id, color }) => `${id}:${color}`), - ].join('|'), - [resolvedTheme, highContrast, allGraphs], - ); - - const groupedData = useMemo( - () => - data.reduce( - (acc, point) => { - if (!selectedPrecisions.includes(point.precision)) return acc; - const key = `${point.date}_${point.hwKey}_${point.precision}`; - if (!acc[key]) acc[key] = []; - acc[key].push(point); - return acc; - }, - {} as Record, - ), - [data, selectedPrecisions], - ); + overlayPoints.forEach(add); + return groups; + }, [data, selectedPrecisions, overlayPoints, seriesIdOf]); // Track which date+GPU combos have actual data points const idsWithData = useMemo(() => { const ids = new Set(); for (const key of Object.keys(groupedData)) { - // key = "date_hwKey_precision" — strip last segment + // key = "seriesId_precision" — strip last segment const lastUnderscore = key.lastIndexOf('_'); ids.add(key.slice(0, lastUnderscore)); } @@ -395,62 +379,90 @@ const GPUGraph = React.memo( }, [groupedData, frontierDirection]); const rooflines = useMemo(() => { + // One path per observed load sweep, never joined across runs or + // topologies (see groupConcurrencySeries). + if (isConcurrencyAxis) { + const result: Record = {}; + for (const [key, points] of Object.entries(groupedData)) { + for (const [segment, sweep] of groupConcurrencySeries(points)) { + result[`${key}__${encodeURIComponent(segment)}`] = sweep; + } + } + return result; + } if (!powerEnvelopeMode) return paretoRooflines; const result: Record = {}; for (const [key, points] of Object.entries(groupedData)) { - result[key] = upperPowerEnvelope(points, chartDefinition.chartType !== 'e2e'); + result[key] = upperPowerEnvelope( + points, + chartDefinition.chartType !== 'e2e', + isPowerGaugeSeries(selectedYAxisMetric, points[0]), + ); } return result; - }, [powerEnvelopeMode, groupedData, paretoRooflines, chartDefinition.chartType]); + }, [ + isConcurrencyAxis, + powerEnvelopeMode, + groupedData, + paretoRooflines, + chartDefinition.chartType, + selectedYAxisMetric, + ]); + const boundaryKeyOf = useCallback( + (p: InferenceData) => `${seriesIdOf(p)}_${p.precision}-${p.x}-${p.y}`, + [seriesIdOf], + ); const boundaryPointKeys = useMemo(() => { const keys = new Set(); - Object.values(rooflines).forEach((pts) => - pts.forEach((p) => keys.add(`${p.date}_${p.hwKey}_${p.precision}-${p.x}-${p.y}`)), - ); + Object.values(rooflines).forEach((pts) => pts.forEach((p) => keys.add(boundaryKeyOf(p)))); return keys; - }, [rooflines]); + }, [rooflines, boundaryKeyOf]); + // Unofficial runs are not date series: the `activeDates` toggles leave them on. const activeData = useMemo( () => Object.values(groupedData) .flat() - .filter((p) => activeDates.has(`${p.date}_${p.hwKey}`)), - [groupedData, activeDates], + .filter((p) => overlayPointSet.has(p) || activeDates.has(`${p.date}_${p.hwKey}`)), + [groupedData, activeDates, overlayPointSet], ); const filteredData = useMemo(() => { if (hideNonOptimal || (powerEnvelopeMode && !showAllMeasurements)) - return activeData.filter((p) => - boundaryPointKeys.has(`${p.date}_${p.hwKey}_${p.precision}-${p.x}-${p.y}`), - ); + return activeData.filter((p) => boundaryPointKeys.has(boundaryKeyOf(p))); return activeData; - }, [activeData, hideNonOptimal, powerEnvelopeMode, showAllMeasurements, boundaryPointKeys]); + }, [ + activeData, + hideNonOptimal, + powerEnvelopeMode, + showAllMeasurements, + boundaryPointKeys, + boundaryKeyOf, + ]); + // Official points join the scatter layer; unofficial ones draw as X markers. + const officialPoints = useMemo( + () => filteredData.filter((point) => !overlayPointSet.has(point)), + [filteredData, overlayPointSet], + ); + const visibleOverlayPoints = useMemo( + () => filteredData.filter((point) => overlayPointSet.has(point)), + [filteredData, overlayPointSet], + ); // Keep domains fixed so revealing off-boundary dots cannot move power curves. const scaleData = powerEnvelopeMode ? activeData : filteredData; - const powerTierCounts = useMemo( - () => ({ - total: countPowerTiers( - data.filter((point) => selectedPrecisions.includes(point.precision)), - ), - visible: countPowerTiers(filteredData), - }), - [data, filteredData, selectedPrecisions], - ); - - // GPU comparison currently renders official DB-backed points only. Unofficial - // overlays have no benchmark_results id or persisted trace, so they cannot - // open the dedicated per-point charts route. + // Only official DB-backed points have a benchmark_results id, a persisted + // trace and logs; unofficial overlays cannot open those routes. const agenticIds = useMemo( () => - filteredData.flatMap((point) => + officialPoints.flatMap((point) => point.benchmark_type === 'agentic_traces' && isPersistedBenchmarkId(point.id) ? [point.id] : [], ), - [filteredData], + [officialPoints], ); const { data: traceAvailability } = useTraceAvailability(agenticIds); const traceAvailabilityRef = useRef(traceAvailability); @@ -459,18 +471,19 @@ const GPUGraph = React.memo( // Log availability applies to every persisted official point in the // comparison, including fixed-sequence runs. const persistedPointIds = useMemo( - () => filteredData.flatMap((point) => (isPersistedBenchmarkId(point.id) ? [point.id] : [])), - [filteredData], + () => officialPoints.flatMap((point) => (isPersistedBenchmarkId(point.id) ? [point.id] : [])), + [officialPoints], ); const { data: logAvailability } = useLogAvailability(persistedPointIds); const logAvailabilityRef = useRef(logAvailability); logAvailabilityRef.current = logAvailability; const [fixedLogPointId, setFixedLogPointId] = useState(null); + const [powerTelemetryPoint, setPowerTelemetryPoint] = useState(null); - // Warning annotations for visible series with known upstream issues — - // same treatment the scatter view gets, applied to the date-comparison view. - // Lines here are colored per (gpu, date) pair, so take the first active - // pair's color as the series swatch. + // Warning annotations for visible series (official and unofficial) with + // known upstream issues — same treatment the scatter view gets. Lines here + // are colored per (gpu, date) pair, so take the first active pair's color + // as the series swatch. Official-preview notices follow official data only. const knownIssueAnnotations = useMemo((): KnownIssueAnnotation[] => { const annotations: KnownIssueAnnotation[] = matchKnownConfigIssues( modelLabel, @@ -490,7 +503,7 @@ const GPUGraph = React.memo( }; }); for (const previewConfig of OFFICIAL_PREVIEW_SERIES) { - const previewPoints = filteredData.filter((point) => + const previewPoints = officialPoints.filter((point) => hardwareKeyMatchesAnyBase(String(point.hwKey), previewConfig.baseGpuKeys), ); if (previewPoints.length === 0) continue; @@ -512,7 +525,16 @@ const GPUGraph = React.memo( }); } return annotations; - }, [modelLabel, filteredData, allGraphs, activeDates, resolveColor, getCssColor, locale]); + }, [ + modelLabel, + filteredData, + officialPoints, + allGraphs, + activeDates, + resolveColor, + getCssColor, + locale, + ]); const knownIssueLayer = useMemo( () => @@ -557,13 +579,16 @@ const GPUGraph = React.memo( return [yMin, yExtent[1] * 1.05] as [number, number]; }, [scaleData, logScale]); + const pointIdentity = useCallback( + (point: InferenceData) => + overlayPointSet.has(point) + ? `overlay:${overlaySeriesId(point)}:${scatterPointConfigId(point)}` + : `${point.date}:${scatterPointConfigId(point)}`, + [overlayPointSet, overlaySeriesId], + ); const dataIdentity = useMemo( - () => - filteredData - .map((point) => `${point.date}:${scatterPointConfigId(point)}`) - .toSorted() - .join('|'), - [filteredData], + () => filteredData.map(pointIdentity).toSorted().join('|'), + [filteredData, pointIdentity], ); // Tooltip-only trace availability is deliberately excluded from chart // identity; the long-lived D3 content callback reads its latest value via @@ -577,9 +602,7 @@ const GPUGraph = React.memo( powerEnvelopeMode ? 'power-envelope' : 'pareto-curves', `linear:${xExtent.join(',')}`, `${logScale ? 'log' : 'linear'}:${yDomain.join(',')}`, - ...filteredData.map( - (point) => `${point.date}:${scatterPointConfigId(point)}:${point.x}:${point.y}`, - ), + ...filteredData.map((point) => `${pointIdentity(point)}:${point.x}:${point.y}`), ] .toSorted() .join('|'), @@ -592,18 +615,20 @@ const GPUGraph = React.memo( logScale, yDomain, filteredData, + pointIdentity, ], ); - // Color resolver for points/rooflines + // Color resolver for points/rooflines; an unofficial run keeps its run color. const getColor = useMemo( () => (d: InferenceData) => { + if (overlayPointSet.has(d)) return overlayRunColor(overlayRunOf(d)); const graphIndex = allGraphs.findIndex( ({ date, hwKey }) => d.date === date && d.hwKey === hwKey, ); return graphIndex === -1 ? '#6b7280' : allGraphs[graphIndex].color; }, - [allGraphs], + [allGraphs, overlayPointSet, overlayRunOf], ); const getRooflineColor = useMemo( @@ -617,9 +642,21 @@ const GPUGraph = React.memo( const isRooflineVisible = useMemo( () => (key: string) => { const point = rooflines[key]?.[0]; - return point !== undefined && activeDates.has(`${point.date}_${point.hwKey}`); + if (point === undefined) return false; + return overlayPointSet.has(point) || activeDates.has(`${point.date}_${point.hwKey}`); + }, + [activeDates, rooflines, overlayPointSet], + ); + + // Unofficial-run curves keep the run's dash, as in ScatterGraph. + const getRooflineDasharray = useCallback( + (key: string) => { + const point = rooflines[key]?.[0]; + return point && overlayPointSet.has(point) + ? overlayRooflineDasharray(overlayRunOf(point)) + : null; }, - [activeDates, rooflines], + [rooflines, overlayPointSet, overlayRunOf], ); // ── Line labels (date along each roofline) ── @@ -635,16 +672,32 @@ const GPUGraph = React.memo( >(); for (const [key, points] of Object.entries(rooflines)) { if (points.length < 2 || !isRooflineVisible(key)) continue; - const graphId = `${points[0].date}_${points[0].hwKey}`; + const graphId = seriesIdOf(points[0]); const previous = bestByGraph.get(graphId); if (!previous || points.length > previous.points.length) { bestByGraph.set(graphId, { key, graphId, points }); } } + // Runs drawing the same hardware need a run tag on their pills. + const overlayRunsByHw = new Map>(); + for (const { points } of bestByGraph.values()) { + if (!overlayPointSet.has(points[0])) continue; + const hwKey = String(points[0].hwKey); + const runs = overlayRunsByHw.get(hwKey) ?? new Set(); + overlayRunsByHw.set(hwKey, runs.add(overlayRunOf(points[0]))); + } + const overlayLabelFor = (points: InferenceData[]) => { + const info = unofficialRunInfos[overlayRunOf(points[0])]; + const hardwareLabel = hardwareLabelFor(points); + const sharesHardware = (overlayRunsByHw.get(String(points[0].hwKey))?.size ?? 0) > 1; + return info ? getOverlayLineLabel(hardwareLabel, info, sharesHardware) : hardwareLabel; + }; return [...bestByGraph.values()].map(({ key, graphId, points }) => ({ key, seriesId: graphId, - label: labelTextFor(points, runNumbering), + label: overlayPointSet.has(points[0]) + ? overlayLabelFor(points) + : labelTextFor(points, runNumbering), color: getRooflineColor(key), points, })); @@ -715,13 +768,22 @@ const GPUGraph = React.memo( chartDefinition.chartType, runNumbering, paletteIdentity, + seriesIdOf, + overlayPointSet, + overlayRunOf, + unofficialRunInfos, ]); - // Dismiss tooltip when pinned point's combo is hidden + // Dismiss tooltip when pinned point's series is hidden useEffect(() => { - const pp = chartRef.current?.getPinnedPoint() as InferenceData | null; - if (pp && !activeDates.has(`${pp.date}_${pp.hwKey}`)) chartRef.current?.dismissTooltip(); - }, [activeDates]); + const handle = chartRef.current; + const pp = handle?.getPinnedPoint() as InferenceData | null; + if (!pp) return; + const visible = handle?.getPinnedPointIsOverlay() + ? activeOverlayHwTypes.has(String(pp.hwKey)) + : activeDates.has(`${pp.date}_${pp.hwKey}`); + if (!visible) handle?.dismissTooltip(); + }, [activeDates, activeOverlayHwTypes]); // Dismiss on filter changes useEffect(() => { @@ -733,6 +795,167 @@ const GPUGraph = React.memo( selectedDates, selectedDateRange, showAllMeasurements, + overlayData, + ]); + + // One legend group per unofficial run (grouped legends split on the first + // word of `name`), one row per chip config the run draws. Runs are + // dismissed from the banner, not the legend. + const overlayLegendItems = useMemo(() => { + const bySeries = new Map(); + for (const point of overlayPoints) { + const id = overlaySeriesId(point); + bySeries.set(id, [...(bySeries.get(id) ?? []), point]); + } + return [...bySeries].map(([id, points]) => { + const runIndex = overlayRunOf(points[0]); + const info = unofficialRunInfos[runIndex]; + const branch = info?.branch || (info ? `run ${info.id}` : id); + return { + name: `unofficial-run-${info?.id ?? runIndex} ${points[0].hwKey}`, + hw: id, + label: getInferenceRunLabel(`✕ ${hardwareLabelFor(points)}`, points), + color: overlayRunColor(runIndex), + title: legendT.unofficialTitle(branch), + isActive: true, + isRemovable: false, + onClick: () => {}, + }; + }); + }, [overlayPoints, overlaySeriesId, overlayRunOf, unofficialRunInfos, legendT]); + + // ── Unofficial-run points: ScatterGraph's X markers in the run color. The + // run curves go through the roofline layer; index keys let repeated + // observations of one config all draw. ── + const attachPowerTraceAction = usePowerTraceAction(chartRef); + const overlayPointsLayer: CustomLayerConfig = useMemo(() => { + const updateLabels = (zoomGroup: d3.Selection) => { + zoomGroup + .selectAll('.unofficial-overlay-pt') + .each(function (d) { + const lines = pointLabelText(d, useAdvancedLabels, showConcurrencyLabels).split('\n'); + d3.select(this) + .selectAll('.overlay-label') + .data([true]) + .join('text') + .attr('class', 'overlay-label') + .attr('text-anchor', 'middle') + .attr('font-size', px(CHART_TYPE.dataLabel)) + .attr('font-weight', '700') + .attr('pointer-events', 'none') + .style('fill', 'var(--foreground)') + .style('display', showPointLabels ? '' : 'none') + .selectAll('tspan') + .data(lines) + .join('tspan') + .attr('x', 0) + .attr('dy', (_line, i) => + i === 0 ? `${-(1 + (lines.length - 1) * 1.1)}em` : '1.1em', + ) + .text((line) => line); + }); + }; + return { + type: 'custom', + key: 'overlay-points', + displayIdentity: `labels:${showPointLabels}`, + render: (zoomGroup, ctx) => { + const xScale = ctx.xScale as ContinuousScale; + const yScale = ctx.yScale as ContinuousScale; + const marks = zoomGroup + .selectAll('.unofficial-overlay-pt') + .data(visibleOverlayPoints, (_d, i) => String(i)) + .join((enter) => { + const g = enter.append('g').attr('class', 'unofficial-overlay-pt'); + g.append('circle') + .attr('r', HIT_AREA_RADIUS) + .attr('fill', 'transparent') + .attr('cursor', 'pointer'); + g.append('path') + .attr('class', 'visible-shape overlay-x') + .attr('d', xMarkerPath(5, 0.7)) + .attr('fill', 'none') + .attr('stroke-width', 2.5) + .attr('stroke-linecap', 'round') + .attr('cursor', 'pointer'); + return g; + }); + marks.attr('transform', (d) => `translate(${xScale(d.x)},${yScale(d.y)})`); + marks.select('.overlay-x').attr('stroke', (d) => overlayRunColor(overlayRunOf(d))); + marks.each(function (d) { + renderOffloadHalo(d3.select(this), d, overlayRunColor(overlayRunOf(d))); + }); + updateLabels(zoomGroup); + + const container = ctx.layout.svg.node()!.parentElement as HTMLDivElement; + const tooltip = d3.select(ctx.tooltipElement); + attachOverlayXMarkerHandlers(marks, { + markerSelector: '.overlay-x', + normalPath: xMarkerPath(5, 0.7), + hoverPath: xMarkerPath(7, 0.7), + tooltip, + handle: chartRef.current, + content: (point, pinned) => + overlayData + ? generateOverlayTooltipContent({ + data: point, + isPinned: pinned, + xLabel, + yLabel, + selectedYAxisMetric, + hardwareConfig: overlayData.hardwareConfig, + overlayData, + locale, + }) + : '', + position: (event) => { + const [mouseX, mouseY] = d3.pointer(event, container); + return computeTooltipPosition(mouseX, mouseY, tooltip, container); + }, + rulers: { + show: (point, marker) => { + const position = overlayMarkerPosition(marker) ?? { + x: xScale(point.x), + y: yScale(point.y), + }; + zoomGroup.select('.ruler-group').style('display', 'block'); + zoomGroup.select('.vertical-ruler').attr('x1', position.x).attr('x2', position.x); + zoomGroup.select('.horizontal-ruler').attr('y1', position.y).attr('y2', position.y); + }, + hide: () => zoomGroup.select('.ruler-group').style('display', 'none'), + }, + onClick: (point) => { + track('gpu_timeseries_data_point_clicked', { + hw: String(point.hwKey), + x: point.x, + y: point.y, + overlay: true, + }); + attachPowerTraceAction(ctx.tooltipElement, point, true); + }, + }); + }, + onDisplayUpdate: updateLabels, + onZoom: (zoomGroup, ctx) => { + const xScale = ctx.newXScale as ContinuousScale; + const yScale = ctx.newYScale as ContinuousScale; + zoomGroup + .selectAll('.unofficial-overlay-pt') + .attr('transform', (d) => `translate(${xScale(d.x)},${yScale(d.y)})`); + }, + }; + }, [ + visibleOverlayPoints, + overlayRunOf, + overlayData, + showPointLabels, + useAdvancedLabels, + showConcurrencyLabels, + xLabel, + yLabel, + selectedYAxisMetric, + locale, + attachPowerTraceAction, ]); // Hover dimming animates via the inline `transition: opacity 150ms ease` @@ -740,29 +963,35 @@ const GPUGraph = React.memo( // d3 `.transition()` here would re-write opacity every animation frame, // each write restarting the CSS transition (transitionrun/cancel per node // per frame). Same rationale as ScatterGraph's hover handlers. - const handleLegendHover = useCallback((seriesId: string) => { - const svg = chartRef.current?.getSvgElement?.(); - if (!svg) return; - const root = d3.select(svg); - root - .selectAll('.dot-group') - .style('opacity', (d) => (`${d.date}_${d.hwKey}` === seriesId ? 1 : 0.15)); - root.selectAll('.roofline-path').style('opacity', function () { - const point = (d3.select(this).datum() as { points: InferenceData[] } | null)?.points[0]; - const series = point ? `${point.date}_${point.hwKey}` : ''; - return series === seriesId ? null : '0.15'; - }); - }, []); + const handleLegendHover = useCallback( + (seriesId: string) => { + const svg = chartRef.current?.getSvgElement?.(); + if (!svg) return; + const root = d3.select(svg); + root + .selectAll('.dot-group') + .style('opacity', (d) => (`${d.date}_${d.hwKey}` === seriesId ? 1 : 0.15)); + root + .selectAll('.unofficial-overlay-pt') + .style('opacity', (d) => (overlaySeriesId(d) === seriesId ? 1 : 0.15)); + root.selectAll('.roofline-path').style('opacity', function () { + const point = (d3.select(this).datum() as { points: InferenceData[] } | null)?.points[0]; + const series = point ? seriesIdOf(point) : ''; + return series === seriesId ? null : '0.15'; + }); + }, + [overlaySeriesId, seriesIdOf], + ); const handleLegendHoverEnd = useCallback(() => { const svg = chartRef.current?.getSvgElement?.(); if (!svg) return; const root = d3.select(svg); - root.selectAll('.dot-group').style('opacity', null); + root.selectAll('.dot-group, .unofficial-overlay-pt').style('opacity', null); root.selectAll('.roofline-path').style('opacity', null); }, []); - if (data.length === 0) { + if (data.length === 0 && overlayPoints.length === 0) { return (
@@ -803,7 +1032,7 @@ const GPUGraph = React.memo( ); } - return ( + const chart = ( ref={chartRef} // Embeds drop the zoom/pan hint line; the host page has its own caption. @@ -812,33 +1041,12 @@ const GPUGraph = React.memo( dataIdentity={dataIdentity} metricIdentity={metricIdentity} displayIdentity={`${showPointLabels}:${paletteIdentity}:${selectedPrecisions.join(',')}`} - data={filteredData} + data={officialPoints} margin={CHART_MARGIN} watermark={getChartWatermark()} testId="gpu-graph" grabCursor={true} - caption={ - isMeasuredEnergyAxis || powerCurveMetric ? ( - <> - {caption} - {isMeasuredEnergyAxis && ( - - )} - {powerCurveMetric && ( -

- {powerEnvelopeMode ? legendT.powerCurves : legendT.powerOptimal} -

- )} - - ) : ( - caption - ) - } + caption={caption} xScale={{ type: 'linear', domain: xExtent, nice: true }} yScale={{ type: logScale ? 'log' : 'linear', domain: yDomain, nice: true }} xAxis={{ @@ -859,13 +1067,15 @@ const GPUGraph = React.memo( config: { getColor: getRooflineColor, isVisible: isRooflineVisible, - curve: d3.curveMonotoneX, + getDasharray: getRooflineDasharray, + // Load sweeps join measured points; they are not fitted curves. + curve: isConcurrencyAxis ? d3.curveLinear : d3.curveMonotoneX, }, }, { type: 'scatter', key: 'points', - data: filteredData, + data: officialPoints, config: { getColor, hideLabels: !showPointLabels, @@ -882,6 +1092,7 @@ const GPUGraph = React.memo( }, keyFn: (point) => `${point.date}:${scatterPointConfigId(point)}`, }, + overlayPointsLayer, lineLabelLayer, knownIssueLayer, ]} @@ -914,6 +1125,7 @@ const GPUGraph = React.memo( yLabel, selectedYAxisMetric, hardwareConfig, + showPowerTelemetry: showPowerTelemetryRef.current, runUrl: d.run_url ? updateRepoUrl(d.run_url) : undefined, hasTrace: isPersistedBenchmarkId(d.id) ? traceAvailabilityRef.current?.[d.id] === true @@ -961,6 +1173,19 @@ const GPUGraph = React.memo( }); }); } + const powerBtn = tooltipEl.querySelector('[data-action="view-power-telemetry"]'); + if (powerBtn && isPersistedBenchmarkId(d.id)) { + powerBtn.addEventListener('click', (event) => { + event.stopPropagation(); + setPowerTelemetryPoint(d); + chartRef.current?.dismissTooltip(); + track('inference_power_telemetry_opened', { + id: d.id, + hwKey: d.hwKey, + conc: d.conc, + }); + }); + } const logsBtn = tooltipEl.querySelector('[data-action="view-logs"]'); if (logsBtn && typeof d.id === 'number') { logsBtn.addEventListener('click', (event) => { @@ -978,6 +1203,7 @@ const GPUGraph = React.memo( }); }); } + attachPowerTraceAction(tooltipEl, d, false); // Pinning updates D3Chart's React state. GPU comparison rebuilds // several inline layer configs on that render, whose cleanup can // briefly hide the otherwise-pinned portal tooltip. Restore its @@ -1018,21 +1244,15 @@ const GPUGraph = React.memo( // CSS transitions for smooth opacity animation on legend hover — // the hover handlers write opacity once and let these animate. ctx.layout.zoomGroup - .selectAll('.dot-group, .roofline-path') + .selectAll('.dot-group, .unofficial-overlay-pt, .roofline-path') .style('transition', 'opacity 150ms ease'); - // Decorations stay inside the point group, so normal zoom transforms - // carry them without a separate update pass. + // The offload halo stays inside the point group, so normal zoom + // transforms carry it without a separate update pass. ctx.layout.zoomGroup .selectAll('.dot-group') .each(function (point) { renderOffloadHalo(d3.select(this), point, 'var(--foreground)'); - renderLegacyPowerRing( - d3.select(this), - point, - isMeasuredEnergyAxis, - 'var(--foreground)', - ); }); }} legendElement={ @@ -1043,20 +1263,23 @@ const GPUGraph = React.memo( onItemHover={handleLegendHover} onItemHoverEnd={handleLegendHoverEnd} onItemRemove={handleLegendRemove} - legendItems={allGraphs - .filter(({ id }) => idsWithData.has(id)) - .map(({ date, color, hwKey, id }) => ({ - name: `${hwKey} ${comparisonEntryLabel(date, runNumbering)}`, - hw: id, - label: comparisonEntryLabel(date, runNumbering), - color, - title: getDisplayLabel(getHardwareConfig(hwKey, modelLabel)), - isActive: activeDates.has(id), - onClick: () => { - toggleActiveDate(id); - track('interactivity_date_toggled', { date, hw: hwKey }); - }, - }))} + legendItems={[ + ...overlayLegendItems, + ...allGraphs + .filter(({ id }) => idsWithData.has(id)) + .map(({ date, color, hwKey, id }) => ({ + name: `${hwKey} ${comparisonEntryLabel(date, runNumbering)}`, + hw: id, + label: comparisonEntryLabel(date, runNumbering), + color, + title: getDisplayLabel(getHardwareConfig(hwKey, modelLabel)), + isActive: activeDates.has(id), + onClick: () => { + toggleActiveDate(id); + track('interactivity_date_toggled', { date, hw: hwKey }); + }, + })), + ]} isLegendExpanded={isLegendExpanded} onExpandedChange={(expanded) => { setIsLegendExpanded(expanded); @@ -1195,6 +1418,21 @@ const GPUGraph = React.memo( } /> ); + + return ( + <> + {powerTelemetryPoint === null ? null : ( + { + if (!open) setPowerTelemetryPoint(null); + }} + /> + )} + {chart} + + ); }, ); diff --git a/packages/app/src/components/inference/ui/InferenceTable.test.ts b/packages/app/src/components/inference/ui/InferenceTable.test.ts index bf9ab1b0c..e40a43589 100644 --- a/packages/app/src/components/inference/ui/InferenceTable.test.ts +++ b/packages/app/src/components/inference/ui/InferenceTable.test.ts @@ -1,7 +1,12 @@ import { describe, it, expect } from 'vitest'; +import { createElement } from 'react'; +import { renderToStaticMarkup } from 'react-dom/server'; import type { ChartDefinition, InferenceData } from '@/components/inference/types'; -import { formatInferenceTableNumber } from '@/components/inference/ui/InferenceTable'; +import InferenceTable, { + formatInferenceTableNumber, +} from '@/components/inference/ui/InferenceTable'; +import { expandPowerCompareSeries } from '../utils/power-compare'; import * as inferenceTableModule from './InferenceTable'; import { chartDefinitions } from '../metric-registry'; @@ -42,6 +47,54 @@ function makePoint(overrides: Partial): InferenceData { } describe('InferenceTable sorting logic', () => { + it.each(['roles', 'boundaries'] as const)( + 'renders and sorts each %s comparison by its plotted value', + (mode) => { + const base = makePoint({ + hwKey: 'gb300_dynamo-trt', + y: 708.1, + measuredAvgPower: { y: 708.1, roof: false }, + measuredPrefillAvgPower: { y: 760.442, roof: false }, + measuredDecodeAvgPower: { y: 690.652, roof: false }, + gpuProvisionedWatts: { y: 1400, roof: false }, + utilityProvisionedWatts: { y: 1920, roof: false }, + }); + const points = expandPowerCompareSeries([base], 'y_measuredAvgPower', mode); + const sorted = sortRowsByYMetric(points, chartDefinitions[0], 'y_measuredAvgPower'); + expect(sorted.map((point) => point.y)).toEqual( + mode === 'roles' ? [690.652, 708.1, 760.442] : [708.1, 1400, 1920], + ); + + const html = renderToStaticMarkup( + createElement(InferenceTable, { + data: points, + chartDefinition: chartDefinitions[0], + selectedYAxisMetric: 'y_measuredAvgPower', + }), + ); + const body = html.split('')[1].split('')[0]; + const cells = [...body.matchAll(/]*>(?.*?)<\/tr>/gu)].map((match) => + [...match.groups!.row.matchAll(/]*>(?.*?)<\/td>/gu)].map( + (cell) => cell.groups!.cell, + ), + ); + expect(cells.map((row) => [row[2], row[3]])).toEqual( + mode === 'roles' + ? [ + ['Decode GPUs', '691'], + ['All GPUs', '708'], + ['Prefill GPUs', '760'], + ] + : [ + ['GPU Level Measured', '708'], + ['GPU Level Provisioned (TDP)', '1,400'], + ['All in Provisioned', '1,920'], + ], + ); + expect(points.every((point) => point.measuredAvgPower?.y === 708.1)).toBe(true); + }, + ); + it('sorts supported modeled estimates by ascending power', () => { const definition = chartDefinitions[0]; const metric = 'y_modeledChassisPowerPerGpu'; diff --git a/packages/app/src/components/inference/ui/InferenceTable.tsx b/packages/app/src/components/inference/ui/InferenceTable.tsx index b6cdafa7e..0d4e8ab74 100644 --- a/packages/app/src/components/inference/ui/InferenceTable.tsx +++ b/packages/app/src/components/inference/ui/InferenceTable.tsx @@ -7,6 +7,7 @@ import { type DataTableColumn, DataTable } from '@/components/ui/data-table'; import { chipCounts } from '@/lib/chip-counts'; import { getNestedYValue, metricLabel, xAxisLabel } from '@/lib/chart-utils'; import { isModeledSystemPowerConfigKey } from '@/components/inference/metric-registry'; +import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare'; import { sortRowsByYMetric } from '@/components/inference/ui/inference-table-sort'; import { type Precision, getPrecisionLabel } from '@/lib/data-mappings'; import { getDisplayLabel } from '@/lib/utils'; @@ -43,6 +44,7 @@ export function inferenceTableHeaderLabels( physicalChips: locale === 'zh' ? '物理芯片数' : 'Physical Chips', configuredChips: locale === 'zh' ? '配置中的芯片数' : 'Configured Chip Count', concurrency: locale === 'zh' ? '并发数' : 'Conc', + series: locale === 'zh' ? '系列' : 'Series', yMetric: metricLabel(chartDefinition, selectedYAxisMetric, locale), xMetric: xAxisLabel(chartDefinition, locale), throughput: locale === 'zh' ? '单芯片吞吐量 (tok/s)' : 'Throughput/Chip (tok/s)', @@ -66,6 +68,9 @@ export default function InferenceTable({ () => sortRowsByYMetric(data, chartDefinition, selectedYAxisMetric), [data, chartDefinition, selectedYAxisMetric], ); + // Boundary / role clones (`i_pcompare`) share every config column with their + // base row; the series column is what tells them apart. + const powerCompare = useMemo(() => inferPowerCompare(data), [data]); const columns = useMemo[]>( () => [ @@ -85,6 +90,19 @@ export default function InferenceTable({ className: 'whitespace-nowrap', importance: 'key', }, + ...(powerCompare === 'none' + ? [] + : [ + { + header: headers.series, + cell: (row: InferenceData) => + powerSeriesLabel(row, selectedYAxisMetric, powerCompare, locale), + sortValue: (row: InferenceData) => + powerSeriesLabel(row, selectedYAxisMetric, powerCompare, locale), + className: 'whitespace-nowrap', + importance: 'key' as const, + }, + ]), { header: headers.tensorParallelism, align: 'right', @@ -129,8 +147,12 @@ export default function InferenceTable({ { header: headers.yMetric, align: 'right', - cell: (row) => formatInferenceTableNumber(yPath ? getNestedYValue(row, yPath) : row.y), - sortValue: (row) => (yPath ? getNestedYValue(row, yPath) : row.y), + // Comparison clones keep the source metrics; y holds the plotted role/boundary. + cell: (row) => + formatInferenceTableNumber( + row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath), + ), + sortValue: (row) => (row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath)), className: 'tabular-nums', importance: 'key', }, @@ -151,7 +173,7 @@ export default function InferenceTable({ importance: 'key', }, ], - [yPath, headers, showModeledPower], + [yPath, headers, showModeledPower, powerCompare, selectedYAxisMetric, locale], ); return ( diff --git a/packages/app/src/components/inference/ui/LegacyPowerLegendKey.tsx b/packages/app/src/components/inference/ui/LegacyPowerLegendKey.tsx deleted file mode 100644 index 3050638f6..000000000 --- a/packages/app/src/components/inference/ui/LegacyPowerLegendKey.tsx +++ /dev/null @@ -1,51 +0,0 @@ -import { POINT_SIZE } from '@/lib/chart-rendering'; -import { useLocale } from '@/lib/use-locale'; - -// Nested outside the KV-offload halo (POINT_SIZE + 4, '3 2' dashes) with a -// dotted dasharray so both decorations stay legible on one point. -export const LEGACY_POWER_RING_RADIUS = POINT_SIZE + 7; -export const LEGACY_POWER_RING_STROKE_WIDTH = 1.5; -export const LEGACY_POWER_RING_DASHARRAY = '1 3'; - -const STRINGS = { - en: 'Historical measurement (not validated under the current method)', - zh: '历史测量(尚未按当前方法验证)', -} as const; - -/** - * Key for the dotted ring drawn around measured-axis points whose power - * telemetry predates the producer validation contract (`power_tier === - * 'legacy'`). Rendered in the axis-metric info footer below the chart (the - * footer is `no-export`, so downloaded PNGs carry only the ring itself). - */ -export function LegacyPowerLegendKey() { - const locale = useLocale(); - - return ( -
- - {STRINGS[locale]} -
- ); -} diff --git a/packages/app/src/components/inference/ui/MeasuredMetricControls.tsx b/packages/app/src/components/inference/ui/MeasuredMetricControls.tsx index 46f0600eb..d371c4db2 100644 --- a/packages/app/src/components/inference/ui/MeasuredMetricControls.tsx +++ b/packages/app/src/components/inference/ui/MeasuredMetricControls.tsx @@ -7,29 +7,41 @@ import { SelectTrigger, SelectValue, } from '@/components/ui/select'; +import { track } from '@/lib/analytics'; +import { + ALL_IN_MEASURED_NOTE, + POWER_BASES, + POWER_BASIS_LABELS, + type PowerBasis, +} from '@/lib/power-basis'; import { useLocale } from '@/lib/use-locale'; import { changeMeasuredMetricConfig, getMeasuredMetricConfig, type MeasuredMetricConfigChange, } from '../measured-metric-config'; +import type { PowerCompare } from '../types'; +import { POWER_COMPARE_MODES, powerCompareAvailable } from '../utils/power-compare'; const STRINGS = { en: { + basis: 'Boundary', + basisHelp: + 'Choose GPU-only or all-in power, including server overhead and PUE. Provisioned values use rated capacity; measured values start from GPU telemetry. Points without the selected value are omitted.', scope: 'Scope', scopeHelp: 'All GPUs measures the whole deployment. Prefill and decode select only GPUs serving that role.', all: 'All GPUs', prefill: 'Prefill GPUs', decode: 'Decode GPUs', - statistic: 'Statistic', + statistic: 'Power statistic', statisticHelp: 'P75 and P90 are time-weighted percentiles of synchronized fleet power, divided by chip count. They are available for all GPUs only.', average: 'Average', - roleHint: 'Prefill and decode power support Average only.', display: 'Display', displayHelp: - 'Power per chip in watts, or average power as a percentage of chip TDP. Percent of TDP is available for the all-GPU average only.', + 'Power per chip in watts, average power as a percentage of chip TDP, or the per-second telemetry timeline behind the average. Percent of TDP and Timeline are available for the all-GPU average only.', + timeline: 'Timeline', denominator: 'Per', denominatorHelp: 'Choose the energy denominator. All-GPU energy per input or output token includes the whole deployment; role energy is selected separately under Scope.', @@ -40,21 +52,30 @@ const STRINGS = { unit: 'Unit', unitHelp: 'Energy is shown in joules. Energy per successful query can also be shown in watt-hours.', + compare: 'Compare', + compareHelp: + 'Compare power boundaries or prefill and decode on the same benchmark points. Role energy uses a common output-token denominator. Available for average W/chip and J per output token.', + compareNone: 'Off', + compareBoundaries: 'All boundaries', + compareRoles: 'Prefill vs decode', }, zh: { + basis: '功耗边界', + basisHelp: + '选择仅统计 GPU,或计入服务器其他组件及 PUE 的整体功耗。预配值按额定容量计算,实测值以 GPU 遥测为基础。缺少所选数值的数据点不绘制。', scope: '统计范围', scopeHelp: '全部 GPU 对应整个部署;预填充和解码仅统计承担相应任务的 GPU。', all: '全部 GPU', prefill: '预填充 GPU', decode: '解码 GPU', - statistic: '统计量', + statistic: '功耗统计量', statisticHelp: 'P75 和 P90 是同步采样的集群总功率按时间加权得到的分位数,再除以芯片数。仅支持全部 GPU。', average: '平均值', - roleHint: '预填充和解码功率仅支持平均值。', display: '显示方式', displayHelp: - '显示单芯片功率(瓦),或平均功率占芯片 TDP 的百分比。TDP 百分比仅支持全部 GPU 的平均功率。', + '显示单芯片功率(瓦)、平均功率占芯片 TDP 的百分比,或平均值背后的逐秒遥测时间线。TDP 百分比和时间线仅支持全部 GPU 的平均功率。', + timeline: '时间线', denominator: '能耗分母', denominatorHelp: '选择能耗的分母。按输入或输出 token 归一化的全部 GPU 能耗仍包含整个部署;预填充或解码能耗需在统计范围中单独选择。', @@ -64,21 +85,39 @@ const STRINGS = { query: '成功请求', unit: '单位', unitHelp: '能耗以焦耳显示;每个成功请求的能耗也可显示为瓦时。', + compare: '对比', + compareHelp: + '在同一批基准测试数据点上对比不同功耗边界,或预填充与解码。各角色的能耗统一按输出 token 归一化。支持平均 W/芯片和每输出 token 能耗。', + compareNone: '关闭', + compareBoundaries: '全部边界', + compareRoles: '预填充 vs 解码', }, } as const; export function MeasuredMetricControls({ metric, onChange, + compare = 'none', + onCompareChange, }: { metric: string; onChange: (metric: string) => void; + /** Comparison series overlaid on the metric (`i_pcompare`). */ + compare?: PowerCompare; + onCompareChange?: (mode: PowerCompare) => void; }) { - const t = STRINGS[useLocale()]; + const locale = useLocale(); + const t = STRINGS[locale]; const config = getMeasuredMetricConfig(metric); if (!config) return null; + const compareLabels: Record = { + none: t.compareNone, + boundaries: t.compareBoundaries, + roles: t.compareRoles, + }; const change = (next: MeasuredMetricConfigChange) => onChange(changeMeasuredMetricConfig(metric, next)); + const basisId = `measured-${config.family}-basis`; const scopeId = `measured-${config.family}-scope`; const roleScope = config.family === 'energy' @@ -91,9 +130,31 @@ export function MeasuredMetricControls({ return (
+
+ + +
{config.family === 'energy' && (
change({ statistic })} @@ -204,12 +265,16 @@ export function MeasuredMetricControls({ > % TDP + + {t.timeline} +
- {config.scope !== 'all' && ( -

{t.roleHint}

- )} ) : (
@@ -237,6 +302,51 @@ export function MeasuredMetricControls({
)} + {onCompareChange && ( +
+ + +
+ )} + {(config.basis === 'utility-modeled' || compare === 'boundaries') && ( +

+ {ALL_IN_MEASURED_NOTE[locale]} +

+ )}
); } diff --git a/packages/app/src/components/inference/ui/MeasuredPowerSummary.tsx b/packages/app/src/components/inference/ui/MeasuredPowerSummary.tsx deleted file mode 100644 index dfbbf68c6..000000000 --- a/packages/app/src/components/inference/ui/MeasuredPowerSummary.tsx +++ /dev/null @@ -1,117 +0,0 @@ -'use client'; - -import type { InferenceData } from '@/components/inference/types'; -import { useLocale } from '@/lib/use-locale'; - -export interface PowerTierCounts { - certified: number; - legacy: number; -} - -export function countPowerTiers(points: readonly InferenceData[]): PowerTierCounts { - const counts: PowerTierCounts = { certified: 0, legacy: 0 }; - for (const point of points) { - if (point.power_tier === 'certified') counts.certified += 1; - else if (point.power_tier === 'legacy') counts.legacy += 1; - } - return counts; -} - -const STRINGS = { - en: { - validated: 'validated', - historical: 'historical', - single: (shown: number, total: number, tier: string) => - `Showing ${shown} of ${total} ${tier} measurements.`, - combined: ( - shown: number, - total: number, - validatedShown: number, - validatedTotal: number, - historicalShown: number, - historicalTotal: number, - ) => - `Showing ${shown} of ${total} measured points: ${validatedShown}/${validatedTotal} validated · ${historicalShown}/${historicalTotal} historical.`, - bestPerSku: 'Best per SKU', - optimalOnly: 'Optimal Only', - controlsEnabled: (controls: string) => - `${controls} ${controls.includes(' and ') ? 'are' : 'is'} enabled.`, - reducedBySelections: 'Chart selections are hiding some points.', - }, - zh: { - validated: '已验证', - historical: '历史', - single: (shown: number, total: number, tier: string) => - `当前显示 ${shown}/${total} 个${tier}测量数据点。`, - combined: ( - shown: number, - total: number, - validatedShown: number, - validatedTotal: number, - historicalShown: number, - historicalTotal: number, - ) => - `当前显示 ${shown}/${total} 个实测数据点:已验证 ${validatedShown}/${validatedTotal} · 历史 ${historicalShown}/${historicalTotal}。`, - bestPerSku: '每个 SKU 仅显示最佳配置', - optimalOnly: '仅最优', - controlsEnabled: (controls: string) => `已启用:${controls}。`, - reducedBySelections: '当前图表选项隐藏了部分数据点。', - }, -} as const; - -function joinControls(controls: string[], locale: 'en' | 'zh'): string { - if (controls.length < 2) return controls[0] ?? ''; - return locale === 'zh' ? controls.join('、') : controls.join(' and '); -} - -export function MeasuredPowerSummary({ - total, - visible, - bestPerSku, - optimalOnly, -}: { - total: PowerTierCounts; - visible: PowerTierCounts; - bestPerSku: boolean; - optimalOnly: boolean; -}) { - const locale = useLocale(); - const t = STRINGS[locale]; - const totalCount = total.certified + total.legacy; - const visibleCount = visible.certified + visible.legacy; - if (totalCount === 0) return null; - - const summary = - total.certified > 0 && total.legacy > 0 - ? t.combined( - visibleCount, - totalCount, - visible.certified, - total.certified, - visible.legacy, - total.legacy, - ) - : total.certified > 0 - ? t.single(visible.certified, total.certified, t.validated) - : t.single(visible.legacy, total.legacy, t.historical); - - const controls: string[] = []; - if (bestPerSku) controls.push(t.bestPerSku); - if (optimalOnly) controls.push(t.optimalOnly); - const reduction = - visibleCount >= totalCount - ? null - : controls.length > 0 - ? t.controlsEnabled(joinControls(controls, locale)) - : t.reducedBySelections; - - return ( -
- {summary} - {reduction && {reduction}} -
- ); -} diff --git a/packages/app/src/components/inference/ui/PowerAnalysisPanels.tsx b/packages/app/src/components/inference/ui/PowerAnalysisPanels.tsx new file mode 100644 index 000000000..d2171eea1 --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerAnalysisPanels.tsx @@ -0,0 +1,142 @@ +'use client'; + +import { useCallback, useEffect, useMemo, useState } from 'react'; + +import { useUnofficialRun } from '@/components/unofficial-run-provider'; +import { useThemeColors } from '@/hooks/useThemeColors'; +import { useUrlState } from '@/hooks/useUrlState'; +import { track } from '@/lib/analytics'; +import { overlayRunColor, overlayRunIndex } from '@/lib/overlay-run-style'; +import { useLocale } from '@/lib/use-locale'; +import type { AggDataEntry, InferenceData } from '../types'; +import { + equalServiceSourceKey, + getEqualServiceSources, + observedPoints, +} from '../utils/equal-service-comparison'; +import PowerFitPanel from './PowerFitPanel'; +import PowerRoleGroup from './PowerRoleGroup'; + +const STRINGS = { + en: { + roles: 'Prefill / decode roles', + fit: 'Power vs output-rate fit', + }, + zh: { + roles: '预填充 / 解码角色', + fit: '功耗与输出速率拟合', + }, +}; + +function PanelToggle({ + checked, + onChange, + label, + testId, + event, +}: { + checked: boolean; + onChange: (checked: boolean) => void; + label: string; + testId: string; + event: string; +}) { + return ( + + ); +} + +export default function PowerAnalysisPanels({ + data, + overlayData = [], + xField, + xLabel, + chartId, + contextLabel, +}: { + data: InferenceData[]; + overlayData?: readonly InferenceData[]; + xField: keyof AggDataEntry; + xLabel: string; + chartId: string; + contextLabel?: string; +}) { + const locale = useLocale(); + const t = STRINGS[locale]; + const { getUrlParam, setUrlParams } = useUrlState(); + const { runIndexByUrl } = useUnofficialRun(); + const [roleShare, setRoleShare] = useState(() => getUrlParam('i_roleshare') === '1'); + const [powerFit, setPowerFit] = useState(() => getUrlParam('i_powerfit') === '1'); + const sources = useMemo(() => getEqualServiceSources(data, locale), [data, locale]); + useEffect(() => { + setUrlParams({ + i_roleshare: roleShare ? '1' : '0', + i_powerfit: powerFit ? '1' : '0', + }); + }, [roleShare, powerFit, setUrlParams]); + const { resolveColor } = useThemeColors({ + highContrast: true, + identifiers: sources.map((s) => s.key), + }); + const overlaySourceKeys = useMemo( + () => new Map(observedPoints(overlayData).map((p) => [equalServiceSourceKey(p), p.run_url])), + [overlayData], + ); + const colorOf = useCallback( + (key: string) => + overlaySourceKeys.has(key) + ? overlayRunColor(overlayRunIndex(overlaySourceKeys.get(key), runIndexByUrl)) + : resolveColor(key), + [overlaySourceKeys, runIndexByUrl, resolveColor], + ); + return ( +
+
+ + +
+ {roleShare && ( + + )} + {powerFit && ( + + )} +
+ ); +} diff --git a/packages/app/src/components/inference/ui/PowerFitPanel.tsx b/packages/app/src/components/inference/ui/PowerFitPanel.tsx new file mode 100644 index 000000000..deab9eb45 --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerFitPanel.tsx @@ -0,0 +1,224 @@ +'use client'; + +import * as d3 from 'd3'; +import { useMemo } from 'react'; + +import { ChartButtons } from '@/components/ui/chart-buttons'; +import { Heading } from '@/components/ui/heading'; +import { exportToCsv } from '@/lib/csv-export'; +import { useLocale } from '@/lib/use-locale'; +import { escapeHtml } from '@/lib/utils'; +import type { InferenceData } from '../types'; +import { buildPowerFits, MIN_FIT_POINTS } from '../utils/power-fit'; +import { PowerPanelPlot, type PanelLine, type PanelMarker } from './PowerPanelPlot'; + +const STRINGS = { + en: { + title: 'Power versus output rate (least-squares fit)', + xAxis: 'Output rate (tok/s per allocated GPU)', + yAxis: 'Mean GPU power (W/GPU)', + source: 'Source', + intercept: 'P₀ (W/GPU)', + tdpShare: 'P₀ ÷ TDP', + slope: 'm (J/output token)', + rSquared: 'R²', + points: 'Points', + range: 'Fitted range (tok/s/GPU)', + tooFew: (count: number) => + `Not fitted: needs ${MIN_FIT_POINTS} distinct output rates, has ${count}.`, + flat: 'Undefined: power did not vary', + empty: 'No validated measured power with a known output rate is available for these filters.', + }, + zh: { + title: '功耗与输出速率(最小二乘拟合)', + xAxis: '输出速率(每个已分配 GPU 的 tok/s)', + yAxis: '平均 GPU 功耗(W/GPU)', + source: '数据源', + intercept: 'P₀(W/GPU)', + tdpShare: 'P₀ ÷ TDP', + slope: 'm(J/输出 token)', + rSquared: 'R²', + points: '点数', + range: '拟合范围(tok/s/GPU)', + tooFew: (count: number) => + `未拟合:需要 ${MIN_FIT_POINTS} 个不同的输出速率,当前只有 ${count} 个。`, + flat: '无定义:功耗没有变化', + empty: '当前筛选条件下没有同时具备有效实测功耗和已知输出速率的数据。', + }, +}; + +const watts = d3.format(',.0f'); +const joules = d3.format(',.3f'); +const rate = d3.format(',.1f'); +const share = d3.format('.0%'); +const rSquared = d3.format('.3f'); +const EXTENSION_DASH = '4 4'; + +export default function PowerFitPanel({ + chartId, + data, + contextLabel, + colorOf, +}: { + chartId: string; + data: InferenceData[]; + contextLabel?: string; + colorOf: (sourceKey: string) => string; +}) { + const locale = useLocale(); + const t = STRINGS[locale]; + const fits = useMemo(() => buildPowerFits(data, locale), [data, locale]); + const sectionId = `${chartId}-power-fit`; + + const markers: PanelMarker[] = fits.flatMap(({ source, observations }) => + observations.map((observation) => ({ + x: observation.x, + y: observation.y, + key: source.key, + tooltip: [ + `${escapeHtml(source.label)}`, + `c${observation.point.conc} · ${rate(observation.x)} tok/s/GPU`, + `${watts(observation.y)} W/GPU`, + ].join('
'), + })), + ); + // Line keys become SVG class names, so they index the source rather than embed its key. + const lines: PanelLine[] = fits.flatMap(({ source, fit }, index) => { + if (!fit) return []; + const at = (x: number) => ({ x, y: fit.intercept + fit.slope * x }); + const color = colorOf(source.key); + return [ + { key: `source${index}-fit`, color, points: [at(fit.xMin), at(fit.xMax)] }, + ...(fit.xMin > 0 + ? [ + { + key: `source${index}-extension`, + color, + dash: EXTENSION_DASH, + points: [at(0), at(fit.xMin)], + }, + ] + : []), + ]; + }); + + const exportCsv = () => { + exportToCsv( + 'InferenceX_power_fit', + [ + 'source', + 'p0_w_per_gpu', + 'tdp_w', + 'p0_over_tdp', + 'm_j_per_output_token', + 'r_squared', + 'n', + 'x_min_tok_s_per_gpu', + 'x_max_tok_s_per_gpu', + 'status', + ], + fits.map((entry) => { + const { fit, tdpWatts } = entry; + return [ + entry.source.label, + fit?.intercept ?? null, + tdpWatts, + fit && tdpWatts ? fit.intercept / tdpWatts : null, + fit?.slope ?? null, + fit?.rSquared ?? null, + entry.observations.length, + fit?.xMin ?? null, + fit?.xMax ?? null, + entry.reason ?? 'fitted', + ]; + }), + contextLabel ? [contextLabel] : [], + ); + }; + + return ( +
+ + {t.title} + +

{contextLabel}

+ {fits.length === 0 ? ( +

{t.empty}

+ ) : ( + <> + +
+ + + + + + + + + + + + + + {fits.map((entry) => { + const { fit, source, tdpWatts } = entry; + return ( + + + {fit ? ( + <> + + + + + + + + ) : ( + + )} + + ); + })} + +
{t.source}{t.intercept}{t.tdpShare}{t.slope}{t.rSquared}{t.points}{t.range}
+ ● + {source.label} + {watts(fit.intercept)} + {tdpWatts + ? `${share(fit.intercept / tdpWatts)} · ${watts(tdpWatts)} W` + : '—'} + {joules(fit.slope)} + {fit.rSquared === null ? t.flat : rSquared(fit.rSquared)} + {fit.n} + {rate(fit.xMin)}–{rate(fit.xMax)} + + {t.tooFew(new Set(entry.observations.map((o) => o.x)).size)} +
+
+
+
+
+ + + )} +
+ ); +} diff --git a/packages/app/src/components/inference/ui/PowerMetricAvailability.tsx b/packages/app/src/components/inference/ui/PowerMetricAvailability.tsx deleted file mode 100644 index a3dd548b0..000000000 --- a/packages/app/src/components/inference/ui/PowerMetricAvailability.tsx +++ /dev/null @@ -1,233 +0,0 @@ -'use client'; - -import { useMemo } from 'react'; -import { useInferenceData, useInferenceFilters } from '../InferenceContext'; -import { isMeasuredEnergyConfigKey, metricOptionTitle, type MetricKey } from '../metric-registry'; -import type { InferenceData } from '../types'; -import { matchesQuickFilters } from '../utils/quickFilters'; -import { - powerMetricAvailability, - powerMetricState, - type PowerAvailabilityState, -} from '../utils/power-metric-availability'; -import { useUnofficialRun } from '@/components/unofficial-run-provider'; -import { hardwareKeyMatchesAnyBase } from '@/lib/constants'; -import { useLocale } from '@/lib/use-locale'; -import { track } from '@/lib/analytics'; - -const STRINGS = { - en: { - title: 'PowerX availability', - scope: - 'Current workload and hardware selection, before Optimal Only. Includes visible unofficial runs.', - loading: 'Updating measurement availability…', - none: 'No benchmark points match this selection.', - summary: (available: number, total: number) => - `${available} of ${total} points have this metric`, - labels: { - strict: 'Validated · schema 2', - validated: 'Validated · other or missing schema', - unverified: 'No validation verdict', - invalid: 'Validation failed', - inapplicable: 'No separate worker pools', - ambiguous: 'Whole-deployment energy schema unavailable', - missing: 'Metric not reported', - }, - note: 'A missing verdict does not establish age or validity. Prefill/decode metrics measure separate worker pools. Missing values are never replaced with zero or TDP estimates.', - all: 'Availability of all measured metrics', - evidence: 'Selected metric: source details', - run: 'Source run', - }, - zh: { - title: 'PowerX 指标可用性', - scope: '按当前工作负载与硬件选择统计,不受“仅最优”影响,包含已显示的非官方运行。', - loading: '正在更新指标可用性…', - none: '当前选择没有匹配的基准测试数据点。', - summary: (available: number, total: number) => - `${total} 个数据点中有 ${available} 个提供此指标`, - labels: { - strict: '已验证 · schema 2', - validated: '已验证 · 其他或未标注 schema', - unverified: '未提供验证结论', - invalid: '验证失败', - inapplicable: '无独立 worker 池', - ambiguous: '缺少整个部署的能耗 schema', - missing: '未提供此指标', - }, - note: '缺少验证结论不能判断数据新旧或有效性。prefill/decode 指标仅衡量独立 worker 池。缺失值不会被替换为零或 TDP 估算值。', - all: '所有实测指标的可用性', - evidence: '当前指标的来源详情', - run: '来源运行', - }, -} as const; - -export function PowerMetricAvailabilityPanel({ - points, - metric, - onSelect, - loading = false, -}: { - points: readonly InferenceData[]; - metric: string; - onSelect: (metric: string) => void; - loading?: boolean; -}) { - const locale = useLocale(); - const t = STRINGS[locale]; - const availability = useMemo(() => powerMetricAvailability(points), [points]); - const selected = availability.find((entry) => entry.metric === metric); - if (!selected) return null; - const sources = new Map< - string, - { point: InferenceData; state: PowerAvailabilityState; count: number } - >(); - for (const point of points) { - const state = powerMetricState(point, metric); - if (state === 'strict') continue; - const key = JSON.stringify([point.hwKey, point.run_url, state, point.power_invalid_reasons]); - const group = sources.get(key); - if (group) group.count++; - else sources.set(key, { point, state, count: 1 }); - } - return ( -
-

- {t.title}:{' '} - {loading - ? t.loading - : points.length > 0 - ? t.summary(selected.available, selected.total) - : t.none} -

- {!loading && ( - <> -

{t.scope}

-

- {Object.entries(selected.counts) - .filter(([, count]) => count > 0) - .map(([state, count]) => `${t.labels[state as PowerAvailabilityState]}: ${count}`) - .join(' · ')} -

-
{ - if (event.currentTarget.open) track('inference_power_availability_opened'); - }} - > - {t.all} -
    - {availability.map((entry) => ( -
  • - -
  • - ))} -
-

{t.note}

-
- {sources.size > 0 && ( -
- {t.evidence} -
    - {[...sources.values()].map(({ point, state, count }, index) => ( -
  • - {point.hwKey}: {t.labels[state]} ({count}) - {point.power_invalid_reasons?.length - ? ` · ${point.power_invalid_reasons.join(', ')}` - : ''} - {point.power_audit?.source ? ` · ${point.power_audit.source}` : ''} - {point.run_url && - /^https:\/\/github\.com\/[\w.-]+\/[\w.-]+\/actions\/runs\/\d+(?:\/attempts\/\d+)?$/u.test( - point.run_url, - ) && ( - <> - {' '} - ·{' '} - - {t.run} - - - )} -
  • - ))} -
-
- )} - - )} -
- ); -} - -/** Reuse pre-metric rows, including overlays: absent power must not erase its own explanation. */ -export function PowerMetricAvailability({ - metric, - onSelect, -}: { - metric: string; - onSelect: (metric: string) => void; -}) { - const { selectionPoints = [], loading, refreshing } = useInferenceData(); - const { - selectedModel, - selectedSequence, - selectedPrecisions, - activeHwTypes, - quickFilters, - compareGpuPair, - } = useInferenceFilters(); - const { getOverlayData, activeOverlayHwTypes, isUnofficialRun, localOfficialOverride } = - useUnofficialRun(); - const points = useMemo(() => { - const officialHw = isUnofficialRun ? (localOfficialOverride ?? activeHwTypes) : activeHwTypes; - const official = selectionPoints.filter( - (point) => - selectedPrecisions.includes(point.precision) && officialHw.has(String(point.hwKey)), - ); - const overlay = getOverlayData?.(selectedModel, selectedSequence, 'interactivity')?.data ?? []; - return [ - ...official, - ...overlay.filter( - (point) => - selectedPrecisions.includes(point.precision) && - activeOverlayHwTypes.has(String(point.hwKey)) && - matchesQuickFilters(point, quickFilters) && - (!compareGpuPair || hardwareKeyMatchesAnyBase(String(point.hwKey), compareGpuPair)), - ), - ]; - }, [ - selectionPoints, - selectedPrecisions, - activeHwTypes, - getOverlayData, - selectedModel, - selectedSequence, - activeOverlayHwTypes, - isUnofficialRun, - localOfficialOverride, - quickFilters, - compareGpuPair, - ]); - if (!isMeasuredEnergyConfigKey(metric)) return null; - return ( - - ); -} diff --git a/packages/app/src/components/inference/ui/PowerPanelPlot.tsx b/packages/app/src/components/inference/ui/PowerPanelPlot.tsx new file mode 100644 index 000000000..595200cd3 --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerPanelPlot.tsx @@ -0,0 +1,196 @@ +'use client'; + +import * as d3 from 'd3'; + +import { D3Chart } from '@/lib/d3-chart/D3Chart'; +import type { CustomLayerConfig, LayerConfig } from '@/lib/d3-chart/D3Chart/types'; +import type { ContinuousScale } from '@/lib/d3-chart/types'; + +/** One observed or derived value drawn as a marker. */ +export interface PanelMarker { + x: number; + y: number; + /** Series identity; the colour comes from `markerColor(key)`. */ + key: string; + selected?: boolean; + /** Tooltip HTML; callers escape any source text. */ + tooltip?: string; +} + +/** + * A polyline drawn beneath the markers (series connections, fit lines). A + * non-finite y breaks the line; it never bridges a missing value. + */ +export interface PanelLine { + key: string; + points: { x: number; y: number }[]; + color: string; + dash?: string; +} + +const MARGIN = { top: 20, right: 18, bottom: 64, left: 65 }; +const compactNumber = d3.format('~g'); +const compact = (value: d3.AxisDomain) => compactNumber(Number(value)); + +/** d3's log labelling: the 1× and 2× steps of each decade get labels, other ticks stay bare. */ +function logTickFormat(domain: [number, number]) { + const format = d3.scaleLog().domain(domain).tickFormat(5, '~g'); + return (value: d3.AxisDomain) => format(Number(value)); +} + +function logDomain(values: number[]): [number, number] { + const positive = values.filter((value) => value > 0 && Number.isFinite(value)); + if (positive.length === 0) return [1, 10]; + const min = Math.min(...positive); + const max = Math.max(...positive); + return min === max ? [min / 2, max * 2] : [min / 1.25, max * 1.25]; +} + +/** Observed x range; a single value widens from zero so the marker is not on an edge. */ +function linearXDomain(values: number[]): [number, number] { + const finite = values.filter(Number.isFinite); + if (finite.length === 0) return [0, 1]; + const min = Math.min(...finite); + const max = Math.max(...finite); + return min === max ? [0, max * 1.05 || 1] : [min, max]; +} + +/** Values plus zero and the reference, padded by a tenth of the span; never below zero for non-negative data. */ +function linearYDomain(values: number[], include: number[]): [number, number] { + const finite = [...values, ...include].filter(Number.isFinite); + const min = Math.min(...finite); + const max = Math.max(...finite); + const pad = (max - min || 1) * 0.1; + return [min >= 0 ? 0 : min - pad, max + pad]; +} + +/** + * Small D3 plot shared by the PowerX comparison panels: markers, optional + * connecting or fitted lines, and one dashed reference value (0 %, 50 %…). + * Log axes drop non-positive values instead of clamping them. + */ +export function PowerPanelPlot({ + chartId, + markers, + lines = [], + markerColor, + xLabel, + yLabel, + reference = null, + xLog = false, + yLog = false, + xDomain, + yDomain, + xTickValues, + height = 320, +}: { + chartId: string; + markers: PanelMarker[]; + lines?: PanelLine[]; + markerColor: (key: string) => string; + xLabel: string; + yLabel: string; + reference?: number | null; + xLog?: boolean; + yLog?: boolean; + xDomain?: [number, number]; + yDomain?: [number, number]; + /** Explicit x ticks, e.g. the observed concurrencies of a load sweep. */ + xTickValues?: number[]; + height?: number; +}) { + const visible = (point: { x: number; y: number }) => + Number.isFinite(point.x) && + Number.isFinite(point.y) && + (!xLog || point.x > 0) && + (!yLog || point.y > 0); + const drawn = markers.filter(visible); + const drawnLines = lines.filter((line) => line.points.filter(visible).length > 1); + const everything = [...drawn, ...drawnLines.flatMap((line) => line.points.filter(visible))]; + const xs = everything.map((point) => point.x); + const ys = everything.map((point) => point.y); + const x = xDomain ?? (xLog ? logDomain(xs) : linearXDomain(xs)); + const y = + yDomain ?? + (yLog ? logDomain(ys) : linearYDomain(ys, reference === null ? [0] : [0, reference])); + const colors = Object.fromEntries(drawnLines.map((line) => [line.key, line.color])); + const dashes = Object.fromEntries(drawnLines.map((line) => [line.key, line.dash ?? ''])); + const drawReference: NonNullable = (group, ctx) => { + const scale = (ctx.renderedYScale ?? ctx.yScale) as ContinuousScale; + group + .selectAll('line') + .data(reference === null ? [] : [reference]) + .join('line') + .attr('x1', 0) + .attr('x2', ctx.width) + .attr('y1', (value) => scale(value)) + .attr('y2', (value) => scale(value)) + .attr('stroke', 'currentColor') + .attr('stroke-dasharray', '5,4') + .attr('opacity', 0.6); + }; + const layers: LayerConfig[] = [ + { type: 'custom', key: 'reference', render: drawReference }, + { + type: 'line', + key: 'lines', + lines: Object.fromEntries(drawnLines.map((line) => [line.key, line.points])), + config: { + curve: d3.curveLinear, + getColor: (key) => colors[key], + getStrokeDasharray: (key) => dashes[key] || 'none', + isDefined: visible, + }, + }, + { + type: 'point', + key: 'markers', + data: drawn, + config: { + getCx: () => 0, + getCy: () => 0, + getX: (point) => point.x, + getY: (point) => point.y, + getRadius: (point) => (point.selected ? 6 : 3), + getColor: (point) => markerColor(point.key), + }, + }, + ]; + const hasTooltips = drawn.some((point) => point.tooltip); + return ( + + chartId={`${chartId}-plot`} + data={drawn} + height={height} + testId={`${chartId}-plot`} + watermark="logo" + zoom={{ enabled: false }} + grabCursor={false} + instructions="" + margin={MARGIN} + xScale={{ type: xLog ? 'log' : 'linear', domain: x, nice: !xLog }} + yScale={{ type: yLog ? 'log' : 'linear', domain: y, nice: !yLog }} + xAxis={{ + label: xLabel, + tickCount: 4, + ...(xTickValues ? { tickValues: xTickValues, tickFormat: compact } : {}), + }} + yAxis={{ + label: yLabel, + tickCount: 5, + ...(yLog ? { tickFormat: logTickFormat(y) } : {}), + }} + layers={layers} + {...(hasTooltips + ? { + tooltip: { + rulerType: 'none' as const, + content: (point: PanelMarker) => + `
${point.tooltip ?? ''}
`, + attachToLayer: 2, + }, + } + : {})} + /> + ); +} diff --git a/packages/app/src/components/inference/ui/PowerRoleGroup.tsx b/packages/app/src/components/inference/ui/PowerRoleGroup.tsx new file mode 100644 index 000000000..e8ce4cef6 --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerRoleGroup.tsx @@ -0,0 +1,321 @@ +'use client'; + +import * as d3 from 'd3'; +import { useMemo } from 'react'; + +import { ChartButtons } from '@/components/ui/chart-buttons'; +import { Heading } from '@/components/ui/heading'; +import { exportToCsv } from '@/lib/csv-export'; +import { useLocale } from '@/lib/use-locale'; +import { escapeHtml } from '@/lib/utils'; +import type { AggDataEntry, InferenceData } from '../types'; +import { + getRolePoints, + type EqualServiceSource, + type RolePoint, +} from '../utils/equal-service-comparison'; +import { powerVariantDash } from '../utils/power-compare'; +import { PowerPanelPlot, type PanelLine, type PanelMarker } from './PowerPanelPlot'; + +type RoleSeriesKey = 'prefill' | 'decode' | 'total'; +type RolePanelId = 'power' | 'local-energy' | 'output-energy' | 'share'; + +const STRINGS = { + en: { + title: 'Prefill and decode roles', + empty: 'No validated prefill/decode telemetry is available for these filters.', + panelEmpty: 'These observations do not report this role figure.', + roles: { prefill: 'Prefill', decode: 'Decode', total: 'Prefill + decode' }, + panels: { + power: { + title: 'Mean GPU power by role', + axis: 'Mean GPU power (W/GPU)', + note: 'Mean board power per GPU inside each pool over the validated window.', + units: { prefill: 'W/GPU', decode: 'W/GPU', total: 'W/GPU' }, + }, + 'local-energy': { + title: 'Role-local GPU energy', + axis: 'GPU energy (J/token, log)', + note: 'Prefill: prefill-pool joules ÷ input tokens. Decode: decode-pool joules ÷ output tokens. The denominators differ, so these two are not added.', + units: { prefill: 'J/input token', decode: 'J/output token', total: 'J/output token' }, + }, + 'output-energy': { + title: 'GPU energy per output token by role', + axis: 'GPU energy (J/output token, log)', + note: 'Prefill energy restated per output token with the same window’s input:output token ratio, so prefill + decode is the request total.', + units: { prefill: 'J/output token', decode: 'J/output token', total: 'J/output token' }, + }, + share: { + title: 'Prefill energy share', + axis: 'Prefill energy share (%)', + note: 'Share = prefill ÷ (prefill + decode), both per output token. The dashed line marks 50%.', + units: { prefill: '%', decode: '%', total: '%' }, + }, + } satisfies Record< + RolePanelId, + { title: string; axis: string; note: string; units: Record } + >, + }, + zh: { + title: '预填充与解码角色', + empty: '当前筛选条件下没有经过验证的预填充 / 解码遥测数据。', + panelEmpty: '这些观测值没有报告该角色指标。', + roles: { prefill: '预填充', decode: '解码', total: '预填充 + 解码' }, + panels: { + power: { + title: '各角色的平均 GPU 功耗', + axis: '平均 GPU 功耗(W/GPU)', + note: '有效测量窗口内各 GPU 池中每个 GPU 的平均板卡功耗。', + units: { prefill: 'W/GPU', decode: 'W/GPU', total: 'W/GPU' }, + }, + 'local-energy': { + title: '各角色按本池 token 计的 GPU 能耗', + axis: 'GPU 能耗(J/token,对数)', + note: '预填充:预填充池焦耳数 ÷ 输入 token 数。解码:解码池焦耳数 ÷ 输出 token 数。两者分母不同,因此不相加。', + units: { prefill: 'J/输入 token', decode: 'J/输出 token', total: 'J/输出 token' }, + }, + 'output-energy': { + title: '各角色每输出 token 的 GPU 能耗', + axis: 'GPU 能耗(J/输出 token,对数)', + note: '用同一窗口的输入:输出 token 比,把预填充能耗换算为每输出 token,因此预填充 + 解码等于整个请求的能耗。', + units: { prefill: 'J/输出 token', decode: 'J/输出 token', total: 'J/输出 token' }, + }, + share: { + title: '预填充能耗占比', + axis: '预填充能耗占比(%)', + note: '占比 = 预填充 ÷(预填充 + 解码),两者都按输出 token 计。虚线标记 50%。', + units: { prefill: '%', decode: '%', total: '%' }, + }, + } satisfies Record< + RolePanelId, + { title: string; axis: string; note: string; units: Record } + >, + }, +}; + +const PANELS: { + id: RolePanelId; + series: { role: RoleSeriesKey; value: (point: RolePoint) => number | null }[]; + yLog?: boolean; + reference?: number; + yDomain?: [number, number]; + /** Share points stay unconnected: each is one observed configuration. */ + connect: boolean; +}[] = [ + { + id: 'power', + series: [ + { role: 'prefill', value: (point) => point.prefillWattsPerGpu }, + { role: 'decode', value: (point) => point.decodeWattsPerGpu }, + ], + connect: true, + }, + { + id: 'local-energy', + series: [ + { role: 'prefill', value: (point) => point.prefillJoulesPerInputToken }, + { role: 'decode', value: (point) => point.decodeJoulesPerOutputToken }, + ], + yLog: true, + connect: true, + }, + { + id: 'output-energy', + series: [ + { role: 'prefill', value: (point) => point.energy?.prefill ?? null }, + { role: 'decode', value: (point) => point.energy?.decode ?? null }, + { role: 'total', value: (point) => point.energy?.total ?? null }, + ], + yLog: true, + connect: true, + }, + { + id: 'share', + series: [{ role: 'prefill', value: (point) => point.energy?.prefillShare ?? null }], + reference: 50, + yDomain: [0, 100], + connect: false, + }, +]; +const ROLE_DASH: Record = { + prefill: powerVariantDash({ kind: 'role', id: 'prefill' }), + decode: powerVariantDash({ kind: 'role', id: 'decode' }), + total: '', +}; +const value = d3.format(',.4~g'); + +export default function PowerRoleGroup({ + chartId, + data, + xField, + xLabel, + contextLabel, + sources, + colorOf, +}: { + chartId: string; + data: InferenceData[]; + xField: keyof AggDataEntry; + xLabel: string; + contextLabel?: string; + sources: EqualServiceSource[]; + colorOf: (sourceKey: string) => string; +}) { + const locale = useLocale(); + const t = STRINGS[locale]; + const points = useMemo(() => getRolePoints(data, xField), [data, xField]); + const label = (key: string) => sources.find((source) => source.key === key)?.label ?? key; + const isConcurrency = xField === 'conc'; + const concurrencies = isConcurrency ? [...new Set(points.map((point) => point.x))] : undefined; + const sectionId = `${chartId}-roles`; + + const panel = (definition: (typeof PANELS)[number]) => { + const copy = t.panels[definition.id]; + const plotId = `${chartId}-role-${definition.id}`; + const markers: PanelMarker[] = definition.series.flatMap(({ role, value: read }) => + points.flatMap((point) => { + const y = read(point); + return y === null + ? [] + : [ + { + x: point.x, + y, + key: point.sourceKey, + tooltip: [ + `${escapeHtml(label(point.sourceKey))}`, + `${escapeHtml(xLabel)}: ${value(point.x)} · c${point.point.conc}`, + `${t.roles[role]}: ${value(y)} ${copy.units[role]}`, + ].join('
'), + }, + ]; + }), + ); + // Line keys become SVG class names, so they index the source rather than embed its key. + const lines: PanelLine[] = definition.connect + ? definition.series.flatMap(({ role, value: read }) => + [...new Set(points.map((point) => point.sourceKey))].map((sourceKey, index) => ({ + key: `source${index}-${role}`, + color: colorOf(sourceKey), + dash: ROLE_DASH[role], + points: points + .filter((point) => point.sourceKey === sourceKey) + .map((point) => ({ x: point.x, y: read(point) ?? Number.NaN })), + })), + ) + : []; + return ( +
+
{copy.title}
+ {markers.length > 0 ? ( + + ) : ( +

+ {t.panelEmpty} +

+ )} +

{copy.note}

+
+ ); + }; + + const exportCsv = () => { + exportToCsv( + 'InferenceX_prefill_decode_roles', + [ + 'source', + String(xField), + 'concurrency', + 'point_id', + 'run_url', + 'prefill_w_per_gpu', + 'decode_w_per_gpu', + 'prefill_j_per_input_token', + 'decode_j_per_output_token', + 'prefill_j_per_output_token', + 'total_j_per_output_token', + 'prefill_energy_share_pct', + ], + points.map((point) => [ + label(point.sourceKey), + point.x, + point.point.conc, + point.point.id ?? null, + point.point.run_url ?? null, + point.prefillWattsPerGpu, + point.decodeWattsPerGpu, + point.prefillJoulesPerInputToken, + point.decodeJoulesPerOutputToken, + point.energy?.prefill ?? null, + point.energy?.total ?? null, + point.energy?.prefillShare ?? null, + ]), + contextLabel ? [contextLabel] : [], + ); + }; + + const shownSources = sources.filter((source) => + points.some((point) => point.sourceKey === source.key), + ); + return ( +
+ + {t.title} + +

{contextLabel}

+ {points.length > 0 ? ( + <> +
+ {shownSources.map((source) => ( + + ● + {source.label} + + ))} + {(['prefill', 'decode', 'total'] as const).map((role) => ( + + + {t.roles[role]} + + ))} +
+
{PANELS.map(panel)}
+
+
+
+ + + ) : ( +

{t.empty}

+ )} +
+ ); +} diff --git a/packages/app/src/components/inference/ui/PowerTimeline.tsx b/packages/app/src/components/inference/ui/PowerTimeline.tsx new file mode 100644 index 000000000..3c6f2bd8f --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerTimeline.tsx @@ -0,0 +1,1659 @@ +'use client'; + +/** + * PowerX "Timeline" display: the per-second GPU power behind each measured + * average, drawn over the whole benchmark job. + * + * ChartDisplay renders this instead of ScatterGraph when the Measured Power + * Display control is `timeline` (`y_measuredPowerTimeline`). The point set is + * the same as the Measured Avg Power axis; each point's `power_audit.source` + * names its `gpu_metrics_*` artifact — or, for disaggregated Slurm / Dynamo + * rows, its validation file inside a `power_audit_*` bundle — which + * `/api/gpu-metrics?series=power` returns as one-second per-GPU buckets + * (`components/gpu-power/power-series.ts`). + * + * One trace per config, coloured by hardware (official) or by run (unofficial + * overlay); in date comparison, official traces take the GPUGraph colour and + * legend toggle of their (date, hardware) series. The validated measurement + * window is emphasized; the rest of the + * job (server start, warmup) is drawn faint. Rated TDP is a dashed reference + * per hardware; the all-in provisioned line is opt-in because it would halve + * the vertical resolution of the traces. + * + * Pool mode (bundle series carry worker roles) sums each prefill / decode pool + * instead, against pool-sized TDP references, so a disaggregated deployment + * reads as two lines on a watts axis. A pinned scatter tooltip can deep-link + * here focused on one config (`requestPowerTraceFocus`). + */ +import * as d3 from 'd3'; +import { useQueries, type UseQueryResult } from '@tanstack/react-query'; +import React, { useCallback, useEffect, useMemo, useRef, useState } from 'react'; +import { HW_REGISTRY } from '@semianalysisai/inferencex-constants'; + +import { + bucketTimeMs, + meanPowerAt, + sumPowerAt, + type GpuPowerSeries, + type GpuPowerSeriesResponse, +} from '@/components/gpu-power/power-series'; +import { useComparisonSeries } from '@/components/inference/hooks/useComparisonSeries'; +import { comparisonEntryLabel } from '@/components/inference/utils/comparisonEntry'; +import ChartLegend, { type LegendSwitchConfig } from '@/components/ui/chart-legend'; +import { SegmentedToggle } from '@/components/ui/segmented-toggle'; +import { useUnofficialRun } from '@/components/unofficial-run-provider'; +import { matchesQuickFilters } from '@/components/inference/utils/quickFilters'; +import { useThemeColors } from '@/hooks/useThemeColors'; +import { useUrlState } from '@/hooks/useUrlState'; +import { track } from '@/lib/analytics'; +import { computeToggle } from '@/lib/toggle-set'; +import { getModelSortIndex } from '@/lib/constants'; +import { D3Chart } from '@/lib/d3-chart/D3Chart'; +import type { LayerConfig, RenderContext, ZoomContext } from '@/lib/d3-chart/D3Chart/types'; +import { lttbDownsample } from '@/lib/d3-chart/downsample'; +import { CHART_FONT_SANS, CHART_TYPE, px } from '@/lib/d3-chart/typography'; +import { overlayRunColor, overlayRunIndex } from '@/lib/overlay-run-style'; +import { useLocale } from '@/lib/use-locale'; +import { getDisplayLabel } from '@/lib/utils'; + +import { + useInferenceActions, + useInferenceData, + useInferenceDisplay, + useInferenceFilters, +} from '../InferenceContext'; +import type { InferenceData, OverlayData } from '../types'; +import { powerVariantDash } from '../utils/power-compare'; +import PowerTimelineSummary from './PowerTimelineSummary'; +import { + allGpuPool, + consumePowerTraceFocus, + groupPoolsBySize, + hasPowerTimelineWindow, + parsePowerTimelineParams, + powerTimelineSampleX, + joinPowerTimeline, + planPowerTimelineRequests, + prioritizeRun, + prioritizeRuns, + referenceLabelSlots, + stackTraceLabels, + runIdFromUrl, + traceConfigLabel, + traceKeyRunId, + tracePools, + windowPhase, + type PoolSizeGroup, + type PowerPool, + type PowerPoolRole, + type PowerTimelineAxis, + type PowerTimelineLines, + type PowerTimelineRequest, + type PowerTimelineTrace, + type WindowPhase, +} from '../utils/powerTimeline'; + +/** Distinct workflow runs fetched per chart; each is one GitHub artifact sweep. */ +export const POWER_TIMELINE_MAX_RUNS = 4; +/** Hover targets per trace after LTTB; paths keep every bucket. */ +const HIT_POINTS_PER_TRACE = 200; +/** Above this many visible traces the `c` end labels would only overlap. */ +const MAX_LABELED_TRACES = 40; +const CHART_HEIGHT = 600; +const MARGIN = { top: 24, right: 84, bottom: 60, left: 64 }; + +const STRINGS = { + en: { + timeAxis: 'Time axis', + concurrency: 'Concurrency', + allConcurrencies: 'All', + wall: 'Wall clock (UTC)', + elapsed: 'Since telemetry start', + serving: 'Since serving start', + xServing: 'Time since serving-window start (s)', + windowOnly: 'Validated window only', + windowOnlyHelp: + 'Show only retained samples inside each recorded validated serving window. Traces without valid window bounds are omitted.', + missingWindow: (count: number) => + `${count} trace${count === 1 ? '' : 's'} omitted: no valid serving-window bounds.`, + missingFocus: + 'The selected trace is unavailable for these filters or has no retained telemetry.', + xWall: 'Time (UTC)', + xElapsed: 'Time since telemetry start (m:ss)', + perGpu: 'One line per GPU', + perGpuHelp: 'Draw every GPU of a config instead of the mean across its GPUs.', + pools: 'Prefill / decode pools', + poolsHelp: + 'One line per worker-role pool: the summed board power of the prefill GPUs and of the decode GPUs of a config. Dashed references are pool size × rated TDP.', + utilityLines: 'All-in provisioned lines', + utilityHelp: + 'Dashed reference at the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model). Off by default because it compresses the traces.', + loading: (runs: number) => + `Loading GPU telemetry for ${runs} run${runs === 1 ? '' : 's'}… (may take a minute)`, + loadError: (runId: string, message: string) => `Run ${runId}: ${message}`, + noTraces: + 'No telemetry traces for the visible hardware. Enable a series in the legend or choose another date.', + noArtifacts: + 'These points predate per-config telemetry artifacts, so no timeline is available for them.', + droppedRuns: (runs: number) => + `Telemetry from ${runs} more run${runs === 1 ? '' : 's'} was not loaded (limit ${POWER_TIMELINE_MAX_RUNS} runs per chart).`, + telemetry: 'Telemetry', + instructions: + 'Shift+Scroll to zoom horizontally · Drag to pan · Double-click to reset · Click a point to pin tooltip', + dismiss: 'Click elsewhere to dismiss', + phase: { + before: 'Before window (startup / warmup)', + window: 'Measurement window', + after: 'After window', + unknown: 'Window not recorded', + } satisfies Record, + meanPerGpu: 'Mean per GPU', + gpus: (count: number) => `${count} GPU${count === 1 ? '' : 's'}`, + min: 'min', + max: 'max', + validated: 'Validated average', + sinceStart: 'since start', + tdp: 'TDP', + allIn: 'all-in', + poolShort: { prefill: 'prefill', decode: 'decode', all: 'all GPUs' } satisfies Record< + PowerPoolRole, + string + >, + yPool: 'GPU pool power (W)', + pool: 'Pool', + poolPower: 'Pool power', + poolTdp: 'pool TDP', + focused: (label: string) => `Focused on ${label}`, + showAll: 'Show all', + unofficialRun: 'Unofficial run', + branch: 'Branch', + viewWorkflow: 'View workflow run', + }, + zh: { + timeAxis: '时间轴', + concurrency: '并发数', + allConcurrencies: '全部', + wall: '实际时刻(UTC)', + elapsed: '距遥测起点', + serving: '距服务窗口起点', + xServing: '距服务窗口开始的时间(秒)', + windowOnly: '仅显示有效测量窗口', + windowOnlyHelp: '仅显示已记录的有效服务窗口内保留的采样点;缺少有效窗口边界的曲线不绘制。', + missingWindow: (count: number) => `${count} 条曲线缺少有效服务窗口边界,未绘制。`, + missingFocus: '所选曲线不符合当前筛选条件,或没有保留的遥测数据。', + xWall: '时间(UTC)', + xElapsed: '距遥测开始的时间(分:秒)', + perGpu: '每个 GPU 一条线', + perGpuHelp: '绘制配置中每个 GPU 的曲线,而不是各 GPU 的平均值。', + pools: '预填充 / 解码 GPU 池', + poolsHelp: + '按 worker 角色分池绘制:每条线是同一配置中预填充 GPU 或解码 GPU 的板卡功耗之和。虚线参考为池内 GPU 数量 × 额定 TDP。', + utilityLines: '全电源配置参考线', + utilityHelp: + '按硬件注册表中每 GPU 的全电源配置(all-in)市电功率绘制虚线参考(SemiAnalysis 数据中心行业模型)。默认关闭,因为它会压缩曲线的纵向分辨率。', + loading: (runs: number) => `正在加载 ${runs} 个运行的 GPU 遥测数据……(可能需要约一分钟)`, + loadError: (runId: string, message: string) => `运行 ${runId}:${message}`, + noTraces: '当前可见硬件没有遥测曲线。请在图例中启用一个系列或选择其他日期。', + noArtifacts: '这些数据点早于按配置上传的遥测产物,因此没有可用的时间线。', + droppedRuns: (runs: number) => + `另有 ${runs} 个运行的遥测数据未加载(每张图表最多 ${POWER_TIMELINE_MAX_RUNS} 个运行)。`, + telemetry: '遥测来源', + instructions: 'Shift+滚轮横向缩放 · 拖动平移 · 双击重置 · 点击数据点固定提示框', + dismiss: '点击其他区域关闭', + phase: { + before: '测量窗口之前(启动 / warmup)', + window: '测量窗口内', + after: '测量窗口之后', + unknown: '未记录测量窗口', + } satisfies Record, + meanPerGpu: '每 GPU 平均', + gpus: (count: number) => `${count} 个 GPU`, + min: '最小', + max: '最大', + validated: '有效平均值', + sinceStart: '距起点', + tdp: 'TDP', + allIn: 'all-in', + poolShort: { prefill: '预填充', decode: '解码', all: '全部 GPU' } satisfies Record< + PowerPoolRole, + string + >, + yPool: 'GPU 池功耗(W)', + pool: 'GPU 池', + poolPower: '池功耗', + poolTdp: '池 TDP', + focused: (label: string) => `聚焦:${label}`, + showAll: '显示全部', + unofficialRun: '非官方运行', + branch: '分支', + viewWorkflow: '查看工作流运行', + }, +} as const; + +type XMode = PowerTimelineAxis; +type LineMode = PowerTimelineLines; + +interface TimelineSample { + trace: PowerTimelineTrace; + color: string; + overlayIndex: number | null; + column: number; + timeMs: number; + /** Data-space x for the active mode: epoch ms (wall) or seconds (elapsed). */ + x: number; + /** Mean watts across the GPUs sampled in the bucket; the pool's summed watts in pool mode. */ + y: number; + min: number; + max: number; + /** GPUs with a sample in the bucket (inside the pool, in pool mode). */ + gpuCount: number; + phase: WindowPhase; + /** The pool this sample sums and its device count, in pool mode. */ + pool?: { role: PowerPoolRole; gpuCount: number }; +} + +interface TracePoint { + x: number; + y: number | null; +} + +interface TracePath { + id: string; + traceKey: string; + hwKey: string; + /** Date comparison: the official trace's `${date}_${hwKey}` legend series. */ + series?: string; + overlayIndex: number | null; + color: string; + segment: 'full' | 'window'; + width: number; + opacity: number; + points: TracePoint[]; + /** Worker-role pool the line sums, in pool mode. */ + pool?: PowerPoolRole; +} + +interface ReferenceLine { + id: string; + watts: number; + label: string; + color: string; + kind: 'tdp' | 'utility'; + /** Pools the line is sized for, in pool mode (roles sharing one GPU count). */ + pools?: PowerPoolRole[]; +} + +/** Vertical distance between stacked reference labels that share a watts value. */ +const REFERENCE_LABEL_ROW = 13; + +interface TraceLabel { + /** Join key: the trace key, plus the pool role in pool mode. */ + id: string; + traceKey: string; + hwKey: string; + series?: string; + pool?: PowerPoolRole; + color: string; + text: string; + x: number; + y: number; +} + +interface DrawModel { + paths: TracePath[]; + labels: TraceLabel[]; +} + +export interface PowerTimelineProps { + chartId: string; + /** Official points of the chart (display-limit clipped points restored). */ + data: InferenceData[]; + overlayData?: OverlayData; + yLabel: string; + caption?: React.ReactNode; + /** Date comparison: official traces follow the per-date legend series. */ + comparison?: boolean; + /** Run numbers shared with the comparison changelog and GPUGraph legend. */ + runNumbering?: Map; +} + +async function fetchPowerSeries( + request: PowerTimelineRequest, + signal: AbortSignal, +): Promise { + const params = new URLSearchParams({ runId: request.runId, series: 'power' }); + if (request.prefix) params.set('prefix', request.prefix); + const response = await fetch(`/api/gpu-metrics?${params.toString()}`, { + method: 'POST', + headers: { 'Content-Type': 'application/json' }, + body: JSON.stringify({ sources: request.sources }), + cache: 'no-store', + signal, + }); + const body = (await response.json()) as GpuPowerSeriesResponse | { error: string }; + if (!response.ok) { + throw new Error('error' in body ? body.error : `HTTP ${response.status}`); + } + return body as GpuPowerSeriesResponse; +} + +function formatElapsed(totalSeconds: number): string { + const seconds = Math.max(0, Math.round(totalSeconds)); + const h = Math.floor(seconds / 3600); + const m = Math.floor((seconds % 3600) / 60); + const s = seconds % 60; + const mm = h > 0 ? String(m).padStart(2, '0') : String(m); + return `${h > 0 ? `${h}:` : ''}${mm}:${String(s).padStart(2, '0')}`; +} + +const formatUtcClock = d3.utcFormat('%H:%M:%S'); +const formatUtcDate = d3.utcFormat('%Y-%m-%d'); +/** Pool sums run to thousands of watts; group the digits. */ +const formatWatts = d3.format(',.0f'); + +/** Rough advance of one label character at the data-label size. */ +const LABEL_CHAR_WIDTH = 6.5; + +function baseHardware(hwKey: string): string { + return hwKey.split('_')[0]; +} + +function runIdsOf(points: readonly InferenceData[]): Set { + return new Set( + points + .map((point) => runIdFromUrl(point.run_url)) + .filter((runId): runId is string => runId !== null), + ); +} + +/** + * The pools a trace draws in pool mode: its worker-role pools, or every GPU as + * one pool when the collector assigned no roles (a single-node trace then shows + * its deployment total on the same axis). + */ +function drawnPools(series: GpuPowerSeries): PowerPool[] { + const pools = tracePools(series); + return pools.length > 0 ? pools : [allGpuPool(series)]; +} + +/** SVG dash of a pool line: per role from the comparison palette; `all` stays solid. */ +function poolDash(pool: PowerPoolRole): string | null { + const dash = powerVariantDash({ kind: 'role', id: pool }); + return dash === '' ? null : dash; +} + +interface TraceRow { + id: string; + pool?: PowerPoolRole; + values: (number | null)[]; +} + +/** One polyline's values per line mode: the GPU mean, each GPU, or each pool's sum. */ +function traceRows(series: GpuPowerSeries, lineMode: LineMode): TraceRow[] { + if (lineMode === 'gpu') { + return series.gpus.map((gpu, row) => ({ id: `gpu${gpu}`, values: series.power[row] })); + } + if (lineMode === 'pool') { + return drawnPools(series).map((pool) => ({ + id: `pool:${pool.role}`, + pool: pool.role, + values: series.t.map((_, column) => sumPowerAt(series, pool.rows, column)), + })); + } + return [{ id: 'mean', values: series.t.map((_, column) => meanPowerAt(series, column)) }]; +} + +/** Builds the mean, per-GPU or per-pool polylines plus the window emphasis for one trace. */ +function tracePaths( + trace: PowerTimelineTrace, + color: string, + overlayIndex: number | null, + xMode: XMode, + lineMode: LineMode, + windowOnly: boolean, +): TracePath[] { + const { series } = trace; + const rows = traceRows(series, lineMode); + const faint = lineMode === 'gpu' ? 0.22 : 0.32; + const strong = lineMode === 'gpu' ? 0.85 : 1; + const widths = lineMode === 'gpu' ? [1, 1.5] : [1.25, 2.25]; + const paths: TracePath[] = []; + for (const row of rows) { + const full: TracePoint[] = []; + const window: TracePoint[] = []; + row.values.forEach((value, column) => { + const x = powerTimelineSampleX(trace, column, xMode, windowOnly); + if (x === null) return; + const point = { x, y: value }; + full.push(point); + if (windowPhase(trace, bucketTimeMs(series, column)) === 'window') window.push(point); + }); + if (!windowOnly && full.length > 1) + paths.push({ + id: `${trace.key}:${row.id}:full`, + traceKey: trace.key, + hwKey: trace.point.hwKey, + overlayIndex, + color, + segment: 'full', + width: widths[0], + opacity: faint, + points: full, + pool: row.pool, + }); + if (window.length > 1) { + paths.push({ + id: `${trace.key}:${row.id}:window`, + traceKey: trace.key, + hwKey: trace.point.hwKey, + overlayIndex, + color, + segment: 'window', + width: widths[1], + opacity: strong, + points: window, + pool: row.pool, + }); + } + } + return paths; +} + +/** + * Hover targets of one trace: one stream over all its GPUs (mean watts), or in + * pool mode one stream per pool (summed watts) so the tooltip can name the pool. + */ +function traceSamples( + trace: PowerTimelineTrace, + color: string, + overlayIndex: number | null, + xMode: XMode, + lineMode: LineMode, + windowOnly: boolean, +): TimelineSample[] { + const { series } = trace; + if (lineMode !== 'pool') { + const rows = series.power.map((_, row) => row); + return sampleRows(trace, color, overlayIndex, xMode, windowOnly, rows, undefined); + } + return drawnPools(series).flatMap((pool) => + sampleRows(trace, color, overlayIndex, xMode, windowOnly, pool.rows, { + role: pool.role, + gpuCount: pool.rows.length, + }), + ); +} + +function sampleRows( + trace: PowerTimelineTrace, + color: string, + overlayIndex: number | null, + xMode: XMode, + windowOnly: boolean, + rows: readonly number[], + pool: TimelineSample['pool'], +): TimelineSample[] { + const { series } = trace; + const samples: TimelineSample[] = []; + for (let column = 0; column < series.t.length; column++) { + const x = powerTimelineSampleX(trace, column, xMode, windowOnly); + if (x === null) continue; + let sum = 0; + let count = 0; + let min = Number.POSITIVE_INFINITY; + let max = Number.NEGATIVE_INFINITY; + for (const row of rows) { + const value = series.power[row]?.[column]; + if (value === null || value === undefined) continue; + sum += value; + count += 1; + if (value < min) min = value; + if (value > max) max = value; + } + // A pool bucket missing a device is a gap, as in `sumPowerAt`, not a dip. + if (count === 0 || (pool && count < rows.length)) continue; + const timeMs = bucketTimeMs(series, column); + samples.push({ + trace, + color, + overlayIndex, + column, + timeMs, + x, + y: pool ? sum : sum / count, + min, + max, + gpuCount: count, + phase: windowPhase(trace, timeMs), + pool, + }); + } + return lttbDownsample( + samples, + HIT_POINTS_PER_TRACE, + (sample) => sample.x, + (sample) => sample.y, + ); +} + +type AnyContinuousScale = d3.ScaleLinear; + +function drawTraces( + group: d3.Selection, + xScale: AnyContinuousScale, + yScale: AnyContinuousScale, + model: DrawModel, + highlight: string | null, +): void { + const line = d3 + .line() + .defined((point) => point.y !== null) + .x((point) => xScale(point.x)) + .y((point) => yScale(point.y ?? 0)) + .curve(d3.curveLinear); + const selection = group + .selectAll('path.power-trace') + .data(model.paths, (path) => path.id); + selection.exit().remove(); + selection + .enter() + .append('path') + .attr('class', 'power-trace') + .attr('fill', 'none') + .attr('stroke-linejoin', 'round') + .attr('stroke-linecap', 'round') + .merge(selection) + .attr('data-trace-key', (path) => path.traceKey) + .attr('data-hw', (path) => path.hwKey) + .attr('data-segment', (path) => path.segment) + .attr('data-run-index', (path) => path.overlayIndex) + .attr('data-pool', (path) => path.pool ?? null) + .attr('stroke-dasharray', (path) => (path.pool ? poolDash(path.pool) : null)) + .attr('stroke', (path) => path.color) + .attr('stroke-width', (path) => path.width) + .attr('opacity', (path) => traceOpacity(path, highlight)) + .attr('d', (path) => line(path.points)); +} + +function isHighlighted(mark: TracePath | TraceLabel, highlight: string): boolean { + return highlight === mark.hwKey || highlight === mark.traceKey || highlight === mark.series; +} + +function traceOpacity(path: TracePath, highlight: string | null): number { + if (highlight === null) return path.opacity; + return isHighlighted(path, highlight) ? Math.min(1, path.opacity + 0.15) : path.opacity * 0.15; +} + +function labelOpacity(label: TraceLabel, highlight: string | null): number { + return highlight === null || isHighlighted(label, highlight) ? 1 : 0.2; +} + +function drawLabels( + group: d3.Selection, + xScale: AnyContinuousScale, + yScale: AnyContinuousScale, + model: DrawModel, + plotWidth: number, + highlight: string | null, +): void { + const selection = group + .selectAll('text.power-trace-label') + .data(model.labels, (label) => label.id); + selection.exit().remove(); + selection + .enter() + .append('text') + .attr('class', 'power-trace-label') + .attr('font-family', CHART_FONT_SANS) + .attr('font-size', px(CHART_TYPE.dataLabel)) + .attr('font-weight', '600') + .attr('dominant-baseline', 'middle') + .attr('pointer-events', 'none') + .merge(selection) + .attr('data-hw', (label) => label.hwKey) + .attr('data-pool', (label) => label.pool ?? null) + .attr('fill', (label) => label.color) + .attr('opacity', (label) => labelOpacity(label, highlight)) + .text((label) => label.text); + // Sit just past the last sample; flip inside the plot near the right edge. + const boxes = model.labels.map((label) => { + const x = xScale(label.x); + const width = label.text.length * LABEL_CHAR_WIDTH; + const flip = x + 6 + width > plotWidth; + const left = flip ? x - 6 - width : x + 6; + return { flip, left, right: left + width, y: yScale(label.y) }; + }); + const rows = stackTraceLabels(boxes, CHART_TYPE.dataLabel + 2); + const byId = new Map(model.labels.map((label, index) => [label.id, index])); + group.selectAll('text.power-trace-label').each(function (label) { + const index = byId.get(label.id)!; + const box = boxes[index]; + d3.select(this) + .attr('text-anchor', box.flip ? 'end' : 'start') + .attr('x', box.flip ? box.right : box.left) + .attr('y', rows[index]); + }); +} + +function drawReferenceLines( + group: d3.Selection, + yScale: AnyContinuousScale, + width: number, + lines: ReferenceLine[], +): void { + group.selectAll('.power-reference').remove(); + const slots = referenceLabelSlots(lines); + lines.forEach((line, index) => { + const y = yScale(line.watts); + const g = group + .append('g') + .attr('class', 'power-reference') + .attr('data-reference', line.kind) + .attr('data-pool', line.pools?.join(' ') ?? null) + .attr('data-watts', line.watts); + g.append('line') + .attr('x1', 0) + .attr('x2', width) + .attr('y1', y) + .attr('y2', y) + .attr('stroke', line.color) + .attr('stroke-width', 1.25) + .attr('stroke-dasharray', line.kind === 'tdp' ? '6,4' : '2,4') + .attr('opacity', 0.9); + g.append('text') + .attr('x', width - 4) + .attr('y', y - 5 - slots[index] * REFERENCE_LABEL_ROW) + .attr('text-anchor', 'end') + .attr('fill', line.color) + .attr('font-family', CHART_FONT_SANS) + .attr('font-size', px(CHART_TYPE.annotation)) + .attr('font-weight', '600') + .text(line.label); + }); +} + +export default function PowerTimeline({ + chartId, + data, + overlayData, + yLabel, + caption, + comparison = false, + runNumbering: providedRunNumbering, +}: PowerTimelineProps) { + const locale = useLocale(); + const t = STRINGS[locale]; + const { hardwareConfig, hwTypesWithData } = useInferenceData(); + const { activeHwTypes, selectedPrecisions, quickFilters, activeDates } = useInferenceFilters(); + const { isLegendExpanded, highContrast } = useInferenceDisplay(); + const { setBestPerSku, toggleHwType, setIsLegendExpanded, toggleActiveDate } = + useInferenceActions(); + const { allGraphs: comparisonSeries, runNumbering } = useComparisonSeries(providedRunNumbering); + const { + unofficialRunInfos, + runIndexByUrl, + activeOverlayHwTypes, + localOfficialOverride, + setUnifiedOverlaySelection, + } = useUnofficialRun(); + + const { getUrlParam, setUrlParams } = useUrlState(); + const [initialView] = useState(() => + parsePowerTimelineParams({ + i_ptaxis: getUrlParam('i_ptaxis'), + i_ptlines: getUrlParam('i_ptlines'), + i_ptwindow: getUrlParam('i_ptwindow'), + i_ptfocus: getUrlParam('i_ptfocus'), + i_ptutility: getUrlParam('i_ptutility'), + i_ptconc: getUrlParam('i_ptconc'), + }), + ); + const [xModeChoice, setXModeChoice] = useState(initialView.axis); + const [lineMode, setLineMode] = useState(initialView.lines); + const [windowOnly, setWindowOnly] = useState(initialView.windowOnly); + const [showUtility, setShowUtility] = useState(initialView.utility); + const [highlight, setHighlight] = useState(null); + /** Trace a "View power trace" deep link asked for, once the join has produced it. */ + const [focusKey, setFocusKey] = useState(initialView.focus); + /** The deep-link request, read once on mount; `undefined` until read, `null` once honoured. */ + const requestedFocusRef = useRef(undefined); + // The deep-link request is read once, before planning, so its run is fetched + // even when the chart spans more runs than the cap. + if (requestedFocusRef.current === undefined) { + requestedFocusRef.current = consumePowerTraceFocus(); + } + /** One load across hardware (e.g. GB200 and GB300 at c4); a deep link names its own trace. */ + const [concurrency, setConcurrency] = useState(() => + requestedFocusRef.current ? null : initialView.concurrency, + ); + + // The chart's point list still carries every precision, quick-filtered rows + // and rows without a validated average (ScatterGraph applies those gates at + // draw time); only rows that plot on the measured-average axis for the + // current selection have a trace to look up. + const plotsHere = useCallback( + (point: InferenceData) => + point.measuredPowerTimeline !== undefined && + selectedPrecisions.includes(point.precision) && + matchesQuickFilters(point, quickFilters), + [selectedPrecisions, quickFilters], + ); + const measuredData = useMemo(() => data.filter(plotsHere), [data, plotsHere]); + const overlayPoints = useMemo( + () => (overlayData?.data ?? []).filter(plotsHere), + [overlayData, plotsHere], + ); + const overlayPointSet = useMemo(() => new Set(overlayPoints), [overlayPoints]); + const concurrencyOptions = useMemo( + () => + [...new Set([...measuredData, ...overlayPoints].map((point) => point.conc))] + .filter((conc) => Number.isSafeInteger(conc) && conc > 0) + .toSorted((a, b) => a - b), + [measuredData, overlayPoints], + ); + // Filtering before planning keeps the run cap for the runs that hold this load. + const allPoints = useMemo( + () => + [...measuredData, ...overlayPoints].filter( + (point) => concurrency === null || point.conc === concurrency, + ), + [measuredData, overlayPoints, concurrency], + ); + + const hwKeysInData = useMemo( + () => + [...new Set(measuredData.map((point) => point.hwKey))].toSorted( + (a, b) => getModelSortIndex(a) - getModelSortIndex(b) || a.localeCompare(b), + ), + [measuredData], + ); + const stableHcKeys = useMemo(() => [...hwTypesWithData], [hwTypesWithData]); + const activeOfficialKeys = useMemo(() => [...activeHwTypes], [activeHwTypes]); + const { resolveColor, getCssColor } = useThemeColors({ + highContrast, + identifiers: hwKeysInData, + activeKeys: activeOfficialKeys, + hcKeys: stableHcKeys, + }); + + const requests = useMemo(() => planPowerTimelineRequests(allPoints), [allPoints]); + const focusRun = traceKeyRunId(requestedFocusRef.current ?? focusKey); + // Same visibility source as ScatterGraph: an overlay session may hold a + // local official selection that has not been written back to the filters. + // Date comparison instead shows the (date, hardware) series toggled on in + // the legend, as GPUGraph does. + const officialHwTypes = localOfficialOverride ?? activeHwTypes; + const comparisonSeriesOf = useCallback( + (point: InferenceData) => + comparison && !overlayPointSet.has(point) ? `${point.date}_${point.hwKey}` : undefined, + [comparison, overlayPointSet], + ); + const isShown = useCallback( + (point: InferenceData) => { + if (overlayPointSet.has(point)) return activeOverlayHwTypes.has(point.hwKey); + const series = comparisonSeriesOf(point); + return series ? activeDates.has(series) : officialHwTypes.has(point.hwKey); + }, + [overlayPointSet, activeOverlayHwTypes, comparisonSeriesOf, activeDates, officialHwTypes], + ); + // Overlay runs were requested explicitly (`?unofficialrun=`), so they take + // the cap's slots before official runs; runs the legend shows go before runs + // it hides entirely, so hiding hardware makes room for the pair compared. + // The deep-linked run still goes first. + const overlayRunIds = useMemo(() => runIdsOf(overlayPoints), [overlayPoints]); + const shownRunIds = useMemo(() => runIdsOf(allPoints.filter(isShown)), [allPoints, isShown]); + const fetchedRequests = useMemo( + () => + prioritizeRun( + prioritizeRuns(prioritizeRuns(requests, overlayRunIds), shownRunIds), + focusRun, + ).slice(0, POWER_TIMELINE_MAX_RUNS), + [requests, overlayRunIds, shownRunIds, focusRun], + ); + const droppedRuns = requests.length - fetchedRequests.length; + // React Query structurally shares the combined result, so `resolved` keeps + // its identity until a run's data actually changes. That gives the response + // map a fixed-shape memo input; a per-query spread would change the deps + // array length whenever runs enter or leave the plot, which React rejects. + const combineQueries = useCallback( + (results: UseQueryResult[]) => ({ + loadingRuns: results.filter((query) => query.isPending).length, + errors: fetchedRequests.flatMap((request, index) => { + const error = results[index]?.error; + return error instanceof Error ? [{ request, error }] : []; + }), + resolved: fetchedRequests.flatMap((request, index) => { + const response = results[index]?.data; + return response ? [[request.runId, response] as const] : []; + }), + }), + [fetchedRequests], + ); + const { loadingRuns, errors, resolved } = useQueries({ + queries: fetchedRequests.map((request) => ({ + queryKey: ['power-timeline', request.runId, request.prefix, request.sources] as const, + queryFn: ({ signal }: { signal: AbortSignal }) => fetchPowerSeries(request, signal), + staleTime: 0, + refetchOnWindowFocus: true, + retry: 1, + })), + combine: combineQueries, + }); + const responses = useMemo(() => new Map(resolved), [resolved]); + + const { traces, missing } = useMemo( + () => joinPowerTimeline(allPoints, responses), + [allPoints, responses], + ); + const hasAnyArtifact = requests.length > 0; + + // Honour the deep link once its trace exists; a disaggregated trace opens in + // pool mode because its prefill / decode split is what the reader came for. + useEffect(() => { + const requested = requestedFocusRef.current; + if (!requested) return; + const trace = traces.find((entry) => entry.key === requested); + if (!trace) return; + requestedFocusRef.current = null; + setFocusKey(trace.key); + if (tracePools(trace.series).length > 0) setLineMode('pool'); + }, [traces]); + + useEffect(() => { + if (loadingRuns > 0 || responses.size === 0) return; + track('inference_power_timeline_loaded', { + traces: traces.length, + missing: missing.length, + runs: responses.size, + }); + }, [loadingRuns, responses.size, traces.length, missing.length]); + + const comparisonColors = useMemo( + () => new Map(comparisonSeries.map((series) => [series.id, series.color])), + [comparisonSeries], + ); + const colorForTrace = useCallback( + (trace: PowerTimelineTrace): { color: string; overlayIndex: number | null } => { + if (overlayPointSet.has(trace.point)) { + const index = overlayRunIndex(trace.point.run_url ?? null, runIndexByUrl); + return { color: overlayRunColor(index), overlayIndex: index }; + } + const series = comparisonSeriesOf(trace.point); + const color = + (series && comparisonColors.get(series)) ?? getCssColor(resolveColor(trace.point.hwKey)); + return { color, overlayIndex: null }; + }, + [ + overlayPointSet, + runIndexByUrl, + comparisonSeriesOf, + comparisonColors, + getCssColor, + resolveColor, + ], + ); + // With an overlay loaded the chart reads localOfficialOverride, so a legend + // click must write the unified selection the way ScatterGraph does; the + // context's toggleHwType would change activeHwTypes with no visible effect. + const handleToggleHwType = useCallback( + (key: string) => { + if (!overlayData) { + toggleHwType(key); + return; + } + setBestPerSku(false, { applySelection: false }); + const official = new Set([...officialHwTypes].filter((hw) => hwTypesWithData.has(hw))); + setUnifiedOverlaySelection( + computeToggle(official, key, hwTypesWithData), + activeOverlayHwTypes, + ); + }, + [ + overlayData, + toggleHwType, + setBestPerSku, + officialHwTypes, + hwTypesWithData, + setUnifiedOverlaySelection, + activeOverlayHwTypes, + ], + ); + const visibleTraces = useMemo( + () => traces.filter((trace) => isShown(trace.point)), + [traces, isShown], + ); + // Focus follows visibility: hiding the focused hardware in the legend lifts + // the dimming and the chip instead of dimming everything with nothing lit. + const focusedTrace = useMemo( + () => visibleTraces.find((trace) => trace.key === focusKey) ?? null, + [visibleTraces, focusKey], + ); + /** Legend hover wins over the deep-link focus while it lasts. */ + const activeHighlight = highlight ?? focusedTrace?.key ?? null; + const visibleRunCount = useMemo( + () => new Set(visibleTraces.map((trace) => trace.runId)).size, + [visibleTraces], + ); + const xMode: XMode = xModeChoice ?? (visibleRunCount <= 1 ? 'wall' : 'elapsed'); + + // Retain a shared pool choice while data loads; keep its switch available even + // if the current selection has no role-tagged trace, so it can be turned off. + const hasPools = useMemo( + () => visibleTraces.some((trace) => tracePools(trace.series).length > 0), + [visibleTraces], + ); + useEffect(() => { + setUrlParams({ + i_ptaxis: xModeChoice ?? '', + i_ptlines: lineMode === 'mean' ? '' : lineMode, + i_ptwindow: windowOnly ? 'window' : '', + i_ptfocus: focusKey ?? '', + i_ptutility: showUtility ? '1' : '', + i_ptconc: concurrency === null ? '' : String(concurrency), + }); + }, [xModeChoice, lineMode, windowOnly, focusKey, showUtility, concurrency, setUrlParams]); + + const missingWindows = + windowOnly || xMode === 'serving' + ? visibleTraces.filter((trace) => !hasPowerTimelineWindow(trace)).length + : 0; + + const hardwareLabel = useCallback( + (point: InferenceData): string => { + const config = overlayPointSet.has(point) + ? overlayData?.hardwareConfig[point.hwKey] + : hardwareConfig[point.hwKey]; + return config ? getDisplayLabel(config) : point.hwKey; + }, + [overlayPointSet, overlayData, hardwareConfig], + ); + const entryLabel = useCallback( + (point: InferenceData) => comparisonEntryLabel(String(point.date), runNumbering), + [runNumbering], + ); + /** Hardware, plus the compared date or run of an official trace in date comparison. */ + const seriesLabel = useCallback( + (point: InferenceData): string => + comparisonSeriesOf(point) + ? `${hardwareLabel(point)} · ${entryLabel(point)}` + : hardwareLabel(point), + [comparisonSeriesOf, hardwareLabel, entryLabel], + ); + + const model = useMemo(() => { + const paths: TracePath[] = []; + const labels: TraceLabel[] = []; + // Two platforms at one load share `c4`; name the hardware so the end labels differ. + const bases = new Set(visibleTraces.map((trace) => baseHardware(trace.point.hwKey))); + const hardwarePrefix = (trace: PowerTimelineTrace) => { + if (bases.size < 2) return ''; + const base = baseHardware(trace.point.hwKey); + return `${HW_REGISTRY[base]?.label ?? base.toUpperCase()} `; + }; + // Several compared dates share `c4` as well; name the date or run too. + const entries = new Set( + visibleTraces + .filter((trace) => comparisonSeriesOf(trace.point)) + .map((trace) => trace.point.date), + ); + const entryPrefix = (trace: PowerTimelineTrace) => + entries.size > 1 && comparisonSeriesOf(trace.point) ? `${entryLabel(trace.point)} ` : ''; + for (const trace of visibleTraces) { + const { color, overlayIndex } = colorForTrace(trace); + const series = comparisonSeriesOf(trace.point); + const tracePathSet = tracePaths(trace, color, overlayIndex, xMode, lineMode, windowOnly).map( + (path) => ({ ...path, series }), + ); + paths.push(...tracePathSet); + if (visibleTraces.length > MAX_LABELED_TRACES) continue; + // One end label per trace; per pool in pool mode, so the role reads off the line. + const groups: { pool?: PowerPoolRole; paths: TracePath[] }[] = + lineMode === 'pool' + ? drawnPools(trace.series).map((pool) => ({ + pool: pool.role, + paths: tracePathSet.filter((path) => path.pool === pool.role), + })) + : [{ paths: tracePathSet }]; + for (const group of groups) { + const anchor = + group.paths.find((path) => path.segment === 'window') ?? + group.paths.find((path) => path.segment === 'full'); + const last = anchor?.points.filter((point) => point.y !== null).at(-1); + if (!last || last.y === null) continue; + labels.push({ + id: group.pool ? `${trace.key}:${group.pool}` : trace.key, + traceKey: trace.key, + hwKey: trace.point.hwKey, + series, + pool: group.pool, + color, + text: `${entryPrefix(trace)}${hardwarePrefix(trace)}${ + group.pool + ? `c${trace.point.conc} · ${t.poolShort[group.pool]}` + : `c${trace.point.conc}` + }`, + x: last.x, + y: last.y, + }); + } + } + return { paths, labels }; + }, [ + visibleTraces, + colorForTrace, + comparisonSeriesOf, + entryLabel, + xMode, + lineMode, + windowOnly, + t, + ]); + + const samples = useMemo( + () => + visibleTraces.flatMap((trace) => { + const { color, overlayIndex } = colorForTrace(trace); + return traceSamples(trace, color, overlayIndex, xMode, lineMode, windowOnly); + }), + [visibleTraces, colorForTrace, xMode, lineMode, windowOnly], + ); + + // Rated references: per hardware in mean / per-GPU modes; per (hardware, + // pool size) in pool mode, scaled to the pool so the summed line and its + // ceiling share the axis. Roles of one hardware that hold the same number + // of GPUs share a ceiling and draw as one line (`prefill / decode ×16`). + const referenceLines = useMemo(() => { + const lines: ReferenceLine[] = []; + const pushLines = (base: string, color: string, pool?: PoolSizeGroup) => { + const specs = HW_REGISTRY[base]; + if (!specs) return; + const label = specs.label ?? base.toUpperCase(); + const size = pool?.size ?? 1; + const id = pool ? `${base}:${pool.roles.join('+')}:${pool.size}` : base; + const name = pool + ? `${label} ${pool.roles.map((role) => t.poolShort[role]).join(' / ')} ×${pool.size}` + : label; + if (specs.tdp > 0) { + lines.push({ + id: `tdp:${id}`, + watts: specs.tdp * size, + label: `${name} ${t.tdp} ${specs.tdp * size} W`, + color, + kind: 'tdp', + pools: pool?.roles, + }); + } + if (showUtility && specs.power > 0) { + const watts = pool ? Math.round(specs.power * 1000) * size : specs.power * 1000; + lines.push({ + id: `utility:${id}`, + watts, + label: `${name} ${t.allIn} ${Math.round(watts)} W`, + color, + kind: 'utility', + pools: pool?.roles, + }); + } + }; + // First trace of a hardware sets the reference colour, as before. + const perBase = new Map(); + for (const trace of visibleTraces) { + const base = baseHardware(trace.point.hwKey); + if (!perBase.has(base)) perBase.set(base, { color: colorForTrace(trace).color, pools: [] }); + if (lineMode === 'pool') perBase.get(base)!.pools.push(...drawnPools(trace.series)); + } + for (const [base, { color, pools }] of perBase) { + if (lineMode !== 'pool') { + pushLines(base, color); + continue; + } + for (const group of groupPoolsBySize(pools)) pushLines(base, color, group); + } + return lines; + }, [visibleTraces, colorForTrace, showUtility, lineMode, t]); + + const xDomain = useMemo<[number, number]>(() => { + let min = Number.POSITIVE_INFINITY; + let max = Number.NEGATIVE_INFINITY; + for (const path of model.paths) { + for (const point of path.points) { + if (point.x < min) min = point.x; + if (point.x > max) max = point.x; + } + } + if (!Number.isFinite(min) || !Number.isFinite(max)) { + return xMode === 'wall' ? [Date.UTC(2026, 0, 1), Date.UTC(2026, 0, 1, 0, 10)] : [0, 600]; + } + if (xMode === 'serving' && windowOnly) min = Math.min(0, min); + return min === max ? [min, max + (xMode === 'wall' ? 60_000 : 60)] : [min, max]; + }, [model.paths, xMode, windowOnly]); + const yDomain = useMemo<[number, number]>(() => { + let max = 0; + for (const path of model.paths) { + for (const point of path.points) if (point.y !== null && point.y > max) max = point.y; + } + for (const line of referenceLines) if (line.watts > max) max = line.watts; + return [0, max > 0 ? max * 1.06 : 100]; + }, [model.paths, referenceLines]); + + const xTickFormat = useMemo(() => { + if (xMode === 'elapsed') return (value: d3.AxisDomain) => formatElapsed(Number(value)); + if (xMode === 'serving') return (value: d3.AxisDomain) => d3.format('~g')(Number(value)); + const span = xDomain[1] - xDomain[0]; + const crossesDate = formatUtcDate(new Date(xDomain[0])) !== formatUtcDate(new Date(xDomain[1])); + const format = d3.utcFormat( + crossesDate ? '%m/%d %H:%M' : span < 3 * 60_000 ? '%H:%M:%S' : '%H:%M', + ); + return (value: d3.AxisDomain) => + format(value instanceof Date ? value : new Date(Number(value))); + }, [xMode, xDomain]); + + const highlightRef = useRef(activeHighlight); + highlightRef.current = activeHighlight; + const layers = useMemo[]>( + () => [ + { + type: 'custom', + key: 'power-reference-lines', + render: (group, ctx) => { + drawReferenceLines(group, ctx.yScale as AnyContinuousScale, ctx.width, referenceLines); + }, + }, + { + type: 'custom', + key: 'power-traces', + render: (group, ctx: RenderContext) => { + drawTraces( + group, + ctx.xScale as AnyContinuousScale, + ctx.yScale as AnyContinuousScale, + model, + highlightRef.current, + ); + }, + onZoom: (group, ctx: ZoomContext) => { + drawTraces( + group, + ctx.newXScale as AnyContinuousScale, + ctx.newYScale as AnyContinuousScale, + model, + highlightRef.current, + ); + }, + }, + { + type: 'point', + key: 'power-hit-points', + data: samples, + config: { + getCx: () => 0, + getCy: () => 0, + getX: (sample) => sample.x, + getY: (sample) => sample.y, + getColor: (sample) => sample.color, + getRadius: () => 2, + // Pool streams of one trace share columns; the role keeps their keys apart. + keyFn: (sample) => + sample.pool + ? `${sample.trace.key}:${sample.pool.role}:${sample.column}` + : `${sample.trace.key}:${sample.column}`, + maxPoints: Number.POSITIVE_INFINITY, + }, + }, + { + type: 'custom', + key: 'power-trace-labels', + render: (group, ctx: RenderContext) => { + drawLabels( + group, + ctx.xScale as AnyContinuousScale, + ctx.yScale as AnyContinuousScale, + model, + ctx.width, + highlightRef.current, + ); + }, + onZoom: (group, ctx: ZoomContext) => { + drawLabels( + group, + ctx.newXScale as AnyContinuousScale, + ctx.newYScale as AnyContinuousScale, + model, + ctx.width, + highlightRef.current, + ); + }, + }, + ], + [referenceLines, model, samples], + ); + + const onDisplayUpdate = useCallback( + (ctx: RenderContext) => { + const root = d3.select(ctx.layout.svg.node() as SVGSVGElement); + root + .selectAll('path.power-trace') + .attr('opacity', (path) => traceOpacity(path, activeHighlight)); + root + .selectAll('text.power-trace-label') + .attr('opacity', (label) => labelOpacity(label, activeHighlight)); + }, + [activeHighlight], + ); + + const tooltipContent = useCallback( + (sample: TimelineSample, isPinned: boolean) => { + const { trace } = sample; + const point = trace.point; + const overlayInfo = + sample.overlayIndex === null ? null : unofficialRunInfos[sample.overlayIndex]; + const tdp = HW_REGISTRY[baseHardware(point.hwKey)]?.tdp ?? 0; + const elapsed = formatElapsed((sample.timeMs - trace.series.startMs) / 1000); + const clock = `${formatUtcClock(new Date(sample.timeMs))} UTC`; + const time = + xMode === 'serving' + ? `${sample.x.toFixed(1)} s · ${t.serving} · ${clock}` + : xMode === 'wall' + ? `${clock} · +${elapsed} ${t.sinceStart}` + : `+${elapsed} · ${clock}`; + const colon = locale === 'zh' ? ':' : ':'; + const validated = point.measuredAvgPower?.y; + const { pool } = sample; + const readings = pool + ? `
${t.pool}${colon} ${t.poolShort[pool.role]} · ${t.gpus(pool.gpuCount)}
+
${t.poolPower}${colon} ${formatWatts(sample.y)} W${ + tdp > 0 + ? ` (${((sample.y / (tdp * pool.gpuCount)) * 100).toFixed(0)}% ${t.poolTdp})` + : '' + }
+
${t.meanPerGpu}${colon} ${(sample.y / sample.gpuCount).toFixed(1)} W · ${t.min} ${sample.min.toFixed(1)} W · ${t.max} ${sample.max.toFixed(1)} W
` + : `
${t.meanPerGpu}${colon} ${sample.y.toFixed(1)} W${ + tdp > 0 + ? ` (${((sample.y / tdp) * 100).toFixed(0)}% ${t.tdp})` + : '' + }
+
${t.gpus(sample.gpuCount)} · ${t.min} ${sample.min.toFixed(1)} W · ${t.max} ${sample.max.toFixed(1)} W
`; + return `
+ ${isPinned ? `
${t.dismiss}
` : ''} +
${seriesLabel(point)} · ${traceConfigLabel(point)}${ + overlayInfo ? ` · ✕ ${overlayInfo.branch || `run ${overlayInfo.id}`}` : '' + }
+
${time}
+ ${readings} +
${t.phase[sample.phase]}
+ ${ + typeof validated === 'number' + ? `
${t.validated}${colon} ${validated.toFixed(1)} W
` + : '' + } +
`; + }, + [unofficialRunInfos, xMode, t, locale, seriesLabel], + ); + + const legendItems = useMemo(() => { + const overlayItems = + overlayData && unofficialRunInfos.length > 0 + ? unofficialRunInfos + .map((info, index) => { + const hasPoints = overlayPoints.some( + (point) => overlayRunIndex(point.run_url ?? null, runIndexByUrl) === index, + ); + if (!hasPoints) return null; + const branch = info.branch || `run ${info.id}`; + return { + name: `✕ unofficial-run-${info.id}`, + label: `✕ ${branch}`, + color: overlayRunColor(index), + // The comparison legend groups on the name's first word: one + // "Unofficial run" group over every run. + title: comparison ? t.unofficialRun : `${t.unofficialRun}: ${branch}`, + isHighlighted: true, + hw: `overlay-run-${info.id}`, + isActive: true, + isRemovable: false, + onClick: () => {}, + tooltip: ( +
+
{t.unofficialRun}
+
+ {t.branch}: {branch} +
+ {info.url && ( + + {t.viewWorkflow} + + )} +
+ ), + }; + }) + .filter((item): item is NonNullable => item !== null) + : []; + if (comparison) { + // GPUGraph's comparison legend: one row per (date, hardware) series with + // a trace candidate, grouped by hardware. + const seriesWithData = new Set(measuredData.map((point) => `${point.date}_${point.hwKey}`)); + const comparisonItems = comparisonSeries + .filter(({ id }) => seriesWithData.has(id)) + .map(({ date, hwKey, id, color }) => ({ + name: `${hwKey} ${comparisonEntryLabel(date, runNumbering)}`, + label: comparisonEntryLabel(date, runNumbering), + color, + title: hardwareConfig[hwKey] ? getDisplayLabel(hardwareConfig[hwKey]) : hwKey, + hw: id, + isActive: activeDates.has(id), + onClick: () => { + toggleActiveDate(id); + track('interactivity_date_toggled', { date, hw: hwKey, view: 'power_timeline' }); + }, + tooltip: null, + })); + return [...overlayItems, ...comparisonItems]; + } + const officialItems = hwKeysInData + .filter((key) => hwTypesWithData.has(key) && hardwareConfig[key]) + .map((key) => { + const config = hardwareConfig[key]; + return { + name: config.name, + label: getDisplayLabel(config), + color: resolveColor(key), + title: config.gpu, + hw: key, + isActive: officialHwTypes.has(key), + onClick: () => { + handleToggleHwType(key); + track('latency_hw_type_toggled', { hw: key }); + }, + tooltip: null, + }; + }); + return [...overlayItems, ...officialItems]; + }, [ + overlayData, + unofficialRunInfos, + overlayPoints, + runIndexByUrl, + hwKeysInData, + hwTypesWithData, + hardwareConfig, + resolveColor, + officialHwTypes, + handleToggleHwType, + t, + comparison, + comparisonSeries, + measuredData, + runNumbering, + activeDates, + toggleActiveDate, + ]); + + // Per-GPU and pools are two views of the same lines, so either switch turns + // the other off; both fall back to the mean. + const chooseLineMode = (next: LineMode) => { + setLineMode(next); + track('inference_power_timeline_lines_changed', { lines: next }); + }; + const switches: LegendSwitchConfig[] = [ + { + id: 'power-timeline-per-gpu', + label: t.perGpu, + checked: lineMode === 'gpu', + onCheckedChange: (checked) => chooseLineMode(checked ? 'gpu' : 'mean'), + infoTooltip: t.perGpuHelp, + }, + ]; + if (hasPools || lineMode === 'pool') { + switches.push({ + id: 'power-timeline-pools', + label: t.pools, + checked: lineMode === 'pool', + onCheckedChange: (checked) => chooseLineMode(checked ? 'pool' : 'mean'), + infoTooltip: t.poolsHelp, + }); + } + switches.push( + { + id: 'power-timeline-window-only', + label: t.windowOnly, + checked: windowOnly, + onCheckedChange: (checked) => { + setWindowOnly(checked); + track('inference_power_timeline_window_changed', { windowOnly: checked }); + }, + infoTooltip: t.windowOnlyHelp, + }, + { + id: 'power-timeline-utility', + label: t.utilityLines, + checked: showUtility, + onCheckedChange: (checked) => { + setShowUtility(checked); + track('inference_power_timeline_utility_toggled', { enabled: checked }); + }, + infoTooltip: t.utilityHelp, + }, + ); + + const legendElement = ( + { + setIsLegendExpanded(expanded); + track('latency_legend_expanded', { expanded }); + }} + onItemHover={(id) => setHighlight(id)} + onItemHoverEnd={() => setHighlight(null)} + hideAtomFootnote + switches={switches} + /> + ); + + const runInfos = useMemo( + () => + fetchedRequests + .map((request) => responses.get(request.runId)?.runInfo) + .filter((info): info is NonNullable => Boolean(info)), + [fetchedRequests, responses], + ); + + const toolbar = ( +
+ + {t.timeAxis} + + value={xMode} + ariaLabel={t.timeAxis} + role="group" + options={[ + { value: 'wall', label: t.wall, testId: 'power-timeline-axis-wall' }, + { value: 'elapsed', label: t.elapsed, testId: 'power-timeline-axis-elapsed' }, + { value: 'serving', label: t.serving, testId: 'power-timeline-axis-serving' }, + ]} + onValueChange={(mode) => { + setXModeChoice(mode); + track('inference_power_timeline_axis_changed', { mode }); + }} + /> +
+ ); + + let emptyMessage: string | null = null; + if (loadingRuns > 0) emptyMessage = t.loading(loadingRuns); + else if (!hasAnyArtifact) emptyMessage = t.noArtifacts; + else if (visibleTraces.length === 0) emptyMessage = t.noTraces; + else if (missingWindows === visibleTraces.length) emptyMessage = t.missingWindow(missingWindows); + + return ( +
+ + key={`${chartId}-${xMode}-${lineMode}-${windowOnly}`} + chartId={chartId} + data={samples} + height={CHART_HEIGHT} + margin={MARGIN} + watermark={overlayData && overlayPoints.length > 0 ? 'unofficial' : 'logo'} + testId="power-timeline-chart-svg" + grabCursor + instructions={t.instructions} + xScale={ + xMode === 'wall' + ? { type: 'time', domain: [new Date(xDomain[0]), new Date(xDomain[1])] } + : { type: 'linear', domain: xDomain } + } + yScale={{ type: 'linear', domain: yDomain, nice: true }} + xAxis={{ + label: xMode === 'wall' ? t.xWall : xMode === 'serving' ? t.xServing : t.xElapsed, + tickValues: (scale) => { + const timeScale = scale as + | d3.ScaleTime + | d3.ScaleLinear; + const [left, right] = timeScale.range(); + return timeScale.ticks(Math.max(2, Math.min(10, Math.floor((right - left) / 80)))); + }, + tickFormat: xTickFormat, + }} + yAxis={{ label: lineMode === 'pool' ? t.yPool : yLabel, tickCount: 8 }} + layers={layers} + displayIdentity={activeHighlight ?? ''} + onDisplayUpdate={onDisplayUpdate} + zoom={{ + enabled: true, + axes: 'x', + scaleExtent: [1, 60], + resetEventName: `power_timeline_zoom_reset_${chartId}`, + }} + tooltip={{ + rulerType: 'crosshair', + content: tooltipContent, + getRulerX: (sample, xScale) => (xScale as AnyContinuousScale)(sample.x), + getRulerY: (sample, yScale) => yScale(sample.y), + onHoverStart: (selection) => { + selection.attr('r', 5).attr('stroke', 'white').attr('stroke-width', 1); + }, + onHoverEnd: (selection) => { + selection.attr('r', 2).attr('stroke', 'none'); + }, + attachToLayer: 2, + }} + legendElement={legendElement} + caption={ + <> + {caption} + {toolbar} + + } + noDataOverlay={ + emptyMessage ? ( +
+ {emptyMessage} +
+ ) : undefined + } + /> +
+ {focusedTrace && ( +

+ + {t.focused( + `${seriesLabel(focusedTrace.point)} · ${traceConfigLabel(focusedTrace.point)}`, + )} + + +

+ )} + {focusKey && !focusedTrace && loadingRuns === 0 && ( +

{t.missingFocus}

+ )} + {errors.map(({ request, error }) => ( +

+ {t.loadError(request.runId, error.message)} +

+ ))} + {droppedRuns > 0 &&

{t.droppedRuns(droppedRuns)}

} + {runInfos.length > 0 && ( +

+ {t.telemetry}:{' '} + {runInfos.map((info, index) => ( + + {index > 0 && ' · '} + + {`run ${info.id}`} + + {info.createdAt ? ` (${formatUtcDate(new Date(info.createdAt))})` : ''} + + ))} +

+ )} +
+ {loadingRuns === 0 && visibleTraces.length > 0 && ( + colorForTrace(trace).color} + hardwareLabel={seriesLabel} + isOverlay={(point) => overlayPointSet.has(point)} + /> + )} +
+ ); +} diff --git a/packages/app/src/components/inference/ui/PowerTimelineSummary.tsx b/packages/app/src/components/inference/ui/PowerTimelineSummary.tsx new file mode 100644 index 000000000..41e49ccef --- /dev/null +++ b/packages/app/src/components/inference/ui/PowerTimelineSummary.tsx @@ -0,0 +1,189 @@ +'use client'; + +import * as d3 from 'd3'; +import { useMemo } from 'react'; +import { HW_REGISTRY } from '@semianalysisai/inferencex-constants'; + +import type { GpuPowerSeriesResponse } from '@/components/gpu-power/power-series'; +import { useLocale } from '@/lib/use-locale'; +import type { InferenceData } from '../types'; +import { + allGpuPool, + runAttemptFromUrl, + summarizeTraceWindow, + telemetryNameForPoint, + traceConfigLabel, + tracePools, + type PowerPoolRole, + type PowerTimelineTrace, +} from '../utils/powerTimeline'; + +const STRINGS = { + en: { + title: 'Validated-window summary', + trace: 'Hardware · config', + run: 'Run · telemetry source', + window: 'Window (s)', + pool: 'Pool · GPUs', + validated: 'Validated average (W/GPU)', + peak: 'Peak 1-s pool power (W)', + poolTdp: 'Pool TDP (W)', + pools: { all: 'All GPUs', prefill: 'Prefill', decode: 'Decode' } satisfies Record< + PowerPoolRole, + string + >, + source: { database: 'database', github: 'GitHub artifact fallback' }, + attempt: (attempt: number) => `attempt ${attempt}`, + unofficial: 'unofficial', + noWindow: 'Not recorded', + }, + zh: { + title: '有效测量窗口汇总', + trace: '硬件 · 配置', + run: '运行 · 遥测来源', + window: '窗口(秒)', + pool: 'GPU 池 · GPU 数', + validated: '有效窗口平均值(W/GPU)', + peak: '池功耗 1 秒峰值(W)', + poolTdp: '池 TDP(W)', + pools: { all: '全部 GPU', prefill: '预填充', decode: '解码' } satisfies Record< + PowerPoolRole, + string + >, + source: { database: '数据库', github: 'GitHub 产物回退' }, + attempt: (attempt: number) => `第 ${attempt} 次尝试`, + unofficial: '非官方', + noWindow: '未记录', + }, +}; + +const watts = d3.format(',.0f'); +const seconds = d3.format(',.0f'); + +const validatedWatts = (point: InferenceData, role: PowerPoolRole) => + (role === 'prefill' + ? point.measuredPrefillAvgPower + : role === 'decode' + ? point.measuredDecodeAvgPower + : point.measuredAvgPower + )?.y ?? null; + +/** + * Per-trace provenance behind the Timeline curves: which run and telemetry file + * each trace came from, the validated window length, and per pool its GPU + * count, the row's validated average, the drawn peak and the rated pool TDP. + */ +export default function PowerTimelineSummary({ + traces, + responses, + colorOf, + hardwareLabel, + isOverlay, +}: { + traces: readonly PowerTimelineTrace[]; + responses: ReadonlyMap; + colorOf: (trace: PowerTimelineTrace) => string; + hardwareLabel: (point: InferenceData) => string; + isOverlay: (point: InferenceData) => boolean; +}) { + const locale = useLocale(); + const t = STRINGS[locale]; + const rows = useMemo( + () => + traces.map((trace) => ({ + trace, + summary: summarizeTraceWindow(trace, [ + allGpuPool(trace.series), + ...tracePools(trace.series), + ]), + })), + [traces], + ); + return ( +
+

{t.title}

+
+ + + + + + + + + + + + + {rows.map(({ trace, summary }) => { + const { point } = trace; + const attempt = runAttemptFromUrl(point.run_url); + const name = telemetryNameForPoint(point); + const source = responses.get(trace.runId)?.source; + const tdp = HW_REGISTRY[point.hwKey.split('_')[0]]?.tdp ?? 0; + return ( + + {summary.pools.map((pool, index) => { + const validated = validatedWatts(point, pool.role); + return ( + + {index === 0 && ( + <> + + + + + )} + + + + + + ); + })} + + ); + })} +
{t.trace}{t.run}{t.window}{t.pool}{t.validated}{t.peak}{t.poolTdp}
+ ● + {hardwareLabel(point)} + + {traceConfigLabel(point)} + {isOverlay(point) ? ` · ${t.unofficial}` : ''} + + + + {`run ${trace.runId}`} + + {attempt === null ? '' : ` · ${t.attempt(attempt)}`} + + {name ? `power_validation_${name}.json` : '—'} + {source ? ` · ${t.source[source]}` : ''} + + + {summary.windowSeconds === null + ? t.noWindow + : seconds(summary.windowSeconds)} + + {t.pools[pool.role]} · {pool.gpuCount} + + {validated === null ? '—' : watts(validated)} + + {pool.peakWatts === null ? '—' : watts(pool.peakWatts)} + + {tdp > 0 ? watts(tdp * pool.gpuCount) : '—'} +
+
+
+ ); +} diff --git a/packages/app/src/components/inference/ui/QuickFiltersDialog.tsx b/packages/app/src/components/inference/ui/QuickFiltersDialog.tsx index 2e9778095..eead845db 100644 --- a/packages/app/src/components/inference/ui/QuickFiltersDialog.tsx +++ b/packages/app/src/components/inference/ui/QuickFiltersDialog.tsx @@ -7,6 +7,7 @@ import { OptionInfo } from '@/components/ui/option-info'; import type { DeploymentMode, SpecMode } from '@/components/inference/types'; import type { PowerTier } from '@/lib/power-tier'; import { FRAMEWORK_FAMILIES } from '@/components/inference/utils/quickFilters'; +import { topologyLabel } from '@/components/inference/utils/topology-filter'; import { useInferenceActions, @@ -56,6 +57,9 @@ const STRINGS = { specHelp: 'MTP groups runs with speculative decoding enabled, including methods such as EAGLE. STP groups standard decoding without speculative decoding. Available only for fixed-sequence benchmarks.', power: 'Measured Power', + topology: 'Topology', + topologyHelp: + 'Keep one GPU allocation and parallelism configuration across its observed concurrency sweep. This does not match software versions or fill missing measurements.', certified: 'Validated', legacyTier: 'Historical', validatedTitle: 'Validated measurement', @@ -97,6 +101,9 @@ const STRINGS = { specHelp: 'MTP 组包含启用投机解码的运行,也包括 EAGLE 等方法;STP 组为未启用投机解码的标准解码运行。该筛选仅适用于固定序列长度基准测试。', power: '实测功耗', + topology: '拓扑', + topologyHelp: + '按 GPU 分配和并行配置保留整条已测并发曲线。此筛选不会匹配软件版本,也不会补造缺失数据。', certified: '已验证', legacyTier: '历史测量', validatedTitle: '已验证测量', @@ -151,6 +158,7 @@ export function QuickFiltersDialog({ setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, } = useInferenceActions(); const { selectedSequence, quickFilters, lockedFrameworks } = useInferenceFilters(); const { availableQuickFilters } = useInferenceData(); @@ -160,6 +168,7 @@ export function QuickFiltersDialog({ framework:

{t.frameworkHelp}

, deployment:

{t.deploymentHelp}

, spec:

{t.specHelp}

, + topology:

{t.topologyHelp}

, power: ( <>
@@ -186,9 +195,12 @@ export function QuickFiltersDialog({ label: framework.label, available: availableQuickFilters.frameworks.includes(framework.key), })); + const topologyKeys = [ + ...new Set([...(availableQuickFilters.topologies ?? []), ...(quickFilters.topologies ?? [])]), + ]; const groups: { - key: 'vendor' | 'framework' | 'deployment' | 'spec' | 'power'; + key: 'vendor' | 'framework' | 'deployment' | 'spec' | 'power' | 'topology'; label: string; options: readonly { value: string; label: string; available: boolean }[]; selected: readonly string[]; @@ -254,6 +266,16 @@ export function QuickFiltersDialog({ })), selected: quickFilters.power, }, + { + key: 'topology', + label: t.topology, + options: topologyKeys.map((value) => ({ + value, + label: topologyLabel(value, locale, topologyKeys), + available: (availableQuickFilters.topologies ?? []).includes(value), + })), + selected: quickFilters.topologies ?? [], + }, ]; const selectedCount = groups.reduce((count, group) => count + group.selected.length, 0); @@ -261,26 +283,29 @@ export function QuickFiltersDialog({ // Every option is a visible toggle: one tap adds or removes it, no dropdown // to open first. Empty selection in a group means "all". const handleToggle = ( - category: 'vendor' | 'framework' | 'deployment' | 'spec' | 'power', + category: 'vendor' | 'framework' | 'deployment' | 'spec' | 'power' | 'topology', value: string, ) => { const previous: readonly string[] = - category === 'vendor' - ? quickFilters.vendors - : category === 'framework' - ? quickFilters.frameworks - : category === 'deployment' - ? quickFilters.deployment - : category === 'power' - ? quickFilters.power - : quickFilters.spec; + category === 'topology' + ? (quickFilters.topologies ?? []) + : category === 'vendor' + ? quickFilters.vendors + : category === 'framework' + ? quickFilters.frameworks + : category === 'deployment' + ? quickFilters.deployment + : category === 'power' + ? quickFilters.power + : quickFilters.spec; const values = toggleValue(previous, value); track('inference_quick_filter_toggled', { category, value, active: values.includes(value), }); - if (category === 'vendor') setQuickFilterVendors(values); + if (category === 'topology') setQuickFilterTopologies(values); + else if (category === 'vendor') setQuickFilterVendors(values); else if (category === 'framework') setQuickFilterFrameworks(values); else if (category === 'deployment') setQuickFilterDeployment(values as DeploymentMode[]); else if (category === 'power') setQuickFilterPower(values as PowerTier[]); @@ -293,6 +318,7 @@ export function QuickFiltersDialog({ setQuickFilterDeployment([]); setQuickFilterSpec([]); setQuickFilterPower([]); + setQuickFilterTopologies([]); track('inference_quick_filters_cleared', { source: 'dialog' }); }; @@ -366,7 +392,10 @@ export function QuickFiltersDialog({ role="group" aria-labelledby={`quick-filter-${group.key}-label`} data-testid={`quick-filter-${group.key}-options`} - className="flex flex-wrap gap-2" + className={cn( + 'flex min-w-0 flex-wrap gap-2', + group.key === 'topology' && 'max-h-56 overflow-y-auto overscroll-contain pr-1', + )} > {group.options.map((option) => { const active = group.selected.includes(option.value); @@ -382,6 +411,8 @@ export function QuickFiltersDialog({ title={disabled ? t.noData : undefined} className={cn( 'rounded-full font-normal', + group.key === 'topology' && + 'h-auto min-h-9 max-w-full whitespace-normal break-words rounded-lg py-2 text-left md:h-auto', active && 'bg-brand hover:bg-brand/90', )} data-testid={`quick-filter-${group.key}-${option.value}`} diff --git a/packages/app/src/components/inference/ui/ScatterGraph.decoration.test.tsx b/packages/app/src/components/inference/ui/ScatterGraph.decoration.test.tsx index 8e0a422cf..19f0543c7 100644 --- a/packages/app/src/components/inference/ui/ScatterGraph.decoration.test.tsx +++ b/packages/app/src/components/inference/ui/ScatterGraph.decoration.test.tsx @@ -4,7 +4,6 @@ import { act } from 'react'; import { describe, expect, it, vi } from 'vitest'; import type { InferenceData } from '@/components/inference/types'; -import { renderLegacyPowerRing } from '@/components/inference/utils/legacy-power-marker'; import { renderOffloadHalo } from '@/components/inference/utils/offload-halo'; import { @@ -130,41 +129,6 @@ describe('ScatterGraph toggle decoration', () => { unmount(); }); - it('rings legacy-power points only on Measured Energy axes', () => { - const legacy = { - ...point('h100', 'fp8', 10, 100, 1), - power_tier: 'legacy', - } as InferenceData; - const certified = { - ...point('h100', 'fp8', 20, 200, 2), - power_tier: 'certified', - } as InferenceData; - const tierless = point('h100', 'fp8', 40, 400, 4); - - inferenceState.current = { - ...baseInferenceState(), - selectedYAxisMetric: 'y_measuredAvgPower', - }; - const measured = mountChart({ data: [legacy, certified, tierless] }); - const groups = dotGroups(measured.container); - - expect(groups[0].querySelector('.legacy-power-ring')).not.toBeNull(); - expect(groups[1].querySelector('.legacy-power-ring')).toBeNull(); - expect(groups[2].querySelector('.legacy-power-ring')).toBeNull(); - // The legend key lives in ChartDisplay's axis-metric footer, not the chart. - expect(measured.container.querySelector('[data-testid="legacy-power-key"]')).toBeNull(); - measured.unmount(); - - // Non-measured axis: the same legacy point renders without a ring. - inferenceState.current = { - ...baseInferenceState(), - selectedYAxisMetric: 'y', - }; - const throughput = mountChart({ data: [legacy, certified, tierless] }); - expect(throughput.container.querySelectorAll('.legacy-power-ring')).toHaveLength(0); - throughput.unmount(); - }); - it('reads current trace availability without changing metric identity', () => { const agenticPoint = { ...point('h100', 'fp8', 20, 200, 2), @@ -466,41 +430,32 @@ function recordCount(root: Node, run: () => void): number { return count; } -const offloaded = { offload_mode: 'on', power_tier: 'legacy' } as InferenceData; +const offloaded = { offload_mode: 'on' } as InferenceData; -describe('point decoration rings', () => { - it('draw the offload halo and legacy-power ring', () => { +describe('offload halo decoration', () => { + it('draws the offload halo', () => { const group = pointGroup(); renderOffloadHalo(group, offloaded, 'red'); - renderLegacyPowerRing(group, offloaded, true, 'red'); expect(group.select('.offload-halo').attr('stroke')).toBe('red'); expect(group.select('.offload-halo').attr('stroke-dasharray')).toBe('3 2'); - expect(group.select('.legacy-power-ring').attr('stroke-dasharray')).toBe('1 3'); }); - it('write nothing when re-rendered with the same state', () => { + it('writes nothing when re-rendered with the same state', () => { const group = pointGroup(); renderOffloadHalo(group, offloaded, 'red'); - renderLegacyPowerRing(group, offloaded, true, 'red'); - expect( - recordCount(group.node()!, () => { - renderOffloadHalo(group, offloaded, 'red'); - renderLegacyPowerRing(group, offloaded, true, 'red'); - }), - ).toBe(0); + expect(recordCount(group.node()!, () => renderOffloadHalo(group, offloaded, 'red'))).toBe(0); }); - it('restyle and remove on real changes', () => { + it('restyles and removes on real changes', () => { const group = pointGroup(); renderOffloadHalo(group, offloaded, 'red'); - renderLegacyPowerRing(group, offloaded, true, 'red'); renderOffloadHalo(group, offloaded, 'blue'); - renderLegacyPowerRing(group, offloaded, false, 'blue'); - expect(group.select('.offload-halo').attr('stroke')).toBe('blue'); - expect(group.select('.legacy-power-ring').empty()).toBe(true); + + renderOffloadHalo(group, { offload_mode: 'off' } as InferenceData, 'blue'); + expect(group.select('.offload-halo').empty()).toBe(true); }); }); diff --git a/packages/app/src/components/inference/ui/ScatterGraph.overlay.test.tsx b/packages/app/src/components/inference/ui/ScatterGraph.overlay.test.tsx index 13f930f08..22d68c48a 100644 --- a/packages/app/src/components/inference/ui/ScatterGraph.overlay.test.tsx +++ b/packages/app/src/components/inference/ui/ScatterGraph.overlay.test.tsx @@ -214,17 +214,7 @@ describe('ScatterGraph unofficial overlays', () => { const visible = showAllMeasurements || datum.x !== 20; expect(group.style.opacity).toBe(visible ? '1' : '0'); expect(group.style.pointerEvents).toBe(visible ? 'auto' : 'none'); - expect(Boolean(group.querySelector('.legacy-power-ring'))).toBe( - datum.power_tier === 'legacy', - ); } - expect( - container.querySelector('[data-testid="measured-power-summary"]')?.textContent, - ).toContain( - showAllMeasurements - ? 'Showing 9 of 9 measured points: 3/3 validated · 6/6 historical.' - : 'Showing 6 of 9 measured points: 3/3 validated · 3/6 historical.', - ); expect(curves.map((curve) => curve.getAttribute('d'))).toEqual(paths); expect(axes.map((axis) => axis.innerHTML)).toEqual(axisGeometry); expect(groups.map((group) => group.getAttribute('transform'))).toEqual(positions); @@ -233,38 +223,6 @@ describe('ScatterGraph unofficial overlays', () => { unmount(); }); - it('includes unofficial measured points in the validated/historical coverage summary', () => { - const runUrl = 'https://github.com/o/r/actions/runs/123'; - const overlayPoints = [ - { ...point('h100', 'fp8', 30, 300, 2), power_tier: 'certified', run_url: runUrl }, - { ...point('h100', 'fp8', 35, 350, 4), power_tier: 'legacy', run_url: runUrl }, - ] as InferenceData[]; - inferenceState.current = { - ...baseInferenceState(), - selectedYAxisMetric: 'y_measuredJPerOutputToken', - }; - overlayState.current = { - ...baseOverlayState(), - isUnofficialRun: true, - activeOverlayHwTypes: new Set(['h100']), - allOverlayHwTypes: new Set(['h100']), - runIndexByUrl: { [runUrl]: 0 }, - unofficialRunInfos: [{ id: '123', branch: 'test-branch', url: runUrl }], - }; - - const { container, unmount } = mountChart({ - overlayData: { - data: overlayPoints, - hardwareConfig: HARDWARE_CONFIG, - } as unknown as Parameters[0]['overlayData'], - }); - - expect( - container.querySelector('[data-testid="measured-power-summary"]')?.textContent, - ).toContain('Showing 2 of 2 measured points: 1/1 validated · 1/1 historical.'); - unmount(); - }); - it('keeps unofficial-run overlay markers rendered through official toggles', () => { const overlayPoints = [point('h100', 'fp8', 30, 300, 2), point('h100', 'fp8', 35, 350, 4)].map( (p) => ({ ...p, run_url: 'https://github.com/o/r/actions/runs/123' }), diff --git a/packages/app/src/components/inference/ui/ScatterGraph.test-harness.tsx b/packages/app/src/components/inference/ui/ScatterGraph.test-harness.tsx index 2cebdba58..5e72fb046 100644 --- a/packages/app/src/components/inference/ui/ScatterGraph.test-harness.tsx +++ b/packages/app/src/components/inference/ui/ScatterGraph.test-harness.tsx @@ -80,6 +80,10 @@ class MockResizeObserver { unobserve() {} disconnect() {} } +const originalGetComputedTextLength = Object.getOwnPropertyDescriptor( + SVGElement.prototype, + 'getComputedTextLength', +); const originalGetBBox = Object.getOwnPropertyDescriptor(SVGElement.prototype, 'getBBox'); export const point = ( @@ -220,6 +224,12 @@ export const rebuildCount = () => vi.mocked(setupChartStructure).mock.calls.leng beforeEach(() => { globalThis.__scatterPathnameState.value = '/inference'; vi.stubGlobal('ResizeObserver', MockResizeObserver); + Object.defineProperty(SVGElement.prototype, 'getComputedTextLength', { + configurable: true, + value(this: SVGElement) { + return (this.textContent?.length ?? 0) * 7; + }, + }); Object.defineProperty(SVGElement.prototype, 'getBBox', { configurable: true, value: () => @@ -259,6 +269,15 @@ beforeEach(() => { afterEach(() => { vi.unstubAllGlobals(); vi.restoreAllMocks(); + if (originalGetComputedTextLength) { + Object.defineProperty( + SVGElement.prototype, + 'getComputedTextLength', + originalGetComputedTextLength, + ); + } else { + Reflect.deleteProperty(SVGElement.prototype, 'getComputedTextLength'); + } if (originalGetBBox) { Object.defineProperty(SVGElement.prototype, 'getBBox', originalGetBBox); } else { diff --git a/packages/app/src/components/inference/ui/ScatterGraph.tsx b/packages/app/src/components/inference/ui/ScatterGraph.tsx index fc860d3ad..e23738e50 100644 --- a/packages/app/src/components/inference/ui/ScatterGraph.tsx +++ b/packages/app/src/components/inference/ui/ScatterGraph.tsx @@ -1,10 +1,13 @@ 'use client'; +import { useFeatureGate } from '@/lib/use-feature-gate'; +import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config'; import { track } from '@/lib/analytics'; import { isPersistedBenchmarkId } from '@/lib/benchmark-id'; import { useEphemeralUrlState } from '@/hooks/useUrlState'; import { rememberChartStateInUrl } from '@/lib/url-state'; import * as d3 from 'd3'; +import { CHART_TYPE } from '@/lib/d3-chart/typography'; import dynamic from 'next/dynamic'; import React, { useCallback, useEffect, useLayoutEffect, useMemo, useRef, useState } from 'react'; @@ -16,6 +19,7 @@ import { useInferenceDisplay, useInferenceFilters, } from '@/components/inference/InferenceContext'; +import { usePerfRulerStore } from '@/components/inference/perf-ruler-store'; import { useTraceAvailability } from '@/hooks/api/use-trace-availability'; import { useLogAvailability } from '@/hooks/api/use-log-availability'; import { computeToggle } from '@/hooks/useTogglableSet'; @@ -47,7 +51,13 @@ import { matchKnownConfigIssues, pointMatchesIssue } from '@/lib/known-issues'; import { useLocale } from '@/lib/use-locale'; import { getLineLabelVendorIcon } from '@/lib/vendor-logos'; import { formatNumber, getDisplayLabel, updateRepoUrl } from '@/lib/utils'; -import { getInferenceHardwareConfig, getInferenceRunLabel } from '@/lib/inference-labels'; +import { + getInferenceHardwareConfig, + getInferenceRunLabel, + getOverlayLineLabel, + OVERLAY_LABEL_MARKER, + overlayRunTag, +} from '@/lib/inference-labels'; import { D3Chart } from '@/lib/d3-chart/D3Chart'; import type { CustomLayerConfig, @@ -76,6 +86,7 @@ import { renderPerfRulers, type PerfRulerEndInput, type PerfRulerGeometry, + type PerfRulerMeasurement, type PerfRulerRenderEntry, type PerfRulerState, } from '@/lib/d3-chart/layers/perf-ruler'; @@ -111,6 +122,7 @@ import { chartFrontier, upperPowerEnvelope, isPowerCurveMetric, + isPowerGaugeSeries, isMeasuredPowerCurveMetric, } from '@/components/inference/utils/powerCurves'; import type { @@ -118,29 +130,45 @@ import type { ClippedInferenceData, InferenceData, ScatterGraphProps, + PowerVariant, } from '@/components/inference/types'; import { generateOverlayTooltipContent, generateTooltipContent, } from '@/components/inference/utils/tooltipUtils'; +import { usePowerTraceAction } from '@/components/inference/hooks/usePowerTraceAction'; import { QuickFiltersDialog } from '@/components/inference/ui/QuickFiltersDialog'; import { ScatterEmptyState } from '@/components/inference/ui/ScatterEmptyState'; +import { ALL_IN_MEASURED_EMPTY } from '@/lib/power-basis'; import { scatterPointConfigId, scatterPointJoinId, + parseScatterSeriesKey, + scatterSeriesKey, } from '@/components/inference/utils/point-identity'; +import { + flatSeriesValue, + inferPowerCompare, + lineLabelHardwareKey, + lineLabelSeriesId, + metricPlotsWatts, + powerCompareBase, + powerLineLabelSuffix, + powerVariantDash, + powerVariantId, + powerVariantLabel, + powerVariantsInData, +} from '@/components/inference/utils/power-compare'; +import FrontierPointsPanel from '@/components/inference/ui/FrontierPointsPanel'; import LegendPointsDialog from '@/components/inference/ui/LegendPointsDialog'; import { renderOffloadHalo } from '@/components/inference/utils/offload-halo'; -import { renderLegacyPowerRing } from '@/components/inference/utils/legacy-power-marker'; -import { - countPowerTiers, - MeasuredPowerSummary, -} from '@/components/inference/ui/MeasuredPowerSummary'; import { + isAllInMeasuredConfigKey, isMeasuredEnergyConfigKey, isRoleLocalMeasuredEnergyConfigKey, } from '@/components/inference/metric-registry'; import { buildLegendPointsRows } from '@/components/inference/utils/legend-points-table'; +import { groupConcurrencySeries } from '@/components/inference/utils/concurrency-series'; import { resolveScatterXAxisScale } from '@/components/inference/utils/x-axis-scale'; import { pointLabelText } from './point-label'; import { @@ -163,6 +191,14 @@ import { fitContinuationLabelBaseline, } from '@/components/inference/utils/overflowContinuations'; +const PowerTelemetryDialog = dynamic( + () => + import('@/components/inference/power-telemetry-dialog').then( + (module) => module.PowerTelemetryDialog, + ), + { ssr: false }, +); + const FixedSequenceLogDialog = dynamic( () => import('@/components/inference/log-viewer/fixed-sequence-log-dialog').then( @@ -221,9 +257,35 @@ const optimalPointKey = (d: InferenceData): string => const EMPTY_OVERLAY_DATA: InferenceData[] = []; const EMPTY_CLIPPED_DATA: ClippedInferenceData[] = []; +/** + * Legend ids of the comparison-series rows (`i_pcompare`), distinct from + * hardware keys so the shared hover / toggle handlers can tell them apart. + */ +const POWER_VARIANT_LEGEND_PREFIX = 'power-variant:'; +/** Comparison clones sit behind the base series they annotate. */ +const POWER_VARIANT_POINT_OPACITY = 0.6; +const pointOpacityForVariant = (d: InferenceData): number => + d.powerVariant ? POWER_VARIANT_POINT_OPACITY : 1; +/** Dash for a series key's variant id (`parseScatterSeriesKey().variant`). */ +const powerVariantDashById = (variantId: string | null | undefined): string => + variantId ? (VARIANT_DASH_BY_ID.get(variantId) ?? '') : ''; +const VARIANT_DASH_BY_ID = new Map( + ( + [ + ['basis', 'gpu-measured'], + ['basis', 'gpu-provisioned'], + ['basis', 'utility-provisioned'], + ['basis', 'utility-modeled'], + ['role', 'all'], + ['role', 'prefill'], + ['role', 'decode'], + ] as const + ).map(([kind, id]) => [id, powerVariantDash({ kind, id } as PowerVariant)]), +); + const LINE_LABEL_RAISE = ['.line-label'] as const; /** Decorations sit above the visible shape, which a precision toggle may replace. */ -const POINT_DECORATION_RAISE = ['.offload-halo', '.legacy-power-ring'] as const; +const POINT_DECORATION_RAISE = ['.offload-halo'] as const; /** Rebuilds and zoom frames re-sync every gradient stop; skip unchanged writes. */ function syncGradientStop(this: SVGStopElement, stop: { offset: number; color: string }): void { @@ -385,6 +447,8 @@ const pointCountEn = (count: number) => `${count} ${count === 1 ? 'point' : 'poi const SCATTER_STRINGS = { en: { + concurrencyCurves: + 'Dots are observed loads. Straight segments connect only matching topology, recipe and run; repeated loads remain separate markers. Concurrency is not a higher-is-better score.', logScale: 'Log Scale', optimalOnly: 'Optimal Only', paretoFrontier: 'Pareto Frontier', @@ -394,10 +458,6 @@ const SCATTER_STRINGS = { optimalInfo: 'Optimal points form the Pareto frontier for the selected axes.', powerBoundaryInfo: 'Show only points on the upper measured power boundary. Turn off to show all measurements; the boundary stays the same. This is a power-load boundary, not an energy-efficiency frontier.', - powerCurves: - 'Smooth lines trace the upper power boundary across tested configurations. Dots are measured; lines are interpolated, not efficiency frontiers.', - powerOptimal: - 'A power Pareto frontier can contain a single point. Turn off Optimal Only to show the upper power boundary.', labels: 'Labels', highContrast: 'High Contrast', parallelismLabels: 'Parallelism Labels', @@ -417,12 +477,16 @@ const SCATTER_STRINGS = { noDataHint: 'Please change the model, sequence, precision, date range or chip selection.', noRoleEnergyDataHint: 'This dataset does not report role-level prefill/decode energy. Choose a different model, scenario, precision, date, or measured-energy metric.', + noMeasuredDataHint: + 'No measured GPU power is reported for this selection. Choose other chip configs, or a different model, scenario, precision or date.', unofficialTitle: (branch: string) => `UNOFFICIAL: ${branch}`, unofficialRun: 'UNOFFICIAL RUN', branch: 'Branch', viewWorkflow: 'View workflow run', }, zh: { + concurrencyCurves: + '点表示实测负载;直线段仅连接相同拓扑、配方和运行的数据,重复负载保留为独立点。并发数不是越高越好的分数。', logScale: '对数缩放', optimalOnly: '仅最优', paretoFrontier: 'Pareto 前沿', @@ -432,9 +496,6 @@ const SCATTER_STRINGS = { optimalInfo: '最优点构成当前所选坐标轴的 Pareto 前沿。', powerBoundaryInfo: '仅显示实测功率上边界上的点。关闭后显示全部测量点,边界曲线保持不变。这是功率负载边界,不是能效前沿。', - powerCurves: - '平滑曲线勾勒各测试配置的功耗上边界。数据点来自实测,曲线通过插值得到,不代表能效 Pareto 前沿。', - powerOptimal: '功耗的 Pareto 前沿可能只有一个点。关闭“仅最优”即可查看功耗上边界。', labels: '标签', highContrast: '高对比度', parallelismLabels: '并行配置标签', @@ -454,6 +515,8 @@ const SCATTER_STRINGS = { noDataHint: '请调整模型、序列长度、精度、日期范围或芯片选项。', noRoleEnergyDataHint: '当前数据集未提供 Prefill/Decode 各角色的能耗数据。请选择其他模型、场景、精度、日期或实测能耗指标。', + noMeasuredDataHint: + '当前选择没有实测 GPU 功耗数据。请选择其他芯片配置,或更换模型、场景、精度或日期。', unofficialTitle: (branch: string) => `非官方:${branch}`, unofficialRun: '非官方运行', branch: '分支', @@ -534,39 +597,57 @@ const ScatterGraph = React.memo( setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, } = useInferenceActions(); - const paretoDirection = chartDefinition[`${selectedYAxisMetric}_roofline`] as - | ParetoDirection - | undefined; + const isConcurrencyAxis = chartDefinition.x_scale_field === 'conc'; + const paretoDirection = ( + isConcurrencyAxis ? undefined : chartDefinition[`${selectedYAxisMetric}_roofline`] + ) as ParetoDirection | undefined; const hideNonOptimal = preferOptimalOnly && Boolean(paretoDirection); const isPowerAxis = isPowerCurveMetric(selectedYAxisMetric); const isMeasuredPowerAxis = isMeasuredPowerCurveMetric(selectedYAxisMetric); // Measured power describes the load sweep. Keep its boundary fixed while // Optimal Only changes marker visibility, as on the other scatter charts. - const showPowerEnvelope = isPowerAxis && (isMeasuredPowerAxis || !hideNonOptimal); + const showPowerEnvelope = + !isConcurrencyAxis && isPowerAxis && (isMeasuredPowerAxis || !hideNonOptimal); const showAllMeasurements = isMeasuredPowerAxis ? !hideNonOptimal : savedShowAllMeasurements; - const supportsGradientLabels = !showPowerEnvelope || isMeasuredPowerAxis; + const supportsGradientLabels = + !isConcurrencyAxis && (!showPowerEnvelope || isMeasuredPowerAxis); const showGradientLabels = preferGradientLabels && supportsGradientLabels; const groupDisplayedPoints = useCallback( (points: InferenceData[]) => { + if (isConcurrencyAxis) return groupConcurrencySeries(points); const groups = groupPointsByDate(points); if (showPowerEnvelope) { for (const [date, samples] of groups) { - groups.set(date, upperPowerEnvelope(samples, chartDefinition.chartType !== 'e2e')); + groups.set( + date, + upperPowerEnvelope( + samples, + chartDefinition.chartType !== 'e2e', + isPowerGaugeSeries(selectedYAxisMetric, samples[0]), + ), + ); } } return groups; }, - [showPowerEnvelope, chartDefinition.chartType], + [isConcurrencyAxis, showPowerEnvelope, chartDefinition.chartType, selectedYAxisMetric], ); const locale = useLocale(); + const featureGateUnlocked = useFeatureGate(); + const showPowerTelemetry = + featureGateUnlocked || getMeasuredMetricConfig(selectedYAxisMetric) !== undefined; const legendT = SCATTER_STRINGS[locale]; + // Comparison series (`i_pcompare`) switched off from the legend. Chart-local, + // like Optimal Only's point set: the URL carries the comparison, not which + // of its rows a reader hid while looking. + const [hiddenPowerVariants, setHiddenPowerVariants] = useState>( + () => new Set(), + ); const ephemeralUrlState = useEphemeralUrlState(); const costLimit = chartDefinition.y_cost_limit ?? 0; const latencyLimit = chartDefinition.y_latency_limit ?? 0; - // Legacy-power rings decorate points only while a Measured Energy y-axis - // is selected (see legacy-power-marker.ts). - const isMeasuredEnergyAxis = isMeasuredEnergyConfigKey(selectedYAxisMetric); const { isUnofficialRun, @@ -801,7 +882,7 @@ const ScatterGraph = React.memo( () => data.reduce( (acc, point) => { - const key = `${point.hwKey}_${point.precision}`; + const key = scatterSeriesKey(point); if (!acc[key]) acc[key] = []; acc[key].push(point); return acc; @@ -835,7 +916,7 @@ const ScatterGraph = React.memo( return result; }, [groupedData, selectedYAxisMetric, chartDefinition]); - const displayedRooflines = showPowerEnvelope ? groupedData : rooflines; + const displayedRooflines = isConcurrencyAxis || showPowerEnvelope ? groupedData : rooflines; const powerEnvelopePointKeys = useMemo(() => { const keys = new Set(); @@ -884,8 +965,13 @@ const ScatterGraph = React.memo( return false; }, [data]); const buildPointId = useCallback( - (point: InferenceData) => scatterPointJoinId(point, distinguishPointDates), - [distinguishPointDates], + (point: InferenceData) => { + const configId = scatterPointJoinId(point, distinguishPointDates); + return isConcurrencyAxis + ? `${configId}|observation-${point.id ?? data.indexOf(point)}|run-${point.run_url ?? ''}` + : configId; + }, + [isConcurrencyAxis, data, distinguishPointDates], ); // filteredData: visible points only (for scale domain calculation) @@ -938,7 +1024,7 @@ const ScatterGraph = React.memo( } const buckets = new Map(); const getBucket = (point: InferenceData) => { - const key = `${point.hwKey}|${point.precision}|${point.date}`; + const key = `${scatterSeriesKey(point)}|${point.date}`; let bucket = buckets.get(key); if (!bucket) { bucket = { @@ -987,7 +1073,7 @@ const ScatterGraph = React.memo( const buckets = new Map(); const getBucket = (point: InferenceData) => { const runIndex = overlayRunIndex(point.run_url ?? null, runIndexByUrl); - const key = `${point.hwKey}|${point.precision}|${point.date}|run${runIndex}`; + const key = `${scatterSeriesKey(point)}|${point.date}|run${runIndex}`; let bucket = buckets.get(key); if (!bucket) { bucket = { @@ -1092,6 +1178,8 @@ const ScatterGraph = React.memo( interface Entry { hwKey: string; runIndex: number; + /** Comparison variant id for boundary / role clones, null for the run's base series. */ + variant: string | null; points: InferenceData[]; } if (processedOverlayData.length === 0) return {} as Record; @@ -1100,8 +1188,15 @@ const ScatterGraph = React.memo( const grouped = processedOverlayData.reduce( (acc, p) => { const runIndex = overlayRunIndex(p.run_url ?? null, runIndexByUrl); - const key = `${p.hwKey}_${p.precision}_run${runIndex}`; - if (!acc[key]) acc[key] = { hwKey: String(p.hwKey), runIndex, points: [] }; + const key = `${scatterSeriesKey(p)}_run${runIndex}`; + if (!acc[key]) { + acc[key] = { + hwKey: String(p.hwKey), + runIndex, + variant: p.powerVariant?.id ?? null, + points: [], + }; + } acc[key].points.push(p); return acc; }, @@ -1124,7 +1219,7 @@ const ScatterGraph = React.memo( [overlayGroups, paretoDirection], ); const displayedOverlayRooflines = useMemo(() => { - if (!showPowerEnvelope) return overlayRooflines; + if (!showPowerEnvelope && !isConcurrencyAxis) return overlayRooflines; return Object.fromEntries( Object.entries(overlayGroups).flatMap(([key, group]) => [...groupDisplayedPoints(group.points)].map(([segment, points]) => [ @@ -1133,7 +1228,13 @@ const ScatterGraph = React.memo( ]), ), ); - }, [showPowerEnvelope, overlayRooflines, overlayGroups, groupDisplayedPoints]); + }, [ + isConcurrencyAxis, + showPowerEnvelope, + overlayRooflines, + overlayGroups, + groupDisplayedPoints, + ]); // Overlay counterpart of `optimalPointKeys`: the points on any overlay // run's drawn roofline (already e2e-restricted for agentic non-e2e modes). @@ -1180,9 +1281,11 @@ const ScatterGraph = React.memo( // its X marker sitting on the dashed roofline and read as a pareto point. const isOverlayPointVisible = useCallback( (d: InferenceData) => + !hiddenPowerVariants.has(powerVariantId(d.powerVariant)) && (!hideNonOptimal || overlayOptimalPoints.has(d)) && (!showPowerEnvelope || showAllMeasurements || overlayEnvelopePoints.has(d)), [ + hiddenPowerVariants, hideNonOptimal, overlayOptimalPoints, showPowerEnvelope, @@ -1216,6 +1319,9 @@ const ScatterGraph = React.memo( ); const { data: persistedLogAvailability } = useLogAvailability(persistedPointIds); const [fixedLogPointId, setFixedLogPointId] = useState(null); + const [powerTelemetryPoint, setPowerTelemetryPoint] = useState(null); + + const attachPowerTraceAction = usePowerTraceAction(chartRef); // --- Legend points table (per-series drill-down opened from the legend) --- const [pointsTableTarget, setPointsTableTarget] = useState(null); @@ -1226,6 +1332,7 @@ const ScatterGraph = React.memo( (lockedFrameworks ? 0 : quickFilters.frameworks.length) + quickFilters.deployment.length + quickFilters.power.length + + (quickFilters.topologies?.length ?? 0) + (selectedSequence === Sequence.AgenticTraces ? 0 : quickFilters.spec.length); const clearQuickFilters = useCallback(() => { setQuickFilterVendors([]); @@ -1233,12 +1340,14 @@ const ScatterGraph = React.memo( setQuickFilterDeployment([]); if (selectedSequence !== Sequence.AgenticTraces) setQuickFilterSpec([]); setQuickFilterPower([]); + setQuickFilterTopologies([]); }, [ setQuickFilterVendors, setQuickFilterFrameworks, setQuickFilterDeployment, setQuickFilterSpec, setQuickFilterPower, + setQuickFilterTopologies, selectedSequence, ]); @@ -1251,6 +1360,7 @@ const ScatterGraph = React.memo( const pts = pointsData.filter( (p) => p.hwKey === hwKey && + !p.powerVariant && selectedPrecisions.includes(p.precision) && (!hideNonOptimal || optimalPointKeys.has(optimalPointKey(p))), ); @@ -1266,6 +1376,7 @@ const ScatterGraph = React.memo( const pts = processedOverlayData.filter( (p) => overlayRunIndex(p.run_url ?? null, runIndexByUrl) === runIndex && + !p.powerVariant && activeOverlayHwTypes.has(p.hwKey as string) && (!hideNonOptimal || overlayOptimalPoints.has(p)), ); @@ -1427,7 +1538,11 @@ const ScatterGraph = React.memo( const metricIdentity = useMemo( () => [ - showPowerEnvelope ? 'power-envelopes' : 'pareto-curves', + isConcurrencyAxis + ? 'observed-load' + : showPowerEnvelope + ? 'power-envelopes' + : 'pareto-curves', useAdvancedLabels ? 'advanced-labels' : 'basic-labels', showConcurrencyLabels ? 'conc-labels' : 'no-conc-labels', selectedYAxisMetric, @@ -1445,6 +1560,7 @@ const ScatterGraph = React.memo( .toSorted() .join('|'), [ + isConcurrencyAxis, showPowerEnvelope, selectedYAxisMetric, useAdvancedLabels, @@ -1486,6 +1602,7 @@ const ScatterGraph = React.memo( (d: InferenceData) => effectiveActiveHwTypes.has(d.hwKey as string) && selectedPrecisions.includes(d.precision) && + !hiddenPowerVariants.has(powerVariantId(d.powerVariant)) && (!hideNonOptimal || optimalPointKeys.has(optimalPointKey(d))) && (!showPowerEnvelope || showAllMeasurements || @@ -1493,6 +1610,7 @@ const ScatterGraph = React.memo( [ effectiveActiveHwTypes, selectedPrecisions, + hiddenPowerVariants, hideNonOptimal, optimalPointKeys, showPowerEnvelope, @@ -1538,6 +1656,25 @@ const ScatterGraph = React.memo( () => frontierHardwareKeys(globalParetoPoints, globalFrontier), [globalParetoPoints, globalFrontier], ); + // PowerX provenance of the drawn frontier: the same points, listed with their runs. + const showFrontierPoints = + showParetoFrontier && + globalFrontier.length > 0 && + !minimalChrome && + getMeasuredMetricConfig(selectedYAxisMetric) !== undefined; + const frontierHardwareLabel = useCallback( + (point: InferenceData) => { + const config = processedOverlayData.includes(point) + ? overlayData?.hardwareConfig[point.hwKey] + : hardwareConfig[point.hwKey]; + return config ? getDisplayLabel(config) : point.hwKey; + }, + [processedOverlayData, overlayData, hardwareConfig], + ); + const frontierHardwareColor = useCallback( + (point: InferenceData) => getCssColor(resolveColor(point.hwKey)), + [getCssColor, resolveColor], + ); const applyParetoFadeRef = useRef<(group: RenderContext['layout']['zoomGroup']) => void>( () => {}, ); @@ -1669,28 +1806,60 @@ const ScatterGraph = React.memo( getCssColor, ]); - const powerTierCounts = useMemo(() => { - const officialTotal = pointsData.filter((point) => - selectedPrecisions.includes(point.precision), - ); - const overlayTotal = processedOverlayData.filter((point) => - selectedPrecisions.includes(point.precision), - ); - const officialVisible = officialTotal.filter(isPointVisible); - const overlayVisible = overlayTotal.filter( - (point) => activeOverlayHwTypes.has(String(point.hwKey)) && isOverlayPointVisible(point), - ); - return { - total: countPowerTiers([...officialTotal, ...overlayTotal]), - visible: countPowerTiers([...officialVisible, ...overlayVisible]), - }; + // The comparison in effect and the base series' identity under it. The + // base is the selected metric's own series; deriving it from which variant + // no official point carries breaks when only an overlay carries the + // comparison, and line labels need the same answer as the legend rows. + const powerCompareMode = useMemo(() => { + const official = inferPowerCompare(pointsData); + return official === 'none' ? inferPowerCompare(processedOverlayData) : official; + }, [pointsData, processedOverlayData]); + const powerCompareBaseId = useMemo( + () => powerVariantId(powerCompareBase(selectedYAxisMetric, powerCompareMode)), + [selectedYAxisMetric, powerCompareMode], + ); + + // One legend row per comparison series present (base first). Rows toggle + // chart-local visibility and hover-highlight that series across hardware. + const powerVariantLegendItems = useMemo(() => { + const allPoints = [...pointsData, ...processedOverlayData]; + const variants = powerVariantsInData(allPoints, selectedYAxisMetric); + const baseId = powerCompareBaseId; + return variants.map((variant) => { + const id = powerVariantId(variant); + const legendId = `${POWER_VARIANT_LEGEND_PREFIX}${id}`; + const isBase = id === baseId; + return { + name: legendId, + hw: legendId, + label: powerVariantLabel(variant, locale), + color: 'var(--foreground)', + // The base series is solid, like its points; siblings carry their dash. + lineDasharray: isBase ? '1 0' : powerVariantDash(variant) || '1 0', + isActive: !hiddenPowerVariants.has(isBase ? '' : id), + isRemovable: false, + onClick: () => { + const key = isBase ? '' : id; + setHiddenPowerVariants((prev) => { + const next = new Set(prev); + if (next.has(key)) next.delete(key); + else next.add(key); + return next; + }); + track('inference_power_compare_series_toggled', { + series: id, + visible: hiddenPowerVariants.has(key), + }); + }, + }; + }); }, [ pointsData, processedOverlayData, - selectedPrecisions, - isPointVisible, - activeOverlayHwTypes, - isOverlayPointVisible, + selectedYAxisMetric, + powerCompareBaseId, + locale, + hiddenPowerVariants, ]); // --- Legend hover highlight --- @@ -1699,9 +1868,13 @@ const ScatterGraph = React.memo( const hw = el.dataset.hwKey; const prec = el.dataset.precision; if (hw === null || hw === undefined || prec === null || prec === undefined) return false; - return effectiveActiveHwTypes.has(hw) && selectedPrecisions.includes(prec); + return ( + effectiveActiveHwTypes.has(hw) && + selectedPrecisions.includes(prec) && + !hiddenPowerVariants.has(el.dataset.powerVariant ?? '') + ); }, - [effectiveActiveHwTypes, selectedPrecisions], + [effectiveActiveHwTypes, selectedPrecisions, hiddenPowerVariants], ); // --- Interaction state ref --- @@ -1718,6 +1891,7 @@ const ScatterGraph = React.memo( isPointVisible, isOverlayPointVisible, effectiveActiveHwTypes, + hiddenPowerVariants, selectedPrecisions, activeOverlayHwTypes, getCssColor, @@ -1725,11 +1899,13 @@ const ScatterGraph = React.memo( knownIssueAnnotations, traceAvailability, logAvailability: persistedLogAvailability, + showPowerTelemetry, }); interactionRef.current = { isPointVisible, isOverlayPointVisible, effectiveActiveHwTypes, + hiddenPowerVariants, selectedPrecisions, activeOverlayHwTypes, getCssColor, @@ -1737,6 +1913,7 @@ const ScatterGraph = React.memo( knownIssueAnnotations, traceAvailability, logAvailability: persistedLogAvailability, + showPowerTelemetry, }; // --- Perf ruler (opt-in: click two curves, drag the ruler to any iso-x) --- @@ -1748,17 +1925,40 @@ const ScatterGraph = React.memo( // the curves' rendered paths at the iso-x — neither end needs to be a // data point. Multiple rulers accumulate (capped in the pure module); // completing one immediately allows starting the next. - const [preferPerfRulerMode, setPerfRulerMode] = useState(false); - const perfRulerMode = preferPerfRulerMode && (!showPowerEnvelope || isMeasuredPowerAxis); - const [perfRulerState, setPerfRulerState] = useState(EMPTY_PERF_RULER_STATE); + // + // The primary chart's rulers live in the InferenceProvider store so they + // ride along in share links (`i_rulers`) and survive a remount (table + // view toggle). Every other instance — the replay chart, which draws the + // same curve classes, and harnesses mounted without the provider — keeps + // component-local state. Both paths share one `[state, setState]` pair + // below, so the reducers, refs, and draw passes are path-agnostic. + const perfRulerStore = usePerfRulerStore(); + const persistedRulers = perfRulerStore?.chartId === chartId ? perfRulerStore : undefined; + // Rulers only render while the mode is on (and the mode-off effect below + // clears them), so restored share-link rulers — pending or already + // committed by a previous mount — switch the mode on for this instance. + const [preferPerfRulerMode, setPerfRulerMode] = useState( + () => + persistedRulers !== undefined && + (persistedRulers.pending !== null || persistedRulers.state.rulers.length > 0), + ); + const perfRulerMode = + !isConcurrencyAxis && preferPerfRulerMode && (!showPowerEnvelope || isMeasuredPowerAxis); + const [localPerfRulerState, setLocalPerfRulerState] = + useState(EMPTY_PERF_RULER_STATE); + const perfRulerState = persistedRulers ? persistedRulers.state : localPerfRulerState; + const setPerfRulerState = persistedRulers ? persistedRulers.setState : setLocalPerfRulerState; // Changing the x- or y-axis metric (including the x percentile, which // `x_scale_field` encodes) clears every ruler: the curves are redrawn // in different units, so a ruler that persisted would measure a ratio // the user never placed. Runs before the draw pass so no stale ruler - // ever paints over the new curves. + // ever paints over the new curves. Render-time adjustment is only legal + // for this component's own state, so the hook targets the local state; + // the store applies the same reset to persisted rulers inside the + // provider (see usePerfRulerStoreValue). usePerfRulerAxisReset( perfRulerAxisMetricKey(chartDefinition.x_scale_field, selectedYAxisMetric), - setPerfRulerState, + setLocalPerfRulerState, ); // Draw passes read mode/state through refs so toggling off clears the // rulers in the same pre-paint layout pass — lines/labels must never @@ -1824,6 +2024,51 @@ const ScatterGraph = React.memo( [], ); + // Share-link rulers commit only once BOTH curve paths are in the DOM — + // otherwise the prune pass would eat them before their data (i_gpus, + // comparison dates, overlay runs) has arrived. Hidden curves (opacity 0) + // count as present, like for prune. The iso-x is clamped to the pair's + // overlap through the drawn paths, so a rounded or since-shifted iso-x + // still renders; a pair with disjoint spans can never be measured on + // these axes and is dropped. This runs from the draw pass rather than a + // React effect: the chart first draws in a D3Chart-local re-render + // (dimensions are measured after mount), which re-renders nothing here, + // so an effect keyed on our props could miss the first draw and leave + // resolvable rulers pending for the rest of the session. The store is + // read through a ref for the same reason the draw passes read the ruler + // state through refs. Nothing commits while the mode is off (forced off + // by the power envelope, or switched off by the user) — the mode-off + // effect discards pending rulers, and the analytics event must not + // report a restore nobody saw. Draw passes can repeat before React has + // applied a commit, so the pending list handed over is remembered by + // identity and skipped until the store replaces it. + const persistedRulersRef = useRef(persistedRulers); + persistedRulersRef.current = persistedRulers; + const committedPendingRef = useRef(null); + const commitPendingPerfRulers = useCallback( + (zoomGroup: d3.Selection) => { + const store = persistedRulersRef.current; + const pending = store?.pending ?? null; + if (!store || !pending || !perfRulerModeRef.current) return; + if (committedPendingRef.current === pending) return; + const curveExists = (cls: string) => !zoomGroup.select(`.${CSS.escape(cls)}`).empty(); + const resolved: PerfRulerMeasurement[] = []; + const remaining: PerfRulerMeasurement[] = []; + for (const ruler of pending) { + if (!curveExists(ruler.curveA) || !curveExists(ruler.curveB)) { + remaining.push(ruler); + continue; + } + const isoX = clampPerfRulerIsoXToOverlap(ruler.curveA, ruler.curveB, ruler.isoX); + if (isoX !== null) resolved.push({ ...ruler, isoX }); + } + if (remaining.length === pending.length) return; + committedPendingRef.current = pending; + store.commitPending(resolved, remaining.length > 0 ? remaining : null); + }, + [clampPerfRulerIsoXToOverlap], + ); + // Curve click (widened hit strokes): iso-x is the click's x pixel // through the CURRENT rendered x scale, stored in data space. const handlePerfRulerCurveClick = useCallback( @@ -1852,7 +2097,7 @@ const ScatterGraph = React.memo( (point: InferenceData, source: 'official' | 'overlay') => { const ctx = perfRulerDrawCtxRef.current; if (!ctx) return; - const series = `${String(point.hwKey)}_${point.precision}`; + const series = scatterSeriesKey(point); const base = source === 'overlay' ? `overlay-roofline-${series}_run${overlayRunIndex(point.run_url ?? null, runIndexByUrl)}` @@ -1882,8 +2127,12 @@ const ScatterGraph = React.memo( // the switch handler also clears synchronously, this covers // programmatic mode changes). `clearPerfRulers` bails out with the same // reference when there is nothing to clear. + // Share-link rulers still waiting for their curves go too — the user + // switched the tool off, so nothing should surface later. useEffect(() => { - if (!perfRulerMode) setPerfRulerState(clearPerfRulers); + if (perfRulerMode) return; + setPerfRulerState(clearPerfRulers); + persistedRulers?.discardPending(); }, [perfRulerMode]); // Invisible widened hit strokes over every rendered roofline path @@ -2033,6 +2282,7 @@ const ScatterGraph = React.memo( ) => { perfRulerDrawCtxRef.current = { zoomGroup, xScale, yScale, width, height }; syncPerfRulerHitPaths(zoomGroup); + commitPendingPerfRulers(zoomGroup); const state = perfRulerStateRef.current; const entries: PerfRulerRenderEntry[] = []; if (perfRulerModeRef.current && state.rulers.length > 0) { @@ -2086,7 +2336,7 @@ const ScatterGraph = React.memo( ); if (!dragHandles.empty()) dragHandles.call(perfRulerDrag); }, - [syncPerfRulerHitPaths, perfRulerDrag], + [syncPerfRulerHitPaths, commitPendingPerfRulers, perfRulerDrag], ); drawPerfRulerRef.current = drawPerfRuler; @@ -2120,16 +2370,31 @@ const ScatterGraph = React.memo( const svg = chartRef.current?.getSvgElement?.(); if (!svg) return; const root = d3.select(svg); + // A comparison-series legend row highlights that boundary / role + // across every hardware instead of one hardware across series. Base + // points and rooflines carry no variant (only siblings are cloned), + // so the empty id maps back to the base legend row's id. + const variantId = hwKey.startsWith(POWER_VARIANT_LEGEND_PREFIX) + ? hwKey.slice(POWER_VARIANT_LEGEND_PREFIX.length) + : null; + const matchesPoint = (d: InferenceData) => + variantId === null + ? String(d.hwKey) === hwKey + : (powerVariantId(d.powerVariant) || powerCompareBaseId) === variantId; root .selectAll('.dot-group') .style('opacity', (d) => - isPointVisible(d) ? (String(d.hwKey) === hwKey ? 1 : 0.15) : 0, + isPointVisible(d) ? (matchesPoint(d) ? pointOpacityForVariant(d) : 0.15) : 0, ); root .selectAll('.roofline-path, .official-overflow-continuation') .style('opacity', function () { if (!isRooflineVisible(this)) return 0; - return this.dataset.hwKey === hwKey ? null : '0.15'; + const matches = + variantId === null + ? this.dataset.hwKey === hwKey + : (this.dataset.powerVariant || powerCompareBaseId) === variantId; + return matches ? null : '0.15'; }); root .selectAll('.parallelism-label, .line-label') @@ -2137,7 +2402,7 @@ const ScatterGraph = React.memo( return labelOpacityForHover((this as SVGGElement).dataset, hwKey); }); }, - [isPointVisible, isRooflineVisible], + [isPointVisible, isRooflineVisible, powerCompareBaseId], ); const handleLegendHoverEnd = useCallback(() => { @@ -2146,7 +2411,7 @@ const ScatterGraph = React.memo( const root = d3.select(svg); root .selectAll('.dot-group') - .style('opacity', (d) => (isPointVisible(d) ? 1 : 0)); + .style('opacity', (d) => (isPointVisible(d) ? pointOpacityForVariant(d) : 0)); root .selectAll('.roofline-path, .official-overflow-continuation') .style('opacity', function () { @@ -2159,9 +2424,16 @@ const ScatterGraph = React.memo( (this as SVGGElement).dataset, effectiveActiveHwTypes, selectedPrecisions, + activeOverlayHwTypes, ); }); - }, [isPointVisible, isRooflineVisible, effectiveActiveHwTypes, selectedPrecisions]); + }, [ + isPointVisible, + isRooflineVisible, + effectiveActiveHwTypes, + selectedPrecisions, + activeOverlayHwTypes, + ]); // --- Zoom config --- const eventPrefix = chartDefinition.chartType === 'e2e' ? 'latency' : 'interactivity'; @@ -2230,6 +2502,7 @@ const ScatterGraph = React.memo( yLabel, selectedYAxisMetric, hardwareConfig, + showPowerTelemetry: interactionRef.current.showPowerTelemetry, runUrl: d.run_url ? updateRepoUrl(d.run_url) : undefined, hasTrace: d.benchmark_type === 'agentic_traces' && isPersistedBenchmarkId(d.id) @@ -2291,6 +2564,15 @@ const ScatterGraph = React.memo( }); }); } + const powerBtn = tooltipEl.querySelector('[data-action="view-power-telemetry"]'); + if (powerBtn && isPersistedBenchmarkId(d.id)) { + powerBtn.addEventListener('click', (event) => { + event.stopPropagation(); + setPowerTelemetryPoint(d); + chartRef.current?.dismissTooltip(); + track('inference_power_telemetry_opened', { id: d.id, hwKey: d.hwKey, conc: d.conc }); + }); + } const logsBtn = tooltipEl.querySelector('[data-action="view-logs"]'); if (logsBtn && typeof d.id === 'number') { logsBtn.addEventListener('click', (btnEvent) => { @@ -2308,10 +2590,12 @@ const ScatterGraph = React.memo( }); }); } + attachPowerTraceAction(tooltipEl, d, false); }, attachToLayer: 1, // scatter layer is index 1 (after rooflines at 0) }), [ + attachPowerTraceAction, xLabel, yLabel, selectedYAxisMetric, @@ -2325,6 +2609,38 @@ const ScatterGraph = React.memo( // --- Layers --- const layers = useMemo((): LayerConfig[] => { + // Observed-load segments join exact measurements; every other curve is smoothed. + const lineCurve = isConcurrencyAxis ? d3.curveLinear : d3.curveMonotoneX; + const curveKind = isConcurrencyAxis + ? 'observed-load' + : showPowerEnvelope + ? 'power-envelope' + : 'pareto'; + // Line-label identity of one drawn series under a power comparison + // (`i_pcompare`): the base series keeps the hardware key, so pinned + // anchors and hover hooks keep working; a sibling is `::`. + const wattsAxis = metricPlotsWatts(selectedYAxisMetric); + const lineLabelIdentity = (hw: string, points: readonly InferenceData[]) => { + const variant = points[0]?.powerVariant; + const variantId = powerVariantId(variant); + const isBase = !variant || variantId === powerCompareBaseId; + return { variant, variantId, isBase, seriesId: lineLabelSeriesId(hw, variant, isBase) }; + }; + // A sibling's label says which series it is; a flat provisioned boundary + // (TDP, all-in) on a watts axis also states its value. + const lineLabelSuffix = ( + identity: ReturnType, + points: readonly InferenceData[], + ) => + powerLineLabelSuffix(identity.variant, { + isBase: identity.isBase, + locale, + flatWatts: + !identity.isBase && wattsAxis && identity.variant?.kind === 'basis' + ? flatSeriesValue(points.map((point) => point.y)) + : null, + }); + // ── Layer 0: Rooflines + gradient labels (custom) ── const rooflineLayer: CustomLayerConfig = { type: 'custom', @@ -2343,7 +2659,7 @@ const ScatterGraph = React.memo( .line() .x((d) => xScale(d.x)) .y((d) => yScale(d.y)) - .curve(d3.curveMonotoneX); + .curve(lineCurve); // Ensure rooflines layer exists before dot-groups let rooflinesLayer = zoomGroup.select('.rooflines-layer'); @@ -2362,6 +2678,8 @@ const ScatterGraph = React.memo( key: string; hw: string; precision: string; + /** Comparison variant id (`i_pcompare`), '' for the base series. */ + variant: string; points: InferenceData[]; stroke: string; visible: boolean; @@ -2370,10 +2688,11 @@ const ScatterGraph = React.memo( const activeGradientIds = new Set(); Object.entries(displayedRooflines).forEach(([key, pts]) => { - const hw = key.split('_').slice(0, -1).join('_'); - const precision = key.split('_').pop()!; + const { hw, precision, variant } = parseScatterSeriesKey(key); const visible = - ir.effectiveActiveHwTypes.has(hw) && ir.selectedPrecisions.includes(precision); + ir.effectiveActiveHwTypes.has(hw) && + ir.selectedPrecisions.includes(precision) && + !ir.hiddenPowerVariants.has(variant ?? ''); const baseStroke = ir.getCssColor(ir.resolveColor(hw)); // Split into per-date sub-paths so the line never crosses dates. @@ -2417,6 +2736,7 @@ const ScatterGraph = React.memo( key: entryKey, hw, precision, + variant: variant ?? '', points: datePoints, stroke, visible, @@ -2441,12 +2761,15 @@ const ScatterGraph = React.memo( ) .join('path') .attr('class', (d) => `roofline-path roofline-${d.key}`) - .attr('data-curve-kind', showPowerEnvelope ? 'power-envelope' : 'pareto') + .attr('data-curve-kind', curveKind) .attr('data-hw-key', (d) => d.hw) .attr('data-precision', (d) => d.precision) + .attr('data-power-variant', (d) => d.variant || null) .attr('fill', 'none') .attr('stroke', (d) => d.stroke) .attr('stroke-width', 2.5) + // Comparison siblings share the hardware colour; the dash tells them apart. + .attr('stroke-dasharray', (d) => powerVariantDashById(d.variant) || null) .attr('d', (d) => lineGen(d.points)) .style('transition', 'opacity 150ms ease') .style('opacity', (d) => (d.visible ? 1 : 0)); @@ -2467,10 +2790,11 @@ const ScatterGraph = React.memo( if (showGradientLabels) { Object.entries(allPointLabelsByKey).forEach(([key, pointLabels]) => { if (pointLabels.length < 2) return; - const hw = key.split('_').slice(0, -1).join('_'); - const precision = key.split('_').pop()!; + const { hw, precision, variant } = parseScatterSeriesKey(key); const visible = - ir.effectiveActiveHwTypes.has(hw) && ir.selectedPrecisions.includes(precision); + ir.effectiveActiveHwTypes.has(hw) && + ir.selectedPrecisions.includes(precision) && + !ir.hiddenPowerVariants.has(variant ?? ''); const segments: { label: string; color: string; points: InferenceData[] }[] = []; let cur = { @@ -2536,7 +2860,7 @@ const ScatterGraph = React.memo( (exit) => exit.remove(), ) .attr('data-seg-key', (d) => d.segKey) - .attr('data-curve-kind', showPowerEnvelope ? 'power-envelope' : 'pareto') + .attr('data-curve-kind', curveKind) .attr('data-hw-key', (d) => d.hw) .attr('data-precision', (d) => d.precision) .attr('transform', (d) => `translate(${d.x},${d.y})`) @@ -2568,12 +2892,22 @@ const ScatterGraph = React.memo( // ── Line labels (run name along each roofline) ── let lineLabels: LineLabelPlacement[] = []; + // Comparison variant and label suffix per label key, for the text + // segments and the `data-power-variant` hook on each pill. + const lineLabelMeta = new Map< + string, + { variantId: string; suffix: string; runTag: string } + >(); if (showLineLabels) { const multiPrecision = ir.selectedPrecisions.length > 1; const officialByGroup = new Map(); for (const entry of entries) { if (!entry.visible) continue; - const groupKey = multiPrecision ? entry.key : entry.hw; + // One label per hardware and, under a power comparison, per + // sibling series: the measured line and its boundary / pool + // lines each say which one they are, instead of the longest + // line taking the hardware's only label. + const groupKey = multiPrecision ? entry.key : `${entry.hw}::${entry.variant}`; const previous = officialByGroup.get(groupKey); if (!previous || entry.points.length > previous.points.length) { officialByGroup.set(groupKey, entry); @@ -2582,39 +2916,58 @@ const ScatterGraph = React.memo( const officialSeries: LineLabelSeries[] = [ ...officialByGroup.values(), - ].map((entry) => ({ - key: entry.key, - seriesId: entry.hw, - label: lineLabelText( - entry.hw, - entry.precision, - multiPrecision, - modelLabel, - entry.points, - ), - color: ir.getCssColor(ir.resolveColor(entry.hw)), - points: entry.points, - keepVisibleOnCollision: entry.points.length === 1, - })); + ].map((entry) => { + const identity = lineLabelIdentity(entry.hw, entry.points); + const suffix = lineLabelSuffix(identity, entry.points); + lineLabelMeta.set(entry.key, { variantId: identity.variantId, suffix, runTag: '' }); + return { + key: entry.key, + seriesId: identity.seriesId, + label: `${lineLabelText( + entry.hw, + entry.precision, + multiPrecision, + modelLabel, + entry.points, + )}${suffix}`, + color: ir.getCssColor(ir.resolveColor(entry.hw)), + points: entry.points, + }; + }); + // Runs drawing the same hardware need a run tag on their pills. + const overlayRunsByHw = new Map>(); + for (const group of Object.values(displayedOverlayRooflines)) { + if (!ir.activeOverlayHwTypes.has(group.hwKey)) continue; + if (!overlayRunsByHw.has(group.hwKey)) overlayRunsByHw.set(group.hwKey, new Set()); + overlayRunsByHw.get(group.hwKey)!.add(group.runIndex); + } const overlaySeries: LineLabelSeries[] = Object.entries( displayedOverlayRooflines, ).flatMap(([overlayKey, group]) => { if (!ir.activeOverlayHwTypes.has(group.hwKey)) return []; const info = unofficialRunInfos[group.runIndex]; const precision = group.points[0]?.precision ?? ''; - const runLabel = info - ? getInferenceRunLabel(`✕ ${info.branch || `run ${info.id}`}`, group.points) - : ''; + const hardwareLabel = lineLabelText( + group.hwKey, + precision, + multiPrecision, + modelLabel, + group.points, + ); + const sharesHardware = (overlayRunsByHw.get(group.hwKey)?.size ?? 0) > 1; + const runTag = info && sharesHardware ? overlayRunTag(info) : ''; const label = info - ? multiPrecision - ? `${runLabel} ${getPrecisionLabel(precision as Precision)}` - : runLabel - : lineLabelText(group.hwKey, precision, multiPrecision, modelLabel, group.points); + ? getOverlayLineLabel(hardwareLabel, info, sharesHardware) + : hardwareLabel; + const identity = lineLabelIdentity(group.hwKey, group.points); + const suffix = lineLabelSuffix(identity, group.points); + const key = `overlay-${overlayKey}`; + lineLabelMeta.set(key, { variantId: identity.variantId, suffix, runTag }); return [ { - key: `overlay-${overlayKey}`, - seriesId: group.hwKey, - label, + key, + seriesId: identity.seriesId, + label: `${label}${suffix}`, color: overlayRunColor(group.runIndex), points: group.points, }, @@ -2637,16 +2990,19 @@ const ScatterGraph = React.memo( const labeledKeys = new Set(lineLabels.map((label) => label.key)); for (const entry of entries) { if (labeledKeys.has(entry.key)) continue; + const identity = lineLabelIdentity(entry.hw, entry.points); + const suffix = lineLabelSuffix(identity, entry.points); + lineLabelMeta.set(entry.key, { variantId: identity.variantId, suffix, runTag: '' }); lineLabels.push({ key: entry.key, - seriesId: entry.hw, - label: lineLabelText( + seriesId: identity.seriesId, + label: `${lineLabelText( entry.hw, entry.precision, multiPrecision, modelLabel, entry.points, - ), + )}${suffix}`, color: ir.getCssColor(ir.resolveColor(entry.hw)), x: xScale(entry.points[0].x), y: yScale(entry.points[0].y), @@ -2663,20 +3019,35 @@ const ScatterGraph = React.memo( } renderLineLabels(zoomGroup, lineLabels, { - seriesAttribute: 'data-hw-key', - iconFor: (label) => getLineLabelVendorIcon(label.seriesId), + seriesAttribute: 'data-series-id', + iconFor: (label) => getLineLabelVendorIcon(lineLabelHardwareKey(label.seriesId)), configureGroup: (labelGroup, label) => { labelGroup .attr('data-visible', label.visible ? '1' : '0') + // Legend hover and filter sync key labels by hardware alone; + // the variant names the comparison sibling ('' for the base). + .attr('data-hw-key', lineLabelHardwareKey(label.seriesId)) + .attr('data-power-variant', lineLabelMeta.get(label.key)?.variantId ?? '') .select('.ll-bg') .attr('opacity', 0.95); }, configureText: (text, label) => { - const config = getHardwareConfig(label.seriesId, modelLabel); + const config = getHardwareConfig(lineLabelHardwareKey(label.seriesId), modelLabel); + // Parse the hardware part without the variant suffix, which gets + // its own segment so the engine is still matched at the end. + const meta = lineLabelMeta.get(label.key); + const suffix = meta?.suffix ?? ''; + const runTag = meta?.runTag ?? ''; + let coreLabel = suffix ? label.label.slice(0, -suffix.length) : label.label; + if (runTag) coreLabel = coreLabel.slice(0, -runTag.length); + // Overlay pills lead with the run marker; the hardware behind it is + // parsed like an official pill so the GPU name stays bold. + const marker = coreLabel.startsWith(OVERLAY_LABEL_MARKER) ? OVERLAY_LABEL_MARKER : ''; + coreLabel = coreLabel.slice(marker.length); const hardwareLabel = getDisplayLabel(config); const isHardwareLabel = - label.label === hardwareLabel || label.label.startsWith(`${config.label} `); - const remainingLabel = isHardwareLabel ? label.label.slice(config.label.length) : ''; + coreLabel === hardwareLabel || coreLabel.startsWith(`${config.label} `); + const remainingLabel = isHardwareLabel ? coreLabel.slice(config.label.length) : ''; // Use this curve's resolved suffix, not the generic hwKey label: // official and overlay curves can share a key but differ by run. const engineLabel = @@ -2686,8 +3057,18 @@ const ScatterGraph = React.memo( engineLabel && remainingLabel.endsWith(engineLabel) ? remainingLabel.slice(0, -engineLabel.length) : remainingLabel; + const markerSegments = marker + ? [{ className: 'll-marker', text: marker, fill: 'white', weight: '600' }] + : []; + const runSegments = runTag + ? [{ className: 'll-run', text: runTag, fill: '#d1d5db', weight: '400' }] + : []; + const variantSegments = suffix + ? [{ className: 'll-variant', text: suffix, fill: 'white', weight: '500' }] + : []; const segments = isHardwareLabel ? [ + ...markerSegments, { className: 'll-gpu', text: config.label, fill: 'white', weight: '700' }, ...(precisionLabel ? [ @@ -2709,14 +3090,19 @@ const ScatterGraph = React.memo( }, ] : []), + ...runSegments, + ...variantSegments, ] : [ + ...markerSegments, { className: 'll-plain', - text: label.label, + text: coreLabel, fill: 'white', weight: '600', }, + ...runSegments, + ...variantSegments, ]; text .selectAll('tspan') @@ -2725,7 +3111,24 @@ const ScatterGraph = React.memo( .attr('class', (segment) => segment.className) .attr('fill', (segment) => segment.fill) .attr('font-weight', (segment) => segment.weight) + .attr('x', null) + .attr('dy', null) .text((segment) => segment.text); + // Keep the framework and role visible when a pill is wider than + // the mobile plot, without shrinking its text or dropping fields. + const textX = Number(text.attr('x') ?? 0); + const maxLineWidth = ctx.width - textX - 10; + let lineWidth = 0; + text.selectAll('tspan').each(function () { + const width = this.getComputedTextLength(); + if (lineWidth > 0 && lineWidth + width > maxLineWidth) { + d3.select(this) + .attr('x', textX) + .attr('dy', CHART_TYPE.lineLabel + 3); + lineWidth = 0; + } + lineWidth += width; + }); }, }); // Labels can be joined independently of the Pareto display pass. @@ -2750,7 +3153,7 @@ const ScatterGraph = React.memo( .line() .x((d) => newXScale(d.x)) .y((d) => newYScale(d.y)) - .curve(d3.curveMonotoneX); + .curve(lineCurve); // Update roofline paths — must split per-date so the zoom redraw // matches the per-date sub-paths created in the initial render. @@ -2826,11 +3229,11 @@ const ScatterGraph = React.memo( { key: string; seriesId: string; points: InferenceData[] } >(); for (const [key, points] of Object.entries(displayedRooflines)) { - const hardware = key.split('_').slice(0, -1).join('_'); - const precision = key.split('_').pop()!; + const { hw: hardware, precision, variant } = parseScatterSeriesKey(key); if ( !ir.effectiveActiveHwTypes.has(hardware) || - !ir.selectedPrecisions.includes(precision) + !ir.selectedPrecisions.includes(precision) || + ir.hiddenPowerVariants.has(variant ?? '') ) { continue; } @@ -2838,12 +3241,12 @@ const ScatterGraph = React.memo( const singleDate = pointsByDate.size === 1; for (const [date, datePoints] of pointsByDate) { const entryKey = singleDate ? key : `${key}__${encodeURIComponent(date)}`; - const groupKey = multiPrecision ? entryKey : hardware; + const groupKey = multiPrecision ? entryKey : `${hardware}::${variant ?? ''}`; const previous = bestByGroup.get(groupKey); if (!previous || datePoints.length > previous.points.length) { bestByGroup.set(groupKey, { key: entryKey, - seriesId: hardware, + seriesId: lineLabelIdentity(hardware, datePoints).seriesId, points: datePoints, }); } @@ -2854,7 +3257,6 @@ const ScatterGraph = React.memo( ...entry, label: '', color: '', - keepVisibleOnCollision: entry.points.length === 1, }), ); const overlaySeries: LineLabelSeries[] = Object.entries( @@ -2864,7 +3266,7 @@ const ScatterGraph = React.memo( ? [ { key: `overlay-${overlayKey}`, - seriesId: group.hwKey, + seriesId: lineLabelIdentity(group.hwKey, group.points).seriesId, label: '', color: '', points: group.points, @@ -2898,7 +3300,8 @@ const ScatterGraph = React.memo( interactionRef.current.getCssColor( interactionRef.current.resolveColor(d.hwKey as string), ), - getOpacity: (d) => (interactionRef.current.isPointVisible(d) ? 1 : 0), + getOpacity: (d) => + interactionRef.current.isPointVisible(d) ? pointOpacityForVariant(d) : 0, getPointerEvents: (d) => (interactionRef.current.isPointVisible(d) ? 'auto' : 'none'), hideLabels: !showPointLabels || showGradientLabels, // Concurrency (C=) is appended only when the advanced @@ -2908,6 +3311,7 @@ const ScatterGraph = React.memo( dataAttrs: { 'hw-key': (d) => String(d.hwKey), precision: (d) => d.precision, + 'power-variant': (d) => d.powerVariant?.id ?? '', // Lets the agentic coach mark pick an anchor out of the DOM // without knowing anything about React state. 'benchmark-type': (d) => d.benchmark_type ?? '', @@ -2979,13 +3383,14 @@ const ScatterGraph = React.memo( .line() .x((d) => xScale(d.x)) .y((d) => yScale(d.y)) - .curve(d3.curveMonotoneX); + .curve(lineCurve); interface OvEntry { key: string; points: InferenceData[]; stroke: string; runIndex: number; + variant: string | null; } const ovEntries: OvEntry[] = []; Object.entries(displayedOverlayRooflines).forEach(([key, group]) => { @@ -2997,6 +3402,7 @@ const ScatterGraph = React.memo( // Color by run — same palette entry the legend uses, so they match. stroke: overlayRunColor(group.runIndex), runIndex: group.runIndex, + variant: group.variant, }); } }); @@ -3010,13 +3416,25 @@ const ScatterGraph = React.memo( .data(ovEntries, (d) => d.key) .join('path') .attr('class', (d) => `overlay-roofline-path overlay-roofline-${d.key}`) - .attr('data-curve-kind', showPowerEnvelope ? 'power-envelope' : 'pareto') + .attr('data-curve-kind', curveKind) .attr('fill', 'none') .attr('stroke', (d) => d.stroke) .attr('stroke-width', 2) - .attr('stroke-dasharray', (d) => overlayRooflineDasharray(d.runIndex)) + .attr('data-power-variant', (d) => d.variant) + // The run keeps its colour; a comparison sibling takes the + // variant dash so it reads like its official counterpart. + .attr('stroke-dasharray', (d) => + d.variant + ? powerVariantDashById(d.variant) + : overlayRooflineDasharray(d.runIndex), + ) .attr('d', (d) => lineGen(d.points)) - .style('filter', null); + .style('filter', null) + // Comparison rows hidden from the legend (the decoration effect + // keeps this in step with later toggles). + .style('opacity', (d) => + interactionRef.current.hiddenPowerVariants.has(d.variant ?? '') ? 0 : null, + ); // Overlay X-shape points — index-keyed so every point renders const overlayPoints = zoomGroup @@ -3052,7 +3470,7 @@ const ScatterGraph = React.memo( overlayPoints.each(function (d) { const visible = interactionRef.current.isOverlayPointVisible(d); d3.select(this) - .style('opacity', visible ? 1 : 0) + .style('opacity', visible ? pointOpacityForVariant(d) : 0) .style('pointer-events', visible ? 'auto' : 'none'); }); overlayPoints @@ -3061,15 +3479,14 @@ const ScatterGraph = React.memo( overlayRunColor(overlayRunIndex(d.run_url ?? null, runIndexByUrl)), ); - // Match official points: KV offload and the measured-axis - // legacy-power ring are the only persistent point decorations. - // Decode method remains in the tooltip. + // Match official points: KV offload is the only persistent point + // decoration. Decode method remains in the tooltip. overlayPoints.each(function (d) { - const overlayStroke = overlayRunColor( - overlayRunIndex(d.run_url ?? null, runIndexByUrl), + renderOffloadHalo( + d3.select(this), + d, + overlayRunColor(overlayRunIndex(d.run_url ?? null, runIndexByUrl)), ); - renderOffloadHalo(d3.select(this), d, overlayStroke); - renderLegacyPowerRing(d3.select(this), d, isMeasuredEnergyAxis, overlayStroke); }); updateOverlayLabels(zoomGroup); @@ -3136,6 +3553,9 @@ const ScatterGraph = React.memo( y: point.y, overlay: true, }); + // The shared helper has just rendered the pinned content into + // this element and pinned it via `handle`. + attachPowerTraceAction(ctx.tooltipElement, point, true); }, }); }, @@ -3149,7 +3569,7 @@ const ScatterGraph = React.memo( .line() .x((d) => newXScale(d.x)) .y((d) => newYScale(d.y)) - .curve(d3.curveMonotoneX); + .curve(lineCurve); Object.entries(displayedOverlayRooflines).forEach(([key, group]) => { if (group.points.length < 2) return; @@ -3237,7 +3657,7 @@ const ScatterGraph = React.memo( .line() .x((point) => xScale(point.x)) .y((point) => yScale(point.y)) - .curve(d3.curveMonotoneX); + .curve(lineCurve); const continuationPath = group .select('.overflow-continuation-line') .attr('d', lineGenerator(entry.points) ?? '') @@ -3378,6 +3798,7 @@ const ScatterGraph = React.memo( // existing DOM when it changes. Only data/structure changes recreate // the layers (and with them, the full chart render). }, [ + isConcurrencyAxis, displayedRooflines, paretoHighlightLayer, groupDisplayedPoints, @@ -3406,10 +3827,11 @@ const ScatterGraph = React.memo( xLabel, yLabel, selectedYAxisMetric, - isMeasuredEnergyAxis, + powerCompareBaseId, chartDefinition, locale, drawPerfRuler, + attachPowerTraceAction, ]); // Layers handle for the decoration effect — lets it re-run individual @@ -3429,10 +3851,8 @@ const ScatterGraph = React.memo( zoomGroup.selectAll('.dot-group').style('transition', 'opacity 150ms ease'); // Offload halo: dashed ring on every point that used KV offload (Pareto or not). - // Legacy-power ring: dotted ring on unvalidated telemetry, measured axes only. zoomGroup.selectAll('.dot-group').each(function (d) { renderOffloadHalo(d3.select(this), d, 'var(--foreground)'); - renderLegacyPowerRing(d3.select(this), d, isMeasuredEnergyAxis, 'var(--foreground)'); }); avoidPointLabelCollisions(zoomGroup); @@ -3462,9 +3882,6 @@ const ScatterGraph = React.memo( optimalPointKeys, getCssColor, resolveColor, - // A metric-only change must re-run the decoration pass so legacy-power - // rings appear/disappear with the Measured Energy axis selection. - isMeasuredEnergyAxis, ], ); @@ -3488,7 +3905,9 @@ const ScatterGraph = React.memo( zoomGroup.selectAll('.dot-group').each(function (d) { const point = d3.select(this); const visible = ir.isPointVisible(d); - point.style('opacity', visible ? 1 : 0).style('pointer-events', visible ? 'auto' : 'none'); + point + .style('opacity', visible ? pointOpacityForVariant(d) : 0) + .style('pointer-events', visible ? 'auto' : 'none'); const color = (showGradientLabels && gradientColorByPoint.get(d)) || ir.getCssColor(ir.resolveColor(d.hwKey as string)); @@ -3507,9 +3926,20 @@ const ScatterGraph = React.memo( zoomGroup.selectAll('.unofficial-overlay-pt').each(function (d) { const visible = ir.isOverlayPointVisible(d); d3.select(this) - .style('opacity', visible ? 1 : 0) + .style('opacity', visible ? pointOpacityForVariant(d) : 0) .style('pointer-events', visible ? 'auto' : 'none'); }); + // Overlay rooflines are only drawn for active overlay hardware; a + // comparison row hidden from the legend is the one visibility toggle + // they answer to here. + zoomGroup.selectAll('.overlay-roofline-path').each(function () { + const roofline = d3.select(this); + if (ir.hiddenPowerVariants.has(this.dataset.powerVariant ?? '')) { + roofline.style('opacity', 0); + } else { + roofline.style('opacity', null); + } + }); // Rooflines: visibility and solid-stroke recolor as direct writes. Keep // gradient url references intact and never touch animated path geometry. @@ -3519,7 +3949,9 @@ const ScatterGraph = React.memo( if (!hw || !precision) return; const roofline = d3.select(this); const visible = - ir.effectiveActiveHwTypes.has(hw) && ir.selectedPrecisions.includes(precision); + ir.effectiveActiveHwTypes.has(hw) && + ir.selectedPrecisions.includes(precision) && + !ir.hiddenPowerVariants.has(this.dataset.powerVariant ?? ''); roofline.style('opacity', visible ? 1 : 0); const stroke = roofline.attr('stroke'); if (stroke && !stroke.startsWith('url(')) { @@ -3553,6 +3985,7 @@ const ScatterGraph = React.memo( (this as SVGGElement).dataset, ir.effectiveActiveHwTypes, ir.selectedPrecisions, + ir.activeOverlayHwTypes, ); }); }, [ @@ -3711,7 +4144,9 @@ const ScatterGraph = React.memo( // brings a hidden ruler back); curves whose paths left the DOM // entirely are truly gone from the data, so prune each ruler (and the // draft) that references one. `prunePerfRulers` bails out with the - // same reference when nothing changed. + // same reference when nothing changed. Share-link rulers still + // pending are not state yet, so prune cannot touch them; drawPerfRuler + // above committed those whose curves now exist. setPerfRulerState((prev) => prunePerfRulers(prev, (cls) => !display.zoomGroup.select(`.${CSS.escape(cls)}`).empty()), ); @@ -3780,8 +4215,14 @@ const ScatterGraph = React.memo( { @@ -3808,7 +4249,11 @@ const ScatterGraph = React.memo(
); @@ -3832,14 +4277,14 @@ const ScatterGraph = React.memo( testId="scatter-graph" grabCursor={true} caption={ - isPowerAxis ? ( + isConcurrencyAxis ? ( <> {caption}

- {showPowerEnvelope ? legendT.powerCurves : legendT.powerOptimal} + {legendT.concurrencyCurves}

) : ( @@ -3973,6 +4418,9 @@ const ScatterGraph = React.memo( ) : null, })), + // Comparison series (`i_pcompare`): one dash-swatch row per + // boundary / role, toggling that series across every hardware. + ...powerVariantLegendItems, ]} disableActiveSort={false} isLegendExpanded={isLegendExpanded} @@ -4111,7 +4559,7 @@ const ScatterGraph = React.memo( track('latency_line_labels_toggled', { enabled: checked }); }, }, - ...(showPowerEnvelope && !isMeasuredPowerAxis + ...(isConcurrencyAxis || (showPowerEnvelope && !isMeasuredPowerAxis) ? [] : [ { @@ -4126,7 +4574,10 @@ const ScatterGraph = React.memo( // the pre-paint decoration effect then removes the rulers // and the curve hit strokes before the next frame (no // lingering lines after toggle-off). - if (!checked) setPerfRulerState(clearPerfRulers); + if (!checked) { + setPerfRulerState(clearPerfRulers); + persistedRulers?.discardPending(); + } track('latency_perf_ruler_toggled', { enabled: checked }); }, }, @@ -4174,6 +4625,7 @@ const ScatterGraph = React.memo( count: perfRulerState.rulers.length, }); setPerfRulerState(clearPerfRulers); + persistedRulers?.discardPending(); }, }, ] @@ -4193,18 +4645,14 @@ const ScatterGraph = React.memo( /> } /> - {isMeasuredEnergyAxis && ( - - )} {pointsTable && ( )} + {showFrontierPoints && ( + + )} + {powerTelemetryPoint === null ? null : ( + { + if (!open) setPowerTelemetryPoint(null); + }} + /> + )} {fixedLogPointId === null ? null : ( -
- - { - setSelectedXAxisMode(mode as XAxisMode); - track('latency_x_axis_mode_selected', { mode }); - }} - groups={[ - { - label: '', - options: options.map(({ value: option, kind, label, labelZh }) => ({ - value: option, - label: locale === 'zh' ? labelZh : label, - testId: `x-axis-mode-${option}`, - help: ( - <> -

- {X_AXIS_EXPLANATIONS[kind].name[locale]( - isAgentic ? selectedPercentile.toUpperCase() : null, +

+
+ + { + setSelectedXAxisMode(mode as XAxisMode); + track('latency_x_axis_mode_selected', { mode }); + }} + groups={[ + { + label: '', + options: options.map(({ value: option, kind, label, labelZh }) => ({ + value: option, + label: locale === 'zh' ? labelZh : label, + testId: `x-axis-mode-${option}`, + help: ( + <> +

+ {X_AXIS_EXPLANATIONS[kind].name[locale]( + isAgentic + ? selectedPercentile.toUpperCase() + : fixedSequenceStatistic === 'mean' + ? 'Mean' + : 'Median', + )} +

+

{X_AXIS_EXPLANATIONS[kind].description[locale]}

+ {isAgenticOnlyXAxisMode(option) && ( + )} -

-

{X_AXIS_EXPLANATIONS[kind].description[locale]}

- {isAgenticOnlyXAxisMode(option) && ( - - )} - - ), - })), - }, - ]} - /> + + ), + })), + }, + ]} + /> +
+ {mounted && !isAgentic && value !== 'concurrency' && ( +
+ + { + setFixedSequenceStatistic(statistic as FixedSequenceStatistic); + track('latency_service_statistic_selected', { statistic }); + }} + groups={[ + { + label: '', + options: (['median', 'mean'] as const).map((statistic) => ({ + value: statistic, + label: t[statistic], + testId: `fixed-sequence-statistic-${statistic}`, + help:

{t.statisticHelp}

, + })), + }, + ]} + /> +
+ )}
); diff --git a/packages/app/src/components/inference/ui/inference-table-sort.ts b/packages/app/src/components/inference/ui/inference-table-sort.ts index d9eb9b8bf..112ec5062 100644 --- a/packages/app/src/components/inference/ui/inference-table-sort.ts +++ b/packages/app/src/components/inference/ui/inference-table-sort.ts @@ -25,8 +25,8 @@ export function sortRowsByYMetric( const yAscending = rooflineDir?.startsWith('lower'); return [...data].toSorted((a, b) => { - const ay = getNestedYValue(a, yPath); - const by = getNestedYValue(b, yPath); + const ay = a.powerVariant ? a.y : getNestedYValue(a, yPath); + const by = b.powerVariant ? b.y : getNestedYValue(b, yPath); return yAscending ? ay - by : by - ay; }); } diff --git a/packages/app/src/components/inference/ui/line-label-layer.test.ts b/packages/app/src/components/inference/ui/line-label-layer.test.ts index 33bb77256..e3eda7273 100644 --- a/packages/app/src/components/inference/ui/line-label-layer.test.ts +++ b/packages/app/src/components/inference/ui/line-label-layer.test.ts @@ -14,17 +14,12 @@ interface Point { y: number; } -const series = ( - key: string, - points: Point[], - keepVisibleOnCollision = false, -): LineLabelSeries => ({ +const series = (key: string, points: Point[]): LineLabelSeries => ({ key, seriesId: key, label: key, color: '#000', points, - keepVisibleOnCollision, }); const identity = (value: number) => value; diff --git a/packages/app/src/components/inference/ui/line-label-layer.ts b/packages/app/src/components/inference/ui/line-label-layer.ts index 6da0a7a23..a3b635dfa 100644 --- a/packages/app/src/components/inference/ui/line-label-layer.ts +++ b/packages/app/src/components/inference/ui/line-label-layer.ts @@ -16,7 +16,6 @@ export interface LineLabelSeries { label: string; color: string; points: readonly TPoint[]; - keepVisibleOnCollision?: boolean; } export interface LineLabelPlacement { @@ -26,6 +25,12 @@ export interface LineLabelPlacement { color: string; x: number; y: number; + /** + * `placeLineLabels` always emits `true`: every series it is given keeps its + * pill, overlapping if it must. The only producer of `false` is the caller's + * de-duplication pass, which keeps a hidden data-join entry for a curve that + * lost the one-label-per-hardware contest (GH #470). + */ visible: boolean; } @@ -146,16 +151,16 @@ interface PillLayoutItem { * its anchor, and both mirrors together. Every candidate is clamped into * `bounds` before the overlap test, so nothing leaves the plot. When every * mirrored candidate collides, nearby rows are tried before the default spot - * is kept: an overlapped label is still - * better than a missing one, and the fallback matches what the anchor pass - * already tolerates for pinned anchors. + * is kept: an overlapped label is still better than a missing one, and the + * fallback matches what the anchor pass already tolerates. * * With no bounds — a chart that clips nothing — the anchor offset is applied * unchanged and the collision pass is skipped, preserving that chart's * existing layout. * * Hidden pills get their default transform and occupy no space, so a label - * that later becomes visible reappears where the anchor pass put it. + * that later becomes visible reappears where the anchor pass put it. Only the + * caller's de-duplication pass hides pills; the anchor pass never does. */ function layoutPills( items: readonly PillLayoutItem[], @@ -319,6 +324,21 @@ function lineCandidates( return candidates; } +/** + * Anchor one pill per series along its line. + * + * Each series tries `ANCHOR_SLOTS` fractions along its own points, rotated by + * its index so converging curves spread out instead of stacking at the + * endpoint. A series that finds a clear slot takes it. A series that finds none + * is deferred and placed afterwards on its least crowded slot, so it never + * steals a clear slot from a series that could have used it. + * + * Every series gets a visible pill. The overlap that survives here is resolved + * by `layoutPills`, which runs later with the pills' real measured boxes; the + * crude nominal box used here is far too small to decide that a label is + * unplaceable — a rendered pill is routinely two to three times + * `collisionWidth`. + */ export function placeLineLabels( series: readonly LineLabelSeries[], xScale: (value: number) => number, @@ -346,6 +366,35 @@ export function placeLineLabels( Math.abs(other.y - y) < collisionHeight && Math.abs(other.x - x) < other.halfW + labelHalfWidth, ); + /** + * Nominal overlap area against the labels already placed. The same crude box + * model as `collides`, scored instead of thresholded, so a slot that clips one + * neighbour is preferred over one that sits on three. + */ + const collisionCost = (x: number, y: number) => + placed.reduce((cost, other) => { + const dx = other.halfW + labelHalfWidth - Math.abs(other.x - x); + const dy = collisionHeight - Math.abs(other.y - y); + return dx > 0 && dy > 0 ? cost + dx * dy : cost; + }, 0); + + const emit = (entry: LineLabelSeries, point: TPoint) => { + const x = xScale(point.x); + const y = yScale(point.y); + placed.push({ x, y, halfW: labelHalfWidth }); + result.push({ + key: entry.key, + seriesId: entry.seriesId, + label: entry.label, + color: entry.color, + x, + y, + visible: true, + }); + }; + + /** Series with no clear slot, deferred to a second pass — see below. */ + const crowded: { entry: LineLabelSeries; candidates: TPoint[] }[] = []; for (const [seriesIndex, entry] of sorted.entries()) { if (entry.points.length === 0) continue; @@ -376,35 +425,37 @@ export function placeLineLabels( const candidate = candidates.find((point) => !collides(xScale(point.x), yScale(point.y))); if (candidate) { - const x = xScale(candidate.x); - const y = yScale(candidate.y); - placed.push({ x, y, halfW: labelHalfWidth }); - result.push({ - key: entry.key, - seriesId: entry.seriesId, - label: entry.label, - color: entry.color, - x, - y, - visible: true, - }); + emit(entry, candidate); continue; } - const fallback = entry.points[0]; - const x = xScale(fallback.x); - const y = yScale(fallback.y); - const visible = entry.keepVisibleOnCollision === true; - if (visible) placed.push({ x, y, halfW: labelHalfWidth }); - result.push({ - key: entry.key, - seriesId: entry.seriesId, - label: entry.label, - color: entry.color, - x, - y, - visible, - }); + // No clear slot. Defer rather than claim one now: a series that is going to + // overlap something must not take a slot a later series could have had to + // itself. + crowded.push({ entry, candidates }); + } + + // Every series keeps a pill. `layoutPills` runs after this with the real + // measured boxes and can still mirror it, shift it a row and clamp it into the + // plot — "an overlapped label is still better than a missing one". Emitting the + // crowded ones last also hands that pass the clean labels first, so the crowded + // ones do the moving. + // + // This is where the chart stopped promising that line labels never overlap + // (#132 introduced the drop as the only way to honour that, #434 restated it). + // The promise was worth less than it cost: a dropped pill is silent, and with + // line labels on, PNG export omits the legend, so the series loses its only + // identifier. An overlapping pill at least announces itself. + for (const { entry, candidates } of crowded) { + emit( + entry, + candidates.reduce((best, point) => + collisionCost(xScale(point.x), yScale(point.y)) < + collisionCost(xScale(best.x), yScale(best.y)) + ? point + : best, + ), + ); } return result; diff --git a/packages/app/src/components/inference/ui/line-label-visibility.test.ts b/packages/app/src/components/inference/ui/line-label-visibility.test.ts index 80e17b049..3b4a075d8 100644 --- a/packages/app/src/components/inference/ui/line-label-visibility.test.ts +++ b/packages/app/src/components/inference/ui/line-label-visibility.test.ts @@ -53,6 +53,26 @@ describe('labelOpacityForActiveState', () => { }); }); +describe('labelOpacityForActiveState with ?unofficialrun= overlays', () => { + const official = new Set(['gb200_dynamo-sglang']); + const precisions = ['fp8']; + + it('hides an overlay label when its overlay hardware row is off, even if the official row is on', () => { + expect( + labelOpacityForActiveState( + { + hwKey: 'gb200_dynamo-sglang', + lineKey: 'overlay-gb200_dynamo-sglang_fp8_run1', + visible: '1', + }, + official, + precisions, + new Set(['gb300_dynamo-sglang']), + ), + ).toBe(0); + }); +}); + describe('labelOpacityForHover', () => { it('lights up the kept label for the hovered hardware', () => { expect(labelOpacityForHover({ hwKey: 'b300_sglang', visible: '1' }, 'b300_sglang')).toBe(1); diff --git a/packages/app/src/components/inference/ui/line-label-visibility.ts b/packages/app/src/components/inference/ui/line-label-visibility.ts index a8b34b810..9dfa8e878 100644 --- a/packages/app/src/components/inference/ui/line-label-visibility.ts +++ b/packages/app/src/components/inference/ui/line-label-visibility.ts @@ -25,6 +25,8 @@ export interface LabelAttrs { /** `data-hw-key` — base hardware key, shared across a hw's curves. */ hwKey?: string; + /** `data-line-key` — `overlay-…` marks an unofficial-run curve's label. */ + lineKey?: string; /** `data-precision` — set on parallelism labels, absent on line labels. */ precision?: string; /** `data-visible` — `'1'`/`'0'`; only line labels set this. */ @@ -50,16 +52,24 @@ export const labelOpacityForHover = (attrs: LabelAttrs, hoveredHwKey: string): 0 * filter-change sync effect. Line labels (no precision) show when their * hardware is active **and** the render kept them; parallelism labels show when * their hardware is active and their precision is selected. + * + * An `?unofficialrun=` overlay curve answers to the overlay legend rows, not + * the official ones: its label follows `activeOverlayHwTypes`, so soloing an + * official hardware no longer hides the overlay pills of every other hardware + * (and hiding an official row keeps its overlay twin labelled). */ export const labelOpacityForActiveState = ( attrs: LabelAttrs, activeHwTypes: ReadonlySet, selectedPrecisions: readonly string[], + activeOverlayHwTypes?: ReadonlySet, ): 0 | 1 => { const { hwKey, precision } = attrs; if (!hwKey) return 0; + const isOverlay = attrs.lineKey?.startsWith('overlay-') ?? false; + const active = isOverlay && activeOverlayHwTypes ? activeOverlayHwTypes : activeHwTypes; if (!precision) { - return activeHwTypes.has(hwKey) && renderKept(attrs) ? 1 : 0; + return active.has(hwKey) && renderKept(attrs) ? 1 : 0; } - return activeHwTypes.has(hwKey) && selectedPrecisions.includes(precision) ? 1 : 0; + return active.has(hwKey) && selectedPrecisions.includes(precision) ? 1 : 0; }; diff --git a/packages/app/src/components/inference/utils.test.ts b/packages/app/src/components/inference/utils.test.ts index f1560a7fb..2b07210fb 100644 --- a/packages/app/src/components/inference/utils.test.ts +++ b/packages/app/src/components/inference/utils.test.ts @@ -226,6 +226,27 @@ describe('processOverlayChartData', () => { expect(result.clippedData).toEqual([]); }); + it('retains exact concurrency in unofficial overlays regardless of latency or optimization stamps', () => { + const points = [1, 4, 128].map((conc) => ({ + ...prefillEnergyPoint(68, 120), + conc, + isOnNormalizedInteractivityFrontier: false, + })); + const result = processOverlayChartDataWithClipping( + points, + 'e2e', + 'y_measuredPrefillJPerInputToken', + 'p90_ttft', + { isAgentic: false, selectedPercentile: 'p90', selectedXAxisMode: 'concurrency' }, + ); + expect(result.data.map((point) => [point.x, point.y])).toEqual([ + [1, 0.2], + [4, 0.2], + [128, 0.2], + ]); + expect(result.clippedData).toEqual([]); + }); + it('uses median TTFT for fixed-sequence overlays in TTFT mode and omits missing measurements', () => { const result = processOverlayChartDataWithClipping( [prefillEnergyPoint(68, 1.5), prefillEnergyPoint(55)], diff --git a/packages/app/src/components/inference/utils.ts b/packages/app/src/components/inference/utils.ts index 5284aac3b..f459b5639 100644 --- a/packages/app/src/components/inference/utils.ts +++ b/packages/app/src/components/inference/utils.ts @@ -6,10 +6,20 @@ import { getGpuSpecs, type TcoBasis } from '@/lib/constants'; */ import chartDefinitions from '@/components/inference/metric-registry'; -import { resolveXAxisField } from '@/components/inference/utils/resolveXAxisField'; +import { + resolveXAxisField, + type FixedSequenceStatistic, +} from '@/components/inference/utils/resolveXAxisField'; import { remapInferencePoint } from '@/lib/chart-utils'; +import { expandPowerCompareSeries } from '@/components/inference/utils/power-compare'; -import type { ChartDefinition, ClippedInferenceData, InferenceData, YAxisMetricKey } from './types'; +import type { + ChartDefinition, + ClippedInferenceData, + InferenceData, + PowerCompare, + YAxisMetricKey, +} from './types'; import type { XAxisMode } from './hooks/useChartData'; /** @@ -96,6 +106,9 @@ export function partitionChartDataByLimits( selectedYAxisMetric: string, options: { isTtftX: boolean; isAgentic: boolean }, ): ProcessedChartData { + if (chartDefinition.x_scale_field === 'conc') { + return { data: data.filter((point) => Number.isFinite(point.x)), clippedData: [] }; + } const costLimitApplies = selectedYAxisMetric.includes('cost') && selectedYAxisMetric !== 'y_costUser' && @@ -143,8 +156,10 @@ export function processOverlayChartData( isAgentic?: boolean; selectedPercentile?: string; selectedXAxisMode?: XAxisMode; + fixedSequenceStatistic?: FixedSequenceStatistic; restrictToNormalizedFrontier?: boolean; tcoBasis?: TcoBasis; + powerCompare?: PowerCompare; }, ): InferenceData[] { return processOverlayChartDataWithClipping( @@ -169,8 +184,11 @@ export function processOverlayChartDataWithClipping( isAgentic?: boolean; selectedPercentile?: string; selectedXAxisMode?: XAxisMode; + fixedSequenceStatistic?: FixedSequenceStatistic; restrictToNormalizedFrontier?: boolean; tcoBasis?: TcoBasis; + /** Sibling boundary / role series, mirroring the official path in useChartData. */ + powerCompare?: PowerCompare; }, ): ProcessedChartData { const chartDef = (chartDefinitions as ChartDefinition[]).find((d) => d.chartType === chartType); @@ -216,15 +234,20 @@ export function processOverlayChartDataWithClipping( isAgentic, percentile: selectedPercentile, xAxisMode: options?.selectedXAxisMode, + fixedSequenceStatistic: options?.fixedSequenceStatistic, }); // The latency limit targets overload outliers on the TTFT axis only; skip it // for the natural axis and for agentic (long TTFTs are normal there). const isTtftX = xAxisField.endsWith('_ttft'); - const processedData = sourceData - .filter((d) => metricKey in d) - .map((d) => remapInferencePoint(d, metricKey, xAxisField)); + const processedData = expandPowerCompareSeries( + sourceData + .filter((d) => metricKey in d) + .map((d) => remapInferencePoint(d, metricKey, xAxisField)), + selectedYAxisMetric, + options?.powerCompare ?? 'none', + ); // The normalized metric is derived from persisted request traces, which an // unofficial overlay does not have. An all-false canonical stamp prevents a @@ -237,8 +260,13 @@ export function processOverlayChartDataWithClipping( } } - return partitionChartDataByLimits(processedData, chartDef, selectedYAxisMetric, { - isTtftX, - isAgentic, - }); + return partitionChartDataByLimits( + processedData, + { ...chartDef, x_scale_field: xAxisField }, + selectedYAxisMetric, + { + isTtftX, + isAgentic, + }, + ); } diff --git a/packages/app/src/components/inference/utils/best-series-per-sku.ts b/packages/app/src/components/inference/utils/best-series-per-sku.ts index 6799b9631..2a35f0b9c 100644 --- a/packages/app/src/components/inference/utils/best-series-per-sku.ts +++ b/packages/app/src/components/inference/utils/best-series-per-sku.ts @@ -40,7 +40,9 @@ export function bestSeriesPerSku(points: InferenceData[], direction: Direction): const bySku = new Map>(); const featured = new Set(); for (const point of points) { - if (!isFrontierEligible(point) || !Number.isFinite(point.y)) continue; + // Comparison clones re-plot the same configs at another boundary or role; + // the best series per SKU is judged on the selected metric alone. + if (point.powerVariant || !isFrontierEligible(point) || !Number.isFinite(point.y)) continue; const sku = baseSku(point); const key = String(point.hwKey); if (point.framework === 'tilert') featured.add(key); diff --git a/packages/app/src/components/inference/utils/concurrency-series.ts b/packages/app/src/components/inference/utils/concurrency-series.ts new file mode 100644 index 000000000..94d46e18a --- /dev/null +++ b/packages/app/src/components/inference/utils/concurrency-series.ts @@ -0,0 +1,36 @@ +import type { InferenceData } from '../types'; +import { pointTopologyKey } from './topology-filter'; + +/** Observed load sweeps, never a Pareto frontier or an envelope. */ +export function groupConcurrencySeries( + points: readonly InferenceData[], +): Map { + const groups = new Map(); + for (const [index, point] of points.entries()) { + // Unknown provenance cannot establish a controlled sweep. Keep its marker. + const run = point.run_url || `unknown-run-${index}`; + const key = JSON.stringify([ + point.hwKey, + point.precision, + point.date, + run, + pointTopologyKey(point), + point.recipe_fingerprint ?? null, + point.powerVariant?.id ?? null, + ]); + const group = groups.get(key); + if (group) group.push(point); + else groups.set(key, [point]); + } + const segments = new Map(); + for (const [key, group] of groups) { + const sorted = group.toSorted((a, b) => a.x - b.x); + if (new Set(sorted.map((point) => point.x)).size === sorted.length) { + segments.set(key, sorted); + } else { + // Repeats at the same load are distinct observations, not a fitted mean. + sorted.forEach((point, index) => segments.set(`${key}:${index}`, [point])); + } + } + return segments; +} diff --git a/packages/app/src/components/inference/utils/equal-service-comparison.test.ts b/packages/app/src/components/inference/utils/equal-service-comparison.test.ts new file mode 100644 index 000000000..a0ebdd036 --- /dev/null +++ b/packages/app/src/components/inference/utils/equal-service-comparison.test.ts @@ -0,0 +1,261 @@ +import { describe, expect, it } from 'vitest'; +import type { AggDataEntry, InferenceData } from '../types'; +import { + buildEqualServiceComparison, + equalServiceSourceKey, + getEqualServiceSources, + getPrefillSharePoints, + getRolePoints, +} from './equal-service-comparison'; + +const metric = (y: number) => ({ y, roof: false }); +function point(overrides: Partial = {}): InferenceData { + return { + x: 20, + y: 500, + hwKey: 'b200_sglang', + date: '2026-09-23', + tp: 4, + physicalChips: 4, + precision: 'fp8', + conc: 8, + run_url: 'https://example.invalid/runs/1', + model: 'qwen3.5', + benchmark_type: 'single_turn', + isl: 8192, + osl: 1024, + decode_tp: 4, + mean_intvty: 20, + output_tput_per_gpu: 50, + measuredAvgPower: metric(400), + measuredJPerOutputToken: metric(10), + tpPerGpu: metric(50), + tpPerMw: metric(50), + costh: metric(1), + costr: metric(1), + costhi: metric(1), + costri: metric(1), + ...overrides, + }; +} +const a = [ + point({ id: 1 }), + point({ + id: 2, + mean_intvty: 60, + conc: 1, + measuredAvgPower: metric(800), + output_tput_per_gpu: 150, + measuredJPerOutputToken: metric(30), + }), +]; +const b = [ + point({ + id: 3, + hwKey: 'b300_sglang', + run_url: 'https://example.invalid/runs/2', + measuredAvgPower: metric(800), + output_tput_per_gpu: 100, + measuredJPerOutputToken: metric(8), + }), + point({ + id: 4, + hwKey: 'b300_sglang', + run_url: 'https://example.invalid/runs/2', + mean_intvty: 60, + conc: 1, + measuredAvgPower: metric(1000), + output_tput_per_gpu: 200, + measuredJPerOutputToken: metric(16), + }), +]; +const options = { + baseline: equalServiceSourceKey(a[0]), + comparator: equalServiceSourceKey(b[0]), + target: 40, + xField: 'mean_intvty' as const, +}; + +describe('equal-service comparison', () => { + it('interpolates each raw quantity first, retains endpoints, and uses one percentage sign convention', () => { + const result = buildEqualServiceComparison([...a, ...b], options); + expect(result.metrics.meanWattsPerGpu).toMatchObject({ + baseline: { value: 600, interpolated: true }, + comparator: { value: 900 }, + changePercent: 50, + }); + expect(result.metrics.outputTokensPerSecond).toMatchObject({ + baseline: { value: 400 }, + comparator: { value: 600 }, + changePercent: 50, + }); + expect(result.metrics.joulesPerOutputToken).toMatchObject({ + baseline: { value: 20 }, + comparator: { value: 12 }, + changePercent: -40, + }); + expect( + result.metrics.meanWattsPerGpu.baseline?.endpoints.map(({ point: p }) => [ + p.id, + p.conc, + p.run_url, + ]), + ).toEqual([ + [1, 8, a[0].run_url], + [2, 1, a[1].run_url], + ]); + // Interpolating the endpoint percentages instead would incorrectly give 62.5%. + expect(result.metrics.meanWattsPerGpu.changePercent).not.toBe(62.5); + }); + + it('does not skip a missing interior measurement or turn it into zero', () => { + const gap = point({ id: 5, mean_intvty: 40, measuredAvgPower: undefined }); + const result = buildEqualServiceComparison([...a, gap, ...b], { ...options, target: 50 }); + expect(result.metrics.meanWattsPerGpu).toMatchObject({ + baseline: null, + changePercent: null, + reason: 'missing-metric', + }); + expect(result.metrics.outputTokensPerSecond.changePercent).not.toBeNull(); + }); + + it('scales aggregate output by physical chips, but PD output by decode GPUs only', () => { + const aggregate = point({ physicalChips: 8, tp: 2 }); + const pd = point({ + hwKey: 'gb200', + disagg: true, + physicalChips: 8, + num_prefill_gpu: 4, + num_decode_gpu: 4, + }); + const result = buildEqualServiceComparison([aggregate, pd], { + ...options, + baseline: equalServiceSourceKey(aggregate), + comparator: equalServiceSourceKey(pd), + target: 20, + }); + expect(result.metrics.outputTokensPerSecond).toMatchObject({ + baseline: { value: 400 }, + comparator: { value: 200 }, + changePercent: -50, + }); + const invalid = { ...pd, num_decode_gpu: 0 }; + expect( + buildEqualServiceComparison([aggregate, invalid], { + ...options, + baseline: equalServiceSourceKey(aggregate), + comparator: equalServiceSourceKey(invalid), + target: 20, + }).metrics.outputTokensPerSecond.reason, + ).toBe('missing-metric'); + }); + + it('labels sources by hardware and date, adding only the details that tell them apart', () => { + const sources = getEqualServiceSources([ + point({ id: 1 }), + point({ id: 2, physicalChips: 8, decode_tp: 8 }), + point({ id: 3, run_url: 'https://example.invalid/runs/3' }), + point({ id: 4, hwKey: 'b300_sglang', run_url: 'https://example.invalid/runs/2' }), + ]); + expect(sources.map((source) => source.label)).toEqual([ + 'B200 (SGLang) · 2026-09-23 · Single-node · GPU4 · TP4 · EP? · Run #1', + 'B200 (SGLang) · 2026-09-23 · Single-node · GPU8 · TP8 · EP?', + 'B200 (SGLang) · 2026-09-23 · Single-node · GPU4 · TP4 · EP? · Run #3', + 'B300 (SGLang) · 2026-09-23', + ]); + }); + + it('keys a stitched append-only curve by its snapshot; rows without one keep their own run', () => { + // B200 TP4: run 35905882425 appended c1–c4 onto run 35843506474's c8–c128. + const snapshot = { curve_workflow_run_id: 35843506474, curve_date: '2026-09-20' }; + const stitched = [ + point({ + id: 1, + conc: 8, + actualDate: '2026-09-20', + power_audit: { producer_sha: 'producer-a', exporter_image_sha256: 'exporter-a' }, + ...snapshot, + }), + point({ + id: 2, + conc: 1, + actualDate: '2026-09-23', + run_url: 'https://example.invalid/runs/35905882425/attempts/1', + power_audit: { producer_sha: 'producer-b', exporter_image_sha256: 'exporter-b' }, + ...snapshot, + }), + ]; + const laterSnapshot = point({ + id: 3, + actualDate: '2026-09-20', + curve_workflow_run_id: 35900000000, + curve_date: '2026-09-20', + }); + const legacy = [ + point({ id: 4, hwKey: 'b300_sglang', run_url: 'https://example.invalid/runs/7' }), + point({ id: 5, hwKey: 'b300_sglang', run_url: 'https://example.invalid/runs/8' }), + ]; + expect(equalServiceSourceKey(stitched[0])).toBe(equalServiceSourceKey(stitched[1])); + expect(equalServiceSourceKey(legacy[0])).not.toBe(equalServiceSourceKey(legacy[1])); + const sources = getEqualServiceSources([...stitched, laterSnapshot, ...legacy]); + expect(sources.map((source) => source.label)).toEqual([ + 'B200 (SGLang) · 2026-09-20 · Run #35843506474', + 'B200 (SGLang) · 2026-09-20 · Run #35900000000', + 'B300 (SGLang) · 2026-09-23 · Run #7', + 'B300 (SGLang) · 2026-09-23 · Run #8', + ]); + }); + + it.each([ + { image: 'another-image' }, + { recipe_fingerprint: 'another-recipe' }, + { decode_tp: 8, physicalChips: 8 }, + { curve_workflow_run_id: 2 }, + ])('keeps distinct snapshot configurations separate: %j', (variant) => { + const original = point({ curve_workflow_run_id: 1, curve_date: '2026-09-20' }); + expect(getEqualServiceSources([original, { ...original, ...variant }])).toHaveLength(2); + }); + + it.each([{ producer_sha: 'another-producer' }, { exporter_image_sha256: 'another-exporter' }])( + 'keeps producer distinctions when no snapshot authorizes stitching: %j', + (power_audit) => { + expect(getEqualServiceSources([point(), point({ power_audit })])).toHaveLength(2); + }, + ); + + it('plots role panels on the trace-derived P75/P90 axes from point.x without interpolating on them', () => { + const role = (overrides: Partial) => + point({ + disagg: true, + num_prefill_gpu: 4, + num_decode_gpu: 4, + power_valid: 1, + power_metric_schema_version: 2, + joules_per_input_token: 1, + joules_per_output_token: 8, + prefill_joules_per_input_token: 0.4, + decode_joules_per_output_token: 4.8, + measuredPrefillAvgPower: metric(300), + measuredDecodeAvgPower: metric(500), + ...overrides, + }); + // The chart and the views API both store the derived value on `x` only. + const rows = [ + role({ id: 1, x: 31.2 }), + role({ id: 2, x: 24.8, conc: 16 }), + role({ id: 3, x: 28, hwKey: 'b300_sglang', run_url: 'https://example.invalid/runs/2' }), + ]; + const derived = 'p90_e2e_norm_intvty' as keyof AggDataEntry; + expect(getRolePoints(rows, derived).map((row) => row.x)).toEqual([24.8, 31.2, 28]); + expect(getPrefillSharePoints(rows, derived).map((row) => row.x)).toEqual([24.8, 31.2, 28]); + expect(getRolePoints(rows, 'p75_e2e_norm_intvty' as keyof AggDataEntry)).toHaveLength(3); + expect( + buildEqualServiceComparison(rows, { + baseline: equalServiceSourceKey(rows[0]), + comparator: equalServiceSourceKey(rows[2]), + target: 28, + xField: derived, + }).reason, + ).toBe('unsupported-axis'); + }); +}); diff --git a/packages/app/src/components/inference/utils/equal-service-comparison.ts b/packages/app/src/components/inference/utils/equal-service-comparison.ts new file mode 100644 index 000000000..c74e061b5 --- /dev/null +++ b/packages/app/src/components/inference/utils/equal-service-comparison.ts @@ -0,0 +1,391 @@ +import type { AggDataEntry, InferenceData } from '../types'; +import { chipCounts } from '@/lib/chip-counts'; +import { getHardwareConfig } from '@/lib/constants'; +import { isPositive, powerBasisNormalization } from '@/lib/power-basis'; +import { getDisplayLabel } from '@/lib/utils'; +import { runIdFromUrl } from './powerTimeline'; +import { reconstructedRoleEnergy, type ReconstructedRoleEnergy } from './role-energy'; +import { pointTopologyKey, topologyLabel } from './topology-filter'; + +export interface EqualServiceSource { + key: string; + label: string; +} +export interface EqualServiceEstimate { + value: number; + interpolated: boolean; + endpoints: { x: number; value: number; point: InferenceData }[]; +} +export type EqualServiceReason = + | 'unsupported-axis' + | 'invalid-target' + | 'same-source' + | 'unknown-source' + | 'out-of-range' + | 'ambiguous-x' + | 'missing-metric'; +export interface EqualServiceMetric { + baseline: EqualServiceEstimate | null; + comparator: EqualServiceEstimate | null; + /** Comparator change relative to baseline; negative energy means less energy. */ + changePercent: number | null; + reason?: EqualServiceReason; +} +export interface EqualServiceOptions { + baseline: string; + comparator: string; + target: number; + xField: keyof AggDataEntry; +} +export interface EqualServiceComparison { + target: number; + xField: keyof AggDataEntry; + baseline: EqualServiceSource | null; + comparator: EqualServiceSource | null; + reason?: EqualServiceReason; + metrics: Record< + 'meanWattsPerGpu' | 'outputTokensPerSecond' | 'joulesPerOutputToken', + EqualServiceMetric + >; +} +const serviceAxis = (field: string) => + field === 'mean_tpot_intvty' || + /^(?:mean|median|p\d+(?:\.\d+)?)_(?:intvty|tpot|ttft|e2el|itl)$/u.test(field); +/** + * Trace-derived agentic axes (`p75_e2e_norm_intvty`, `p90_e2e_norm_intvty`) + * live only on `point.x`; no row field carries that name. The role panels + * plot observations, so they accept them; equal-service interpolation keeps + * its observed-field policy and reports `unsupported-axis`. + */ +const derivedAxis = (field: string) => /^p\d+_e2e_norm_intvty$/u.test(field); +const roleAxisValue = (point: InferenceData, xField: keyof AggDataEntry) => + derivedAxis(xField) ? point.x : point[xField]; +/** A positive finite reading, or null: a missing value is never zero. */ +export const positiveOrNull = (value: unknown): number | null => (isPositive(value) ? value : null); +/** Measured rows only: hidden rows and power-comparison clones are never sources. */ +export const observedPoints = (points: readonly InferenceData[]) => + points.filter((point) => !point.hidden && !point.powerVariant); + +/** + * The curve snapshot a row belongs to, or null when it carries none. An + * append-only run stitches new points onto an older run's curve; the chart + * draws them as one series, so the panels treat them as one source. + */ +const curveSnapshotId = (point: InferenceData): number | null => + point.curve_workflow_run_id ?? null; +/** Snapshot date of a stitched curve, else the row's own measured date. */ +const sourceDate = (point: InferenceData): string => + (curveSnapshotId(point) === null ? undefined : point.curve_date) ?? + point.actualDate ?? + point.date; + +/** + * Concurrency is excluded; unknown run identity must never join distinct rows. + * Rows with a curve snapshot key by that snapshot instead of their own run. + * A snapshot can retain multiple telemetry producers; their hashes stay on the + * observations, while recipe, image and topology still distinguish configurations. + */ +export function equalServiceSourceKey(point: InferenceData): string { + const snapshot = curveSnapshotId(point); + const attempt = 'run_attempt' in point ? point.run_attempt : null; + return JSON.stringify([ + point.hwKey, + point.model ?? null, + point.framework ?? null, + point.precision, + point.benchmark_type ?? null, + point.isl ?? null, + point.osl ?? null, + sourceDate(point), + snapshot === null + ? point.run_url || `unknown-run-point-${point.id ?? JSON.stringify(point)}` + : `curve-${snapshot}`, + snapshot === null ? attempt : null, + pointTopologyKey(point), + point.recipe_fingerprint ?? null, + point.image ?? null, + point.mtp ?? null, + point.spec_decoding ?? null, + point.kv_offloading ?? null, + point.kv_offload_backend ?? null, + point.kv_offload_backend_version ?? null, + point.kv_p2p_transfer ?? null, + point.router_name ?? null, + point.router_version ?? null, + snapshot === null ? (point.power_audit?.producer_sha ?? null) : null, + snapshot === null ? (point.power_audit?.exporter_image_sha256 ?? null) : null, + ]); +} + +const SOURCE_LABEL_WORDS = { + en: { + point: 'Point', + run: (id: string) => `Run #${id}`, + attempt: (attempt: unknown) => `Attempt ${attempt}`, + recipe: 'Recipe', + }, + zh: { + point: '数据点', + run: (id: string) => `运行 #${id}`, + attempt: (attempt: unknown) => `第 ${attempt} 次尝试`, + recipe: '配方', + }, +}; + +/** + * Hardware and snapshot date, plus only the details that tell otherwise + * identical sources apart, in this order; the opaque key stays the exact + * identity. A stitched curve is named by its snapshot run, not each row's run. + */ +function sourceLabels(points: readonly InferenceData[], locale: 'en' | 'zh'): string[] { + const words = SOURCE_LABEL_WORDS[locale]; + const details: ((point: InferenceData, group: readonly InferenceData[]) => string | null)[] = [ + (point) => point.precision.toUpperCase(), + (point, group) => topologyLabel(pointTopologyKey(point), locale, group.map(pointTopologyKey)), + (point) => { + const snapshot = curveSnapshotId(point); + const runId = snapshot === null ? runIdFromUrl(point.run_url) : String(snapshot); + return runId ? words.run(runId) : null; + }, + (point) => + curveSnapshotId(point) === null && 'run_attempt' in point + ? words.attempt(point.run_attempt) + : null, + (point) => + point.recipe_fingerprint ? `${words.recipe} ${point.recipe_fingerprint.slice(0, 8)}` : null, + (point) => point.image ?? null, + (point) => `${words.point} ${point.id ?? '?'}`, + ]; + const labels = points.map((point) => { + const hardware = getHardwareConfig(point.hwKey); + return [ + hardware.name === 'unknown' ? point.hwKey : getDisplayLabel(hardware), + sourceDate(point), + ].join(' · '); + }); + for (const detail of details) { + const groups = new Map(); + labels.forEach((label, index) => groups.set(label, [...(groups.get(label) ?? []), index])); + for (const indices of groups.values()) { + if (indices.length < 2) continue; + const group = indices.map((index) => points[index]); + const values = group.map((point) => detail(point, group)); + if (new Set(values).size < 2) continue; + indices.forEach((index, position) => { + if (values[position]) labels[index] += ` · ${values[position]}`; + }); + } + } + return labels; +} + +/** Labels are English by default: the read-only API has no locale. */ +export function getEqualServiceSources( + points: readonly InferenceData[], + locale: 'en' | 'zh' = 'en', +): EqualServiceSource[] { + const sources = [ + ...new Map(observedPoints(points).map((point) => [equalServiceSourceKey(point), point])), + ].sort(([a], [b]) => a.localeCompare(b)); + const labels = sourceLabels( + sources.map(([, point]) => point), + locale, + ); + return sources.map(([key], index) => ({ key, label: labels[index] })); +} + +function deploymentOutput(point: InferenceData): number | undefined { + if (!isPositive(point.output_tput_per_gpu)) return undefined; + if (point.disagg) { + return ( + powerBasisNormalization({ + output_tput_per_gpu: point.output_tput_per_gpu, + disagg: true, + benchmark_type: point.benchmark_type, + num_prefill_gpu: point.num_prefill_gpu ?? 0, + num_decode_gpu: point.num_decode_gpu ?? 0, + }).totalOutputTokPerSec ?? undefined + ); + } + const count = chipCounts(point, false).physical; + return isPositive(count) && Number.isSafeInteger(count) + ? point.output_tput_per_gpu * count + : undefined; +} +/** The measured quantities every service panel compares, read from one point. */ +export const serviceMetricValue = { + meanWattsPerGpu: (point: InferenceData) => point.measuredAvgPower?.y, + outputTokensPerSecond: deploymentOutput, + joulesPerOutputToken: (point: InferenceData) => point.measuredJPerOutputToken?.y, +}; + +function estimate( + points: readonly InferenceData[], + field: keyof AggDataEntry, + target: number, + quantity: (point: InferenceData) => number | undefined, +): { estimate: EqualServiceEstimate | null; reason?: EqualServiceReason } { + const rows = points + .flatMap((point) => { + const x = point[field]; + return isPositive(x) ? [{ point, x, value: quantity(point) }] : []; + }) + .sort((a, b) => a.x - b.x || (a.point.id ?? 0) - (b.point.id ?? 0)); + const xs = [...new Set(rows.map((row) => row.x))]; + if (xs.length === 0 || target < xs[0] || target > xs.at(-1)!) + return { estimate: null, reason: 'out-of-range' }; + const upper = xs.findIndex((x) => x >= target); + const brackets = xs[upper] === target ? [target] : [xs[upper - 1], xs[upper]]; + const endpoints: EqualServiceEstimate['endpoints'] = []; + for (const x of brackets) { + const matches = rows.filter((row) => row.x === x); + if (matches.some((row) => !Object.is(row.value, matches[0].value))) + return { estimate: null, reason: 'ambiguous-x' }; + for (const row of matches) { + if (!isPositive(row.value)) return { estimate: null, reason: 'missing-metric' }; + endpoints.push({ ...row, value: row.value }); + } + } + const left = endpoints[0], + right = endpoints.at(-1)!; + const value = + brackets.length === 1 + ? left.value + : left.value + (right.value - left.value) * ((target - left.x) / (right.x - left.x)); + return isPositive(value) + ? { estimate: { value, interpolated: brackets.length === 2, endpoints } } + : { estimate: null, reason: 'missing-metric' }; +} + +/** Numerical linear interpolation of raw quantities, then ratios; never extrapolation. */ +export function buildEqualServiceComparison( + points: readonly InferenceData[], + options: EqualServiceOptions, +): EqualServiceComparison { + const { baseline, comparator, target, xField } = options; + const sources = getEqualServiceSources(points); + const sourceA = sources.find((source) => source.key === baseline) ?? null; + const sourceB = sources.find((source) => source.key === comparator) ?? null; + let reason: EqualServiceReason | undefined; + if (!serviceAxis(xField)) reason = 'unsupported-axis'; + else if (!isPositive(target)) reason = 'invalid-target'; + else if (baseline === comparator) reason = 'same-source'; + else if (!sourceA || !sourceB) reason = 'unknown-source'; + const rows = observedPoints(points); + const a = rows.filter((point) => equalServiceSourceKey(point) === baseline); + const b = rows.filter((point) => equalServiceSourceKey(point) === comparator); + const compare = (quantity: (point: InferenceData) => number | undefined): EqualServiceMetric => { + if (reason) return { baseline: null, comparator: null, changePercent: null, reason }; + const left = estimate(a, xField, target, quantity); + const right = estimate(b, xField, target, quantity); + const change = + left.estimate && right.estimate + ? 100 * (right.estimate.value / left.estimate.value - 1) + : null; + return { + baseline: left.estimate, + comparator: right.estimate, + changePercent: change !== null && Number.isFinite(change) ? change : null, + ...(left.reason || right.reason ? { reason: left.reason ?? right.reason } : {}), + }; + }; + return { + target, + xField, + baseline: sourceA, + comparator: sourceB, + ...(reason ? { reason } : {}), + metrics: { + meanWattsPerGpu: compare(serviceMetricValue.meanWattsPerGpu), + outputTokensPerSecond: compare(serviceMetricValue.outputTokensPerSecond), + joulesPerOutputToken: compare(serviceMetricValue.joulesPerOutputToken), + }, + }; +} + +/** Knots are observed X values within both source ranges, not a fitted hardware model. */ +export function getEqualServiceComparisonCurve( + points: readonly InferenceData[], + options: Omit, +): EqualServiceComparison[] { + if (!serviceAxis(options.xField) || options.baseline === options.comparator) return []; + const xs = (key: string) => + observedPoints(points) + .filter((point) => equalServiceSourceKey(point) === key) + .map((point) => point[options.xField]) + .filter(isPositive); + const a = xs(options.baseline), + b = xs(options.comparator); + if (a.length === 0 || b.length === 0) return []; + const lower = Math.max(Math.min(...a), Math.min(...b)); + const upper = Math.min(Math.max(...a), Math.max(...b)); + return [...new Set([...a, ...b])] + .filter((x) => x >= lower && x <= upper) + .sort((x, y) => x - y) + .map((target) => buildEqualServiceComparison(points, { ...options, target })); +} + +export function getPrefillSharePoints( + points: readonly InferenceData[], + xField: keyof AggDataEntry, +) { + if (!serviceAxis(xField) && xField !== 'conc' && !derivedAxis(xField)) return []; + return observedPoints(points) + .flatMap((point) => { + const x = roleAxisValue(point, xField); + const energy = reconstructedRoleEnergy(point); + return isPositive(x) && energy + ? [{ x, sourceKey: equalServiceSourceKey(point), point, ...energy }] + : []; + }) + .sort((a, b) => a.sourceKey.localeCompare(b.sourceKey) || a.x - b.x); +} + +export interface RolePoint { + x: number; + sourceKey: string; + point: InferenceData; + /** Mean board W/GPU inside each pool (role-local). */ + prefillWattsPerGpu: number | null; + decodeWattsPerGpu: number | null; + /** Role-local energy: each pool's joules over its own token count. */ + prefillJoulesPerInputToken: number | null; + decodeJoulesPerOutputToken: number | null; + /** Both pools on the output-token denominator, when reconstructable. */ + energy: ReconstructedRoleEnergy | null; +} + +/** + * Validated disaggregated rows with any role measurement (PowerX Figures + * 12–14). Pool-local denominators stay separate from the additive + * output-token reconstruction; a missing figure is null, never zero. + */ +export function getRolePoints( + points: readonly InferenceData[], + xField: keyof AggDataEntry, +): RolePoint[] { + if (!serviceAxis(xField) && xField !== 'conc' && !derivedAxis(xField)) return []; + return observedPoints(points) + .flatMap((point): RolePoint[] => { + const x = roleAxisValue(point, xField); + if (!point.disagg || !isPositive(x)) return []; + const role: RolePoint = { + x, + sourceKey: equalServiceSourceKey(point), + point, + prefillWattsPerGpu: positiveOrNull(point.measuredPrefillAvgPower?.y), + decodeWattsPerGpu: positiveOrNull(point.measuredDecodeAvgPower?.y), + prefillJoulesPerInputToken: positiveOrNull(point.measuredPrefillJPerInputToken?.y), + decodeJoulesPerOutputToken: positiveOrNull(point.measuredDecodeJPerOutputToken?.y), + energy: reconstructedRoleEnergy(point) ?? null, + }; + const measured = + role.prefillWattsPerGpu !== null || + role.decodeWattsPerGpu !== null || + role.prefillJoulesPerInputToken !== null || + role.decodeJoulesPerOutputToken !== null || + role.energy !== null; + return measured ? [role] : []; + }) + .sort((a, b) => a.sourceKey.localeCompare(b.sourceKey) || a.x - b.x); +} diff --git a/packages/app/src/components/inference/utils/frontier-points.ts b/packages/app/src/components/inference/utils/frontier-points.ts new file mode 100644 index 000000000..8af690d21 --- /dev/null +++ b/packages/app/src/components/inference/utils/frontier-points.ts @@ -0,0 +1,44 @@ +import type { InferenceData } from '../types'; +import { runAttemptFromUrl, runIdFromUrl } from './powerTimeline'; +import { pointTopologyKey } from './topology-filter'; + +/** + * CSV export of a cross-platform frontier (PowerX Figure 16): the chart's own + * global non-dominated set, one row per observation with its provenance. + * Stable export columns: identity first, then plotted coordinates and provenance. + */ +export const FRONTIER_EXPORT_HEADERS = [ + 'hardware', + 'framework', + 'precision', + 'topology', + 'concurrency', + 'x', + 'y', + 'date', + 'run_url', + 'run_id', + 'run_attempt', + 'recipe_fingerprint', + 'image', + 'point_id', +] as const; + +export function frontierExportRow(point: InferenceData): (string | number | null)[] { + return [ + point.hwKey, + point.framework ?? null, + point.precision, + pointTopologyKey(point), + point.conc, + point.x, + point.y, + point.actualDate ?? point.date, + point.run_url ?? null, + runIdFromUrl(point.run_url), + runAttemptFromUrl(point.run_url), + point.recipe_fingerprint ?? null, + point.image ?? null, + point.id ?? null, + ]; +} diff --git a/packages/app/src/components/inference/utils/legacy-power-marker.ts b/packages/app/src/components/inference/utils/legacy-power-marker.ts deleted file mode 100644 index 43be6f4c8..000000000 --- a/packages/app/src/components/inference/utils/legacy-power-marker.ts +++ /dev/null @@ -1,38 +0,0 @@ -import type * as d3 from 'd3'; - -import type { InferenceData } from '@/components/inference/types'; -import { setAttrIfChanged } from '@/lib/d3-chart/chart-update'; -import { - LEGACY_POWER_RING_DASHARRAY, - LEGACY_POWER_RING_RADIUS, - LEGACY_POWER_RING_STROKE_WIDTH, -} from '@/components/inference/ui/LegacyPowerLegendKey'; - -/** - * Dotted ring flagging measured-power telemetry without a producer validation - * verdict (`power_tier === 'legacy'`). Drawn only while a Measured Energy - * y-axis is selected; the join-on-empty-data pattern removes the ring when the - * axis changes or the point's tier is not legacy. - */ -export function renderLegacyPowerRing( - group: d3.Selection, - point: InferenceData, - isMeasuredAxis: boolean, - stroke: string, -): void { - group - .selectAll('.legacy-power-ring') - .data(isMeasuredAxis && point.power_tier === 'legacy' ? [true] : []) - .join('circle') - // Every chart render re-syncs every point; skip unchanged writes. - .each(function () { - setAttrIfChanged(this, 'class', 'legacy-power-ring'); - setAttrIfChanged(this, 'r', String(LEGACY_POWER_RING_RADIUS)); - setAttrIfChanged(this, 'fill', 'none'); - setAttrIfChanged(this, 'stroke', stroke); - setAttrIfChanged(this, 'stroke-width', String(LEGACY_POWER_RING_STROKE_WIDTH)); - setAttrIfChanged(this, 'stroke-dasharray', LEGACY_POWER_RING_DASHARRAY); - setAttrIfChanged(this, 'opacity', '0.9'); - setAttrIfChanged(this, 'pointer-events', 'none'); - }); -} diff --git a/packages/app/src/components/inference/utils/matched-concurrency.test.ts b/packages/app/src/components/inference/utils/matched-concurrency.test.ts new file mode 100644 index 000000000..e6fb61279 --- /dev/null +++ b/packages/app/src/components/inference/utils/matched-concurrency.test.ts @@ -0,0 +1,74 @@ +import { describe, expect, it } from 'vitest'; +import type { InferenceData } from '../types'; +import { equalServiceSourceKey } from './equal-service-comparison'; +import { buildMatchedConcurrencyTable } from './matched-concurrency'; + +const metric = (y: number) => ({ y, roof: false }); +function point(overrides: Partial = {}): InferenceData { + return { + x: 100, + y: 1, + hwKey: 'b200_sglang', + date: '2026-09-18', + tp: 4, + physicalChips: 4, + precision: 'fp8', + conc: 1, + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/35317697106/attempts/1', + model: 'Qwen-3.5-397B-A17B', + benchmark_type: 'single_turn', + isl: 8192, + osl: 1024, + decode_tp: 4, + output_tput_per_gpu: 50, + tpPerGpu: metric(50), + tpPerMw: metric(50), + costh: metric(1), + costr: metric(1), + costhi: metric(1), + costri: metric(1), + ...overrides, + }; +} +// PowerX Figure 6 values (B200 / MI355X, TP4, three-dispatch means). +const b200 = (conc: number, joules: number, watts: number, speed: number, id: number) => + point({ + id, + conc, + mean_tpot_intvty: speed, + measuredJPerOutputToken: metric(joules), + measuredAvgPower: metric(watts), + }); +const mi355x = (conc: number, joules: number, watts: number, speed: number, id: number) => + point({ + id, + conc, + hwKey: 'mi355x_sglang', + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/33348766792/attempts/2', + mean_tpot_intvty: speed, + measuredJPerOutputToken: metric(joules), + measuredAvgPower: metric(watts), + }); +const baseline = [b200(1, 9.065, 402, 186.9, 1), b200(4, 3.87, 507, 141.2, 2)]; +const comparator = [mi355x(1, 13.271, 412, 128.1, 11), mi355x(128, 1.311, 869, 22.1, 12)]; +const options = { + baseline: equalServiceSourceKey(baseline[0]), + comparator: equalServiceSourceKey(comparator[0]), + interactivityField: 'mean_tpot_intvty' as const, +}; + +describe('matched-concurrency table', () => { + it('pairs observed rows by concurrency with one signed change convention', () => { + const table = buildMatchedConcurrencyTable([...baseline, ...comparator], options); + expect(table.reason).toBeUndefined(); + expect(table.rows.map((row) => row.concurrency)).toEqual([1, 4, 128]); + const [first] = table.rows; + expect(first.baseline).toMatchObject({ + status: 'observed', + values: { joulesPerOutputToken: 9.065, meanWattsPerGpu: 402, interactivity: 186.9 }, + }); + expect(first.changePercent.joulesPerOutputToken).toBeCloseTo(46.4, 1); + expect(first.changePercent.meanWattsPerGpu).toBeCloseTo(2.49, 2); + expect(first.changePercent.interactivity).toBeCloseTo(-31.46, 2); + }); +}); diff --git a/packages/app/src/components/inference/utils/matched-concurrency.ts b/packages/app/src/components/inference/utils/matched-concurrency.ts new file mode 100644 index 000000000..41d97e868 --- /dev/null +++ b/packages/app/src/components/inference/utils/matched-concurrency.ts @@ -0,0 +1,128 @@ +import type { AggDataEntry, InferenceData } from '../types'; +import { + equalServiceSourceKey, + getEqualServiceSources, + observedPoints, + positiveOrNull, + serviceMetricValue, + type EqualServiceSource, +} from './equal-service-comparison'; + +/** + * Pairs two sources' observations at the same concurrency (PowerX Figures 6, + * 7 and 9). Matched load is a diagnostic next to the equal-service comparison, + * not a replacement: the two sources usually serve different speeds at one + * concurrency. Nothing is interpolated, averaged or carried between rows. + */ +export const MATCHED_CONCURRENCY_METRICS = [ + 'joulesPerOutputToken', + 'meanWattsPerGpu', + 'interactivity', +] as const; +export type MatchedConcurrencyMetric = (typeof MATCHED_CONCURRENCY_METRICS)[number]; +export type MatchedConcurrencyValues = Record; + +export type MatchedConcurrencySide = + | { status: 'observed'; point: InferenceData; values: MatchedConcurrencyValues } + | { status: 'missing' } + /** Several observations of one source at one load disagree; none is chosen. */ + | { status: 'ambiguous'; points: InferenceData[] }; + +export interface MatchedConcurrencyRow { + concurrency: number; + baseline: MatchedConcurrencySide; + comparator: MatchedConcurrencySide; + /** `100 × (comparator ÷ baseline − 1)`; null unless both values were observed. */ + changePercent: MatchedConcurrencyValues; +} + +export type MatchedConcurrencyReason = 'same-source' | 'unknown-source'; + +export interface MatchedConcurrencyTable { + baseline: EqualServiceSource | null; + comparator: EqualServiceSource | null; + /** Streaming-speed field read for the `interactivity` column (selected statistic). */ + interactivityField: keyof AggDataEntry; + reason?: MatchedConcurrencyReason; + rows: MatchedConcurrencyRow[]; +} + +const NO_CHANGE: MatchedConcurrencyValues = { + joulesPerOutputToken: null, + meanWattsPerGpu: null, + interactivity: null, +}; + +function values(point: InferenceData, field: keyof AggDataEntry): MatchedConcurrencyValues { + return { + joulesPerOutputToken: positiveOrNull(serviceMetricValue.joulesPerOutputToken(point)), + meanWattsPerGpu: positiveOrNull(serviceMetricValue.meanWattsPerGpu(point)), + interactivity: positiveOrNull(point[field]), + }; +} + +function side(points: readonly InferenceData[], field: keyof AggDataEntry): MatchedConcurrencySide { + if (points.length === 0) return { status: 'missing' }; + const first = values(points[0], field); + const agree = points.every((point) => { + const other = values(point, field); + return MATCHED_CONCURRENCY_METRICS.every((key) => Object.is(other[key], first[key])); + }); + return agree + ? { status: 'observed', point: points[0], values: first } + : { status: 'ambiguous', points: [...points] }; +} + +export function buildMatchedConcurrencyTable( + points: readonly InferenceData[], + options: { baseline: string; comparator: string; interactivityField: keyof AggDataEntry }, +): MatchedConcurrencyTable { + const { baseline, comparator, interactivityField } = options; + const sources = getEqualServiceSources(points); + const sourceA = sources.find((source) => source.key === baseline) ?? null; + const sourceB = sources.find((source) => source.key === comparator) ?? null; + const table = { baseline: sourceA, comparator: sourceB, interactivityField }; + if (baseline === comparator) return { ...table, reason: 'same-source', rows: [] }; + if (!sourceA || !sourceB) return { ...table, reason: 'unknown-source', rows: [] }; + + const byConcurrency = (key: string) => { + const groups = new Map(); + for (const point of observedPoints(points)) { + if (equalServiceSourceKey(point) !== key || !Number.isSafeInteger(point.conc)) continue; + if (point.conc <= 0) continue; + const group = groups.get(point.conc); + if (group) group.push(point); + else groups.set(point.conc, [point]); + } + return groups; + }; + const a = byConcurrency(baseline); + const b = byConcurrency(comparator); + const rows = [...new Set([...a.keys(), ...b.keys()])] + .toSorted((x, y) => x - y) + .map((concurrency): MatchedConcurrencyRow => { + const left = side(a.get(concurrency) ?? [], interactivityField); + const right = side(b.get(concurrency) ?? [], interactivityField); + if (left.status !== 'observed' || right.status !== 'observed') { + return { concurrency, baseline: left, comparator: right, changePercent: NO_CHANGE }; + } + const change = (key: MatchedConcurrencyMetric) => { + const before = left.values[key]; + const after = right.values[key]; + if (before === null || after === null) return null; + const percent = 100 * (after / before - 1); + return Number.isFinite(percent) ? percent : null; + }; + return { + concurrency, + baseline: left, + comparator: right, + changePercent: { + joulesPerOutputToken: change('joulesPerOutputToken'), + meanWattsPerGpu: change('meanWattsPerGpu'), + interactivity: change('interactivity'), + }, + }; + }); + return { ...table, rows }; +} diff --git a/packages/app/src/components/inference/utils/point-identity.ts b/packages/app/src/components/inference/utils/point-identity.ts index 2c975593e..0cc8bbed8 100644 --- a/packages/app/src/components/inference/utils/point-identity.ts +++ b/packages/app/src/components/inference/utils/point-identity.ts @@ -23,9 +23,47 @@ export function scatterPointConfigId(point: InferenceData): string { // Agentic series omit spec decoding from hwKey so one curve can mix methods. // It remains point identity to avoid collapsing overlapping MTP/STP results. key += agenticSpecDecodingKeySuffix(point); + // Comparison clones share every config field with their base point. + if (point.powerVariant) key += `|variant-${point.powerVariant.id}`; return key; } +/** + * Comparison-series suffix inside a scatter series key. Letters, digits and + * dashes only, so the key stays a valid CSS class token (the perf ruler and + * `i_rulers` address rooflines by class) and needs no escaping. + */ +const SERIES_VARIANT_DELIMITER = '-v-'; + +/** + * Identity of one drawn series: hardware key, precision and, on a power + * comparison, the boundary or role variant. Rooflines, frontiers, line labels + * and the perf ruler all key on this string. + */ +export function scatterSeriesKey( + point: Pick, +): string { + const base = `${point.hwKey}_${point.precision}`; + return point.powerVariant ? `${base}${SERIES_VARIANT_DELIMITER}${point.powerVariant.id}` : base; +} + +export interface ScatterSeriesIdentity { + hw: string; + precision: string; + /** Comparison variant id (`gpu-provisioned`, `prefill`, …) or null for the base series. */ + variant: string | null; +} + +/** Inverse of `scatterSeriesKey`; hardware keys may themselves contain underscores. */ +export function parseScatterSeriesKey(key: string): ScatterSeriesIdentity { + const delimiter = key.indexOf(SERIES_VARIANT_DELIMITER); + const core = delimiter === -1 ? key : key.slice(0, delimiter); + const variant = delimiter === -1 ? null : key.slice(delimiter + SERIES_VARIANT_DELIMITER.length); + const parts = core.split('_'); + const precision = parts.pop() ?? ''; + return { hw: parts.join('_'), precision, variant }; +} + /** * Stable D3 join key for an official scatter point. * diff --git a/packages/app/src/components/inference/utils/power-compare.ts b/packages/app/src/components/inference/utils/power-compare.ts new file mode 100644 index 000000000..d0dea0b73 --- /dev/null +++ b/packages/app/src/components/inference/utils/power-compare.ts @@ -0,0 +1,326 @@ +/** + * Comparison series for the gated measured-power charts (`i_pcompare`). + * + * The selected metric stays the chart's base series. A comparison adds sibling + * series drawn from the SAME points — one per other power boundary + * (`boundaries`, PowerX Figures 2/3) or per worker role (`roles`, Figures 6/7) + * — so ScatterGraph can draw them with the hardware's colour and a per-variant + * dash through its ordinary series pipeline (frontier, Optimal Only, tooltip, + * table, CSV, `?unofficialrun=` overlay). A variant point is a clone of its + * base point with `y` remapped and `powerVariant` set; base points are left + * untouched so a chart without comparison is byte-identical to before. + * + * Variants exist only where the metric names a whole-deployment average + * (W per chip) or J per output token, because those are the only quantities + * every boundary and role publishes on a common axis; everywhere else the + * comparison yields no series and the control says so. + */ +import type { Locale } from '@/lib/i18n'; +import { + POWER_BASES, + POWER_BASIS_FIELDS, + POWER_BASIS_LABELS, + type PowerBasis, +} from '@/lib/power-basis'; + +import { getMeasuredMetricConfig } from '../measured-metric-config'; +import type { InferenceData, PowerCompare, PowerRole, PowerVariant } from '../types'; + +export const POWER_COMPARE_MODES = [ + 'none', + 'boundaries', + 'roles', +] as const satisfies readonly PowerCompare[]; + +export function parsePowerCompare(value: string | null | undefined): PowerCompare { + return value === 'boundaries' || value === 'roles' ? value : 'none'; +} + +/** Point fields a comparison series can plot; each is a `{ y, roof }` pair. */ +export type PowerSeriesField = keyof Pick< + InferenceData, + | 'measuredAvgPower' + | 'measuredPrefillAvgPower' + | 'measuredDecodeAvgPower' + | 'measuredJPerOutputToken' + | 'measuredDecodeJPerOutputToken' + | 'reconstructedPrefillJPerOutputToken' + | 'gpuProvisionedWatts' + | 'gpuProvisionedJPerOutputToken' + | 'utilityProvisionedWatts' + | 'utilityProvisionedJPerOutputToken' + | 'utilityModeledWatts' + | 'utilityModeledJPerOutputToken' +>; + +export interface PowerCompareSeries { + variant: PowerVariant; + field: PowerSeriesField; +} + +const ROLES = ['all', 'prefill', 'decode'] as const satisfies readonly PowerRole[]; + +const ROLE_WATT_FIELDS: Record = { + all: 'measuredAvgPower', + prefill: 'measuredPrefillAvgPower', + decode: 'measuredDecodeAvgPower', +}; +// Energy per output token per role. The prefill pool's own figure is per input +// token; `reconstructedPrefillJPerOutputToken` carries it onto the output-token +// axis (utils/role-energy.ts). +const ROLE_ENERGY_FIELDS: Record = { + all: 'measuredJPerOutputToken', + prefill: 'reconstructedPrefillJPerOutputToken', + decode: 'measuredDecodeJPerOutputToken', +}; + +const ROLE_LABELS: Record = { + all: { en: 'All GPUs', zh: '全部 GPU' }, + prefill: { en: 'Prefill GPUs', zh: '预填充 GPU' }, + decode: { en: 'Decode GPUs', zh: '解码 GPU' }, +}; + +/** SVG dash per variant; the base series and `all` stay solid. */ +const VARIANT_DASH: Record = { + 'gpu-measured': '', + 'gpu-provisioned': '8 4', + 'utility-provisioned': '3 3', + 'utility-modeled': '10 3 2 3', + all: '', + prefill: '7 3', + decode: '2 3', +}; + +/** + * The quantity a metric key plots on the comparison's common axis, or null + * when the key is not a whole-deployment average W/chip or J per output token. + */ +function comparableQuantity(metric: string): 'watts' | 'energy' | null { + const config = getMeasuredMetricConfig(metric); + if (!config) return null; + if (config.family === 'power') { + return config.scope === 'all' && config.statistic === 'average' && config.display === 'watts' + ? 'watts' + : null; + } + return config.scope === 'all' && config.denominator === 'output' && config.unit === 'joules' + ? 'energy' + : null; +} + +/** + * The base series' own identity under a comparison — the boundary or role the + * selected metric already plots — or null when the metric admits no comparison + * of that kind. + */ +export function powerCompareBase(metric: string, mode: PowerCompare): PowerVariant | null { + const config = getMeasuredMetricConfig(metric); + if (!config || mode === 'none') return null; + if (mode === 'boundaries') { + return comparableQuantity(metric) ? { kind: 'basis', id: config.basis } : null; + } + if (config.basis !== 'gpu-measured') return null; + if (config.family === 'power') { + return config.statistic === 'average' && config.display === 'watts' + ? { kind: 'role', id: config.scope } + : null; + } + // Role energy compares J per output token; a prefill J per input token axis + // has no decode counterpart. + return config.denominator === 'output' && config.unit === 'joules' && config.scope !== 'prefill' + ? { kind: 'role', id: config.scope } + : null; +} + +/** The sibling series a comparison adds to the selected metric (never the base itself). */ +export function powerCompareVariants(metric: string, mode: PowerCompare): PowerCompareSeries[] { + const base = powerCompareBase(metric, mode); + const config = getMeasuredMetricConfig(metric); + if (!base || !config) return []; + if (base.kind === 'basis') { + const quantity = comparableQuantity(metric); + if (!quantity) return []; + return POWER_BASES.filter((basis) => basis !== base.id).map((basis) => ({ + variant: { kind: 'basis', id: basis }, + field: + basis === 'gpu-measured' + ? quantity === 'watts' + ? 'measuredAvgPower' + : 'measuredJPerOutputToken' + : POWER_BASIS_FIELDS[basis][quantity], + })); + } + const fields = config.family === 'power' ? ROLE_WATT_FIELDS : ROLE_ENERGY_FIELDS; + return ROLES.filter((role) => role !== base.id).map((role) => ({ + variant: { kind: 'role', id: role }, + field: fields[role], + })); +} + +/** Whether choosing `mode` on `metric` draws anything. `none` is always available. */ +export function powerCompareAvailable(metric: string, mode: PowerCompare): boolean { + return mode === 'none' || powerCompareVariants(metric, mode).length > 0; +} + +/** + * Appends one clone per comparison series to `points` (already remapped onto + * the selected metric). A point lacking a variant's field contributes nothing + * to that series — never a 0 — so the availability rules of each boundary and + * role carry through unchanged. + */ +export function expandPowerCompareSeries( + points: readonly InferenceData[], + metric: string, + mode: PowerCompare, +): InferenceData[] { + const series = powerCompareVariants(metric, mode); + if (series.length === 0) return [...points]; + const result: InferenceData[] = [...points]; + for (const { variant, field } of series) { + for (const point of points) { + const value = point[field]; + if (!value || !Number.isFinite(value.y)) continue; + result.push({ ...point, y: value.y, roof: value.roof, powerVariant: variant }); + } + } + return result; +} + +/** Comparison mode implied by the variants present in a rendered point set. */ +export function inferPowerCompare(points: readonly InferenceData[]): PowerCompare { + for (const point of points) { + if (point.powerVariant?.kind === 'basis') return 'boundaries'; + if (point.powerVariant?.kind === 'role') return 'roles'; + } + return 'none'; +} + +/** Distinct variants in draw order: the base first, then siblings in canonical order. */ +export function powerVariantsInData( + points: readonly InferenceData[], + metric: string, +): PowerVariant[] { + const mode = inferPowerCompare(points); + if (mode === 'none') return []; + const present = new Set(points.map((point) => point.powerVariant?.id).filter(Boolean)); + const base = powerCompareBase(metric, mode); + const ordered: PowerVariant[] = + mode === 'boundaries' + ? POWER_BASES.map((id) => ({ kind: 'basis', id })) + : ROLES.map((id) => ({ kind: 'role', id })); + return ordered.filter((variant) => variant.id === base?.id || present.has(variant.id)); +} + +export function powerVariantId(variant: PowerVariant | null | undefined): string { + return variant?.id ?? ''; +} + +export function powerVariantDash(variant: PowerVariant | null | undefined): string { + return variant ? VARIANT_DASH[variant.id] : ''; +} + +export function powerVariantLabel(variant: PowerVariant, locale: Locale): string { + return variant.kind === 'basis' + ? POWER_BASIS_LABELS[variant.id][locale] + : ROLE_LABELS[variant.id][locale]; +} + +/** Short boundary names for in-chart line labels; the legend keeps the full names. */ +const BASIS_SHORT_LABELS: Record = { + 'gpu-measured': { en: 'Measured', zh: '实测' }, + 'gpu-provisioned': { en: 'TDP', zh: 'TDP' }, + 'utility-provisioned': { en: 'All-in', zh: '全站' }, + 'utility-modeled': { en: 'PUE modeled', zh: 'PUE 建模' }, +}; + +/** Suffix text a line label carries for one comparison series. */ +export function powerVariantShortLabel(variant: PowerVariant, locale: Locale): string { + return variant.kind === 'basis' + ? BASIS_SHORT_LABELS[variant.id][locale] + : ROLE_LABELS[variant.id][locale]; +} + +/** `700 W`, `1.37 kW`, `19.2 kW`: kilowatts from 1000 W, at most two decimals, zeros trimmed. */ +export function formatWatts(watts: number): string { + if (Math.abs(watts) >= 1000) { + return `${(watts / 1000).toFixed(2).replace(/\.?0+$/u, '')} kW`; + } + return `${Math.round(watts)} W`; +} + +/** + * The value a series holds at every point, or null when it varies. A + * provisioned boundary (TDP, all-in) is one number per hardware, so its line + * label can state it instead of sending the reader to the axis. Non-finite + * values are ignored; an empty series is null. + */ +export function flatSeriesValue(values: readonly number[], relTolerance = 0.005): number | null { + const finite = values.filter((value) => Number.isFinite(value)); + if (finite.length === 0) return null; + const first = finite[0]; + const tolerance = Math.abs(first) * relTolerance; + return finite.every((value) => Math.abs(value - first) <= tolerance) ? first : null; +} + +export interface PowerLineLabelOptions { + /** The selected metric's own series keeps the plain hardware label. */ + isBase: boolean; + locale: Locale; + /** Shared watts of a flat series (`flatSeriesValue`), appended after the name. */ + flatWatts?: number | null; +} + +const LINE_LABEL_SUFFIX_SEPARATOR = ' · '; + +/** Suffix appended to a comparison sibling's line label; '' for the base series. */ +export function powerLineLabelSuffix( + variant: PowerVariant | null | undefined, + opts: PowerLineLabelOptions, +): string { + if (opts.isBase || !variant) return ''; + const watts = + typeof opts.flatWatts === 'number' && Number.isFinite(opts.flatWatts) + ? ` ${formatWatts(opts.flatWatts)}` + : ''; + return `${LINE_LABEL_SUFFIX_SEPARATOR}${powerVariantShortLabel(variant, opts.locale)}${watts}`; +} + +const LINE_LABEL_SERIES_DELIMITER = '::'; + +/** + * Line-label series id: the hardware key for the base series (existing pinned + * anchors and hover hooks key on it) and `::` for a sibling. + */ +export function lineLabelSeriesId( + hw: string, + variant: PowerVariant | null | undefined, + isBase: boolean, +): string { + return isBase || !variant ? hw : `${hw}${LINE_LABEL_SERIES_DELIMITER}${variant.id}`; +} + +/** Inverse of `lineLabelSeriesId`: the hardware key behind a line-label series id. */ +export function lineLabelHardwareKey(seriesId: string): string { + const index = seriesId.indexOf(LINE_LABEL_SERIES_DELIMITER); + return index === -1 ? seriesId : seriesId.slice(0, index); +} + +/** Whether `metric` plots watts per chip, so a flat boundary's label can state its value. */ +export function metricPlotsWatts(metric: string): boolean { + const config = getMeasuredMetricConfig(metric); + return config?.family === 'power' && config.display === 'watts'; +} + +/** + * Series label for a table or CSV row: the point's variant, or the base + * series' identity when the chart is comparing and this is a base point. + */ +export function powerSeriesLabel( + point: Pick, + metric: string, + mode: PowerCompare, + locale: Locale, +): string { + const variant = point.powerVariant ?? powerCompareBase(metric, mode); + return variant ? powerVariantLabel(variant, locale) : ''; +} diff --git a/packages/app/src/components/inference/utils/power-fit.test.ts b/packages/app/src/components/inference/utils/power-fit.test.ts new file mode 100644 index 000000000..ce5571872 --- /dev/null +++ b/packages/app/src/components/inference/utils/power-fit.test.ts @@ -0,0 +1,93 @@ +import { describe, expect, it } from 'vitest'; +import type { InferenceData } from '../types'; +import { buildPowerFits, ordinaryLeastSquares } from './power-fit'; + +const metric = (y: number) => ({ y, roof: false }); +function point(overrides: Partial = {}): InferenceData { + return { + x: 100, + y: 1, + hwKey: 'h200_sglang', + date: '2026-09-18', + tp: 8, + physicalChips: 8, + precision: 'fp8', + conc: 1, + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28719990565/attempts/1', + model: 'Qwen-3.5-397B-A17B', + benchmark_type: 'single_turn', + isl: 8192, + osl: 1024, + decode_tp: 8, + output_tput_per_gpu: 50, + tpPerGpu: metric(50), + tpPerMw: metric(50), + costh: metric(1), + costr: metric(1), + costhi: metric(1), + costri: metric(1), + ...overrides, + }; +} +const ladder = (throughputs: number[], watts: (x: number) => number, base = {}) => + throughputs.map((x, index) => + point({ + id: index + 1, + conc: 2 ** index, + output_tput_per_gpu: x, + measuredAvgPower: metric(watts(x)), + ...base, + }), + ); + +describe('ordinaryLeastSquares', () => { + it('reports R² below one for scattered observations', () => { + const fit = ordinaryLeastSquares([ + { x: 1, y: 2 }, + { x: 2, y: 3 }, + { x: 3, y: 7 }, + ]); + // Least squares: slope 2.5, intercept -1, residuals (0.5, -1, 0.5) → SSres 1.5, SStot 14. + expect(fit!.slope).toBeCloseTo(2.5, 12); + expect(fit!.intercept).toBeCloseTo(-1, 12); + expect(fit!.rSquared).toBeCloseTo(1 - 1.5 / 14, 12); + }); +}); + +describe('buildPowerFits', () => { + it('fits disaggregated sources on output per allocated GPU', () => { + const gb200 = ladder([100, 300, 500], () => 0, { + hwKey: 'gb200_dynamo-sglang', + disagg: true, + num_prefill_gpu: 4, + num_decode_gpu: 4, + }).map((entry) => ({ + ...entry, + measuredAvgPower: metric(307 + 0.7 * (entry.output_tput_per_gpu! / 2)), + })); + const [fit] = buildPowerFits(gb200); + expect(fit.observations.map((observation) => observation.x)).toEqual([50, 150, 250]); + expect(fit.fit!.intercept).toBeCloseTo(307, 9); + expect(fit.fit!.slope).toBeCloseTo(0.7, 9); + }); + + it('fits a stitched append-only curve as one full ladder, not one fit per producing run', () => { + const snapshot = { curve_workflow_run_id: 35843506474, curve_date: '2026-09-18' }; + const stitched = ladder([100, 200, 300, 400, 500], (x) => 300 + 0.5 * x, snapshot).map( + (entry, index) => + index < 2 + ? { + ...entry, + actualDate: '2026-09-23', + run_url: + 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/35905882425/attempts/1', + power_audit: { producer_sha: 'appended-producer', exporter_image_sha256: 'exporter' }, + } + : entry, + ); + const fits = buildPowerFits(stitched); + expect(fits.map((fit) => [fit.source.label, fit.fit?.n])).toEqual([ + ['H200 (SGLang) · 2026-09-18', 5], + ]); + }); +}); diff --git a/packages/app/src/components/inference/utils/power-fit.ts b/packages/app/src/components/inference/utils/power-fit.ts new file mode 100644 index 000000000..80877c0ba --- /dev/null +++ b/packages/app/src/components/inference/utils/power-fit.ts @@ -0,0 +1,125 @@ +import type { InferenceData } from '../types'; +import { getGpuSpecs } from '@/lib/constants'; +import { isPositive, powerBasisNormalization } from '@/lib/power-basis'; +import { + equalServiceSourceKey, + getEqualServiceSources, + observedPoints, + type EqualServiceSource, +} from './equal-service-comparison'; + +/** + * Serving-regime power model per source (PowerX Figure 15): an ordinary + * least-squares line `W/GPU = P0 + m × output tok/s per GPU` through one + * source's measured load points. `P0` is the fit's zero-throughput intercept, + * not measured idle power; `m` is marginal GPU energy per output token in + * joules. The fit describes only the observed throughput range. + */ + +/** Distinct throughputs a line needs before its intercept and R² mean anything. */ +export const MIN_FIT_POINTS = 3; + +export interface LinearFit { + /** `P0`, W/GPU at zero output rate (extrapolated). */ + intercept: number; + /** `m`, J per output token (W per output tok/s/GPU). */ + slope: number; + /** Null when every observation draws the same power (no variance to explain). */ + rSquared: number | null; + n: number; + xMin: number; + xMax: number; +} + +export function ordinaryLeastSquares(pairs: readonly { x: number; y: number }[]): LinearFit | null { + const n = pairs.length; + if (n === 0) return null; + const xMean = pairs.reduce((sum, pair) => sum + pair.x, 0) / n; + const yMean = pairs.reduce((sum, pair) => sum + pair.y, 0) / n; + let sxx = 0; + let sxy = 0; + let syy = 0; + for (const { x, y } of pairs) { + sxx += (x - xMean) ** 2; + sxy += (x - xMean) * (y - yMean); + syy += (y - yMean) ** 2; + } + if (!(sxx > 0)) return null; + const slope = sxy / sxx; + const intercept = yMean - slope * xMean; + const residual = pairs.reduce((sum, { x, y }) => sum + (y - (intercept + slope * x)) ** 2, 0); + return { + intercept, + slope, + rSquared: syy > 0 ? 1 - residual / syy : null, + n, + xMin: Math.min(...pairs.map((pair) => pair.x)), + xMax: Math.max(...pairs.map((pair) => pair.x)), + }; +} + +/** + * Whole-deployment output tokens/s divided by every allocated GPU, so a + * disaggregated deployment's prefill GPUs share the output they enable and its + * mean W/GPU (averaged over all GPUs) sits on the same basis. + */ +export function outputRatePerAllocatedGpu(point: InferenceData): number | undefined { + if (!isPositive(point.output_tput_per_gpu)) return undefined; + const { allocatedGpus, totalOutputTokPerSec } = powerBasisNormalization({ + output_tput_per_gpu: point.output_tput_per_gpu, + disagg: Boolean(point.disagg), + benchmark_type: point.benchmark_type, + num_prefill_gpu: point.num_prefill_gpu ?? 0, + num_decode_gpu: point.num_decode_gpu ?? 0, + }); + return isPositive(allocatedGpus) && isPositive(totalOutputTokPerSec) + ? totalOutputTokPerSec / allocatedGpus + : undefined; +} + +export interface PowerFitObservation { + /** Output tok/s per allocated GPU. */ + x: number; + /** Measured mean GPU-board W/GPU. */ + y: number; + point: InferenceData; +} + +export interface PowerFit { + source: EqualServiceSource; + /** Rated board TDP per GPU from the hardware registry, to read `P0 ÷ TDP`; null when unknown. */ + tdpWatts: number | null; + observations: PowerFitObservation[]; + fit: LinearFit | null; + reason?: 'too-few-points'; +} + +/** One fit per equal-service source that has any measured observation. */ +export function buildPowerFits( + points: readonly InferenceData[], + locale: 'en' | 'zh' = 'en', +): PowerFit[] { + const rows = observedPoints(points); + return getEqualServiceSources(points, locale).flatMap((source): PowerFit[] => { + const observations = rows + .filter((point) => equalServiceSourceKey(point) === source.key) + .flatMap((point) => { + const x = outputRatePerAllocatedGpu(point); + const y = point.measuredAvgPower?.y; + return isPositive(x) && isPositive(y) ? [{ x, y, point }] : []; + }) + .toSorted((a, b) => a.x - b.x || (a.point.id ?? 0) - (b.point.id ?? 0)); + if (observations.length === 0) return []; + const tdp = getGpuSpecs(observations[0].point.hwKey).tdp; + const tdpWatts = isPositive(tdp) ? tdp : null; + const fit = + new Set(observations.map((observation) => observation.x)).size >= MIN_FIT_POINTS + ? ordinaryLeastSquares(observations) + : null; + return [ + fit + ? { source, tdpWatts, observations, fit } + : { source, tdpWatts, observations, fit: null, reason: 'too-few-points' }, + ]; + }); +} diff --git a/packages/app/src/components/inference/utils/power-metric-availability.test.ts b/packages/app/src/components/inference/utils/power-metric-availability.test.ts deleted file mode 100644 index d63c225ce..000000000 --- a/packages/app/src/components/inference/utils/power-metric-availability.test.ts +++ /dev/null @@ -1,61 +0,0 @@ -import { describe, expect, it } from 'vitest'; -import type { InferenceData } from '../types'; -import { powerMetricAvailability, powerMetricState } from './power-metric-availability'; -import { createMockInferenceData } from '../../../../cypress/support/mock-data'; - -function point(overrides: Partial = {}) { - return createMockInferenceData(overrides); -} -describe('PowerX availability', () => { - it('separates strict, validated unversioned, no verdict, invalid and missing without hiding zero', () => { - const values = { measuredAvgPower: { y: 0, roof: false } }; - const rows = [ - point({ ...values, power_valid: 1, power_metric_schema_version: 2 }), - point({ ...values, power_valid: 1 }), - point(values), - point({ power_valid: 0, ...values }), - point(), - ]; - expect( - powerMetricAvailability(rows).find((row) => row.metric === 'y_measuredAvgPower'), - ).toMatchObject({ - available: 3, - total: 5, - counts: { strict: 1, validated: 1, unverified: 1, invalid: 1, missing: 1 }, - }); - }); - it('retains real role metrics while distinguishing shared pools and ambiguous old deployment energy', () => { - expect(powerMetricState(point(), 'y_measuredPrefillAvgPower')).toBe('inapplicable'); - expect( - powerMetricState( - point({ - disagg: true, - power_valid: 1, - power_metric_schema_version: 2, - measuredPrefillJPerInputToken: { y: 2, roof: false }, - }), - 'y_measuredPrefillJPerInputToken', - ), - ).toBe('strict'); - expect( - powerMetricState(point({ disagg: true, power_valid: 1 }), 'y_measuredJPerInputToken'), - ).toBe('ambiguous'); - }); - it('counts the actual conversion and percentile fields independently', () => { - const rows = [ - point({ - power_valid: 1, - power_metric_schema_version: 2, - measuredWhPerSuccessfulQuery: { y: 1, roof: false }, - measuredPowerPercentTdp: { y: 40, roof: false }, - }), - ]; - const byMetric = Object.fromEntries( - powerMetricAvailability(rows).map((row) => [row.metric, row.available]), - ); - expect(byMetric.y_measuredWhPerSuccessfulQuery).toBe(1); - expect(byMetric.y_measuredPowerPercentTdp).toBe(1); - expect(byMetric.y_measuredP75Power).toBe(0); - expect(byMetric.y_measuredP90Power).toBe(0); - }); -}); diff --git a/packages/app/src/components/inference/utils/power-metric-availability.ts b/packages/app/src/components/inference/utils/power-metric-availability.ts deleted file mode 100644 index 6a2e7882f..000000000 --- a/packages/app/src/components/inference/utils/power-metric-availability.ts +++ /dev/null @@ -1,68 +0,0 @@ -import { MEASURED_ENERGY_METRIC_CONFIG_KEYS, isMetricKey } from '../metric-registry'; -import type { InferenceData } from '../types'; - -export const POWER_AVAILABILITY_STATES = [ - 'strict', - 'validated', - 'unverified', - 'invalid', - 'inapplicable', - 'ambiguous', - 'missing', -] as const; -export type PowerAvailabilityState = (typeof POWER_AVAILABILITY_STATES)[number]; -export type PowerAvailabilityCounts = Record; -const ROLE_METRICS = new Set([ - 'y_measuredPrefillAvgPower', - 'y_measuredDecodeAvgPower', - 'y_measuredPrefillJPerInputToken', - 'y_measuredDecodeJPerOutputToken', -]); -const WHOLE_ENERGY_METRICS = new Set([ - 'y_measuredJPerInputToken', - 'y_measuredJPerOutputToken', - 'y_measuredJPerTotalToken', - 'y_measuredJPerSuccessfulQuery', - 'y_measuredWhPerSuccessfulQuery', -]); - -/** Uses the chart's actual admitted field, so conversions and semantic gates agree. */ -export function powerMetricState(point: InferenceData, configKey: string): PowerAvailabilityState { - if (ROLE_METRICS.has(configKey) && !point.disagg) return 'inapplicable'; - if (point.power_valid === 0) return 'invalid'; - const key = configKey.replace(/^y_/u, ''); - const value = isMetricKey(key) ? point[key] : undefined; - if ( - value && - typeof value === 'object' && - 'y' in value && - typeof value.y === 'number' && - Number.isFinite(value.y) - ) { - if (point.power_valid === 1) - return point.power_metric_schema_version === 2 ? 'strict' : 'validated'; - return 'unverified'; - } - if ( - WHOLE_ENERGY_METRICS.has(configKey) && - point.disagg && - point.power_metric_schema_version !== 2 - ) - return 'ambiguous'; - return 'missing'; -} - -export function powerMetricAvailability(points: readonly InferenceData[]) { - return [...MEASURED_ENERGY_METRIC_CONFIG_KEYS].map((metric) => { - const counts = Object.fromEntries( - POWER_AVAILABILITY_STATES.map((state) => [state, 0]), - ) as PowerAvailabilityCounts; - for (const point of points) counts[powerMetricState(point, metric)]++; - return { - metric, - counts, - available: counts.strict + counts.validated + counts.unverified, - total: points.length, - }; - }); -} diff --git a/packages/app/src/components/inference/utils/powerCurves.test.ts b/packages/app/src/components/inference/utils/powerCurves.test.ts index a71f8bde2..dc0125649 100644 --- a/packages/app/src/components/inference/utils/powerCurves.test.ts +++ b/packages/app/src/components/inference/utils/powerCurves.test.ts @@ -94,8 +94,12 @@ describe('upper power envelope', () => { it('mirrors the boundary for latency and resolves tied coordinates deterministically', () => { const fast = point(1, 1, 350); const middle = point(8, 2, 700); + const plateau = point(16, 3, 700); const slow = point(32, 4, 950); - const samples = [slow, point(16, 3, 700), point(4, 2, 500), middle, { ...middle }, fast]; + // Measured watts: a tie at the running maximum is a repeat marker, so the + // plateau leaves the boundary and Optimal Only can collapse it; a repeated + // X keeps only its first vertex. + const samples = [slow, plateau, point(4, 2, 500), middle, { ...middle }, fast]; expect(upperPowerEnvelope(samples, false)).toEqual([fast, middle, slow]); expect( upperPowerEnvelope( @@ -103,6 +107,8 @@ describe('upper power envelope', () => { true, ).map((p) => p.y), ).toEqual([950, 700, 350]); + // A gauge keeps the plateau: it is part of the outer edge it draws. + expect(upperPowerEnvelope(samples, false, true)).toEqual([fast, middle, plateau, slow]); }); it('uses only finite positive coordinates and preserves singleton boundaries', () => { diff --git a/packages/app/src/components/inference/utils/powerCurves.ts b/packages/app/src/components/inference/utils/powerCurves.ts index 35146640a..7c251acbc 100644 --- a/packages/app/src/components/inference/utils/powerCurves.ts +++ b/packages/app/src/components/inference/utils/powerCurves.ts @@ -1,3 +1,4 @@ +import { isPowerBasisConfigKey } from '@/components/inference/metric-registry'; import type { InferenceData } from '@/components/inference/types'; import { isFrontierEligible, @@ -15,16 +16,39 @@ const POWER_CURVE_METRICS: ReadonlySet = new Set([ 'y_measuredDecodeAvgPower', 'y_measuredPowerPercentTdp', 'y_modeledChassisPowerPerGpu', + // Provisioned / modelled boundary gauges: same upper-envelope curve as measured watts. + 'y_gpuProvisionedWatts', + 'y_utilityProvisionedWatts', + 'y_utilityModeledWatts', ]); export function isPowerCurveMetric(metric: string): boolean { return POWER_CURVE_METRICS.has(metric); } +/** + * Power gauges whose curve is always the upper envelope and whose Optimal Only + * toggle only hides off-envelope markers: measured watts plus the provisioned / + * modelled boundary gauges that sit beside them. A Pareto corner would collapse + * a flat TDP series to one marker. The modelled chassis axis keeps its legacy + * Pareto behaviour. + */ export function isMeasuredPowerCurveMetric(metric: string): boolean { return isPowerCurveMetric(metric) && metric !== 'y_modeledChassisPowerPerGpu'; } +/** + * Whether a drawn series is a provisioned or modelled gauge rather than + * telemetry: a boundary axis, or a boundary comparison clone of one. Gauges + * keep envelope ties (see `upperPowerEnvelope`); measured series, the measured + * boundary clone and role clones stay on the strict envelope. + */ +export function isPowerGaugeSeries(metric: string, sample: InferenceData | undefined): boolean { + const variant = sample?.powerVariant; + if (variant) return variant.kind === 'basis' && variant.id !== 'gpu-measured'; + return isPowerBasisConfigKey(metric); +} + /** No declared direction means there is no Pareto frontier to draw or filter by. */ export function chartFrontier( points: InferenceData[], @@ -45,14 +69,23 @@ export function chartFrontier( export function upperPowerEnvelope( points: readonly InferenceData[], maximizeX: boolean, + keepTies = false, ): InferenceData[] { const sorted = points .filter((point) => isFrontierEligible(point) && Number.isFinite(point.y) && point.y > 0) .sort((a, b) => (maximizeX ? b.x - a.x : a.x - b.x) || b.y - a.y); + // Measured telemetry keeps the strict envelope: a tie at the running maximum + // is a repeat marker that Optimal Only collapses. A provisioned or modelled + // gauge (`keepTies`) is flat by construction, so its ties stay on the + // boundary and the series draws across its tested x-range instead of one + // marker. Repeated X keeps only its first vertex so the smoothing never + // backtracks. let maxY = -Infinity; + let lastX = Number.NaN; const envelope = sorted.filter((point) => { - if (point.y <= maxY) return false; + if (point.y < maxY || (point.y === maxY && (!keepTies || point.x === lastX))) return false; maxY = point.y; + lastX = point.x; return true; }); return maximizeX ? envelope.toReversed() : envelope; diff --git a/packages/app/src/components/inference/utils/powerTimeline.test.ts b/packages/app/src/components/inference/utils/powerTimeline.test.ts new file mode 100644 index 000000000..083ed100d --- /dev/null +++ b/packages/app/src/components/inference/utils/powerTimeline.test.ts @@ -0,0 +1,88 @@ +import { describe, expect, it } from 'vitest'; + +import type { InferenceData } from '@/components/inference/types'; + +import { + prioritizeRuns, + summarizeTraceWindow, + tracePools, + type PowerTimelineTrace, +} from './powerTimeline'; + +const RUN_URL = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34716669498'; +const NAME_B = + 'qwen3.5_8k1k_fp8_dynamo-sglang_prefill-tp4-pp1-dcp1-pcp1-ep1-dpfalse-nw1_decode-tp4-5b35252b29d80b104d11_sa-bench_isl_8192_osl_1024_conc1_gpus_8_ctx_4_gen_4'; + +function point(overrides: Partial): InferenceData { + return { + x: 1, + y: 1, + hwKey: 'b200_sglang', + tp: 4, + conc: 16, + precision: 'fp8', + date: '2026-09-12', + run_url: RUN_URL, + ...overrides, + } as InferenceData; +} + +describe('summarizeTraceWindow', () => { + // Two prefill and two decode GPUs; the window holds buckets 1–3 (t = 11..13 s). + const pooled: PowerTimelineTrace = { + key: `34716669498:${NAME_B}`, + point: point({ disagg: true, conc: 4 }), + runId: '34716669498', + series: { + artifact: `power_audit_${NAME_B}`, + startMs: 1_000_000, + bucketSeconds: 1, + gpus: [0, 1, 2, 3], + t: [10, 11, 12, 13, 14], + power: [ + [900, 250, 260, 270, 900], + [900, 250, null, 270, 900], + [900, 400, 410, 420, 900], + [900, 400, 410, 420, 900], + ], + devices: [ + { id: 'a/0', role: 'prefill' }, + { id: 'a/1', role: 'prefill' }, + { id: 'b/0', role: 'decode' }, + { id: 'b/1', role: 'decode' }, + ], + }, + windowStartMs: 1_011_000, + windowEndMs: 1_013_500, + }; + + it('reports window length and each pool’s peak summed bucket inside the window', () => { + const summary = summarizeTraceWindow(pooled, tracePools(pooled.series)); + expect(summary.windowSeconds).toBe(2.5); + expect(summary.bucketSeconds).toBe(1); + // The 900 W buckets sit outside the window; bucket 2 lacks one prefill GPU, + // so it is a gap, not a 260 W dip. + expect(summary.pools).toEqual([ + { role: 'prefill', gpuCount: 2, peakWatts: 540 }, + { role: 'decode', gpuCount: 2, peakWatts: 840 }, + ]); + }); +}); + +describe('prioritizeRuns', () => { + const requests = ['1', '2', '3', '4', '5'].map((runId) => ({ + runId, + prefix: '', + sources: [], + })); + + it('moves overlay runs ahead of official runs and keeps both orders', () => { + expect(prioritizeRuns(requests, new Set(['5', '3'])).map((request) => request.runId)).toEqual([ + '3', + '5', + '1', + '2', + '4', + ]); + }); +}); diff --git a/packages/app/src/components/inference/utils/powerTimeline.ts b/packages/app/src/components/inference/utils/powerTimeline.ts new file mode 100644 index 000000000..0a7a22254 --- /dev/null +++ b/packages/app/src/components/inference/utils/powerTimeline.ts @@ -0,0 +1,482 @@ +/** + * Joins chart points to the per-second GPU telemetry behind their measured + * average power (the PowerX "Timeline" display). + * + * Every validated row records its window in `power_audit`, whose `source` is + * `power_validation_.json`. Two collectors publish the telemetry: + * - single-node runners upload one `gpu_metrics_` CSV artifact per + * config, so the source names the artifact exactly; + * - Slurm / Dynamo runners upload one `power_audit_` bundle per + * sweep whose `LOGS/power/samples.csv` covers every concurrency; the API cuts + * it into one series per `power_validation_*.json` it contains and labels + * each with that `source`, so the same file name joins it to the row. + * Nothing is matched by hardware or concurrency; points whose telemetry is + * missing (expired artifact, another collector) are reported, not guessed. + */ +import { + bucketTimeMs, + sumPowerAt, + type GpuPowerRole, + type GpuPowerSeries, + type GpuPowerSeriesResponse, +} from '@/components/gpu-power/power-series'; +import type { InferenceData } from '@/components/inference/types'; + +export const POWER_TIMELINE_METRIC_KEY = 'y_measuredPowerTimeline'; +const ARTIFACT_PREFIX = 'gpu_metrics_'; +const SOURCE_PATTERN = /^(?:.*\/)?power_validation_(?.+)\.json$/u; + +interface AuditedPoint { + power_audit?: { source?: string } | null; +} + +/** The `` of a point's `power_validation_.json` audit source. */ +export function telemetryNameForPoint(point: AuditedPoint): string | null { + const source = point.power_audit?.source; + if (!source) return null; + return SOURCE_PATTERN.exec(source)?.groups?.name ?? null; +} + +/** `gpu_metrics_` for a point, from its power-audit source. */ +export function telemetryArtifactForPoint(point: AuditedPoint): string | null { + const name = telemetryNameForPoint(point); + return name ? `${ARTIFACT_PREFIX}${name}` : null; +} + +/** + * Stable identity of a point's trace: run id plus audit name. Unique within a + * chart because the audit name carries the config and the concurrency. + */ +export function traceKeyForPoint(point: AuditedPoint & { run_url?: string }): string | null { + const runId = runIdFromUrl(point.run_url); + const name = telemetryNameForPoint(point); + return runId && name ? `${runId}:${name}` : null; +} + +/** Workflow run id from a GitHub Actions run URL. */ +export function runIdFromUrl(url: string | null | undefined): string | null { + return url?.match(/\/runs\/(?\d+)/u)?.groups?.runId ?? null; +} + +/** Attempt number from a `/runs//attempts/` URL; null when the URL names none. */ +export function runAttemptFromUrl(url: string | null | undefined): number | null { + const attempt = url?.match(/\/runs\/\d+\/attempts\/(?\d+)/u)?.groups?.attempt; + return attempt ? Number(attempt) : null; +} + +export function longestCommonPrefix(values: readonly string[]): string { + if (values.length === 0) return ''; + let prefix = values[0]; + for (const value of values) { + let end = 0; + while (end < prefix.length && end < value.length && prefix[end] === value[end]) end++; + prefix = prefix.slice(0, end); + if (prefix === '') break; + } + return prefix; +} + +export interface PowerTimelineRequest { + runId: string; + /** RESULT_FILENAME prefix shared by every wanted artifact of the run. */ + prefix: string; + /** Sorted, unique validation basenames expected by the displayed points. */ + sources: string[]; +} + +/** + * One request per workflow run, narrowed to the common RESULT_FILENAME prefix + * of the points' artifacts so a sweep of other models is not downloaded. + * Points without a run URL or an audit source plan nothing. + */ +export function planPowerTimelineRequests( + points: readonly InferenceData[], +): PowerTimelineRequest[] { + const byRun = new Map>(); + for (const point of points) { + const runId = runIdFromUrl(point.run_url); + const artifact = telemetryArtifactForPoint(point); + if (!runId || !artifact) continue; + if (!byRun.has(runId)) byRun.set(runId, new Set()); + byRun.get(runId)!.add(artifact); + } + return [...byRun.entries()] + .toSorted(([a], [b]) => a.localeCompare(b)) + .map(([runId, artifacts]) => { + const names = [...artifacts].toSorted(); + return { + runId, + prefix: longestCommonPrefix(names.map((name) => name.slice(ARTIFACT_PREFIX.length))), + sources: names.map((name) => `power_validation_${name.slice(ARTIFACT_PREFIX.length)}.json`), + }; + }); +} + +/** Run id half of a trace key (`traceKeyForPoint`). */ +export function traceKeyRunId(key: string | null | undefined): string | null { + const runId = key?.split(':')[0]; + return runId || null; +} + +/** + * Moves the request for `runId` to the front so it survives the per-chart run + * cap; the order of the other requests is kept. Returns the same array when + * nothing needs moving. + */ +export function prioritizeRun( + requests: PowerTimelineRequest[], + runId: string | null, +): PowerTimelineRequest[] { + const index = runId ? requests.findIndex((request) => request.runId === runId) : -1; + if (index <= 0) return requests; + return [requests[index], ...requests.slice(0, index), ...requests.slice(index + 1)]; +} + +/** + * Moves every request whose run is in `runIds` ahead of the others, keeping + * the relative order inside both groups. Unofficial-run overlays are loaded + * on purpose, so their telemetry must survive the per-chart run cap before + * official rows compete for the remaining slots. Returns the same array when + * nothing needs moving. + */ +export function prioritizeRuns( + requests: PowerTimelineRequest[], + runIds: ReadonlySet, +): PowerTimelineRequest[] { + if (runIds.size === 0) return requests; + const first = requests.filter((request) => runIds.has(request.runId)); + if (first.length === 0 || first.length === requests.length) return requests; + const rest = requests.filter((request) => !runIds.has(request.runId)); + const moved = first.some((request, index) => requests[index] !== request); + return moved ? [...first, ...rest] : requests; +} + +export interface PowerTimelineTrace { + /** `traceKeyForPoint(point)`. */ + key: string; + point: InferenceData; + runId: string; + series: GpuPowerSeries; + /** Validated measurement window (UTC ms) from `power_audit`, when recorded. */ + windowStartMs: number | null; + windowEndMs: number | null; +} + +export type PowerTimelineAxis = 'wall' | 'elapsed' | 'serving'; +export type PowerTimelineLines = 'mean' | 'gpu' | 'pool'; + +/** URL input is optional: an omitted axis retains the existing per-run default. */ +export function parsePowerTimelineParams(params: { + i_ptaxis?: string; + i_ptlines?: string; + i_ptwindow?: string; + i_ptfocus?: string; + i_ptutility?: string; + i_ptconc?: string; +}) { + return { + axis: (['wall', 'elapsed', 'serving'].includes(params.i_ptaxis ?? '') + ? params.i_ptaxis + : null) as PowerTimelineAxis | null, + lines: (params.i_ptlines === 'gpu' || params.i_ptlines === 'pool' + ? params.i_ptlines + : 'mean') as PowerTimelineLines, + windowOnly: params.i_ptwindow === 'window', + focus: + params.i_ptfocus && /^[1-9]\d*:[\w.-]+$/u.test(params.i_ptfocus) ? params.i_ptfocus : null, + utility: params.i_ptutility === '1', + concurrency: /^[1-9]\d{0,5}$/u.test(params.i_ptconc ?? '') ? Number(params.i_ptconc) : null, + }; +} + +export function hasPowerTimelineWindow(trace: PowerTimelineTrace): boolean { + return ( + trace.windowStartMs !== null && + trace.windowEndMs !== null && + Number.isFinite(trace.windowStartMs) && + Number.isFinite(trace.windowEndMs) && + trace.windowEndMs > trace.windowStartMs + ); +} + +/** Filter and position retained buckets only; never infer a missing serving origin. */ +export function powerTimelineSampleX( + trace: PowerTimelineTrace, + column: number, + axis: PowerTimelineAxis, + windowOnly: boolean, +): number | null { + const timeMs = bucketTimeMs(trace.series, column); + if ((axis === 'serving' || windowOnly) && !hasPowerTimelineWindow(trace)) return null; + if (windowOnly && windowPhase(trace, timeMs) !== 'window') return null; + if (axis === 'wall') return timeMs; + const origin = axis === 'serving' ? trace.windowStartMs! : bucketTimeMs(trace.series, 0); + return (timeMs - origin) / 1000; +} + +/** + * Why a validated point has no trace: + * - `no-source`: the row predates `power_audit.source` (older schema), so no + * artifact can be named; + * - `no-run`: no workflow run URL to look in; + * - `run-not-fetched`: its run is not among the loaded responses (over the + * per-chart run limit, still loading, or the request failed); + * - `not-in-run`: the run was loaded but holds neither a matching + * `gpu_metrics_*` artifact nor a power-audit bundle with the point's + * validation file (expired, over the download cap, or another collector). + */ +export type MissingTraceReason = 'no-source' | 'no-run' | 'run-not-fetched' | 'not-in-run'; + +export interface MissingTrace { + point: InferenceData; + reason: MissingTraceReason; +} + +export interface PowerTimelineJoin { + traces: PowerTimelineTrace[]; + /** Points with a validated average but no telemetry trace, with the reason. */ + missing: MissingTrace[]; +} + +/** Attaches fetched series to points; order follows `points`. */ +export function joinPowerTimeline( + points: readonly InferenceData[], + responses: ReadonlyMap, +): PowerTimelineJoin { + const traces: PowerTimelineTrace[] = []; + const missing: MissingTrace[] = []; + for (const point of points) { + const runId = runIdFromUrl(point.run_url); + const name = telemetryNameForPoint(point); + if (!name) { + missing.push({ point, reason: 'no-source' }); + continue; + } + if (!runId) { + missing.push({ point, reason: 'no-run' }); + continue; + } + const response = responses.get(runId); + if (!response) { + missing.push({ point, reason: 'run-not-fetched' }); + continue; + } + // A bundle-cut series names the point's validation file; a per-config + // CSV artifact names the config itself. + const source = `power_validation_${name}.json`; + const artifact = `${ARTIFACT_PREFIX}${name}`; + const series = + response.series.find((entry) => entry.source === source) ?? + response.series.find((entry) => entry.source === undefined && entry.artifact === artifact); + if (!series) { + missing.push({ point, reason: 'not-in-run' }); + continue; + } + const audit = point.power_audit; + traces.push({ + key: `${runId}:${name}`, + point, + runId, + series, + windowStartMs: unixSecondsToMs(audit?.window_start_unix), + windowEndMs: unixSecondsToMs(audit?.window_end_unix), + }); + } + return { traces, missing }; +} + +function unixSecondsToMs(seconds: number | undefined): number | null { + return typeof seconds === 'number' && Number.isFinite(seconds) ? seconds * 1000 : null; +} + +/** Sample position relative to the validated window. */ +export type WindowPhase = 'before' | 'window' | 'after' | 'unknown'; + +export function windowPhase(trace: PowerTimelineTrace, timeMs: number): WindowPhase { + if (!hasPowerTimelineWindow(trace)) return 'unknown'; + if (timeMs < trace.windowStartMs!) return 'before'; + if (timeMs > trace.windowEndMs!) return 'after'; + return 'window'; +} + +/** Short config label: `TP8 · c64`, plus `PD` for disaggregated points. */ +export function traceConfigLabel(point: InferenceData): string { + const parts: string[] = []; + if (point.disagg) parts.push('PD'); + if (typeof point.tp === 'number' && point.tp > 0) parts.push(`TP${point.tp}`); + parts.push(`c${point.conc}`); + return parts.join(' · '); +} + +export type PowerPoolRole = GpuPowerRole | 'all'; + +export interface PowerPool { + role: PowerPoolRole; + /** Row indices into `series.power`, in series order. */ + rows: number[]; +} + +const POOL_ORDER: readonly PowerPoolRole[] = ['all', 'prefill', 'decode']; + +/** + * The GPU pools of a series by worker role, in `prefill`, `decode` order — + * only the roles that have at least one device. A series whose collector + * assigns no roles yields no pools; callers fall back to `allGpuPool`. + */ +export function tracePools(series: Pick): PowerPool[] { + const rows = new Map(); + series.devices?.forEach((device, row) => { + if (!device.role) return; + if (!rows.has(device.role)) rows.set(device.role, []); + rows.get(device.role)!.push(row); + }); + return POOL_ORDER.filter((role) => rows.has(role)).map((role) => ({ + role, + rows: rows.get(role)!, + })); +} + +export interface PoolSizeGroup { + size: number; + roles: PowerPoolRole[]; +} + +/** + * Pools that hold the same number of GPUs share one rated ceiling, so their + * reference draws once, labelled `prefill / decode ×16`, instead of two labels + * printed over each other. Sizes ascending, roles in pool order, deduplicated. + */ +export function groupPoolsBySize( + pools: readonly Pick[], +): PoolSizeGroup[] { + const bySize = new Map>(); + for (const pool of pools) { + if (!bySize.has(pool.rows.length)) bySize.set(pool.rows.length, new Set()); + bySize.get(pool.rows.length)!.add(pool.role); + } + return [...bySize.entries()] + .toSorted(([a], [b]) => a - b) + .map(([size, roles]) => ({ + size, + roles: POOL_ORDER.filter((role) => roles.has(role)), + })); +} + +/** + * Label row for each reference line: lines at the same watts (different + * hardware with an equal pool ceiling) stack their labels upward, slot 0 on + * the line and slot n `n` rows above, instead of overprinting. Input order. + */ +export function referenceLabelSlots(lines: readonly { watts: number }[]): number[] { + const used = new Map(); + return lines.map((line) => { + const slot = used.get(line.watts) ?? 0; + used.set(line.watts, slot + 1); + return slot; + }); +} + +/** + * Rows for end-of-line trace labels in pixel space: a label whose box overlaps + * one already placed (both ranges intersect) moves one `rowHeight` below it, + * top anchors first, so pools ending at the same watts do not overprint. + * Returns each label's y in input order. + */ +export function stackTraceLabels( + labels: readonly { left: number; right: number; y: number }[], + rowHeight: number, +): number[] { + const ys = labels.map((label) => label.y); + const placed: { left: number; right: number; y: number }[] = []; + const order = labels.map((_, index) => index).toSorted((a, b) => labels[a].y - labels[b].y); + for (const index of order) { + const { left, right } = labels[index]; + let y = ys[index]; + // Each move strictly lowers the label, so this ends within `placed.length` passes. + let moved = true; + while (moved) { + moved = false; + for (const other of placed) { + if (left < other.right && other.left < right && Math.abs(y - other.y) < rowHeight) { + y = other.y + rowHeight; + moved = true; + } + } + } + ys[index] = y; + placed.push({ left, right, y }); + } + return ys; +} + +/** Every GPU of the series as one pool. */ +export function allGpuPool(series: Pick): PowerPool { + return { role: 'all', rows: series.power.map((_, row) => row) }; +} + +export interface TracePoolSummary { + role: PowerPoolRole; + gpuCount: number; + /** Highest drawn one-second pool sum inside the window; null when none was complete. */ + peakWatts: number | null; +} + +export interface TraceWindowSummary { + /** Recorded validated serving-window length, or null when bounds are missing. */ + windowSeconds: number | null; + bucketSeconds: number; + pools: TracePoolSummary[]; +} + +/** + * What a Timeline reader needs next to the curves: the validated window's + * length and, per pool, its GPU count and the peak of the drawn pool line + * inside the window. A bucket missing a pool device is a gap in the line, so + * it is skipped. Averages are not recomputed: the benchmark row's validated + * figures stay the only means. + */ +export function summarizeTraceWindow( + trace: PowerTimelineTrace, + pools: readonly PowerPool[], +): TraceWindowSummary { + const hasWindow = hasPowerTimelineWindow(trace); + const columns = hasWindow + ? trace.series.t + .map((_, column) => column) + .filter((column) => windowPhase(trace, bucketTimeMs(trace.series, column)) === 'window') + : []; + return { + windowSeconds: hasWindow ? (trace.windowEndMs! - trace.windowStartMs!) / 1000 : null, + bucketSeconds: trace.series.bucketSeconds, + pools: pools.map((pool) => { + const sums = columns + .map((column) => sumPowerAt(trace.series, pool.rows, column)) + .filter((value): value is number => value !== null); + return { + role: pool.role, + gpuCount: pool.rows.length, + peakWatts: sums.length > 0 ? Math.max(...sums) : null, + }; + }), + }; +} + +// ── Deep link from a pinned scatter tooltip ───────────────────────────────── +// +// "View power trace" on a pinned tooltip switches the metric to the Timeline +// display; the timeline mounts afterwards and reads the requested trace here +// so it can emphasise that config. The one-shot gesture takes precedence over +// restored URL focus; the mounted timeline then persists it with its view state. + +let pendingFocus: string | null = null; + +export function requestPowerTraceFocus(key: string): void { + pendingFocus = key; +} + +/** The pending focus request, cleared on read. */ +export function consumePowerTraceFocus(): string | null { + const key = pendingFocus; + pendingFocus = null; + return key; +} diff --git a/packages/app/src/components/inference/utils/quick-filter-summary.ts b/packages/app/src/components/inference/utils/quick-filter-summary.ts index ba6b3bbac..e843dc9af 100644 --- a/packages/app/src/components/inference/utils/quick-filter-summary.ts +++ b/packages/app/src/components/inference/utils/quick-filter-summary.ts @@ -1,5 +1,6 @@ import type { QuickFilters } from '../types'; import { FRAMEWORK_FAMILIES } from './quickFilters'; +import { topologyLabel } from './topology-filter'; const LABELS = { en: { @@ -8,6 +9,7 @@ const LABELS = { deployment: 'Deployment', spec: 'Spec Decoding', power: 'Measured Power', + topologies: 'Topology', 'single-node': 'Single-node', 'multi-node': 'Multi-node', disagg: 'Disaggregated', @@ -22,6 +24,7 @@ const LABELS = { deployment: '部署模式', spec: '投机解码', power: '实测功耗', + topologies: '拓扑', 'single-node': '单节点', 'multi-node': '多节点聚合', disagg: '分离式', @@ -37,14 +40,16 @@ export function quickFilterSummary(filters: QuickFilters, locale: 'en' | 'zh', a const labels = LABELS[locale]; return (Object.keys(filters) as (keyof QuickFilters)[]).flatMap((category) => { if (agentic && category === 'spec') return []; - return filters[category].map((value) => ({ + return (filters[category] ?? []).map((value) => ({ category, value, categoryLabel: labels[category], label: - category === 'frameworks' - ? (FRAMEWORK_FAMILIES.find((family) => family.key === value)?.label ?? value) - : (labels[value as keyof typeof labels] ?? value), + category === 'topologies' + ? topologyLabel(value, locale) + : category === 'frameworks' + ? (FRAMEWORK_FAMILIES.find((family) => family.key === value)?.label ?? value) + : (labels[value as keyof typeof labels] ?? value), })); }); } diff --git a/packages/app/src/components/inference/utils/quickFilters.test.ts b/packages/app/src/components/inference/utils/quickFilters.test.ts index 82490d6f0..15c4ae01a 100644 --- a/packages/app/src/components/inference/utils/quickFilters.test.ts +++ b/packages/app/src/components/inference/utils/quickFilters.test.ts @@ -72,7 +72,7 @@ describe('computeAvailableQuickFilters', () => { power_tier: 'certified', }), ]; - expect(computeAvailableQuickFilters(points)).toEqual({ + expect(computeAvailableQuickFilters(points)).toMatchObject({ vendors: ['NVIDIA', 'AMD'], frameworks: ['vllm', 'trt', 'atom'], deployment: ['single-node', 'multi-node', 'disagg'], @@ -85,7 +85,7 @@ describe('computeAvailableQuickFilters', () => { const points = [ point({ hwKey: 'h100_vllm', framework: 'vllm', disagg: false, spec_decoding: 'none' }), ]; - expect(computeAvailableQuickFilters(points)).toEqual({ + expect(computeAvailableQuickFilters(points)).toMatchObject({ vendors: ['NVIDIA'], frameworks: ['vllm'], deployment: ['single-node'], @@ -112,6 +112,7 @@ describe('computeAvailableQuickFilters', () => { deployment: [], spec: [], power: [], + topologies: [], }); }); }); diff --git a/packages/app/src/components/inference/utils/quickFilters.ts b/packages/app/src/components/inference/utils/quickFilters.ts index 479b06d9b..50513bb82 100644 --- a/packages/app/src/components/inference/utils/quickFilters.ts +++ b/packages/app/src/components/inference/utils/quickFilters.ts @@ -9,6 +9,7 @@ import type { } from '@/components/inference/types'; import { frameworkFamily } from '@/lib/framework-family'; import type { PowerTier } from '@/lib/power-tier'; +import { pointTopologyKey } from './topology-filter'; export type { AvailableQuickFilters, DeploymentMode, PowerTier, QuickFilters, SpecMode }; @@ -64,6 +65,7 @@ export function computeAvailableQuickFilters( ): AvailableQuickFilters { const vendors = new Set(); const frameworks = new Set(); + const topologies = new Set(); let hasSingleNode = false; let hasMultiNode = false; let hasDisagg = false; @@ -72,6 +74,7 @@ export function computeAvailableQuickFilters( let hasCertified = false; let hasLegacy = false; for (const p of points) { + topologies.add(pointTopologyKey(p)); const vendor = pointVendor(String(p.hwKey)); if (vendor) vendors.add(vendor); const fam = frameworkFamily(p.framework); @@ -98,6 +101,7 @@ export function computeAvailableQuickFilters( deployment, spec, power: POWER_TIER_ORDER.filter((tier) => (tier === 'certified' ? hasCertified : hasLegacy)), + topologies: [...topologies].toSorted(), }; } @@ -108,7 +112,8 @@ export function quickFiltersActive(f: QuickFilters): boolean { f.frameworks.length > 0 || f.deployment.length > 0 || f.spec.length > 0 || - f.power.length > 0 + f.power.length > 0 || + (f.topologies?.length ?? 0) > 0 ); } @@ -159,6 +164,7 @@ export function parsePowerTiers(values: readonly string[]): PowerTier[] { /** Whether a single data point satisfies every active quick-filter category. */ export function matchesQuickFilters(point: InferenceData, f: QuickFilters): boolean { + if (f.topologies?.length && !f.topologies.includes(pointTopologyKey(point))) return false; if (f.vendors.length > 0) { const vendor = pointVendor(String(point.hwKey)); if (!vendor || !f.vendors.includes(vendor)) return false; diff --git a/packages/app/src/components/inference/utils/resolveXAxisField.ts b/packages/app/src/components/inference/utils/resolveXAxisField.ts index 282b778a7..06b48aa10 100644 --- a/packages/app/src/components/inference/utils/resolveXAxisField.ts +++ b/packages/app/src/components/inference/utils/resolveXAxisField.ts @@ -17,8 +17,11 @@ import { withPercentile } from '@/lib/benchmark-transform'; import type { AggDataEntry, ChartDefinition } from '../types'; import type { XAxisMode } from '../hooks/useChartData'; +export type FixedSequenceStatistic = 'mean' | 'median'; + /** Which rung of the branch ladder chose the x field (drives label choice). */ export type XAxisBranch = + | 'concurrency' | 'natural' | 'user-input-override' | 'config-input-override' @@ -34,15 +37,31 @@ export interface ResolvedXAxis { branch: XAxisBranch; } +/** + * A service-metric field at the selected statistic: the percentile for agentic + * rows, mean or median for fixed sequences. Fixed-sequence mean interactivity + * is reciprocal mean TPOT, never the raw arithmetic-mean interactivity field. + */ +export function resolveServiceField( + field: string, + opts: { isAgentic: boolean; percentile: string; fixedSequenceStatistic?: FixedSequenceStatistic }, +): keyof AggDataEntry { + const { isAgentic, percentile, fixedSequenceStatistic = 'median' } = opts; + const resolved = withPercentile(field, isAgentic ? percentile : fixedSequenceStatistic); + return ( + !isAgentic && resolved === 'mean_intvty' ? 'mean_tpot_intvty' : resolved + ) as keyof AggDataEntry; +} + /** * Resolve the x-axis data field for a chart definition + metric selection. * * Rules, in order: * - The global x-axis mode takes precedence over legacy per-input-metric - * overrides. TTFT uses median for fixed-sequence runs. + * overrides. Fixed-sequence service axes use the selected mean/median statistic. * - Natural x = the chart's latency metric at the selected percentile for - * agentic, forced to median for fixed-seq, where percentile-specific - * columns are not guaranteed to exist. + * agentic, mean or median for fixed-sequence. Mean interactivity uses + * reciprocal mean TPOT, never the raw arithmetic-mean interactivity field. * - Without a global mode, input metrics on the interactivity chart override x to a TTFT column: * the user-picked metric for fixed-seq (the manual dropdown is hidden in * agentic mode), else the config default. @@ -58,13 +77,16 @@ export function resolveXAxisField( chartDef: ChartDefinition, selectedYAxisMetric: string, effectiveXMetric: string | null, - opts: { isAgentic: boolean; percentile: string; xAxisMode?: XAxisMode }, + opts: { + isAgentic: boolean; + percentile: string; + xAxisMode?: XAxisMode; + fixedSequenceStatistic?: FixedSequenceStatistic; + }, ): ResolvedXAxis { const { isAgentic, percentile, xAxisMode } = opts; - const naturalX = withPercentile( - chartDef.x, - isAgentic ? percentile : 'median', - ) as keyof AggDataEntry; + const serviceField = (field: string) => resolveServiceField(field, opts); + const naturalX = serviceField(chartDef.x); const metricTitle = (chartDef[`${selectedYAxisMetric}_title` as keyof ChartDefinition] as string) || ''; @@ -77,11 +99,11 @@ export function resolveXAxisField( let xAxisField: keyof AggDataEntry = naturalX; let branch: XAxisBranch = 'natural'; if (xAxisMode !== undefined) { - if (xAxisMode === 'ttft' && chartDef.chartType === 'e2e') { - xAxisField = withPercentile( - 'median_ttft', - isAgentic ? percentile : 'median', - ) as keyof AggDataEntry; + if (xAxisMode === 'concurrency') { + xAxisField = 'conc'; + branch = 'concurrency'; + } else if (xAxisMode === 'ttft' && chartDef.chartType === 'e2e') { + xAxisField = serviceField('median_ttft'); branch = 'e2e-ttft-override'; } } else if ( @@ -101,7 +123,7 @@ export function resolveXAxisField( branch = 'e2e-ttft-override'; } - if (isAgentic) { + if (isAgentic && xAxisMode !== 'concurrency') { xAxisField = withPercentile(xAxisField, percentile) as keyof AggDataEntry; } diff --git a/packages/app/src/components/inference/utils/role-energy.test.ts b/packages/app/src/components/inference/utils/role-energy.test.ts new file mode 100644 index 000000000..102703cc0 --- /dev/null +++ b/packages/app/src/components/inference/utils/role-energy.test.ts @@ -0,0 +1,26 @@ +import { describe, expect, it } from 'vitest'; + +import { reconstructedRoleEnergy } from './role-energy'; + +// A disaggregated 8K/1K run: the deployment's energy over the window divided +// by 8× more input than output tokens, so J/out ÷ J/in = 7.9 = the served ratio. +const entry = { + disagg: true, + power_valid: 1, + power_metric_schema_version: 2, + joules_per_input_token: 1, + joules_per_output_token: 7.9, + prefill_joules_per_input_token: 0.25, + decode_joules_per_output_token: 6, +}; + +describe('reconstructedRoleEnergy', () => { + it('expresses prefill energy per output token with the served token ratio and sums the roles', () => { + expect(reconstructedRoleEnergy(entry)).toEqual({ + prefill: 1.975, + decode: 6, + total: 7.975, + prefillShare: (100 * 1.975) / 7.975, + }); + }); +}); diff --git a/packages/app/src/components/inference/utils/role-energy.ts b/packages/app/src/components/inference/utils/role-energy.ts new file mode 100644 index 000000000..53d719dd0 --- /dev/null +++ b/packages/app/src/components/inference/utils/role-energy.ts @@ -0,0 +1,63 @@ +import { isPositive } from '@/lib/power-basis'; + +/** + * Reconstructs how a disaggregated deployment's request energy splits between + * its prefill and decode pools (PowerX Figure 7). + * + * Schema-2 aggregate energy has one numerator: the deployment's energy over the + * validated window is divided by input tokens for `joules_per_input_token` and + * by output tokens for `joules_per_output_token`. Their ratio is therefore the + * input:output token ratio the benchmark actually served. Multiplying the + * prefill pool's J per input token by that ratio expresses the prefill energy + * per output token, on the same axis as the decode pool's J per output token. + * The two add up to the whole request's J per output token, which equals the + * deployment figure whenever the role energies partition the deployment energy. + * + * Nothing is estimated: every input is a same-window telemetry figure, and the + * result is `undefined` whenever one is missing, not validated, or not from a + * disaggregated deployment. + */ +export interface RoleEnergyInput { + disagg?: boolean; + power_valid?: number; + power_metric_schema_version?: number; + joules_per_input_token?: number; + joules_per_output_token?: number; + prefill_joules_per_input_token?: number; + decode_joules_per_output_token?: number; +} + +export interface ReconstructedRoleEnergy { + /** Prefill pool energy per output token (J). */ + prefill: number; + /** Decode pool energy per output token (J). */ + decode: number; + /** Prefill + decode (J per output token). */ + total: number; + /** Prefill share of the reconstructed total, in percent. */ + prefillShare: number; +} + +export function reconstructedRoleEnergy( + entry: RoleEnergyInput, +): ReconstructedRoleEnergy | undefined { + if (!entry.disagg || entry.power_valid !== 1 || entry.power_metric_schema_version !== 2) { + return undefined; + } + const input = entry.joules_per_input_token; + const output = entry.joules_per_output_token; + const prefill = entry.prefill_joules_per_input_token; + const decode = entry.decode_joules_per_output_token; + if (!isPositive(input) || !isPositive(output) || !isPositive(prefill) || !isPositive(decode)) { + return undefined; + } + const prefillPerOutputToken = prefill * (output / input); + const total = prefillPerOutputToken + decode; + if (!isPositive(prefillPerOutputToken) || !isPositive(total)) return undefined; + return { + prefill: prefillPerOutputToken, + decode, + total, + prefillShare: (100 * prefillPerOutputToken) / total, + }; +} diff --git a/packages/app/src/components/inference/utils/tooltip-utils.power-trace.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.power-trace.test.ts new file mode 100644 index 000000000..45157d317 --- /dev/null +++ b/packages/app/src/components/inference/utils/tooltip-utils.power-trace.test.ts @@ -0,0 +1,94 @@ +import { afterEach, describe, expect, it, vi } from 'vitest'; + +import type { HardwareConfig, InferenceData, OverlayData } from '@/components/inference/types'; +import { + generateOverlayTooltipContent, + type OverlayTooltipConfig, + type TooltipConfig, +} from '@/components/inference/utils/tooltipUtils'; + +// "View power trace" on pinned tooltips: the deep link from a measured-power +// scatter point to its per-second telemetry on the Timeline display. + +const RUN_URL = 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34716669498'; +const AUDIT_NAME = + 'dsv4_8k1k_fp4_sglang_tp8-pp1-dcp1-pcp1-ep1-dpafalse_disagg-false_spec-none_conc64_b200-host-0123'; +const ACTION = 'data-action="view-power-trace"'; + +const hardwareConfig = { + b200: { + name: 'b200', + label: 'B200', + suffix: '', + gpu: 'B200', + color: 'blue', + power: 1000, + costh: 5, + costr: 1.25, + }, +} as unknown as HardwareConfig; + +function measuredPoint(overrides: Partial = {}): InferenceData { + return { + id: 980001, + date: '2026-09-01', + x: 60, + y: 600, + tp: 8, + conc: 64, + hwKey: 'b200', + precision: 'fp4', + benchmark_type: 'single_turn', + run_url: RUN_URL, + power_audit: { source: `power_validation_${AUDIT_NAME}.json` }, + tpPerGpu: { y: 400, roof: false }, + tpPerMw: { y: 50, roof: false }, + costh: { y: 1, roof: false }, + costr: { y: 1, roof: false }, + costhi: { y: 1, roof: false }, + costri: { y: 1, roof: false }, + ...overrides, + } as InferenceData; +} + +function config(overrides: Partial = {}): TooltipConfig { + return { + data: measuredPoint(), + isPinned: true, + xLabel: 'Interactivity (tok/s/user)', + yLabel: 'Measured Power per Chip (W)', + selectedYAxisMetric: 'y_measuredAvgPower', + hardwareConfig, + ...overrides, + }; +} + +function overlayConfig(overrides: Partial = {}): OverlayTooltipConfig { + return { + ...config({ data: measuredPoint({ id: 0 }) }), + overlayData: { + label: 'powerx-timeline', + hardwareConfig, + data: [], + runUrl: RUN_URL, + } as unknown as OverlayData, + ...overrides, + }; +} + +describe('View power trace tooltip action', () => { + afterEach(() => { + vi.unstubAllGlobals(); + }); + + it('renders on pinned overlay tooltips whose points carry id 0', () => { + const html = generateOverlayTooltipContent(overlayConfig()); + expect(html).toContain(` { expect(html).toContain('Click elsewhere to dismiss'); }); + it('caps pinned tooltip height so stacked actions stay inside the mobile viewport', () => { + const html = generateTooltipContent( + tooltipConfig({ + isPinned: true, + hasTrace: true, + hasLog: true, + showPowerTelemetry: true, + selectedYAxisMetric: 'y_measuredAvgPower', + data: pt({ + id: 42, + benchmark_type: 'agentic_traces', + power_audit: { source: 'power_validation_h100_conc8.json' }, + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/1', + }), + }), + ); + expect(html).toContain('max-height: min(70vh, calc(100dvh - 16px))'); + expect(html).toContain('overflow-y: auto'); + expect(html).toContain('max-width: min(320px, calc(100vw - 16px))'); + expect(html).toContain('data-action="view-power-trace"'); + }); + it('does not show dismiss text when isPinned is false', () => { const html = generateTooltipContent(tooltipConfig({ isPinned: false })); expect(html).not.toContain('Click elsewhere to dismiss'); diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts index 314941a89..a4877ffe4 100644 --- a/packages/app/src/components/inference/utils/tooltipUtils.ts +++ b/packages/app/src/components/inference/utils/tooltipUtils.ts @@ -6,6 +6,7 @@ import { isPersistedBenchmarkId } from '@/lib/benchmark-id'; import { frameworkFamily } from '@/lib/framework-family'; import type { Locale } from '@/lib/i18n'; import { isKvOffloadEnabled } from '@/lib/kv-offload'; +import { chartStateHref } from '@/lib/url-state'; import { chipCounts } from '@/lib/chip-counts'; import type { SystemPowerUnsupportedReason } from '@/lib/modeled-system-power'; @@ -14,6 +15,13 @@ import { isMeasuredEnergyConfigKey, isModeledSystemPowerConfigKey, } from '@/components/inference/metric-registry'; +import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config'; +import { powerVariantLabel } from '@/components/inference/utils/power-compare'; +import { + POWER_TIMELINE_METRIC_KEY, + traceKeyForPoint, +} from '@/components/inference/utils/powerTimeline'; +import { reconstructedRoleEnergy } from '@/components/inference/utils/role-energy'; import { meaningfulParallelismSize, parallelismLabel, @@ -49,6 +57,8 @@ export interface TooltipConfig { hasTrace?: boolean; /** Whether this official DB-backed point has a linked `server_logs` row. */ hasLog?: boolean; + /** Opt in only when the host handles the dialog action. */ + showPowerTelemetry?: boolean; /** Page locale for tooltip metadata labels. Defaults to English. */ locale?: Locale; } @@ -157,6 +167,8 @@ const TOOLTIP_STRINGS = { powerCertified: 'Validated (current PowerX method)', powerLegacy: 'Historical (not validated under the current method)', powerWithheld: 'Measured power withheld', + series: 'Series', + roleEnergyShare: 'Share of request energy', }, zh: { dismiss: '点击其他区域关闭', @@ -176,9 +188,39 @@ const TOOLTIP_STRINGS = { powerCertified: '已验证(采用当前 PowerX 方法)', powerLegacy: '历史测量(尚未按当前方法验证)', powerWithheld: '实测功耗未采信', + series: '系列', + roleEnergyShare: '在请求能耗中的占比', }, } as const; +/** + * Which comparison series (`i_pcompare`) a point belongs to, for the clones a + * boundary / role comparison appends. On the energy axis a role clone also + * reports its share of the reconstructed request energy (utils/role-energy.ts). + */ +const powerVariantHTML = ( + d: InferenceData, + selectedYAxisMetric: string, + locale: Locale, +): string => { + const variant = d.powerVariant; + if (!variant) return ''; + const t = TOOLTIP_STRINGS[locale]; + let html = tooltipLine(t.series, powerVariantLabel(variant, locale)); + if ( + variant.kind === 'role' && + variant.id !== 'all' && + getMeasuredMetricConfig(selectedYAxisMetric)?.family === 'energy' + ) { + const energy = reconstructedRoleEnergy(d); + if (energy) { + const share = variant.id === 'prefill' ? energy.prefillShare : 100 - energy.prefillShare; + html += tooltipLine(t.roleEnergyShare, `${share.toFixed(1)}%`); + } + } + return html; +}; + const totalChipsHTML = (d: InferenceData, selectedYAxisMetric: string, locale: Locale): string => { const t = TOOLTIP_STRINGS[locale]; const { physical, configured } = chipCounts( @@ -228,6 +270,8 @@ const SYSTEM_POWER_STRINGS = { : `${chassis} eight-GPU chassis · ${measured} of ${modeled} GPUs measured, extrapolated to full chassis`, extrapolation: 'Unmeasured chassis GPUs are assumed to run the same workload at the measured per-GPU power; deployment values are the measured GPUs’ share.', + uniformHosts: + 'No per-host telemetry for this multinode deployment; every chassis is modeled at the deployment-mean GPU power.', normalization: 'AC power is divided by all modeled chassis GPUs, including prefill and decode.', boundary: 'Includes GPU chassis CPUs; excludes separate CPU-only frontend/router hosts.', model: 'Power model source', @@ -257,6 +301,7 @@ const SYSTEM_POWER_STRINGS = { : `${chassis} 个八卡机箱 · 实测 ${measured}/${modeled} 张 GPU,按满机箱外推`, extrapolation: '假设机箱内未实测的 GPU 运行相同负载、功耗与实测每卡功耗相同;部署数值为实测 GPU 所占份额。', + uniformHosts: '该多节点部署没有逐主机功耗数据;每个机箱按部署平均每卡功耗建模。', normalization: '交流功耗按所有建模机箱的 GPU 总数分摊,包括 Prefill 与 Decode。', boundary: '计入 GPU 机箱内的 CPU;不计入独立的纯 CPU 前端或路由主机。', model: '功耗模型来源', @@ -303,7 +348,7 @@ const modeledSystemPowerHTML = ( ? ` ${tooltipLine(t.deploymentAc, `${fmt(estimate.deploymentAcWatts)} W`)} ${tooltipLine(`${t.facility} (PUE ${fmt(estimate.pue)})`, `${fmt(estimate.deploymentFacilityWatts)} W`)} -
${t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount)}${estimate.chassisBasis === 'extrapolated' ? `
${t.extrapolation}` : ''}
${t.assumptions}
${t.platformAssumptions}
${t.normalization}
${t.boundary}
+
${t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount)}${estimate.chassisBasis === 'extrapolated' ? `
${t.extrapolation}` : ''}${estimate.topologyBasis === 'uniform-hosts' ? `
${t.uniformHosts}` : ''}
${t.assumptions}
${t.platformAssumptions}
${t.normalization}
${t.boundary}
${tooltipLine(t.model, `
${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.slice(0, 12))}`)} ${t.sweep} ` @@ -493,39 +538,136 @@ const generateAgenticHTML = (d: InferenceData, locale: Locale): string => { }; const ACTION_STRINGS = { - en: { charts: 'View charts', logs: 'View logs' }, - zh: { charts: '查看图表', logs: '查看日志' }, + en: { + charts: 'View charts', + logs: 'View logs', + powerTelemetry: 'View PowerX', + powerTrace: 'View power trace', + }, + zh: { + charts: '查看图表', + logs: '查看日志', + powerTelemetry: '查看 PowerX', + powerTrace: '查看功耗曲线', + }, } as const; -const pointDetailActionLink = (action: 'view-charts' | 'view-logs', href: string, label: string) => +type TooltipAction = 'view-charts' | 'view-logs' | 'view-power-trace'; + +const pointDetailActionLink = (action: TooltipAction, href: string, label: string) => `${label} →`; -/** Point-detail links rendered only for persisted, pinned official points. */ -const viewActionsHTML = ( - isPinned: boolean, - hasTraceData: boolean, - hasLogData: boolean, - pointId: number | undefined, - benchmarkType: string | undefined, - locale: Locale, -): string => { - const isAgentic = benchmarkType === 'agentic_traces'; - const showCharts = isAgentic && hasTraceData; - if (!isPinned || !isPersistedBenchmarkId(pointId) || (!showCharts && !hasLogData)) return ''; - const prefix = locale === 'zh' ? '/zh' : ''; - const agenticHref = agenticDetailHref(pointId, locale); - const logHref = isAgentic - ? `${agenticHref}${agenticHref.includes('?') ? '&' : '?'}view=logs` - : `${prefix}/inference/logs/${pointId}`; +/** + * Outer shell for pinned/hover chart tooltips. Caps width and height to the + * viewport so stacked action buttons (logs / PowerX / power trace) stay + * reachable on narrow mobile viewports — `computeTooltipPosition` can only + * clamp top/left, so a taller-than-viewport shell would still overflow. + */ +export const tooltipShellStyle = (opts: { + isPinned: boolean; + /** Hex/CSS color; omit for the default border token. */ + border?: string; +}): string => { + const border = opts.border ?? '1px solid var(--border)'; + return [ + 'background: var(--popover)', + `border: ${border}`, + 'border-radius: 8px', + 'padding: 12px', + 'box-shadow: 0 4px 6px -1px rgb(0 0 0 / 0.1)', + `user-select: ${opts.isPinned ? 'text' : 'none'}`, + 'max-width: min(320px, calc(100vw - 16px))', + 'max-height: min(70vh, calc(100dvh - 16px))', + 'overflow-x: hidden', + 'overflow-y: auto', + 'overscroll-behavior: contain', + 'box-sizing: border-box', + ].join('; '); +}; + +/** + * Whether a point on the measured-power / energy scatter can jump to its + * per-second telemetry on the Timeline display. Overlay points qualify too: + * the trace is keyed by run id and audit name, not by a persisted row id. + */ +export const showsPowerTraceAction = ( + point: Pick, + selectedYAxisMetric: string, +): boolean => + selectedYAxisMetric !== POWER_TIMELINE_METRIC_KEY && + getMeasuredMetricConfig(selectedYAxisMetric) !== undefined && + traceKeyForPoint(point) !== null; + +/** + * Same-tab click is intercepted by the chart (in-page metric switch); the href + * keeps open-in-new-tab landing on the Timeline display of THIS chart. The + * address bar is stripped of chart state after load, so the share-link store + * is layered over the live location first (`chartStateHref`). + */ +const powerTraceHref = (): string => + typeof window === 'undefined' ? '#' : chartStateHref({ i_metric: POWER_TIMELINE_METRIC_KEY }); + +interface ViewActionsInput { + isPinned: boolean; + hasTraceData: boolean; + hasLogData: boolean; + point: InferenceData; + /** + * Metric the chart currently plots, when that chart can switch to the + * Timeline display in place (the scatter). Omitted by charts that cannot, + * so they never render a "View power trace" link nothing would handle. + */ + powerTraceMetric?: string; + showPowerTelemetry?: boolean; + locale: Locale; +} + +/** + * Point-detail links rendered only on pinned tooltips. "View charts" and + * "View logs" need a persisted row id (overlay points have none); "View power + * trace" needs only the run and audit source, so it works for overlays too. + */ +const viewActionsHTML = ({ + isPinned, + hasTraceData, + hasLogData, + point, + powerTraceMetric, + showPowerTelemetry, + locale, +}: ViewActionsInput): string => { + if (!isPinned) return ''; const t = ACTION_STRINGS[locale]; - const actions = [ - showCharts ? pointDetailActionLink('view-charts', agenticHref, t.charts) : '', - hasLogData ? pointDetailActionLink('view-logs', logHref, t.logs) : '', - ].filter(Boolean); + const actions: string[] = []; + const pointId = point.id; + const isAgentic = point.benchmark_type === 'agentic_traces'; + const showCharts = isAgentic && hasTraceData; + if (isPersistedBenchmarkId(pointId) && (showCharts || hasLogData)) { + const prefix = locale === 'zh' ? '/zh' : ''; + const agenticHref = agenticDetailHref(pointId, locale); + if (showCharts) { + actions.push(pointDetailActionLink('view-charts', agenticHref, t.charts)); + } + if (hasLogData) { + const logHref = isAgentic + ? `${agenticHref}${agenticHref.includes('?') ? '&' : '?'}view=logs` + : `${prefix}/inference/logs/${pointId}`; + actions.push(pointDetailActionLink('view-logs', logHref, t.logs)); + } + } + if (showPowerTelemetry && isPersistedBenchmarkId(pointId)) { + actions.push( + ``, + ); + } + if (powerTraceMetric !== undefined && showsPowerTraceAction(point, powerTraceMetric)) { + actions.push(pointDetailActionLink('view-power-trace', powerTraceHref(), t.powerTrace)); + } + if (actions.length === 0) return ''; return `
${actions.join('')}
`; }; @@ -681,7 +823,7 @@ export const generateTooltipContent = (config: TooltipConfig): string => { const t = TOOLTIP_STRINGS[locale]; return ` -
+
${isPinned ? `
${t.dismiss}
` : ''}
${hardwareConfig[d.hwKey] ? getDisplayLabel(getPointHardwareConfig(d, hardwareConfig[d.hwKey])) : d.hwKey} @@ -708,6 +850,7 @@ export const generateTooltipContent = (config: TooltipConfig): string => { : '' } ${powerTierHTML(d, selectedYAxisMetric, locale)} + ${powerVariantHTML(d, selectedYAxisMetric, locale)} ${modeledSystemPowerHTML(d, selectedYAxisMetric, isPinned, locale)} ${totalChipsHTML(d, selectedYAxisMetric, locale)} ${generateParallelismHTML(d, locale)} @@ -718,7 +861,15 @@ export const generateTooltipContent = (config: TooltipConfig): string => { ${generateAgenticHTML(d, locale)} ${generateWorkerPowerHTML(d, isPinned, locale)} ${runLinkHTML(runUrl, locale)} - ${viewActionsHTML(isPinned, Boolean(hasTrace), Boolean(config.hasLog), d.id, d.benchmark_type, locale)} + ${viewActionsHTML({ + isPinned, + hasTraceData: Boolean(hasTrace), + hasLogData: Boolean(config.hasLog), + point: d, + showPowerTelemetry: config.showPowerTelemetry, + powerTraceMetric: selectedYAxisMetric, + locale, + })}
`; }; @@ -739,7 +890,7 @@ export const generateOverlayTooltipContent = (config: OverlayTooltipConfig): str const branch = perRow?.branch ?? overlayData.label; return ` -
+
${isPinned ? `
${t.dismiss}
` : ''}
${t.unofficialRun} @@ -752,6 +903,7 @@ export const generateOverlayTooltipContent = (config: OverlayTooltipConfig): str ${tooltipLine(xLabel, fmt(d.x))} ${tooltipLine(yLabel, fmt(d.y))} ${powerTierHTML(d, selectedYAxisMetric, locale)} + ${powerVariantHTML(d, selectedYAxisMetric, locale)} ${modeledSystemPowerHTML(d, selectedYAxisMetric, isPinned, locale)} ${totalChipsHTML(d, selectedYAxisMetric, locale)} ${generateParallelismHTML(d, locale)} @@ -761,6 +913,14 @@ export const generateOverlayTooltipContent = (config: OverlayTooltipConfig): str ${powerWithheldHTML(d, locale)} ${generateAgenticHTML(d, locale)} ${generateWorkerPowerHTML(d, isPinned, locale)} + ${viewActionsHTML({ + isPinned, + hasTraceData: false, + hasLogData: false, + point: d, + powerTraceMetric: selectedYAxisMetric, + locale, + })}
`; }; @@ -788,7 +948,7 @@ export const generateGPUGraphTooltipContent = (config: TooltipConfig): string => const t = TOOLTIP_STRINGS[locale]; return ` -
+
${isPinned ? `
${t.dismiss}
` : ''} ${tooltipLine(t.date, `${formatTooltipDate(d.date, locale)}${d.actualDate && d.actualDate !== d.date ? ` ${t.dataFrom(formatTooltipDate(d.actualDate, locale))}` : ''}`)} ${tooltipLine(t.chipConfig, `${hardwareConfig[d.hwKey] ? getDisplayLabel(getPointHardwareConfig(d, hardwareConfig[d.hwKey])) : d.hwKey}`)} @@ -813,6 +973,7 @@ export const generateGPUGraphTooltipContent = (config: TooltipConfig): string => : '' } ${powerTierHTML(d, selectedYAxisMetric, locale)} + ${powerVariantHTML(d, selectedYAxisMetric, locale)} ${modeledSystemPowerHTML(d, selectedYAxisMetric, isPinned, locale)} ${totalChipsHTML(d, selectedYAxisMetric, locale)} ${generateParallelismHTML(d, locale)} @@ -823,7 +984,15 @@ export const generateGPUGraphTooltipContent = (config: TooltipConfig): string => ${generateAgenticHTML(d, locale)} ${generateWorkerPowerHTML(d, isPinned, locale)} ${runLinkHTML(runUrl, locale)} - ${viewActionsHTML(isPinned, Boolean(hasTrace), Boolean(hasLog), d.id, d.benchmark_type, locale)} + ${viewActionsHTML({ + isPinned, + hasTraceData: Boolean(hasTrace), + hasLogData: Boolean(hasLog), + point: d, + showPowerTelemetry: config.showPowerTelemetry, + powerTraceMetric: selectedYAxisMetric, + locale, + })}
`; }; diff --git a/packages/app/src/components/inference/utils/topology-filter.test.ts b/packages/app/src/components/inference/utils/topology-filter.test.ts new file mode 100644 index 000000000..abd9eb195 --- /dev/null +++ b/packages/app/src/components/inference/utils/topology-filter.test.ts @@ -0,0 +1,18 @@ +import { describe, expect, it } from 'vitest'; + +import type { InferenceData } from '@/components/inference/types'; + +import { pointTopologyKey, topologyLabel } from './topology-filter'; + +describe('topologyLabel', () => { + it('spells out offload modes beside compact parallelism tokens', () => { + const base = { physicalChips: 8, decode_tp: 8, decode_ep: 1, offload_mode: 'off' }; + const keys = [base, { ...base, offload_mode: 'on' }].map((point) => + pointTopologyKey(point as InferenceData), + ); + expect(keys.map((key) => topologyLabel(key, 'en', keys))).toEqual([ + 'Single-node · GPU8 · TP8 · EP1 · offload off', + 'Single-node · GPU8 · TP8 · EP1 · offload on', + ]); + }); +}); diff --git a/packages/app/src/components/inference/utils/topology-filter.ts b/packages/app/src/components/inference/utils/topology-filter.ts new file mode 100644 index 000000000..23a4c5e19 --- /dev/null +++ b/packages/app/src/components/inference/utils/topology-filter.ts @@ -0,0 +1,83 @@ +import type { InferenceData } from '../types'; +import { meaningfulParallelismSize } from './parallelism-label'; + +function topologyValue(value: number | boolean | undefined) { + return value === undefined || value === null || !Number.isFinite(Number(value)) + ? '?' + : Number(value); +} + +/** Load is deliberately excluded: a topology filter retains its whole concurrency sweep. */ +export function pointTopologyKey(point: InferenceData): string { + const fields: [string, number | boolean | undefined][] = [ + ['GPU', point.physicalChips ?? point.tp], + ['DP', point.dp], + ['TP', point.decode_tp], + ['EP', point.decode_ep ?? point.ep], + ['PP', point.decode_pp ?? point.pp], + [ + 'DCP', + point.disagg + ? point.decode_dcp_size + : (meaningfulParallelismSize(point.prefill_dcp_size, point.decode_dcp_size) ?? + point.decode_dcp_size ?? + point.prefill_dcp_size), + ], + [ + 'PCP', + point.disagg + ? point.decode_pcp_size + : (meaningfulParallelismSize(point.prefill_pcp_size, point.decode_pcp_size) ?? + point.decode_pcp_size ?? + point.prefill_pcp_size), + ], + ['DPA', point.decode_dp_attention ?? point.dp_attention], + ]; + if (point.disagg) + fields.push( + ['P-GPU', point.num_prefill_gpu], + ['D-GPU', point.num_decode_gpu], + ['P-workers', point.prefill_num_workers], + ['D-workers', point.decode_num_workers], + ['P-TP', point.prefill_tp], + ['P-EP', point.prefill_ep], + ['P-PP', point.prefill_pp], + ['P-DCP', point.prefill_dcp_size], + ['P-PCP', point.prefill_pcp_size], + ['P-DPA', point.prefill_dp_attention], + ); + const mode = point.disagg ? 'PD' : point.is_multinode ? 'Multi' : 'Single'; + return [ + mode, + ...fields.map(([key, v]) => `${key}=${topologyValue(v)}`), + `offload=${point.offload_mode ?? '?'}`, + ].join('|'); +} + +/** Values stay human-readable in share links and exported filter metadata. */ +export function topologyLabel( + key: string, + locale: 'en' | 'zh', + availableKeys?: readonly string[], +): string { + const [mode, ...parts] = key.split('|'); + const modes = + locale === 'zh' + ? { PD: '分离式', Multi: '多节点', Single: '单节点' } + : { PD: 'Disaggregated', Multi: 'Multi-node', Single: 'Single-node' }; + // Hide shared detail, never a difference: defaults and unknowns distinguish options too. + const peers = availableKeys?.filter((sibling) => sibling.split('|')[0] === mode); + const visible = parts.filter((part) => { + const name = part.split('=')[0]; + return ( + ['GPU', 'TP', 'EP', 'P-GPU', 'D-GPU', 'P-TP', 'P-EP'].includes(name) || + !peers?.length || + !peers.every((sibling) => sibling.split('|').includes(part)) + ); + }); + // Numbers and unknowns join their name (TP4, DP?); word values need a space (offload off). + return [ + modes[mode as keyof typeof modes] ?? mode, + ...visible.map((part) => part.replace(/=(?=\p{L})/u, ' ').replace('=', '')), + ].join(' · '); +} diff --git a/packages/app/src/components/inference/utils/x-axis-scale.ts b/packages/app/src/components/inference/utils/x-axis-scale.ts index 205c803ae..9b089a60b 100644 --- a/packages/app/src/components/inference/utils/x-axis-scale.ts +++ b/packages/app/src/components/inference/utils/x-axis-scale.ts @@ -18,6 +18,7 @@ export function resolveScatterXAxisScale({ xAxisField, scaleType, }: ResolveScatterXAxisScaleOptions): ScatterXAxisScale { + if (xAxisField === 'conc') return 'linear'; if (selectedYAxisMetric !== 'y_inputTputPerGpu') return 'linear'; if (scaleType === 'linear') return 'linear'; if (scaleType === 'log') return extent[0] > 0 ? 'log' : 'linear'; diff --git a/packages/app/src/lib/api-route-catalog.ts b/packages/app/src/lib/api-route-catalog.ts index 92f1d162f..c85b34fea 100644 --- a/packages/app/src/lib/api-route-catalog.ts +++ b/packages/app/src/lib/api-route-catalog.ts @@ -139,7 +139,7 @@ export const apiRouteCatalog = [ method: 'GET', classification: 'published-read', operationId: 'get-inference-view', - sourceSha256: 'de9192086b27e530ad6b1ece082b4989b3a9771a78195ae9f6d1c3899519588f', + sourceSha256: 'bafac08dbe6e6e4b75dd15987e79f84b514f3e93351149ad44b296f6dd2a9662', }, { source: 'src/app/api/v1/views/options/route.ts', @@ -147,7 +147,7 @@ export const apiRouteCatalog = [ method: 'GET', classification: 'published-read', operationId: 'get-view-options', - sourceSha256: '79571c3e6faf5af977c1afc9c29ebd6c1bfeb54ccf389711fa49668466fb58f4', + sourceSha256: 'ceba9504cfc50cb37748c9455f22dca9813789eb0148da740b24678bb540fb27', }, { source: 'src/app/api/v1/views/overview/route.ts', @@ -859,6 +859,54 @@ export const apiContractSourceDigests = [ zh: 'PowerX 实时产物选择、下载限制、单产物故障隔离、CSV context 规范化及 bundle 解码。', }, }, + { + source: 'src/components/inference/utils/resolveXAxisField.ts', + sourceSha256: '4783579c7b3c1a21b91968cb03e9c35a57a85251f4992cb5667a855ba7c77497', + reviewArea: { + en: 'Shared service-axis resolution: fixed-sequence mean/median, reciprocal mean TPOT, and AgentX percentile isolation.', + zh: '共用服务轴解析:固定长度工作负载 mean/median、mean TPOT 的倒数,以及 AgentX 独立的分位数选择。', + }, + }, + { + source: 'src/components/inference/utils/equal-service-comparison.ts', + sourceSha256: 'ff06f0eb908b6912706bb8dc67d931698d9cbd356546f20850ea8221fedc9354', + reviewArea: { + en: 'Source-scoped equal-service interpolation, comparator-relative changes, endpoint provenance, prefill share projection and role points.', + zh: '按完整来源限定的同等服务插值、相对基准变化、端点来源、prefill 占比视图以及各角色数据点。', + }, + }, + { + source: 'src/components/inference/utils/matched-concurrency.ts', + sourceSha256: 'db81deeff9e6ff64f607b71751d18acc59626f43fecbe037145e3a55c4fe959f', + reviewArea: { + en: 'Same-concurrency pairing of two exact sources: missing and conflicting observations, signed comparator-relative changes, no interpolation.', + zh: '两个完整来源在相同并发下的配对:缺失与冲突观测、相对基准的带符号变化,不做插值。', + }, + }, + { + source: 'src/components/inference/utils/power-fit.ts', + sourceSha256: 'a1e30d09648b74897e283b67a4e63e2591a485ed0f92e00acbe606b75967444a', + reviewArea: { + en: 'Per-source least-squares power fit on output per allocated GPU: intercept, marginal J/token, R², fitted range and registry TDP.', + zh: '按来源对每个已分配 GPU 的输出做最小二乘功耗拟合:截距、边际 J/token、R²、拟合范围与注册表 TDP。', + }, + }, + { + source: 'src/components/inference/utils/role-energy.ts', + sourceSha256: '36b1ceda97af99731482ed815167a660f1619ab94c2956afe11cf72d58fa5380', + reviewArea: { + en: 'Validated prefill/decode energy reconstruction and shares on a common output-token denominator.', + zh: '使用统一 output token 分母的已验证 prefill/decode 能耗重建与占比。', + }, + }, + { + source: 'src/lib/benchmark-transform.ts', + sourceSha256: '53fb9102b1ddd0c597b3b2a414894564d2deda3da5ea2cf0059f997b2db852c8', + reviewArea: { + en: 'Raw benchmark means and derived reciprocal mean-TPOT interactivity used by Dashboard and read-only views.', + zh: '仪表板和只读视图共用的原始 benchmark 均值与 mean TPOT 倒数形式的 interactivity。', + }, + }, { source: '../db/src/etl/power-audit-validations.ts', sourceSha256: '54f37f2bc06835b6cfbbe306d77acd3d5677e10997a9764a340f448f59e953a2', @@ -869,7 +917,7 @@ export const apiContractSourceDigests = [ }, { source: 'src/components/gpu-power/power-audit-bundle.ts', - sourceSha256: '96aa2ed60692116d5b5219b90b562fdc8143b63359bd917bbc91812ffaf274a8', + sourceSha256: '21607a0bc3795c545d7f84bb7af2558624a8d1e054e7c92b9203ae94825d1804', reviewArea: { en: 'Artifact Timeline validation windows, strict nested AgentX result identity, adjacent context selection, timezone normalization and device identity semantics.', zh: '产物 Timeline 验证窗口、严格匹配的嵌套 AgentX result 身份、相邻 context 选择、时区规范化及设备身份语义。', @@ -885,7 +933,7 @@ export const apiContractSourceDigests = [ }, { source: 'src/components/gpu-power/types.ts', - sourceSha256: 'e8c5460821f5d8228bcf8b7dd087abe14e29a8e220e1d8bcd75112fb5ead1baa', + sourceSha256: '9f4e90196e912af9e02ade8521321f35d804c25c02eca847a1b5717581740a42', reviewArea: { en: 'GPU telemetry units, missing values, timestamp deduplication, and full-record live statistics definitions.', zh: 'GPU 遥测单位、缺失值、时间戳去重,以及实时产物全记录统计的定义。', @@ -947,7 +995,7 @@ export const apiContractSourceDigests = [ { source: 'src/components/inference/hooks/chart-data-core.ts', - sourceSha256: '0c86987ed025557172a8d020144ca462be1ecac880d53b7d929f93a617086c33', + sourceSha256: '8601c33f6979541786362276697e4d4a21278de1d57b50d31f705391b5da041d', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', @@ -983,7 +1031,7 @@ export const apiContractSourceDigests = [ { source: 'src/lib/views-api/calculator-extensions.ts', - sourceSha256: '74c15d1dc1bca33de2ee74e22e9bf1186521c7623c6fd609d34677960c5d9633', + sourceSha256: 'c585d3827fb84e1142bb7ebc4ddbdb905098d7c427352b7809ddab66cb01ff19', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', @@ -1001,7 +1049,7 @@ export const apiContractSourceDigests = [ { source: 'src/lib/views-api/series.ts', - sourceSha256: 'b5ebddee9d9ea6b85a8fdab2cf8fcc50bbf05e0fa10564987698abd38a6d4488', + sourceSha256: '9974fe5166ac847f4b9284eda8a901ee0f9c8b432567ef184890453c715d2f4f', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', @@ -1010,7 +1058,7 @@ export const apiContractSourceDigests = [ { source: 'src/lib/views-api/registry.ts', - sourceSha256: '1958ff6a851aa32a167c4acafbdaf69c7ef18bac9ca59851d2153300dbf88a59', + sourceSha256: '5836360c07ce714b78ca3a93b541f704cffce63567f645cf2ad8199687e48c16', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', diff --git a/packages/app/src/lib/benchmark-transform.test.ts b/packages/app/src/lib/benchmark-transform.test.ts index 59fef0b45..c9b1ae00e 100644 --- a/packages/app/src/lib/benchmark-transform.test.ts +++ b/packages/app/src/lib/benchmark-transform.test.ts @@ -69,6 +69,18 @@ function makeRow(overrides: Partial = {}): BenchmarkRow { } describe('rowToAggDataEntry', () => { + it('carries the curve snapshot identity of append-only rows and leaves legacy rows without one', () => { + const stitched = rowToAggDataEntry( + makeRow({ curve_date: '2026-09-20', curve_workflow_run_id: 35843506474 }), + ); + expect([stitched.curve_date, stitched.curve_workflow_run_id]).toEqual([ + '2026-09-20', + 35843506474, + ]); + const legacy = rowToAggDataEntry(makeRow()); + expect([legacy.curve_date, legacy.curve_workflow_run_id]).toEqual([undefined, undefined]); + }); + it.each([1, undefined])( 'labels run 35879254139 in the official/overlay legend without changing data (DB id %s)', (id) => { diff --git a/packages/app/src/lib/benchmark-transform.ts b/packages/app/src/lib/benchmark-transform.ts index f1712e306..3c7263a56 100644 --- a/packages/app/src/lib/benchmark-transform.ts +++ b/packages/app/src/lib/benchmark-transform.ts @@ -199,6 +199,9 @@ export function rowToAggDataEntry(row: BenchmarkRow): AggDataEntry { p99_tpot: m.p99_tpot ?? 0, 'p99.9_tpot': m['p99.9_tpot'] ?? 0, mean_intvty: m.mean_intvty ?? 0, + ...(typeof m.mean_tpot === 'number' && Number.isFinite(m.mean_tpot) && m.mean_tpot > 0 + ? { mean_tpot_intvty: 1 / m.mean_tpot } + : {}), median_intvty: m.median_intvty ?? 0, std_intvty: m.std_intvty ?? 0, p75_intvty: m.p75_intvty ?? 0, @@ -312,6 +315,8 @@ export function rowToAggDataEntry(row: BenchmarkRow): AggDataEntry { date: row.date, actualDate: (row as any).actualDate ?? row.date, run_url: row.run_url ?? undefined, + curve_date: row.curve_date, + curve_workflow_run_id: row.curve_workflow_run_id, benchmark_type: row.benchmark_type, isl: row.isl, osl: row.osl, diff --git a/packages/app/src/lib/chart-utils.test.ts b/packages/app/src/lib/chart-utils.test.ts index 9d5203b66..a23699602 100644 --- a/packages/app/src/lib/chart-utils.test.ts +++ b/packages/app/src/lib/chart-utils.test.ts @@ -928,6 +928,89 @@ describe('createChartDataPoint energy fields', () => { }); }); +// =========================================================================== +// createChartDataPoint — power-boundary fields (B2 GPU provisioned, B3 utility +// provisioned, B4 utility modeled). Mock specs: tdp 700 W, power 700 "kW". +// =========================================================================== +const boundaryPoint = (e: AggDataEntry) => + createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); + +describe('createChartDataPoint power-boundary fields', () => { + // Eight measured GPUs (2 prefill + 6 decode) on two partially filled chassis: + // the model evaluates 16 GPUs, so only deploymentFacilityWatts ÷ gpuCount yields 877.5 W. + // Every other numerator/denominator pairing gives a different number. + const supportedModel = { + status: 'supported' as const, + hardware: 'h100', + modelRevision: 'test', + modelPath: 'test', + gpuCount: 8, + chassisCount: 2, + modeledGpuCount: 16, + measuredGpuWattsPerGpu: 500, + chassisAcWatts: 10_400, + chassisAcWattsPerGpu: 650, + facilityWatts: 13_520, + deploymentAcWatts: 5400, + deploymentFacilityWatts: 7020, + pue: 1.3, + telemetryBasis: 'validated-v2' as const, + topologyBasis: 'worker-hosts' as const, + chassisBasis: 'extrapolated' as const, + }; + const validated = { + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 500, + joules_per_output_token: 10, + modeledSystemPower: supportedModel, + }; + it('emits all six boundary fields for a validated official row', () => { + const p = boundaryPoint( + entry({ output_tput_per_gpu: 400, benchmark_type: 'single_turn', ...validated }), + ); + expect(p.gpuProvisionedWatts).toEqual({ y: 700, roof: false }); + expect(p.gpuProvisionedJPerOutputToken?.y).toBeCloseTo(700 / 400, 10); + expect(p.utilityProvisionedWatts).toEqual({ y: 700_000, roof: false }); + expect(p.utilityProvisionedJPerOutputToken?.y).toBeCloseTo(700_000 / 400, 10); + // B4 W = deployment facility ÷ measured GPUs; J scales B1 by B4 W ÷ B1 W. + expect(p.utilityModeledWatts).toEqual({ y: 877.5, roof: false }); + expect(p.utilityModeledJPerOutputToken?.y).toBeCloseTo((10 * 877.5) / 500, 10); + // Existing measured (B1) fields are untouched by the new boundaries. + expect(p.measuredAvgPower).toEqual({ y: 500, roof: false }); + expect(p.measuredJPerOutputToken).toEqual({ y: 10, roof: false }); + }); + + it('normalizes fixed-sequence disaggregated energy by all GPUs while jOutput stays per decode GPU', () => { + const p = boundaryPoint( + entry({ + output_tput_per_gpu: 400, + disagg: true, + benchmark_type: 'single_turn', + num_prefill_gpu: 4, + num_decode_gpu: 4, + }), + ); + expect(p.gpuProvisionedJPerOutputToken?.y).toBeCloseTo((700 * 8) / (400 * 4), 10); + expect(p.utilityProvisionedJPerOutputToken?.y).toBeCloseTo((700_000 * 8) / (400 * 4), 10); + expect(p.jOutput?.y).toBeCloseTo(700_000 / 400, 10); + expect(p.utilityProvisionedJPerOutputToken?.y).toBeCloseTo(2 * p.jOutput!.y, 10); + + const agentic = boundaryPoint( + entry({ + output_tput_per_gpu: 400, + disagg: true, + benchmark_type: 'agentic_traces', + num_prefill_gpu: 4, + num_decode_gpu: 4, + }), + ); + expect(agentic.gpuProvisionedWatts?.y).toBe(700); + expect(agentic.gpuProvisionedJPerOutputToken).toBeUndefined(); + expect(agentic.utilityProvisionedJPerOutputToken).toBeUndefined(); + }); +}); + // =========================================================================== // createChartDataPoint — measured power / energy fields (from runner telemetry) // =========================================================================== @@ -952,20 +1035,6 @@ describe('createChartDataPoint measured power fields', () => { expect(missing.measuredP75Power).toBeUndefined(); expect(missing.measuredP90Power).toBeUndefined(); }); - it('emits measuredAvgPower when avg_power_w is present on the entry', () => { - const e = entry({ avg_power_w: 685.5 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredAvgPower).toBeDefined(); - expect(point.measuredAvgPower!.y).toBe(685.5); - expect(point.measuredAvgPower!.roof).toBe(false); - }); - - it('emits measuredJPerOutputToken when joules_per_output_token is present', () => { - const e = entry({ joules_per_output_token: 8.4 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredJPerOutputToken).toBeDefined(); - expect(point.measuredJPerOutputToken!.y).toBe(8.4); - }); it('derives J/query, Wh/query, and percent TDP from validated source fields', () => { const e = entry({ avg_power_w: 560, joules_per_successful_query: 1800 }); @@ -1019,14 +1088,6 @@ describe('createChartDataPoint measured power fields', () => { expect(point.measuredAvgPower!.y).toBe(0); }); - it('emits measuredJPerTotalToken when joules_per_total_token is present', () => { - const e = entry({ joules_per_total_token: 0.93 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredJPerTotalToken).toBeDefined(); - expect(point.measuredJPerTotalToken!.y).toBe(0.93); - expect(point.measuredJPerTotalToken!.roof).toBe(false); - }); - it('emits J/output and J/total independently — different denominators', () => { // 8k1k workload: J/output ≈ 9 × J/total (input is ~8x output, so output/total ≈ 1/9). const e = entry({ joules_per_output_token: 2.04, joules_per_total_token: 0.23 }); @@ -1051,30 +1112,6 @@ describe('createChartDataPoint measured power fields', () => { // createChartDataPoint — per-stage measured power / energy (disagg prefill/decode) // =========================================================================== describe('createChartDataPoint per-stage measured power fields', () => { - it('emits measuredPrefillAvgPower when prefill_avg_power_w is present', () => { - const e = entry({ prefill_avg_power_w: 920.3 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredPrefillAvgPower).toBeDefined(); - expect(point.measuredPrefillAvgPower!.y).toBe(920.3); - expect(point.measuredPrefillAvgPower!.roof).toBe(false); - }); - - it('emits measuredDecodeAvgPower when decode_avg_power_w is present', () => { - const e = entry({ decode_avg_power_w: 612.1 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredDecodeAvgPower).toBeDefined(); - expect(point.measuredDecodeAvgPower!.y).toBe(612.1); - expect(point.measuredDecodeAvgPower!.roof).toBe(false); - }); - - it('emits measuredJPerInputToken when joules_per_input_token is present', () => { - const e = entry({ joules_per_input_token: 0.27 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredJPerInputToken).toBeDefined(); - expect(point.measuredJPerInputToken!.y).toBe(0.27); - expect(point.measuredJPerInputToken!.roof).toBe(false); - }); - it('omits all per-stage fields on legacy rows predating per-stage attribution', () => { // Single-node / pre-disagg runs emit avg_power_w only, no prefill/decode split. const e = entry({ avg_power_w: 685.5 }); @@ -1084,15 +1121,6 @@ describe('createChartDataPoint per-stage measured power fields', () => { expect(point.measuredJPerInputToken).toBeUndefined(); }); - it('emits prefill and decode independently — the disagg per-stage split', () => { - // GB300 disagg: prefill GPUs run compute-bound (higher W) than decode GPUs. - const e = entry({ prefill_avg_power_w: 948, decode_avg_power_w: 631 }); - const point = createChartDataPoint('2025-01-01', e, 'median_e2el', 'tput_per_gpu', 'h100'); - expect(point.measuredPrefillAvgPower!.y).toBe(948); - expect(point.measuredDecodeAvgPower!.y).toBe(631); - expect(point.measuredPrefillAvgPower!.y).toBeGreaterThan(point.measuredDecodeAvgPower!.y); - }); - it('preserves a zero per-stage power value (not falsy-coerced away)', () => { // Same typeof===number gate as total power — 0 W must survive, not be dropped. const e = entry({ prefill_avg_power_w: 0, decode_avg_power_w: 0 }); diff --git a/packages/app/src/lib/chart-utils.ts b/packages/app/src/lib/chart-utils.ts index a8e974d41..340e748c1 100644 --- a/packages/app/src/lib/chart-utils.ts +++ b/packages/app/src/lib/chart-utils.ts @@ -21,6 +21,8 @@ import { import { DEFAULT_TCO_BASIS, getGpuSpecs, isKnownGpu, type TcoBasis } from '@/lib/constants'; import { getVendor, type Vendor } from '@/lib/dynamic-colors'; import type { Locale } from '@/lib/i18n'; +import { buildPowerBasisChartFields, type PowerBasisChartFields } from '@/lib/power-basis'; +import { reconstructedRoleEnergy } from '@/components/inference/utils/role-energy'; // --------------------------------------------------------------------------- // High-contrast color generation (iwanthue — k-means in CIELab) @@ -292,7 +294,8 @@ export function buildAvailabilityHwKey( return hwKey; } -export type DerivedMetricKey = BenchmarkMetricKey; +// The reconstructed prefill energy is a comparison-only series (never an axis). +export type DerivedMetricKey = BenchmarkMetricKey | 'reconstructedPrefillJPerOutputToken'; export type DerivedChartFields = Pick; const chartMetric = (y: number): { y: number; roof: boolean } => ({ y, roof: false }); @@ -414,6 +417,11 @@ export function buildDerivedChartFields( hardwarePower && tputPerGpu ? (hardwarePower * 1000) / tputPerGpu : 0, ); } + // jOutput keeps the historical per-GPU normalization: for disaggregated rows + // output_tput_per_gpu is per decode GPU, so this is all-in W of one decode GPU + // per output token and ignores the prefill pool. The power-boundary field + // utilityProvisionedJPerOutputToken uses the same all-in W but counts every + // allocated GPU, so the two differ on disaggregated rows by (P + D) / D. if (hardwarePower > 0 && wants('jOutput') && outputTputPerGpu) { fields.jOutput = chartMetric(hardwarePower ? (hardwarePower * 1000) / outputTputPerGpu : 0); } @@ -429,6 +437,14 @@ export function buildDerivedChartFields( if (wants(key)) fields[key] = value; } + const powerBasis = buildPowerBasisChartFields(entry, specs); + for (const [key, value] of Object.entries(powerBasis) as [ + keyof PowerBasisChartFields, + { y: number; roof: boolean }, + ][]) { + if (wants(key)) fields[key] = value; + } + if (wants('modeledChassisPowerPerGpu') && entry.modeledSystemPower?.status === 'supported') { fields.modeledChassisPowerPerGpu = chartMetric(entry.modeledSystemPower.chassisAcWattsPerGpu); } @@ -524,6 +540,8 @@ type MeasuredPowerChartFields = Partial< | 'measuredJPerSuccessfulQuery' | 'measuredWhPerSuccessfulQuery' | 'measuredPowerPercentTdp' + | 'measuredPowerTimeline' + | 'reconstructedPrefillJPerOutputToken' > >; @@ -532,9 +550,17 @@ function buildMeasuredPowerChartFields( entry: AggDataEntry, tdpWatts: number, ): MeasuredPowerChartFields { + // Prefill energy on the output-token axis, so the roles comparison can + // stack it against the decode pool (PowerX Figure 7). + const roleEnergy = reconstructedRoleEnergy(entry); return { + // The timeline axis aliases the validated average: the point set (and + // its table row) is the same, only the chart body changes. ...(typeof entry.avg_power_w === 'number' - ? { measuredAvgPower: chartMetric(entry.avg_power_w) } + ? { + measuredAvgPower: chartMetric(entry.avg_power_w), + measuredPowerTimeline: chartMetric(entry.avg_power_w), + } : {}), ...(typeof entry.p75_power_w === 'number' && Number.isFinite(entry.p75_power_w) ? { measuredP75Power: chartMetric(entry.p75_power_w) } @@ -565,6 +591,7 @@ function buildMeasuredPowerChartFields( ...(typeof entry.decode_joules_per_output_token === 'number' ? { measuredDecodeJPerOutputToken: chartMetric(entry.decode_joules_per_output_token) } : {}), + ...(roleEnergy ? { reconstructedPrefillJPerOutputToken: chartMetric(roleEnergy.prefill) } : {}), ...(typeof entry.joules_per_successful_query === 'number' ? { measuredJPerSuccessfulQuery: chartMetric(entry.joules_per_successful_query), @@ -589,14 +616,21 @@ export function remapInferencePoint( const metric = point[metricKey]; const xCandidate = (point as Partial)[xAxisField]; // Absent TTFT values are zero-filled by the row transform. Neither that - // sentinel nor an unrelated fallback coordinate is a latency measurement. - const missingTtft = - xAxisField.endsWith('_ttft') && (typeof xCandidate !== 'number' || xCandidate <= 0); + // sentinel nor an unrelated fallback coordinate is a latency measurement; + // the same holds for concurrency and the mean service fields. + const requiresMeasuredValue = + xAxisField === 'conc' || + xAxisField.endsWith('_ttft') || + xAxisField === 'mean_e2el' || + xAxisField === 'mean_tpot_intvty'; + const missingMeasuredValue = + requiresMeasuredValue && + (typeof xCandidate !== 'number' || !Number.isFinite(xCandidate) || xCandidate <= 0); return { ...point, - x: missingTtft ? NaN : typeof xCandidate === 'number' ? xCandidate : point.x, + x: missingMeasuredValue ? NaN : typeof xCandidate === 'number' ? xCandidate : point.x, y: metric?.y ?? point.y, - roof: metric?.roof ?? false, + roof: xAxisField === 'conc' ? false : (metric?.roof ?? false), }; } diff --git a/packages/app/src/lib/csv-export-helpers.test.ts b/packages/app/src/lib/csv-export-helpers.test.ts index 5fe616931..316f6af1f 100644 --- a/packages/app/src/lib/csv-export-helpers.test.ts +++ b/packages/app/src/lib/csv-export-helpers.test.ts @@ -8,6 +8,7 @@ import { historicalTrendToCsv, } from './csv-export-helpers'; import type { InferenceData } from '@/components/inference/types'; +import { expandPowerCompareSeries } from '@/components/inference/utils/power-compare'; const makePoint = (overrides: Partial = {}): InferenceData => ({ x: 100, @@ -752,3 +753,39 @@ describe('historicalTrendToCsv (mirrors HistoricalTrendsDisplay export)', () => expect(rows[1][headers.indexOf('Date')]).toBe('2025-01-15'); }); }); + +describe('inferenceChartToCsv power comparison', () => { + it.each([false, true])('exports plotted role values with overlay=%s', (overlay) => { + const base = makePoint({ + hwKey: 'gb300_dynamo-trt', + y: 708.1, + measuredAvgPower: { y: 708.1, roof: false }, + measuredPrefillAvgPower: { y: 760.442, roof: false }, + measuredDecodeAvgPower: { y: 690.652, roof: false }, + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/35532106109', + }); + const points = expandPowerCompareSeries([base], 'y_measuredAvgPower', 'roles'); + const { headers, rows } = inferenceChartToCsv( + overlay ? [] : points, + 'Kimi-K3', + 'agentic-traces', + overlay ? points : [], + { + yHeader: 'Measured Power per Chip (W)', + yPath: 'measuredAvgPower.y', + xHeader: 'Interactivity (tok/s/user)', + }, + ); + expect( + rows.map((row) => [ + row[headers.indexOf('Power Series')], + row[headers.indexOf('Measured Power per Chip (W)')], + ]), + ).toEqual([ + ['All GPUs', 708.1], + ['Prefill GPUs', 760.442], + ['Decode GPUs', 690.652], + ]); + expect(points.every((point) => point.measuredAvgPower?.y === 708.1)).toBe(true); + }); +}); diff --git a/packages/app/src/lib/csv-export-helpers.ts b/packages/app/src/lib/csv-export-helpers.ts index 576569af8..a58ba948f 100644 --- a/packages/app/src/lib/csv-export-helpers.ts +++ b/packages/app/src/lib/csv-export-helpers.ts @@ -9,6 +9,7 @@ import { METRIC_REGISTRY } from '@/components/inference/metric-registry'; import type { InferenceData, TrendDataPoint } from '@/components/inference/types'; +import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare'; import { chipCounts } from '@/lib/chip-counts'; import type { SubmissionVolumeRow } from '@/lib/submissions-types'; @@ -57,6 +58,12 @@ export function inferenceChartToCsv( const islOsl = sequenceToIslOsl(sequence); const showModeledPower = displayedMetrics?.yPath === METRIC_REGISTRY.modeledChassisPowerPerGpu.field; + // A power comparison (`i_pcompare`) appends boundary / role clones of the + // plotted points; name each row's series so the export stays unambiguous. + const allPoints = [...data, ...overlayData]; + const powerCompare = inferPowerCompare(allPoints); + const showPowerSeries = powerCompare !== 'none'; + const plottedMetric = displayedMetrics ? `y_${displayedMetrics.yPath.split('.')[0]}` : ''; const headers = [ 'Model', 'ISL', @@ -111,13 +118,15 @@ export function inferenceChartToCsv( 'Physical Chips', 'DP', ...(showModeledPower ? ['Configured Chip Count'] : []), + ...(showPowerSeries ? ['Power Series'] : []), ]; const displayedColumns = displayedMetrics ? [ { header: displayedMetrics.yHeader, - value: (point: InferenceData) => nestedMetric(point, displayedMetrics.yPath), + value: (point: InferenceData) => + point.powerVariant ? point.y : nestedMetric(point, displayedMetrics.yPath), }, { header: displayedMetrics.xHeader, value: (point: InferenceData) => point.x }, ].filter( @@ -128,7 +137,7 @@ export function inferenceChartToCsv( : []; headers.splice(10, 0, ...displayedColumns.map((column) => column.header)); - const rows = [...data, ...overlayData] + const rows = allPoints .filter((d) => !d.hidden) .map((d) => { const chips = chipCounts(d, showModeledPower); @@ -177,6 +186,7 @@ export function inferenceChartToCsv( chips.physical, d.dp ?? '', ...(showModeledPower ? [chips.configured] : []), + ...(showPowerSeries ? [powerSeriesLabel(d, plottedMetric, powerCompare, 'en')] : []), ]; row.splice(10, 0, ...displayedColumns.map((column) => column.value(d))); return row; diff --git a/packages/app/src/lib/d3-chart/layers/perf-ruler.test.ts b/packages/app/src/lib/d3-chart/layers/perf-ruler.test.ts index f1a3484f6..843a970e8 100644 --- a/packages/app/src/lib/d3-chart/layers/perf-ruler.test.ts +++ b/packages/app/src/lib/d3-chart/layers/perf-ruler.test.ts @@ -16,10 +16,12 @@ import { isPerfRulerCurveVisible, movePerfRulerIsoX, nextPerfRulerState, + parsePerfRulers, pathXExtent, perfRulerCurveSet, prunePerfRulers, renderPerfRulers, + serializePerfRulers, type PerfRulerEndInput, type PerfRulerGeometry, type PerfRulerLabelLayoutOptions, @@ -634,6 +636,47 @@ describe('perfRulerCurveSet', () => { }); }); +// ── serializePerfRulers / parsePerfRulers (share links) ───────────── + +describe('serializePerfRulers / parsePerfRulers', () => { + const OFFICIAL_A = 'roofline-b200_trt_fp8'; + const OFFICIAL_B = 'roofline-mi355x_sglang_fp4'; + // Overlay curves carry the unofficial run index; power-envelope curves are + // split per date with the encoded date appended (`%2F` from a slash). + const OVERLAY = 'overlay-roofline-h100_vllm_fp8_run1__2026-09%2F11'; + + it('round-trips completed rulers, including overlay and date-scoped curve ids', () => { + const state = complete( + complete(EMPTY_PERF_RULER_STATE, OFFICIAL_A, OFFICIAL_B, 41.5), + OFFICIAL_A, + OVERLAY, + 120, + ); + const encoded = serializePerfRulers(state); + expect(encoded).toBe(`41.5|${OFFICIAL_A}|${OFFICIAL_B};120|${OFFICIAL_A}|${OVERLAY}`); + const parsed = parsePerfRulers(encoded); + expect(parsed.rulers).toEqual([ + { id: 1, curveA: OFFICIAL_A, curveB: OFFICIAL_B, isoX: 41.5 }, + { id: 2, curveA: OFFICIAL_A, curveB: OVERLAY, isoX: 120 }, + ]); + expect(parsed.draft).toBeNull(); + expect(parsed.nextId).toBe(3); + }); + + it('round-trips run-specific date-comparison curve ids that contain ~', () => { + // GPUGraph series ids stamp the comparison entry onto point.date, so a + // run-qualified selection yields `roofline-~r__`. + const RUN_A = 'roofline-2026-09-09~r27489075807_b200_fp8'; + const RUN_B = 'roofline-2026-09-09~r27489075808_b200_fp8'; + const state = complete(EMPTY_PERF_RULER_STATE, RUN_A, RUN_B, 55.25); + const encoded = serializePerfRulers(state); + expect(encoded).toBe(`55.25|${RUN_A}|${RUN_B}`); + expect(parsePerfRulers(encoded).rulers).toEqual([ + { id: 1, curveA: RUN_A, curveB: RUN_B, isoX: 55.25 }, + ]); + }); +}); + // ── pathXExtent ───────────────────────────────────────────── describe('pathXExtent', () => { diff --git a/packages/app/src/lib/d3-chart/layers/perf-ruler.ts b/packages/app/src/lib/d3-chart/layers/perf-ruler.ts index c44143529..16c7284f3 100644 --- a/packages/app/src/lib/d3-chart/layers/perf-ruler.ts +++ b/packages/app/src/lib/d3-chart/layers/perf-ruler.ts @@ -323,6 +323,77 @@ export function prunePerfRulers( return { ...prev, rulers, draft }; } +const PERF_RULER_URL_RULER_SEPARATOR = ';'; +const PERF_RULER_URL_FIELD_SEPARATOR = '|'; +/** + * Shape of a curve id the link may reference: one roofline path's identity + * class (`roofline-` / `overlay-roofline-`), never the shared + * `roofline-path` / `overlay-roofline-path` marker classes or any other node + * inside the zoom group — those match many paths, so a hand-edited link + * would draw a ruler between whichever two come first in DOM order. + * + * Date-comparison curves include the comparison entry in the series id + * (`roofline->__`), so `~` is part of the + * allowed alphabet alongside the encoded-date `__` suffix. + */ +const PERF_RULER_CURVE_ID = /^(?:overlay-)?roofline-(?!path$)[\w%.~-]+$/u; + +/** + * Share-link encoding of the COMPLETED rulers (`i_rulers`). One ruler per + * `;`, fields joined by `|`: `isoX|curveA|curveB`. Curve ids are the rendered + * roofline path identity classes (`roofline-_`, + * `overlay-roofline-__run`, date-comparison + * `roofline-]>__`, optionally `__`), whose alphabet is `[A-Za-z0-9_%.~-]`, so neither separator can + * appear inside one; `URLSearchParams` percent-encodes both on the wire. + * The iso-x is rounded to four significant digits to keep links short — a + * 0.05% shift on the x metric is far below the ruler's visual resolution. + * The draft is never serialized: it is an unfinished click, not a + * measurement. Empty state serializes to '' so the param strips as default. + */ +export function serializePerfRulers(state: PerfRulerState): string { + return state.rulers + .map((ruler) => + [Number(ruler.isoX.toPrecision(4)), ruler.curveA, ruler.curveB].join( + PERF_RULER_URL_FIELD_SEPARATOR, + ), + ) + .join(PERF_RULER_URL_RULER_SEPARATOR); +} + +/** + * Inverse of {@link serializePerfRulers}. Malformed entries (wrong field + * count, non-numeric iso-x, identical curve ids, or ids that are not + * roofline identity classes) are dropped silently — a hand-edited or + * truncated link degrades to fewer rulers, never to an error. The list is capped at + * {@link MAX_PERF_RULERS} keeping the NEWEST (last-serialized) entries, the + * same end the click reducer drops from. Ids are reassigned 1..n with + * `nextId = n + 1`, so parsed rulers are valid D3 join keys and a ruler + * placed afterwards never collides. Returns {@link EMPTY_PERF_RULER_STATE} + * (same reference) for '', null, or an all-malformed value. + */ +export function parsePerfRulers(raw: string | null | undefined): PerfRulerState { + if (!raw) return EMPTY_PERF_RULER_STATE; + const parsed: Omit[] = []; + for (const entry of raw.split(PERF_RULER_URL_RULER_SEPARATOR)) { + const fields = entry.split(PERF_RULER_URL_FIELD_SEPARATOR); + if (fields.length !== 3) continue; + const [isoXField, curveA, curveB] = fields; + if (isoXField.trim() === '' || curveA === curveB) continue; + if (!PERF_RULER_CURVE_ID.test(curveA) || !PERF_RULER_CURVE_ID.test(curveB)) continue; + const isoX = Number(isoXField); + if (!Number.isFinite(isoX)) continue; + parsed.push({ curveA, curveB, isoX }); + } + if (parsed.length === 0) return EMPTY_PERF_RULER_STATE; + const kept = parsed.slice(-MAX_PERF_RULERS); + return { + rulers: kept.map((ruler, index) => ({ id: index + 1, ...ruler })), + draft: null, + nextId: kept.length + 1, + }; +} + /** Every curve referenced by any ruler or the draft (hit-halo styling). */ export function perfRulerCurveSet(state: PerfRulerState): Set { const curves = new Set(); diff --git a/packages/app/src/lib/d3-chart/layers/rooflines.test.ts b/packages/app/src/lib/d3-chart/layers/rooflines.test.ts index 409cc1f50..170158964 100644 --- a/packages/app/src/lib/d3-chart/layers/rooflines.test.ts +++ b/packages/app/src/lib/d3-chart/layers/rooflines.test.ts @@ -19,7 +19,12 @@ vi.mock('d3', async () => { }; }); -import { renderRooflines, updateRooflinesOnZoom, type RooflineConfig } from './rooflines'; +import { + renderRooflines, + updateRooflinesForDisplay, + updateRooflinesOnZoom, + type RooflineConfig, +} from './rooflines'; // ── Fixtures ───────────────────────────────────────────────────────── @@ -201,6 +206,31 @@ describe('renderRooflines', () => { } }); + it('dashes each curve from getDasharray and keeps the dashes on display updates', () => { + const group = createMockGroup(); + const { xScale, yScale } = makeScales(); + // A null dash is a solid curve even when a chart-wide dash is also set. + const config = makeConfig({ + strokeDasharray: '5,3', + getDasharray: (key) => (key === 'modelB' ? '2 3' : null), + }); + const dashes = () => + Object.fromEntries( + group + .selectAll('.roofline-path') + .elements.map((el) => [el.attrs['class'], el.attrs['stroke-dasharray']]), + ); + const expected = { + 'roofline-path roofline-modelA': null, + 'roofline-path roofline-modelB': '2 3', + }; + + renderRooflines(group as any, SAMPLE_ROOFLINES, xScale, yScale, config); + expect(dashes()).toEqual(expected); + updateRooflinesForDisplay(group as any, config); + expect(dashes()).toEqual(expected); + }); + it('does not set strokeDasharray when not specified', () => { const group = createMockGroup(); const { xScale, yScale } = makeScales(); diff --git a/packages/app/src/lib/d3-chart/layers/rooflines.ts b/packages/app/src/lib/d3-chart/layers/rooflines.ts index c97b880cc..b58d0adb7 100644 --- a/packages/app/src/lib/d3-chart/layers/rooflines.ts +++ b/packages/app/src/lib/d3-chart/layers/rooflines.ts @@ -8,9 +8,15 @@ export interface RooflineConfig { isVisible?: (key: string) => boolean; strokeWidth?: number; strokeDasharray?: string; + /** Per-curve dash pattern; `null` draws a solid line. Takes precedence over `strokeDasharray`. */ + getDasharray?: (key: string) => string | null; curve?: d3.CurveFactory; } +function dasharrayFor(config: RooflineConfig, key: string): string | null { + return config.getDasharray ? config.getDasharray(key) : (config.strokeDasharray ?? null); +} + interface RooflineEntry { key: string; points: T[]; @@ -27,7 +33,14 @@ export function renderRooflines( yScale: ContinuousScale, config: RooflineConfig, ): void { - const { getColor, getOpacity, isVisible, strokeWidth = 2, strokeDasharray } = config; + const { + getColor, + getOpacity, + isVisible, + strokeWidth = 2, + strokeDasharray, + getDasharray, + } = config; const lineGenerator = d3 .line() .x((d) => xScale(d.x)) @@ -60,7 +73,9 @@ export function renderRooflines( .attr('stroke-width', strokeWidth) .attr('d', (d) => lineGenerator(d.points) ?? ''); - if (strokeDasharray) { + if (getDasharray) { + merged.attr('stroke-dasharray', (d) => getDasharray(d.key)); + } else if (strokeDasharray) { merged.attr('stroke-dasharray', strokeDasharray); } @@ -78,12 +93,12 @@ export function updateRooflinesForDisplay( zoomGroup: d3.Selection, config: RooflineConfig, ): void { - const { getColor, getOpacity, strokeWidth = 2, strokeDasharray } = config; + const { getColor, getOpacity, strokeWidth = 2 } = config; zoomGroup .selectAll>('.roofline-path') .attr('stroke', (d) => getColor(d.key)) .attr('stroke-width', strokeWidth) - .attr('stroke-dasharray', strokeDasharray ?? null) + .attr('stroke-dasharray', (d) => dasharrayFor(config, d.key)) .each(function (d) { const opacity = getOpacity?.(d.key); if (opacity !== undefined) d3.select(this).style('opacity', opacity); diff --git a/packages/app/src/lib/inference-labels.ts b/packages/app/src/lib/inference-labels.ts index c95883095..b5cbccba4 100644 --- a/packages/app/src/lib/inference-labels.ts +++ b/packages/app/src/lib/inference-labels.ts @@ -40,6 +40,48 @@ export function inferenceFrameworkLabelOverride( return override && Date.now() < override.expiresAt ? override.label : undefined; } +/** Leads every unofficial-run label, in line labels and legend rows alike. */ +export const OVERLAY_LABEL_MARKER = '✕ '; +const RUN_TAG_MAX = 20; +const RUN_TAG_TAIL = 17; + +export interface OverlayRunIdentity { + id: number | string; + branch?: string | null; +} + +/** + * Short, still recognisable name for a run: the whole branch when it is short, + * else its last path segment, else the branch tail (klaud nightlies end in + * `-`). Falls back to the run id when the branch is unknown. + */ +export function shortRunTag(run: OverlayRunIdentity): string { + const branch = run.branch?.trim() || `run ${run.id}`; + if (branch.length <= RUN_TAG_MAX) return branch; + const segment = branch.slice(branch.lastIndexOf('/') + 1); + if (segment.length > 0 && segment.length <= RUN_TAG_MAX) return segment; + return `…${branch.slice(-RUN_TAG_TAIL)}`; +} + +/** ` · ` appended to an overlay line label when other runs draw the same hardware. */ +export function overlayRunTag(run: OverlayRunIdentity): string { + return ` · ${shortRunTag(run)}`; +} + +/** + * Line-label text for an unofficial-run curve. Pills name the hardware, not the + * branch: branch names run to 70+ characters and the legend already carries + * them. The run tag is added only when several overlay runs draw the same + * hardware, so the pills stay distinguishable. + */ +export function getOverlayLineLabel( + hardwareLabel: string, + run: OverlayRunIdentity, + sharesHardware: boolean, +): string { + return `${OVERLAY_LABEL_MARKER}${hardwareLabel}${sharesHardware ? overlayRunTag(run) : ''}`; +} + /** Keep unofficial-run identity/markers while making its special engine visible. */ export function getInferenceRunLabel( label: string, diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts index c968537ae..c8fb4c26e 100644 --- a/packages/app/src/lib/modeled-system-power.test.ts +++ b/packages/app/src/lib/modeled-system-power.test.ts @@ -222,6 +222,89 @@ describe('modeled system power admission and accounting', () => { expect(modelSystemPower(source)).toMatchObject({ reason: 'role-power' }); }); + it('models an aggregate multinode deployment without worker telemetry at the deployment mean', () => { + // Kimi K3 B200 dynamo-vLLM TP8/PP2 (prod rows, 2026-09-17): two eight-GPU hosts, + // aggregate producer, no per-worker array. The K3 H200 vLLM row below is + // TP16 × 2 DP replicas across four hosts. + const b200 = row({ + is_multinode: true, + num_prefill_gpu: 16, + num_decode_gpu: 16, + decode_num_workers: 1, + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 715.095, + avg_total_gpu_power_w: 11441.513, + decode_pp: 2, + }, + }); + const perChassis = estimateChassisPower('b200', 11441.513 / 2, 1.3)!; + const estimate = modelSystemPower(b200); + expect(estimate).toMatchObject({ + status: 'supported', + topologyBasis: 'uniform-hosts', + chassisBasis: 'full', + gpuCount: 16, + chassisCount: 2, + modeledGpuCount: 16, + }); + if (estimate.status !== 'supported') throw new Error('unreachable'); + expect(estimate.chassisAcWatts).toBeCloseTo(perChassis.chassisAcWatts * 2, 6); + expect(estimate.deploymentFacilityWatts).toBe(estimate.facilityWatts); + expect(estimate.chassisAcWattsPerGpu).toBeCloseTo(perChassis.chassisAcWatts / 8, 6); + + const h200 = row({ + hardware: 'h200', + is_multinode: true, + prefill_tp: 16, + decode_tp: 16, + decode_ep: 32, + decode_dp_attention: true, + decode_num_workers: 2, + num_prefill_gpu: 32, + num_decode_gpu: 32, + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 167.357, + avg_total_gpu_power_w: 5355.413, + decode_pp: 1, + }, + }); + expect(modelSystemPower(h200)).toMatchObject({ + status: 'supported', + topologyBasis: 'uniform-hosts', + chassisCount: 4, + gpuCount: 32, + }); + + // A replica count that does not explain the telemetry width is not guessed around. + expect(modelSystemPower({ ...h200, decode_num_workers: 1 })).toMatchObject({ + reason: 'gpu-count', + }); + // Twelve GPUs cannot fill whole eight-GPU hosts; placement is unknown. + const twelve = row({ is_multinode: true, decode_tp: 12, prefill_tp: 12 }); + twelve.metrics.avg_total_gpu_power_w = twelve.metrics.avg_power_w * 12; + expect(modelSystemPower(twelve)).toMatchObject({ reason: 'topology' }); + // Per-worker telemetry, when present, keeps the more exact worker path. + const withWorkers = { + ...b200, + workers: ['host-a', 'host-b'].map((host, worker_idx) => ({ + role: 'agg', + worker_idx, + hosts: [host], + num_gpus: 8, + avg_power_w: 715.095, + })), + }; + expect(modelSystemPower(withWorkers)).toMatchObject({ + status: 'supported', + topologyBasis: 'worker-hosts', + chassisCount: 2, + }); + }); + it('preserves meaningful aggregate PP and PCP aliases before checking physical width', () => { for (const widths of [ { decode_pp: 1, prefill_pp: 2 }, diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts index 550fe23fc..aadb737f6 100644 --- a/packages/app/src/lib/modeled-system-power.ts +++ b/packages/app/src/lib/modeled-system-power.ts @@ -43,7 +43,13 @@ export type SystemPowerEstimate = deploymentFacilityWatts: number; pue: number; telemetryBasis: 'validated-v2' | 'validated-unversioned-single-node'; - topologyBasis: 'single-node' | 'worker-hosts'; + /** + * 'single-node': one host, one chassis. 'worker-hosts': one chassis per + * measured worker, each at its own telemetry. 'uniform-hosts': an + * aggregate multinode deployment whose producer emitted no per-worker + * telemetry; every eight-GPU chassis is modeled at the deployment mean. + */ + topologyBasis: 'single-node' | 'worker-hosts' | 'uniform-hosts'; /** * 'full': every chassis had all eight GPUs measured. 'extrapolated': at least * one chassis was partially allocated; its model input is the measured per-GPU @@ -138,7 +144,7 @@ export function modelSystemPower( } const chassis: MeasuredChassis[] = []; - let topologyBasis: 'single-node' | 'worker-hosts'; + let topologyBasis: 'single-node' | 'worker-hosts' | 'uniform-hosts'; if (row.disagg === false && row.is_multinode === false) { // One host cannot hold more than one chassis. if (gpuCount > CHASSIS_GPU_COUNT) return unavailable('topology'); @@ -173,6 +179,36 @@ export function modelSystemPower( ? m.avg_total_gpu_power_w : m.avg_power_w * CHASSIS_GPU_COUNT, }); + } else if (row.disagg === false && (!Array.isArray(row.workers) || row.workers.length === 0)) { + // Aggregate multinode producers emit no per-worker telemetry. Symmetric + // TP/PP/DP shards load every host alike, so each full eight-GPU chassis is + // modeled at the deployment mean; the supported hardware only ships in + // eight-GPU hosts, so the count must fill whole chassis on several hosts. + // Disaggregated roles differ in load and stay on the worker path. + const hostCount = gpuCount / CHASSIS_GPU_COUNT; + if (!count(hostCount) || hostCount < 2) return unavailable('topology'); + const tp = row.decode_tp > 0 ? row.decode_tp : row.prefill_tp; + const pp = Math.max(m.pp ?? 1, m.decode_pp ?? 1, m.prefill_pp ?? 1); + const pcp = Math.max(m.pcp_size ?? 1, m.decode_pcp_size ?? 1, m.prefill_pcp_size ?? 1); + // Data-parallel replicas widen the deployment beyond one TP×PP×PCP group. + const replicas = Math.max(1, row.decode_num_workers); + if ( + !count(tp) || + !count(pp) || + !count(pcp) || + !count(replicas) || + tp * pp * pcp * replicas !== gpuCount + ) { + return unavailable('gpu-count'); + } + topologyBasis = 'uniform-hosts'; + for (let host = 0; host < hostCount; host++) { + chassis.push({ + measuredGpus: CHASSIS_GPU_COUNT, + // Partition the producer's exact total so the chassis inputs sum back to it. + modelInputWatts: m.avg_total_gpu_power_w / hostCount, + }); + } } else { // A role average across several hosts is insufficient for nonlinear // fan/PSU evaluation. Require one chassis per measured worker and a diff --git a/packages/app/src/lib/power-basis.test.ts b/packages/app/src/lib/power-basis.test.ts new file mode 100644 index 000000000..0aa67a4f5 --- /dev/null +++ b/packages/app/src/lib/power-basis.test.ts @@ -0,0 +1,73 @@ +import { describe, expect, it } from 'vitest'; + +import type { BenchmarkRow } from '@/lib/api'; +import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform'; +import { buildDerivedChartFields, getHardwareKey } from '@/lib/chart-utils'; +import { POWER_BASIS_FIELDS } from '@/lib/power-basis'; + +// Qwen3.5 B200 c1, run 34175132645: actual rounded telemetry, eight GPUs +// (same fixture as modeled-system-power.test.ts) plus an output rate. +function row(overrides: Partial = {}): BenchmarkRow { + return { + id: 441192, + model: 'qwen3.5', + hardware: 'b200', + framework: 'sglang', + precision: 'fp8', + spec_method: 'none', + disagg: false, + is_multinode: false, + prefill_tp: 8, + prefill_ep: 1, + prefill_dp_attention: false, + prefill_num_workers: 0, + decode_tp: 8, + decode_ep: 1, + decode_dp_attention: false, + decode_num_workers: 0, + num_prefill_gpu: 8, + num_decode_gpu: 8, + benchmark_type: 'single_turn', + isl: 8192, + osl: 1024, + conc: 1, + offload_mode: 'off', + image: 'lmsysorg/sglang:v0.5.19-cu130', + date: '2026-09-08', + run_url: 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34175132645/attempts/1', + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 349.859, + avg_total_gpu_power_w: 2798.868, + total_gpu_energy_j: 120361.299, + joules_per_output_token: 12.937902, + tput_per_gpu: 270, + output_tput_per_gpu: 30, + pp: 1, + pcp_size: 1, + median_intvty: 100, + }, + ...overrides, + }; +} + +/** Full official-path derivation for one row with real HW_REGISTRY specs. */ +function derive(source: BenchmarkRow) { + const entry = rowToAggDataEntry(source); + const hwKey = getHardwareKey(entry); + return { entry, hwKey, fields: buildDerivedChartFields(entry, hwKey) }; +} + +describe('power boundaries through the derived-field builder', () => { + it('serves the same fields to ?unofficialrun= overlays through transformBenchmarkRows', () => { + const { chartData } = transformBenchmarkRows([row()], 'median', 'external'); + const point = chartData[0][0]; + const official = derive(row()).fields; + for (const basis of Object.values(POWER_BASIS_FIELDS)) { + expect(point[basis.watts]).toEqual(official[basis.watts]); + expect(point[basis.energy]).toEqual(official[basis.energy]); + expect(point[basis.watts]?.roof).toBe(false); + } + }); +}); diff --git a/packages/app/src/lib/power-basis.ts b/packages/app/src/lib/power-basis.ts new file mode 100644 index 000000000..5b5f6e61b --- /dev/null +++ b/packages/app/src/lib/power-basis.ts @@ -0,0 +1,223 @@ +/** + * Power boundaries for one benchmark point, from the GPU board out to the + * utility meter. Pure numbers only — no React, DOM, or registry imports — so + * the derived-field builder, overlays, and tests share one formula set. + * + * | Basis | W / GPU | J / output token | + * | ---------------------- | ------------------------------------------------ | ----------------------------------------- | + * | B1 gpu-measured | `avg_power_w` (existing `measuredAvgPower`) | `joules_per_output_token` (existing) | + * | B2 gpu-provisioned | `HW_REGISTRY.tdp` | W × N_alloc ÷ total output tok/s | + * | B3 utility-provisioned | `HW_REGISTRY.power` × 1000 | W × N_alloc ÷ total output tok/s | + * | B4 utility-modeled | modeled `deploymentFacilityWatts` ÷ `gpuCount` | B1 J/out × (B4 W ÷ B1 W) | + * + * N_alloc counts every allocated GPU (prefill + decode for disaggregation); + * total output tok/s is the whole deployment's. B4 reuses the estimate that + * `modelSystemPower` already attached to the entry — PUE is applied exactly + * once inside that model, never here — and is withheld wherever B1 is. + * Unavailable values are `null`; callers omit the field rather than plotting 0. + */ +import type { AggDataEntry, InferenceData, PowerBasisFieldKey } from '@/components/inference/types'; + +export const POWER_BASES = [ + 'gpu-measured', + 'gpu-provisioned', + 'utility-provisioned', + 'utility-modeled', +] as const; +export type PowerBasis = (typeof POWER_BASES)[number]; +export type PowerQuantity = 'watts' | 'energy'; + +export const POWER_BASIS_LABELS: Record = { + 'gpu-measured': { en: 'GPU Level Measured', zh: 'GPU 实测功耗' }, + 'gpu-provisioned': { en: 'GPU Level Provisioned (TDP)', zh: 'GPU 额定功耗(TDP)' }, + 'utility-provisioned': { en: 'All in Provisioned', zh: '整体预配功耗' }, + 'utility-modeled': { en: 'All in Measured', zh: '整体实测功耗' }, +}; + +export const ALL_IN_MEASURED_NOTE = { + en: 'GPU power is measured; unmeasured components are modeled, with PUE included.', + zh: 'GPU 功耗来自实测;未实测的组件功耗由模型估算,并计入数据中心 PUE。', +}; + +export const ALL_IN_MEASURED_EMPTY = { + en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K, validated GPU telemetry, and hardware covered by the chassis power model (not NVL72 systems). Choose another boundary to keep the points.', + zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 场景、已验证的 GPU 遥测,且硬件在机箱功耗模型覆盖范围内(不含 NVL72 系统)。可切换到其他功耗边界以保留数据点。', +}; + +/** InferenceData keys per derived basis and quantity. B1 lives on the measured* fields. */ +export const POWER_BASIS_FIELDS: Record< + Exclude, + Record +> = { + 'gpu-provisioned': { + watts: 'gpuProvisionedWatts', + energy: 'gpuProvisionedJPerOutputToken', + }, + 'utility-provisioned': { + watts: 'utilityProvisionedWatts', + energy: 'utilityProvisionedJPerOutputToken', + }, + 'utility-modeled': { + watts: 'utilityModeledWatts', + energy: 'utilityModeledJPerOutputToken', + }, +}; + +export interface PowerBasisInput { + /** B2 W/GPU: HW_REGISTRY tdp. 0 means the spec is not yet available. */ + tdpWatts: number | null; + /** B3 W/GPU: HW_REGISTRY all-in power, already in watts. */ + utilityWatts: number | null; + /** Every GPU the deployment occupies (prefill + decode for disaggregation). */ + allocatedGpus: number | null; + /** Whole-deployment successful output tokens per second. */ + totalOutputTokPerSec: number | null; + /** B1 W/GPU from validated telemetry. */ + measuredWatts: number | null; + /** B1 J/output token from the same telemetry window. */ + measuredJPerOutputToken: number | null; + /** + * B4 W/GPU: modeled facility watts (PUE already applied) per measured GPU. + * B4 is a scaling of B1, so it is withheld whenever `measuredWatts` is null. + */ + modeledFacilityWattsPerGpu: number | null; +} + +export type PowerBasisValues = Record; + +export const isPositive = (value: unknown): value is number => + typeof value === 'number' && Number.isFinite(value) && value > 0; +const count = (value: unknown): value is number => isPositive(value) && Number.isSafeInteger(value); +const orNull = (value: number): number | null => (isPositive(value) ? value : null); + +/** + * Derives the B2–B4 boundary values from plain numbers. Any unavailable input + * yields `null` for the values that depend on it and leaves the rest intact. + */ +export function computePowerBasisFields(input: PowerBasisInput): PowerBasisValues { + const tdp = isPositive(input.tdpWatts) ? input.tdpWatts : null; + const utility = isPositive(input.utilityWatts) ? input.utilityWatts : null; + const measuredWatts = isPositive(input.measuredWatts) ? input.measuredWatts : null; + const measuredJ = isPositive(input.measuredJPerOutputToken) + ? input.measuredJPerOutputToken + : null; + // B4 is B1 carried out to the utility meter, so it follows B1's availability: + // no measured watts, no modeled boundary (B3 ≥ B4 ≥ B1 needs its anchor). + const modeled = + measuredWatts !== null && isPositive(input.modeledFacilityWattsPerGpu) + ? input.modeledFacilityWattsPerGpu + : null; + + // Provisioned energy: GPU-seconds spent per output token by the whole + // deployment (N_alloc ÷ total tok/s) × W per GPU = J per output token. + const gpuSecondsPerOutputToken = + isPositive(input.allocatedGpus) && isPositive(input.totalOutputTokPerSec) + ? input.allocatedGpus / input.totalOutputTokPerSec + : null; + const provisionedEnergy = (watts: number | null) => + watts !== null && gpuSecondsPerOutputToken !== null + ? orNull(watts * gpuSecondsPerOutputToken) + : null; + + // Modeled energy scales the producer's same-window E/N by modeled ÷ measured W, + // so it inherits B1's token denominator instead of re-deriving one. + const modeledEnergy = + modeled !== null && measuredWatts !== null && measuredJ !== null + ? orNull((measuredJ * modeled) / measuredWatts) + : null; + + return { + gpuProvisionedWatts: tdp, + gpuProvisionedJPerOutputToken: provisionedEnergy(tdp), + utilityProvisionedWatts: utility, + utilityProvisionedJPerOutputToken: provisionedEnergy(utility), + utilityModeledWatts: modeled, + utilityModeledJPerOutputToken: modeledEnergy, + }; +} + +type PowerBasisEntry = Pick< + AggDataEntry, + | 'output_tput_per_gpu' + | 'disagg' + | 'benchmark_type' + | 'num_prefill_gpu' + | 'num_decode_gpu' + | 'avg_power_w' + | 'joules_per_output_token' + | 'modeledSystemPower' +>; + +/** + * Whole-deployment normalization for the provisioned energies. Aggregate rows + * already report output per allocated GPU, so N_alloc cancels and the ratio + * 1 GPU : per-GPU throughput is exact without trusting display counts (legacy + * ingest can encode TP × EP twice). Fixed-sequence disaggregated rows report + * output per decode GPU while the deployment also powers the prefill pool, so + * total output = per-GPU × decode GPUs and N_alloc = prefill + decode GPUs. + * Other disaggregated benchmark types are left out: whether AgentX throughput + * already divides by all GPUs is not verifiable in-app. + */ +export function powerBasisNormalization( + entry: Pick< + PowerBasisEntry, + 'output_tput_per_gpu' | 'disagg' | 'benchmark_type' | 'num_prefill_gpu' | 'num_decode_gpu' + >, +): Pick { + const perGpu = entry.output_tput_per_gpu; + const unavailable = { allocatedGpus: null, totalOutputTokPerSec: null }; + if (!isPositive(perGpu)) return unavailable; + if (!entry.disagg) return { allocatedGpus: 1, totalOutputTokPerSec: perGpu }; + if (entry.benchmark_type !== 'single_turn') return unavailable; + const prefill = entry.num_prefill_gpu; + const decode = entry.num_decode_gpu; + if (!count(prefill) || !count(decode)) return unavailable; + return { allocatedGpus: prefill + decode, totalOutputTokPerSec: perGpu * decode }; +} + +/** + * B4 W/GPU from the estimate `rowToAggDataEntry` attached. The model owns + * telemetry admission: `modelSystemPower` requires `power_valid === 1` plus + * schema v2, or the validated unversioned single-node producer it records as + * `telemetryBasis: 'validated-unversioned-single-node'`. That is the same + * population the app plots as B1 (`measuredAvgPower`) and as + * `modeledChassisPowerPerGpu`, so B4 renders exactly where they do. The public + * API's stricter `strictV2` row filter is not re-applied here; it is not + * applied to the chart's B1 either. + */ +export function modeledFacilityWattsPerGpu( + entry: Pick, +): number | null { + const model = entry.modeledSystemPower; + if (model?.status !== 'supported') return null; + if (!isPositive(model.deploymentFacilityWatts) || !count(model.gpuCount)) return null; + return orNull(model.deploymentFacilityWatts / model.gpuCount); +} + +export type PowerBasisChartFields = Partial>; + +/** + * Chart-shaped B2–B4 fields for one entry. Keys are present only for finite, + * positive values: the metric filters drop a point by `metricKey in point`, + * and the coordinate remap falls back to raw throughput when a key exists + * with an unusable value. + */ +export function buildPowerBasisChartFields( + entry: PowerBasisEntry, + specs: { tdp?: number; power?: number }, +): PowerBasisChartFields { + const values = computePowerBasisFields({ + tdpWatts: specs.tdp ?? null, + utilityWatts: isPositive(specs.power) ? specs.power * 1000 : null, + ...powerBasisNormalization(entry), + measuredWatts: entry.avg_power_w ?? null, + measuredJPerOutputToken: entry.joules_per_output_token ?? null, + modeledFacilityWattsPerGpu: modeledFacilityWattsPerGpu(entry), + }); + const fields: PowerBasisChartFields = {}; + for (const key of Object.keys(values) as PowerBasisFieldKey[]) { + const y = values[key]; + if (y !== null) fields[key] = { y, roof: false }; + } + return fields; +} diff --git a/packages/app/src/lib/url-state.test.ts b/packages/app/src/lib/url-state.test.ts index 39a1937e0..c10e39e74 100644 --- a/packages/app/src/lib/url-state.test.ts +++ b/packages/app/src/lib/url-state.test.ts @@ -99,6 +99,22 @@ describe('PARAM_DEFAULTS', () => { expect(PARAM_DEFAULTS.i_advlabel).toBe(''); }); + it('drops retired comparison controls from old share links while preserving analysis toggles', async () => { + setupWindow( + '?i_servicecompare=1&i_servicebase=baseline&i_servicepeer=comparator&i_servicetarget=8&i_roleshare=1&i_powerfit=1', + ); + const { readUrlParams, buildShareUrl } = await import('@/lib/url-state'); + const params = readUrlParams(); + const shared = new URL(buildShareUrl()).searchParams; + expect(params).toMatchObject({ i_roleshare: '1', i_powerfit: '1' }); + expect(shared.get('i_roleshare')).toBe('1'); + expect(shared.get('i_powerfit')).toBe('1'); + for (const key of ['i_servicecompare', 'i_servicebase', 'i_servicepeer', 'i_servicetarget']) { + expect(params).not.toHaveProperty(key); + expect(shared.has(key)).toBe(false); + } + }); + it('strips the normalized revenue source but preserves OpenRouter as explicit state', async () => { const { PARAM_DEFAULTS } = await import('@/lib/url-state'); expect(PARAM_DEFAULTS.i_revenue).toBe('normalized'); diff --git a/packages/app/src/lib/url-state.ts b/packages/app/src/lib/url-state.ts index 590a80edc..ef62e9674 100644 --- a/packages/app/src/lib/url-state.ts +++ b/packages/app/src/lib/url-state.ts @@ -30,6 +30,9 @@ const URL_STATE_KEYS = [ // Token-revenue sale-price source: normalized $1/M or live OpenRouter catalog. 'i_revenue', 'i_pctl', + 'i_mstat', + 'i_roleshare', + 'i_powerfit', 'i_xmetric', 'i_e2e_xmetric', 'i_xmode', @@ -63,6 +66,20 @@ const URL_STATE_KEYS = [ 'i_spec', // Measured-power certification tiers ('certified' / 'legacy', comma-joined). 'i_power', + 'i_topology', + 'i_ptlines', + 'i_ptaxis', + 'i_ptwindow', + 'i_ptfocus', + 'i_ptutility', + // Power Timeline concurrency filter: one positive integer, empty = every load. + 'i_ptconc', + // Completed Perf Rulers on the primary inference chart: `isoX|curveA|curveB` + // entries joined by `;` (see serializePerfRulers in d3-chart/layers/perf-ruler). + 'i_rulers', + // Comparison series overlaid on a gated power metric: `boundaries` (every + // power boundary) or `roles` (prefill / decode pools). Empty = the metric alone. + 'i_pcompare', // Exact serving-envelope pair behind an Overview 30-day comparison cell. 'i_overview_current', 'i_overview_baseline', @@ -156,6 +173,9 @@ export const PARAM_DEFAULTS: Record = { i_metric: DEFAULT_Y_AXIS_METRIC, i_revenue: 'normalized', i_pctl: 'p90', + i_mstat: 'median', + i_roleshare: '0', + i_powerfit: '0', i_xmetric: 'p90_ttft', i_e2e_xmetric: 'p90_ttft', i_xmode: '', @@ -183,6 +203,15 @@ export const PARAM_DEFAULTS: Record = { i_disagg: '', i_spec: '', i_power: '', + i_topology: '', + i_ptlines: '', + i_ptaxis: '', + i_ptwindow: '', + i_ptfocus: '', + i_ptutility: '', + i_ptconc: '', + i_rulers: '', + i_pcompare: '', i_overview_current: '', i_overview_baseline: '', e_rundate: '', @@ -500,6 +529,26 @@ export function rememberChartStateInUrl(): string { return chartParams.toString(); } +/** + * The current page's URL carrying its chart state plus `overrides`, + * canonicalised like `rememberChartStateInUrl`: chart params and both + * unofficial-run spellings are dropped from the live address bar before the + * store's state (and the overrides) are layered on. For anchors that must + * work with open-in-new-tab, where the in-memory state would otherwise be lost. + */ +export function chartStateHref(overrides: Record): string { + const { origin, pathname, hash, search } = window.location; + const merged = new URLSearchParams(search); + for (const key of URL_STATE_KEYS) merged.delete(key); + // Collected first: deleting while iterating the params would skip entries. + const staleRunKeys = [...merged.keys()].filter((key) => UNOFFICIAL_RUN_PARAM_RE.test(key)); + for (const key of staleRunKeys) merged.delete(key); + for (const [key, value] of collectTabParams()) merged.set(key, value); + for (const [key, value] of Object.entries(overrides)) merged.set(key, value); + const query = merged.toString(); + return `${origin}${pathname}${query ? `?${query}` : ''}${hash}`; +} + /** * Append the current chart state to an outbound in-app href, so the page it * opens can link back to the chart the user left. Used for the agentic diff --git a/packages/app/src/lib/views-api/calculator-extensions.ts b/packages/app/src/lib/views-api/calculator-extensions.ts index 55262f9cf..0c4624d0f 100644 --- a/packages/app/src/lib/views-api/calculator-extensions.ts +++ b/packages/app/src/lib/views-api/calculator-extensions.ts @@ -18,6 +18,7 @@ import { fetchOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing'; import { cachedJson } from '@/lib/api-cache'; import { getGpuSpecs } from '@/lib/constants'; import { getOpenRouterModelId, Sequence, type Model } from '@/lib/data-mappings'; +import { POWER_BASIS_LABELS } from '@/lib/power-basis'; import type { NextRequest } from 'next/server'; import { runViewsRoute, ViewsApiParamError } from './errors'; import { @@ -217,8 +218,8 @@ export function calculatorExtension(view: CalculatorExtension, request: NextRequ powerBasis, target, { - provisioned: 'Provisioned', - modeled: 'Measured + modeled', + provisioned: POWER_BASIS_LABELS['utility-provisioned'].en, + modeled: POWER_BASIS_LABELS['utility-modeled'].en, extrapolated: 'Full-chassis extrapolation', }, ); diff --git a/packages/app/src/lib/views-api/docs/extensions.ts b/packages/app/src/lib/views-api/docs/extensions.ts index 34ce40379..e05e59e3f 100644 --- a/packages/app/src/lib/views-api/docs/extensions.ts +++ b/packages/app/src/lib/views-api/docs/extensions.ts @@ -130,13 +130,17 @@ const PARAMETER_NOTES: Record = { '模型许可或收入分成百分比,范围 0 至 100,默认值随模型变化。', ], powerBasis: [ - 'provisioned (default), modeled or compare. Modeled power requires eligible measured source rows; estimates extrapolated from partial-GPU measurements to a full chassis are identified by powerLabel. Missing coverage is not zero.', - 'provisioned(默认)、modeled 或 compare。建模功耗需要符合条件的实测数据行;由部分 GPU 的实测数据外推到整机的估算,会通过 powerLabel 标明。缺失数据不按零处理。', + 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. All in Measured uses measured GPU power plus modeled unmeasured components and PUE; it is not measured wall power. Eligible measured source rows are required; powerLabel identifies paired estimates and full-chassis extrapolation. Missing coverage is not zero.', + 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。整体实测功耗采用 GPU 实测值,加上未实测组件的功耗估算和 PUE,并非墙上电表读数。该估算需要符合条件的实测数据行;powerLabel 标明对比方式和整机外推。缺失数据不按零处理。', ], power: [ 'Comma-separated certified and/or legacy power tiers. Omit for all tiers.', '以逗号分隔的 certified、legacy 功率数据等级。省略时选择全部等级。', ], + topologies: [ + 'Comma-separated exact topology keys from the Dashboard topology selector or view point topologyKey. Filters allocation and parallelism without filtering concurrency; omitted means all. Preserve the keys verbatim. Unknown keys return no matching points.', + '逗号分隔的精确拓扑键,取自仪表盘拓扑选择器或数据点的 topologyKey。按 GPU 分配和并行配置筛选,保留全部已测并发;省略时不限制拓扑。原样使用这些键,未知键不会匹配数据点。', + ], allPoints: [ 'Boolean, default false. Include points clipped by dashboard limits; optimal and best still apply independently.', '布尔值,默认为 false。包括被图表范围裁剪的数据点;optimal 和 best 仍独立生效。', @@ -452,14 +456,14 @@ export const operations: ApiOperation[] = Object.entries(NEW_VIEWS).map( view === 'video' ? ' Only already published artifacts are read. Cell, phase, slot and GPU-basis choices select result evidence and normalized serving rates; x/y/cost/workload filters produce computed tradeoff points. Local bundles and arbitrary URLs are excluded. Responses are no-store.' : view === 'gpu-metrics' - ? ' Telemetry reads stored DB series first and falls back to GitHub artifacts when storage is absent. Responses use private, no-store; an upstream database failure remains HTTP 503 rather than an empty dataset. File/host series retain separate identities. Full-record statistics include startup and warmup and use current-version stored per-GPU digests; a current empty or missing metric digest stays empty. Outdated or unversioned digests are recomputed read-only from retained DB samples. Statistics cover every chip in the selected series, while raw rows and charts respect selected GPU indices. Chart downsampling does not alter statistics. Statistics expose the selected metric values without the storage-only metric column. These sample-weighted statistics are separate from serving-window power, J/token and selected-time-window calculations. Line time remains relative to the first sample across all chips; missing metric readings are omitted rather than zero-filled. Correlations require both readings.' + ? ' Telemetry reads stored DB series first and falls back to GitHub artifacts when storage is absent. Responses use private, no-store; an upstream database failure remains HTTP 503 rather than an empty dataset. File/host series retain separate identities. Full-record statistics include startup and warmup and use current-version stored per-GPU digests; a current empty or missing metric digest stays empty. Outdated or unversioned digests are recomputed read-only from retained DB samples. Statistics cover every chip in the selected series, while raw rows and charts respect selected GPU indices. Chart downsampling does not alter statistics. Statistics expose the selected metric values without the storage-only metric column. These sample-weighted statistics are separate from serving-window power, J/token and selected-time-window calculations. Line time remains relative to the first sample across all chips; missing metric readings are omitted rather than zero-filled. Correlations require both readings. This endpoint projects the raw GPU-metrics explorer, not the inference Power Timeline. Timeline role/window evidence comes from its existing series=power read plus benchmark power_audit. Timeline i_ptaxis, i_ptlines, i_ptwindow, i_ptfocus and i_ptutility share fields control client rendering only and are not accepted here; they do not change raw telemetry or full-record statistics.' : '' }`, `只读${zh},使用仪表板的数据读取和计算函数。未知或重复查询键返回 400;响应包含解析后的参数,保留缺失数据。仅影响样式的控件不作为 API 参数。${ view === 'video' ? ' 仅读取已发布产物。cell、阶段、slot 和 GPU 口径选择对应结果证据,并计算 serving 归一化速率;x/y、成本及工作负载筛选生成权衡图数据点。不读取本地数据包或任意 URL。响应不缓存。' : view === 'gpu-metrics' - ? ' 遥测优先读取数据库中已存储的序列;缺少存储数据时回退到 GitHub 产物。响应使用 private, no-store;上游数据库故障保留 HTTP 503,不作为空数据返回。各文件、主机的序列身份独立保留。全记录统计包含服务启动与 warmup,已有数据使用当前算法版本的每 GPU 统计摘要;当前摘要为空或缺少所选指标时仍返回空统计数组。旧版本或无版本摘要从保留的 DB 样本只读重算。统计覆盖所选序列的全部芯片,原始数据行和图表则按芯片索引筛选;图表降采样不改变统计。统计项只包含所选指标的数值,不返回数据库内部的 metric 字段。这里按样本计算的统计与 serving-window 功率、J/token 及用户所选时间窗口的计算分别处理。折线时间以所有芯片的首个采样为起点;缺失指标读数会被跳过,不补零。相关性图要求两个指标均有读数。' + ? ' 遥测优先读取数据库中已存储的序列;缺少存储数据时回退到 GitHub 产物。响应使用 private, no-store;上游数据库故障保留 HTTP 503,不作为空数据返回。各文件、主机的序列身份独立保留。全记录统计包含服务启动与 warmup,已有数据使用当前算法版本的每 GPU 统计摘要;当前摘要为空或缺少所选指标时仍返回空统计数组。旧版本或无版本摘要从保留的 DB 样本只读重算。统计覆盖所选序列的全部芯片,原始数据行和图表则按芯片索引筛选;图表降采样不改变统计。统计项只包含所选指标的数值,不返回数据库内部的 metric 字段。这里按样本计算的统计与 serving-window 功率、J/token 及用户所选时间窗口的计算分别处理。折线时间以所有芯片的首个采样为起点;缺失指标读数会被跳过,不补零。相关性图要求两个指标均有读数。 本接口投影原始 GPU-metrics 浏览器,不是推理页的 Power Timeline。时间线通过现有 series=power 读取和基准测试 power_audit 获取角色与窗口证据;i_ptaxis、i_ptlines、i_ptwindow、i_ptfocus 和 i_ptutility 分享字段仅控制客户端显示,本接口不接受这些参数。它们不改变原始遥测或全记录统计。' : '' }`, ), diff --git a/packages/app/src/lib/views-api/docs/inference.ts b/packages/app/src/lib/views-api/docs/inference.ts index eeee1abaf..9fc1e6bf1 100644 --- a/packages/app/src/lib/views-api/docs/inference.ts +++ b/packages/app/src/lib/views-api/docs/inference.ts @@ -21,7 +21,13 @@ import { const VIEWS_GROUP: ApiOperation['group'] = 'views'; -const X_MODE_ENUM = ['interactivity', 'ttft', 'e2e', 'e2e-normalized-interactivity'] as const; +const X_MODE_ENUM = [ + 'interactivity', + 'ttft', + 'e2e', + 'e2e-normalized-interactivity', + 'concurrency', +] as const; const parameters: readonly ApiParameter[] = [ { @@ -82,12 +88,24 @@ const parameters: readonly ApiParameter[] = [ required: false, type: 'enum', description: text( - 'X-axis mode. e2e-normalized-interactivity uses persisted derived AgentX metrics; points without eligible derived values are omitted.', - 'X 轴模式。e2e-normalized-interactivity 使用已持久化的 AgentX 派生指标;没有合格派生值的数据点不参与此视图。', + 'X-axis mode. concurrency uses observed load levels, with no interpolation or optimization ranking; optimal and best resolve to false. e2e-normalized-interactivity uses persisted derived AgentX metrics; points without eligible derived values are omitted.', + 'X 轴模式。concurrency 使用实测并发值,不插值、不作优化排名;optimal 和 best 均解析为 false。e2e-normalized-interactivity 使用已持久化的 AgentX 派生指标;没有合格派生值的数据点不参与此视图。', ), schema: { type: 'string', enum: X_MODE_ENUM, default: 'interactivity' }, example: 'e2e', }, + { + name: 'xstat', + location: 'query', + required: false, + type: 'enum', + description: text( + 'Fixed-sequence service-axis statistic: median (default) or mean. Mean streaming speed is 1 / mean TPOT, not the arithmetic mean of per-request speeds. Mean TTFT and E2E use their recorded mean fields. Missing values are omitted, never replaced with median. Ignored for AgentX and concurrency; params.xstat then resolves to null and xAxis.statistic records the effective percentile or null.', + '固定长度工作负载服务轴的统计量:median(默认)或 mean。Mean streaming speed 为 1 / mean TPOT,不是各请求速度的算术平均;mean TTFT 和 E2E 使用各自记录的均值。缺失时不回退到 median。AgentX 和 concurrency 不使用此参数,params.xstat 为 null,xAxis.statistic 返回实际分位数或 null。', + ), + schema: { type: 'string', enum: ['mean', 'median'], default: 'median' }, + example: 'mean', + }, { name: 'xmetric', location: 'query', @@ -197,14 +215,86 @@ const parameters: readonly ApiParameter[] = [ schema: { type: 'boolean', default: true }, example: 'true', }, + { + name: 'serviceCompare', + location: 'query', + required: false, + type: 'boolean', + description: text( + 'API-only analysis: include source options, equal-service percentage curves, an optional target comparison and the same-concurrency diagnostic table (matchedConcurrency). Uses scoped observed points before frontier/best pruning. No corresponding dashboard control or panel. JSON only.', + '仅通过 API 提供的分析:返回来源选项、同等服务条件下的百分比对比曲线、可选目标值对比,以及相同并发下的诊断表(matchedConcurrency)。使用筛选后、前沿和 best 筛选前的实测点。仪表板不提供对应控件或面板。仅支持 JSON。', + ), + schema: { type: 'boolean', default: false }, + example: 'true', + }, + { + name: 'serviceBaseline', + location: 'query', + required: false, + type: 'string', + description: text( + 'Exact baseline key from serviceSources. Omitted selects the first deterministic source; an unknown explicit key remains unavailable. Percent change is 100 × (comparator / baseline − 1).', + 'serviceSources 中的完整基准来源键。省略时按确定性顺序选择首项;显式未知键保持不可用。变化百分比为 100 ×(对比值 / 基准值 − 1)。', + ), + schema: stringSchema, + example: 'Exact key returned in serviceSources', + }, + { + name: 'serviceComparator', + location: 'query', + required: false, + type: 'string', + description: text( + 'Exact comparator key from serviceSources. Omitted selects the second deterministic source. Keys retain hardware, curve snapshot, recipe, image, topology and workload identity. Stitched telemetry producer/exporter hashes do not split a snapshot; rows without one retain their own run, date and hashes. Never substitute a hardware name.', + 'serviceSources 中的完整对比来源键。省略时按确定性顺序选择第二项。来源键保留硬件、曲线快照、测试配置指纹、镜像、拓扑和工作负载标识。同一快照内不同 telemetry producer/exporter 的 hash 不拆分来源;没有快照的行仍按原运行、日期和 hash 区分。不能用硬件名称代替。', + ), + schema: stringSchema, + example: 'Exact key returned in serviceSources', + }, + { + name: 'serviceTarget', + location: 'query', + required: false, + type: 'number', + description: text( + 'Positive finite service-axis target: tok/s/user for streaming speed, seconds for TTFT/E2E. Omitted returns the curve and null target comparison. Numerical linear interpolation is bounded by each exact source; no extrapolation or interpolation across missing metric endpoints. Concurrency is unsupported.', + '有限正数服务轴目标:streaming speed 单位为 tok/s/user,TTFT/E2E 单位为秒。省略时返回曲线,目标值对比为 null。仅在各完整来源的实测范围内做数值线性插值,不外推、不跨越缺失指标端点。并发轴不适用。', + ), + schema: { type: 'number', minimum: Number.MIN_VALUE }, + example: 40, + }, + { + name: 'roleShare', + location: 'query', + required: false, + type: 'boolean', + description: text( + 'Include validated disaggregated prefill/decode energy shares on one output-token denominator, plus rolePoints with each pool’s W/GPU and role-local energy. Uses same-window aggregate J/output ÷ J/input to convert prefill J/input; share denominator is reconstructed prefill + decode energy, not pool-local token counts. Missing/invalid data is omitted or null. JSON only.', + '返回通过验证的分离式 prefill/decode 能耗占比,统一使用 output token 分母;rolePoints 另含各池的 W/GPU 和按本池 token 计的能耗。用同窗口总 J/output ÷ J/input 将 prefill J/input 转换为 J/output;占比分母为重建的 prefill + decode 能耗,不使用各池独立的 token 数。缺失或无效的数据会被省略或返回 null,仅支持 JSON。', + ), + schema: { type: 'boolean', default: false }, + example: 'true', + }, + { + name: 'powerFit', + location: 'query', + required: false, + type: 'boolean', + description: text( + 'Include one ordinary least-squares line per source: measured mean W/GPU = P0 + m × output tok/s per allocated GPU (disaggregated output spread over prefill and decode GPUs). P0 is the zero-output intercept, not measured idle power; m is marginal J/output token. Sources with fewer than 3 distinct output rates return observations with fit null. JSON only.', + '按数据源分别返回普通最小二乘拟合:实测平均 W/GPU = P0 + m × 每个已分配 GPU 的输出 tok/s(分离式部署的输出量均摊到 prefill 和 decode GPU)。P0 是零输出截距,不是实测空载功耗;m 是每个输出 token 的边际能耗(J)。不同输出速率少于 3 个的数据源返回观测点,fit 为 null。仅支持 JSON。', + ), + schema: { type: 'boolean', default: false }, + example: 'true', + }, { name: 'format', location: 'query', required: false, type: 'enum', description: text( - 'Response encoding. csv returns one flat row per point.', - '响应编码。csv 为每个数据点返回一行平面数据。', + 'Response encoding. csv returns one flat row per plotted point; serviceCompare, roleShare and powerFit analytical results require JSON and return 400 with CSV.', + '响应编码。csv 为每个图表点返回一行平面数据;serviceCompare、roleShare 和 powerFit 分析结果仅支持 JSON,与 CSV 同用时返回 400。', ), schema: { type: 'string', enum: ['json', 'csv'], default: 'json' }, example: 'csv', @@ -218,6 +308,7 @@ const pointSchema = objectSchema( x: numberSchema, y: numberSchema, concurrency: numberSchema, + topologyKey: stringSchema, tp: numberSchema, date: { type: 'string', format: 'date' }, runId: integerSchema, @@ -225,7 +316,7 @@ const pointSchema = objectSchema( bestPerSku: booleanSchema, metrics: { type: 'object', additionalProperties: numberSchema }, }, - ['x', 'y', 'concurrency', 'tp', 'date', 'frontier', 'bestPerSku', 'metrics'], + ['x', 'y', 'concurrency', 'topologyKey', 'tp', 'date', 'frontier', 'bestPerSku', 'metrics'], ); const seriesSchema = objectSchema( @@ -254,6 +345,79 @@ const seriesSchema = objectSchema( ], ); +const sourceSchema = objectSchema({ key: stringSchema, label: stringSchema }, ['key', 'label']); +const identitySchema = objectSchema({ + id: { type: ['integer', 'null'] }, + sourceKey: stringSchema, + hwKey: stringSchema, + precision: stringSchema, + concurrency: numberSchema, + topologyKey: stringSchema, + date: stringSchema, + runUrl: { type: ['string', 'null'] }, + recipeFingerprint: { type: ['string', 'null'] }, + image: { type: ['string', 'null'] }, +}); +const estimateSchema = { + ...objectSchema({ + value: numberSchema, + interpolated: booleanSchema, + endpoints: arraySchema( + objectSchema({ x: numberSchema, value: numberSchema, point: identitySchema }), + ), + }), + type: ['object', 'null'] as const, +}; +const serviceMetricSchema = objectSchema({ + baseline: estimateSchema, + comparator: estimateSchema, + changePercent: { type: ['number', 'null'] }, + reason: stringSchema, +}); +const comparisonSchema = objectSchema({ + target: numberSchema, + xField: stringSchema, + baseline: { ...sourceSchema, type: ['object', 'null'] }, + comparator: { ...sourceSchema, type: ['object', 'null'] }, + reason: stringSchema, + metrics: objectSchema({ + meanWattsPerGpu: { ...serviceMetricSchema, description: 'Mean measured GPU board W/GPU.' }, + outputTokensPerSecond: { + ...serviceMetricSchema, + description: + 'Whole-deployment output tokens/s; disaggregated GPU-count normalization preserves the workload denominator.', + }, + joulesPerOutputToken: { + ...serviceMetricSchema, + description: 'Validated measured GPU joules per output token.', + }, + }), +}); + +const nullableNumber = (description: string) => ({ + type: ['number', 'null'] as const, + description, +}); +const matchedValuesSchema = objectSchema({ + joulesPerOutputToken: { type: ['number', 'null'] }, + meanWattsPerGpu: { type: ['number', 'null'] }, + interactivity: { type: ['number', 'null'] }, +}); +const matchedSideSchema = objectSchema( + { + status: { + type: 'string', + enum: ['observed', 'missing', 'ambiguous'], + description: + 'observed: one agreeing observation; missing: none at this concurrency; ambiguous: conflicting observations, none chosen.', + }, + values: matchedValuesSchema, + point: identitySchema, + points: arraySchema(identitySchema), + }, + ['status'], +); + const responseSchema = objectSchema( { view: { type: 'string', enum: ['inference'] }, @@ -274,12 +438,17 @@ const responseSchema = objectSchema( }, ['key', 'configKey', 'label', 'labelZh'], ), - xAxis: objectSchema({ mode: stringSchema, field: stringSchema, label: stringSchema }), + xAxis: objectSchema({ + mode: stringSchema, + field: stringSchema, + label: stringSchema, + statistic: { type: ['string', 'null'] }, + }), frontier: objectSchema({ direction: { type: ['string', 'null'], description: - 'Selected boundary direction. Measured-power gauges use upper_right for interactivity or upper_left for latency, independently of metric.direction.', + 'Selected boundary direction; null for observed concurrency, which has no preferred direction. Measured-power gauges use upper_right for interactivity or upper_left for latency, independently of metric.direction.', }, points: integerSchema, }), @@ -291,6 +460,91 @@ const responseSchema = objectSchema( ), series: arraySchema(seriesSchema), count: integerSchema, + serviceSources: arraySchema(sourceSchema), + equalServiceComparison: { ...comparisonSchema, type: ['object', 'null'] }, + equalServiceCurve: arraySchema(comparisonSchema), + roleEnergyShares: arraySchema( + objectSchema({ + x: numberSchema, + sourceKey: stringSchema, + point: identitySchema, + prefill: { ...numberSchema, description: 'Prefill energy in J/output token.' }, + decode: { ...numberSchema, description: 'Decode energy in J/output token.' }, + total: { + ...numberSchema, + description: 'Reconstructed prefill + decode energy in J/output token.', + }, + prefillShare: { + ...numberSchema, + description: 'Prefill percentage of reconstructed total.', + }, + decodeShare: { ...numberSchema, description: 'Decode percentage of reconstructed total.' }, + }), + ), + matchedConcurrency: objectSchema({ + baseline: { ...sourceSchema, type: ['object', 'null'] }, + comparator: { ...sourceSchema, type: ['object', 'null'] }, + interactivityField: stringSchema, + reason: { type: 'string', enum: ['same-source', 'unknown-source'] }, + rows: arraySchema( + objectSchema({ + concurrency: integerSchema, + baseline: matchedSideSchema, + comparator: matchedSideSchema, + changePercent: { + ...matchedValuesSchema, + description: + '100 × (comparator / baseline − 1); null unless both values were observed.', + }, + }), + ), + }), + rolePoints: arraySchema( + objectSchema({ + x: numberSchema, + sourceKey: stringSchema, + point: identitySchema, + prefillWattsPerGpu: nullableNumber('Mean W/GPU inside the prefill pool.'), + decodeWattsPerGpu: nullableNumber('Mean W/GPU inside the decode pool.'), + prefillJoulesPerInputToken: nullableNumber('Prefill-pool joules per input token.'), + decodeJoulesPerOutputToken: nullableNumber('Decode-pool joules per output token.'), + energy: { + ...objectSchema({ + prefill: numberSchema, + decode: numberSchema, + total: numberSchema, + prefillShare: numberSchema, + }), + type: ['object', 'null'], + description: 'Both pools on the output-token denominator (J/output token, share %).', + }, + }), + ), + powerFits: arraySchema( + objectSchema({ + source: sourceSchema, + tdpWatts: nullableNumber('Rated board TDP per GPU; P0 ÷ TDP = fit.intercept / tdpWatts.'), + fit: { + ...objectSchema({ + intercept: { ...numberSchema, description: 'P0, W/GPU at zero output (extrapolated).' }, + slope: { ...numberSchema, description: 'm, J per output token.' }, + rSquared: { type: ['number', 'null'] }, + n: integerSchema, + xMin: numberSchema, + xMax: numberSchema, + }), + type: ['object', 'null'], + }, + reason: { type: 'string', enum: ['too-few-points'] }, + observations: arraySchema( + objectSchema({ + x: { ...numberSchema, description: 'Output tok/s per allocated GPU.' }, + y: { ...numberSchema, description: 'Measured mean W/GPU.' }, + point: identitySchema, + }), + ), + }), + ), pricing: { type: ['object', 'null'], additionalProperties: true }, comparisons: arraySchema({ type: 'object', additionalProperties: true }), overlays: arraySchema({ type: 'object', additionalProperties: true }), @@ -307,6 +561,7 @@ const responseExample = { precisions: ['fp8'], metric: 'y_tpPerGpu', xmode: 'interactivity', + xstat: 'median', xmetric: 'p90_ttft', percentile: 'p90', date: null, @@ -316,9 +571,16 @@ const responseExample = { frameworks: [], deployment: [], spec: [], + topologies: [], optimal: true, best: true, format: 'json', + serviceCompare: false, + serviceBaseline: null, + serviceComparator: null, + serviceTarget: null, + roleShare: false, + powerFit: false, }, metric: { key: 'tpPerGpu', @@ -332,6 +594,7 @@ const responseExample = { xAxis: { mode: 'interactivity', field: 'median_intvty', + statistic: 'median', label: 'Median Interactivity (tok/s/user)', }, frontier: { direction: 'upper_left', points: 14 }, @@ -352,6 +615,7 @@ const responseExample = { x: 12.5, y: 450.5, concurrency: 64, + topologyKey: 'Single|GPU=8|DP=?|TP=8|EP=1|PP=?|DCP=?|PCP=?|DPA=0|offload=off', tp: 8, date: '2026-08-20', runId: 12345678, @@ -380,7 +644,7 @@ const responses: readonly ApiResponse[] = [ mediaType: 'text/csv', schema: stringSchema, example: - 'hwKey,gpu,framework,specMethod,label,vendor,deployment,kvOffload,x,y,concurrency,tp,date,runId,frontier,bestPerSku,metric_tpPerGpu\r\nh200_trt,h200,trt,none,H200 (TRTLLM),NVIDIA,single-node,false,12.5,450.5,64,8,2026-08-20,12345678,true,true,450.5', + 'hwKey,gpu,framework,specMethod,label,vendor,deployment,kvOffload,x,y,concurrency,topologyKey,tp,date,runId,frontier,bestPerSku,metric_tpPerGpu\r\nh200_trt,h200,trt,none,H200 (TRTLLM),NVIDIA,single-node,false,12.5,450.5,64,Single|GPU=8|DP=?|TP=8|EP=1|PP=?|DCP=?|PCP=?|DPA=0|offload=off,8,2026-08-20,12345678,true,true,450.5', }, ], }, diff --git a/packages/app/src/lib/views-api/docs/options.ts b/packages/app/src/lib/views-api/docs/options.ts index bfd72e0be..6d7b24ec3 100644 --- a/packages/app/src/lib/views-api/docs/options.ts +++ b/packages/app/src/lib/views-api/docs/options.ts @@ -78,6 +78,7 @@ const responseSchema = objectSchema( specMethods: arraySchema(stringSchema), percentiles: arraySchema(stringSchema), xAxisModes: arraySchema(stringSchema), + fixedSequenceStatistics: arraySchema(stringSchema), scaleModes: arraySchema(stringSchema), metrics: arraySchema( objectSchema({ @@ -157,6 +158,7 @@ const responseExample = { specMethods: ['mtp', 'none'], percentiles: ['p75', 'p90'], xAxisModes: ['interactivity', 'ttft', 'e2e', 'e2e-normalized-interactivity'], + fixedSequenceStatistics: ['median', 'mean'], scaleModes: ['auto', 'linear', 'log'], metrics: [ { @@ -183,6 +185,10 @@ const responseExample = { metric: 'y_tokensPerDollarH', percentile: 'p90', xmode: 'interactivity', + xstat: 'median', + serviceCompare: false, + roleShare: false, + powerFit: false, }, }; diff --git a/packages/app/src/lib/views-api/registry.ts b/packages/app/src/lib/views-api/registry.ts index 0cf343e93..5621e8955 100644 --- a/packages/app/src/lib/views-api/registry.ts +++ b/packages/app/src/lib/views-api/registry.ts @@ -140,6 +140,7 @@ export const VIEW_QUERY_PARAMS = { 'vendors', ], inference: [ + 'topologies', 'allPoints', 'best', 'date', @@ -167,6 +168,13 @@ export const VIEW_QUERY_PARAMS = { 'vendors', 'xmetric', 'xmode', + 'xstat', + 'serviceCompare', + 'serviceBaseline', + 'serviceComparator', + 'serviceTarget', + 'roleShare', + 'powerFit', ], options: ['format'], overview: ['compare', 'engine', 'format', 'hwrows', 'models', 'ref', 'rows', 'tier'], diff --git a/packages/app/src/lib/views-api/series.ts b/packages/app/src/lib/views-api/series.ts index acdb88abe..54cc033e4 100644 --- a/packages/app/src/lib/views-api/series.ts +++ b/packages/app/src/lib/views-api/series.ts @@ -30,7 +30,11 @@ import { upperPowerEnvelope, } from '@/components/inference/utils/powerCurves'; import { pointDeploymentMode, type QuickFilters } from '@/components/inference/utils/quickFilters'; -import { resolveXAxisField } from '@/components/inference/utils/resolveXAxisField'; +import { pointTopologyKey } from '@/components/inference/utils/topology-filter'; +import { + resolveXAxisField, + type FixedSequenceStatistic, +} from '@/components/inference/utils/resolveXAxisField'; import type { DerivedAgenticMetricMap } from '@/hooks/api/use-derived-agentic-metrics'; import type { BenchmarkRow } from '@/lib/api'; import { transformBenchmarkRows } from '@/lib/benchmark-transform'; @@ -61,7 +65,12 @@ import { hardwareLegendLabel, unitFromLabel } from '@/lib/views-api/legend'; /** X-axis modes the API serves. `e2e-normalized-interactivity` is trace-derived * client-side and has no server-side data source, so it is not accepted here. */ -export type SeriesXMode = 'interactivity' | 'ttft' | 'e2e' | 'e2e-normalized-interactivity'; +export type SeriesXMode = + | 'interactivity' + | 'ttft' + | 'e2e' + | 'e2e-normalized-interactivity' + | 'concurrency'; export interface InferenceSeriesOptions { readonly sequence: Sequence; @@ -71,6 +80,8 @@ export interface InferenceSeriesOptions { readonly precisions: readonly string[]; readonly metricConfigKey: MetricConfigKey; readonly xmode: SeriesXMode; + /** Fixed-sequence service-axis statistic; ignored for Agentic and concurrency. */ + readonly fixedSequenceStatistic?: FixedSequenceStatistic; /** TTFT x metric override (e.g. `p90_ttft`); used by `ttft` mode and input metrics. */ readonly xmetric: string; /** Explicit hwKey / bare-GPU selection; empty = all. */ @@ -94,6 +105,7 @@ export interface InferenceSeriesPoint { readonly x: number; readonly y: number; readonly concurrency: number; + readonly topologyKey: string; readonly tp: number; readonly date: string; readonly runId?: number; @@ -130,7 +142,9 @@ export interface InferenceSeriesResult { readonly hardware: readonly { key: string; label: string; vendor?: string }[]; readonly frontier: { direction: ParetoDirection | null; points: number }; readonly metric: InferenceSeriesMetricMeta; - readonly xAxis: { mode: SeriesXMode; field: string; label: string }; + readonly xAxis: { mode: SeriesXMode; field: string; label: string; statistic: string | null }; + /** Scoped observed points before frontier/best-only pruning; not a public raw-row payload. */ + readonly observedPoints: readonly InferenceData[]; readonly count: number; } @@ -171,16 +185,19 @@ function resolveXAxisLabel( effectiveXMetric: string | null, isAgentic: boolean, percentile: string, + xAxisField: string, ): string { + if (branch === 'concurrency') return 'Concurrency'; let label = chartDef.x_label; if (branch === 'e2e-ttft-override') { const pctl = (effectiveXMetric ?? 'p90_ttft').replace(/_ttft$/u, ''); const pctlWord = pctl === 'median' ? 'Median' : pctl.toUpperCase(); label = `${pctlWord} Time To First Token (s)`; } - if (isAgentic) { - label = applyAgenticPercentileToXLabel(label, percentile.toUpperCase()); - } + label = applyAgenticPercentileToXLabel( + label, + isAgentic ? percentile.toUpperCase() : xAxisField.startsWith('mean_') ? 'Mean' : 'Median', + ); return label; } @@ -221,6 +238,7 @@ export function buildInferenceSeries( isAgentic, percentile, xAxisMode: xmode, + fixedSequenceStatistic: options.fixedSequenceStatistic, }); // 4. Precision + scope filters (GPU picks and vendor/framework/deployment/spec pills). @@ -259,10 +277,15 @@ export function buildInferenceSeries( const remapped = scoped .filter((point) => metricKey in point && supportsPointTokenMetric(point, tokenType)) .map((point) => remapInferencePoint(point, metricKey, resolved.xAxisField)); - const partition = partitionChartDataByLimits(remapped, chartDef, metricConfigKey, { - isTtftX: resolved.xAxisField.endsWith('_ttft'), - isAgentic, - }); + const partition = partitionChartDataByLimits( + remapped, + { ...chartDef, x_scale_field: resolved.xAxisField }, + metricConfigKey, + { + isTtftX: resolved.xAxisField.endsWith('_ttft'), + isAgentic, + }, + ); let mapped = options.allPoints ? remapped : partition.data; if (xmode === 'e2e-normalized-interactivity') mapped = mapped.flatMap((point) => { @@ -280,19 +303,24 @@ export function buildInferenceSeries( resolved.xAxisField !== resolved.naturalX && !(chartDef.chartType === 'e2e' && resolved.isTtftOverride); const direction = - configuredDirection && xAxisFlipped - ? flipRooflineDirection(configuredDirection) - : (configuredDirection ?? null); + xmode === 'concurrency' + ? null + : configuredDirection && xAxisFlipped + ? flipRooflineDirection(configuredDirection) + : (configuredDirection ?? null); // 7. Frontier flags, scoped per (hwKey, precision, date) like ScatterGraph. // Measured power represents load demand, so retain its upper boundary. const isMeasuredPower = isMeasuredPowerCurveMetric(metricConfigKey); const maximizePowerX = chartDef.chartType !== 'e2e'; - const frontierDirection = isMeasuredPower - ? maximizePowerX - ? 'upper_right' - : 'upper_left' - : direction; + const frontierDirection = + xmode === 'concurrency' + ? null + : isMeasuredPower + ? maximizePowerX + ? 'upper_right' + : 'upper_left' + : direction; const frontierPoints = new Set(); if (direction) { const frontierFn = paretoFrontForDirection(direction); @@ -334,7 +362,10 @@ export function buildInferenceSeries( if (best && bestHwKeys.size > 0 && !seriesIsBest) continue; const allPoints = byHwKey.get(hwKey)!; - const kept = optimal ? allPoints.filter((point) => frontierPoints.has(point)) : allPoints; + const kept = + optimal && xmode !== 'concurrency' + ? allPoints.filter((point) => frontierPoints.has(point)) + : allPoints; if (kept.length === 0) continue; const sample = kept[0]; @@ -363,6 +394,7 @@ export function buildInferenceSeries( x: point.x, y: point.y, concurrency: point.conc ?? 0, + topologyKey: pointTopologyKey(point), tp: point.tp ?? 0, date: point.date ?? '', ...(runIdFromUrl(point.run_url) === undefined @@ -396,6 +428,12 @@ export function buildInferenceSeries( }, xAxis: { mode: xmode, + statistic: + xmode === 'concurrency' + ? null + : isAgentic + ? percentile + : (options.fixedSequenceStatistic ?? 'median'), field: xmode === 'e2e-normalized-interactivity' ? `${percentile}_e2e_norm_intvty` @@ -403,8 +441,16 @@ export function buildInferenceSeries( label: xmode === 'e2e-normalized-interactivity' ? `${percentile.toUpperCase()} E2E Normalized Interactivity (tok/s/user)` - : resolveXAxisLabel(chartDef, resolved.branch, effectiveXMetric, isAgentic, percentile), + : resolveXAxisLabel( + chartDef, + resolved.branch, + effectiveXMetric, + isAgentic, + percentile, + String(resolved.xAxisField), + ), }, + observedPoints: mapped, count, }; } diff --git a/packages/app/timings.json b/packages/app/timings.json index 2979c5996..df735201a 100644 --- a/packages/app/timings.json +++ b/packages/app/timings.json @@ -236,6 +236,14 @@ "spec": "cypress/e2e/performance.cy.ts", "duration": 2377 }, + { + "spec": "cypress/e2e/powerx-timeline.cy.ts", + "duration": 4000 + }, + { + "spec": "cypress/e2e/powerx-compare.cy.ts", + "duration": 4000 + }, { "spec": "cypress/e2e/profit-estimator.cy.ts", "duration": 39450 diff --git a/packages/db/src/backfill-gpu-metrics.ts b/packages/db/src/backfill-gpu-metrics.ts index 8e5ebdde7..bb73647cb 100644 --- a/packages/db/src/backfill-gpu-metrics.ts +++ b/packages/db/src/backfill-gpu-metrics.ts @@ -81,6 +81,7 @@ import { repositoryFromRunUrl } from './lib/runtime-metadata-artifacts.js'; const DEFAULT_REPO = 'SemiAnalysisAI/InferenceX'; const GITHUB_RETENTION_DAYS = 90; +const GITHUB_RETENTION_MS = GITHUB_RETENTION_DAYS * 24 * 60 * 60 * 1000; const sql = createAdminSql({ noSsl: hasNoSslFlag(), max: 4, onnotice: () => {} }); interface CandidateRun { @@ -146,8 +147,7 @@ function parseFlags(): BackfillFlags { } function isWithinGithubRetention(date: string): boolean { - const ageMs = Date.now() - new Date(date).getTime(); - return ageMs <= GITHUB_RETENTION_DAYS * 24 * 60 * 60 * 1000; + return Date.now() - new Date(date).getTime() <= GITHUB_RETENTION_MS; } async function loadCandidateRuns( @@ -155,9 +155,7 @@ async function loadCandidateRuns( limit: number | null, force: boolean, ): Promise { - const cutoff = new Date(Date.now() - GITHUB_RETENTION_DAYS * 24 * 60 * 60 * 1000) - .toISOString() - .slice(0, 10); + const cutoff = new Date(Date.now() - GITHUB_RETENTION_MS).toISOString().slice(0, 10); const since = flags.since ?? cutoff; const rows = await sql` select wr.id, wr.github_run_id, wr.run_attempt, wr.html_url, wr.date::text as date, @@ -241,13 +239,13 @@ async function processPair( powerAudit?: RecoveredPowerAudit; }[] = []; for (const row of mappedRows) { + const identity = benchmarkPublicationIdentity(row); const ids = await findBenchmarkResultIds(sql, run, [row], (id) => - uniqueFallbacks.set(stablePowerPointIdentity(benchmarkPublicationIdentity(row)), id), + uniqueFallbacks.set(stablePowerPointIdentity(identity), id), ); if (ids.length === 0) throw new Error(`${pair.gpuMetrics.name}: no matching benchmark rows`); matchedIds.push(...ids); if (row.benchmarkType === 'agentic_traces') { - const identity = benchmarkPublicationIdentity(row); for (const id of ids) agenticPoints.push({ id, benchmarkType: row.benchmarkType, conc: row.conc, identity }); } @@ -303,16 +301,14 @@ async function processPair( auditsWritten: written.length, }; } catch (error) { + const message = error instanceof Error ? error.message : String(error); if (pointKeys.length === 0) expectationErrors.push({ benchmarkArtifact: pair.benchmarks.name, artifactNames: [pair.gpuMetrics.name], - error: error instanceof Error ? error.message : String(error), + error: message, }); - for (const key of pointKeys) { - const observation = observations.get(key)!; - observation.error = error instanceof Error ? error.message : String(error); - } + for (const key of pointKeys) observations.get(key)!.error = message; console.error(` ✗ run ${run.github_run_id} artifact ${pair.gpuMetrics.name}:`, error); return { kind: 'failed' }; } finally { diff --git a/packages/db/src/etl/gpu-metrics-artifacts.ts b/packages/db/src/etl/gpu-metrics-artifacts.ts index 3c50827ad..2873db95f 100644 --- a/packages/db/src/etl/gpu-metrics-artifacts.ts +++ b/packages/db/src/etl/gpu-metrics-artifacts.ts @@ -85,19 +85,20 @@ function isGpuMetricsCsvName(fileName: string): boolean { return !lower.includes('_identity') && !lower.includes('_energy_'); } -/** Recursively list every telemetry CSV under an extracted artifact root. */ -export function listGpuMetricsCsvFiles(root: string): GpuMetricsCsvFile[] { +/** Recursively list the files under `root` whose POSIX-relative name passes `matches`. */ +function listFiles( + root: string, + matches: (fileName: string, baseName: string) => boolean, +): GpuMetricsCsvFile[] { if (!fs.existsSync(root) || !fs.statSync(root).isDirectory()) return []; const files: GpuMetricsCsvFile[] = []; const visit = (directory: string): void => { for (const entry of fs.readdirSync(directory, { withFileTypes: true })) { const pathname = path.join(directory, entry.name); if (entry.isDirectory()) visit(pathname); - else if (entry.isFile() && isGpuMetricsCsvName(entry.name)) { - files.push({ - fileName: path.relative(root, pathname).split(path.sep).join('/'), - path: pathname, - }); + else if (entry.isFile()) { + const fileName = path.relative(root, pathname).split(path.sep).join('/'); + if (matches(fileName, entry.name)) files.push({ fileName, path: pathname }); } } }; @@ -105,30 +106,19 @@ export function listGpuMetricsCsvFiles(root: string): GpuMetricsCsvFile[] { return files.toSorted((a, b) => a.fileName.localeCompare(b.fileName)); } +/** Recursively list every telemetry CSV under an extracted artifact root. */ +export function listGpuMetricsCsvFiles(root: string): GpuMetricsCsvFile[] { + return listFiles(root, (_fileName, baseName) => isGpuMetricsCsvName(baseName)); +} + /** Every multinode power CSV under an extracted `power_audit_` root. */ export function listMultinodePowerSampleFiles(root: string): GpuMetricsCsvFile[] { - if (!fs.existsSync(root) || !fs.statSync(root).isDirectory()) return []; - const files: GpuMetricsCsvFile[] = []; - const visit = (directory: string): void => { - for (const entry of fs.readdirSync(directory, { withFileTypes: true })) { - const pathname = path.join(directory, entry.name); - if (entry.isDirectory()) visit(pathname); - else if (entry.isFile()) { - const fileName = path.relative(root, pathname).split(path.sep).join('/'); - if (isMultinodePowerSamplesPath(fileName)) files.push({ fileName, path: pathname }); - } - } - }; - visit(root); - return files.toSorted((a, b) => a.fileName.localeCompare(b.fileName)); + return listFiles(root, isMultinodePowerSamplesPath); } /** The producer manifest next to `samples.csv`; null when absent or malformed. */ export function readMultinodePowerManifest(samplesPath: string): Record | null { - const parsed = readJsonIfPresent(path.join(path.dirname(samplesPath), 'manifest.json')); - return parsed && typeof parsed === 'object' && !Array.isArray(parsed) - ? (parsed as Record) - : null; + return readJsonObjectIfPresent(path.join(path.dirname(samplesPath), 'manifest.json')); } /** Preserve window boundaries and role overrides before GitHub artifact expiry. */ @@ -175,6 +165,14 @@ function readJsonIfPresent(pathname: string): unknown | null { } } +/** Like `readJsonIfPresent`, but only a plain JSON object counts as present. */ +function readJsonObjectIfPresent(pathname: string): Record | null { + const parsed = readJsonIfPresent(pathname); + return parsed && typeof parsed === 'object' && !Array.isArray(parsed) + ? (parsed as Record) + : null; +} + /** `gpu,total_energy_consumption` two-column CSV → { "": joules }. */ export function parseEnergyCsv(text: string): Record | null { const lines = text @@ -225,11 +223,8 @@ export function readGpuMetricsSidecars(csvPath: string): GpuMetricsSidecars { ) .sort(); for (const entry of contextFiles) { - const parsed = readJsonIfPresent(path.join(dir, entry)); - if (parsed && typeof parsed === 'object' && !Array.isArray(parsed)) { - context = parsed as Record; - break; - } + context = readJsonObjectIfPresent(path.join(dir, entry)); + if (context) break; } const energyStartPath = path.join(dir, 'gpu_metrics_energy_start.csv'); const energyEndPath = path.join(dir, 'gpu_metrics_energy_end.csv'); diff --git a/packages/db/src/etl/gpu-metrics-csv.ts b/packages/db/src/etl/gpu-metrics-csv.ts index 35ec33e57..135fa5cef 100644 --- a/packages/db/src/etl/gpu-metrics-csv.ts +++ b/packages/db/src/etl/gpu-metrics-csv.ts @@ -139,7 +139,7 @@ export function parseAmdTimestamp(raw: string): number | null { return Number.isFinite(iso) ? iso : null; } -function emptySample(timestampMs: number, gpuIndex: number): GpuMetricSample { +export function emptySample(timestampMs: number, gpuIndex: number): GpuMetricSample { return { timestampMs, gpuIndex, diff --git a/packages/db/src/etl/gpu-metrics-ingest.ts b/packages/db/src/etl/gpu-metrics-ingest.ts index 7bdd58b2e..e137acc50 100644 --- a/packages/db/src/etl/gpu-metrics-ingest.ts +++ b/packages/db/src/etl/gpu-metrics-ingest.ts @@ -23,9 +23,6 @@ import { export { statMetricColumn } from '../lib/gpu-metric-stats.js'; import type { Sql } from './db-utils.js'; - -/** Either a pooled client or the transaction handle passed to `sql.begin` callbacks. */ -type TxLike = Sql | postgres.TransactionSql; import { computeGpuMetricStats, parseGpuMetricsCsv, @@ -47,6 +44,9 @@ import { } from './gpu-metrics-artifacts.js'; import { multinodePowerVendor, parseMultinodePowerSamples } from './multinode-power-samples.js'; +/** Either a pooled client or the transaction handle passed to `sql.begin` callbacks. */ +type TxLike = Sql | postgres.TransactionSql; + /** Samples are streamed to Postgres in unnest batches of this many rows. */ const SAMPLE_BATCH_SIZE = 5000; @@ -80,6 +80,25 @@ export interface GpuMetricsIngestResult { seriesSkipped: number; } +/** Dedupe, window and digest one CSV's samples; null when nothing usable remains. */ +function digestSeries( + base: Pick, + rawSamples: readonly GpuMetricSample[], +): PreparedGpuMetricSeries | null { + const samples = uniqueSamples(rawSamples); + const summary = summarizeGpuMetricSamples(samples); + if (!summary) return null; + return { + ...base, + samples, + stats: computeGpuMetricStats(samples), + sampleIntervalS: summary.sampleIntervalS, + gpuCount: summary.gpuCount, + startedAtMs: summary.startedAtMs, + endedAtMs: summary.endedAtMs, + }; +} + /** * One series per host from the multinode power bundle. The deployment-wide * CSV is hashed once, so every host series of one upload shares its source sha. @@ -97,31 +116,26 @@ function prepareMultinodePowerSeries(artifact: GpuMetricsArtifact): PreparedGpuM const vendor = multinodePowerVendor(manifest); const csvSha256 = createHash('sha256').update(csvText).digest('hex'); for (const host of hosts) { - const samples = uniqueSamples(host.samples); - const summary = summarizeGpuMetricSamples(samples); - if (!summary) continue; - prepared.push({ - fileName: `${file.fileName}#${host.hostname}`, - vendor, - csvSha256, - samples, - stats: computeGpuMetricStats(samples), - sampleIntervalS: summary.sampleIntervalS, - gpuCount: summary.gpuCount, - startedAtMs: summary.startedAtMs, - endedAtMs: summary.endedAtMs, - sidecars: { - context: manifest, - validations, - identity: Object.entries(host.gpuUuids).map(([index, uuid]) => ({ - hostname: host.hostname, - gpu_index: Number(index), - gpu_uuid: uuid, - })), - energyStart: null, - energyEnd: null, + const series = digestSeries( + { + fileName: `${file.fileName}#${host.hostname}`, + vendor, + csvSha256, + sidecars: { + context: manifest, + validations, + identity: Object.entries(host.gpuUuids).map(([index, uuid]) => ({ + hostname: host.hostname, + gpu_index: Number(index), + gpu_uuid: uuid, + })), + energyStart: null, + energyEnd: null, + }, }, - }); + host.samples, + ); + if (series) prepared.push(series); } } return prepared; @@ -155,28 +169,22 @@ export function prepareGpuMetricsArtifact(artifact: GpuMetricsArtifact): Prepare const parsed = parseGpuMetricsCsv(csvText, { nvidiaUtcOffsetMinutes: contextUtcOffsetMinutes(sidecars.context), }); - if (!parsed) { - unreadable.push(file.fileName); - continue; - } - const samples = uniqueSamples(parsed.samples); - const summary = summarizeGpuMetricSamples(samples); - if (!summary) { + const series = + parsed && + digestSeries( + { + fileName: file.fileName, + vendor: parsed.vendor, + csvSha256: createHash('sha256').update(csvText).digest('hex'), + sidecars, + }, + parsed.samples, + ); + if (!series) { unreadable.push(file.fileName); continue; } - prepared.push({ - fileName: file.fileName, - vendor: parsed.vendor, - csvSha256: createHash('sha256').update(csvText).digest('hex'), - samples, - stats: computeGpuMetricStats(samples), - sampleIntervalS: summary.sampleIntervalS, - gpuCount: summary.gpuCount, - startedAtMs: summary.startedAtMs, - endedAtMs: summary.endedAtMs, - sidecars, - }); + prepared.push(series); } if (prepared.length > 0 && unreadable.length > 0) { throw new Error( diff --git a/packages/db/src/etl/multinode-power-samples.ts b/packages/db/src/etl/multinode-power-samples.ts index 9c711c4de..552588057 100644 --- a/packages/db/src/etl/multinode-power-samples.ts +++ b/packages/db/src/etl/multinode-power-samples.ts @@ -14,7 +14,12 @@ * staging ("one CSV per node"). Pure module: no I/O. */ -import { splitCsvLine, type GpuMetricSample, type GpuMetricsVendor } from './gpu-metrics-csv.js'; +import { + emptySample, + splitCsvLine, + type GpuMetricSample, + type GpuMetricsVendor, +} from './gpu-metrics-csv.js'; export interface MultinodePowerHost { hostname: string; @@ -32,27 +37,6 @@ export function isMultinodePowerSamplesPath(relativePath: string): boolean { return /^LOGS\/(?:[^/]+\/)*samples\.csv$/u.test(posix); } -function powerOnlySample(timestampMs: number, gpuIndex: number, powerW: number): GpuMetricSample { - return { - timestampMs, - gpuIndex, - powerW, - temperatureC: null, - smClockMhz: null, - memClockMhz: null, - gpuUtilPct: null, - memUtilPct: null, - edgeTempC: null, - memTempC: null, - gfxVoltageMv: null, - socVoltageMv: null, - memVoltageMv: null, - fclkMhz: null, - socclkMhz: null, - mmActivityPct: null, - }; -} - /** * Group the deployment-wide CSV by host. Returns null when the header is not * the multinode power schema; malformed rows are skipped. Hosts sort by name @@ -90,7 +74,7 @@ export function parseMultinodePowerSamples(csvText: string): MultinodePowerHost[ host = { hostname, samples: [], gpuUuids: {} }; hosts.set(hostname, host); } - host.samples.push(powerOnlySample(Math.round(seconds * 1000), gpuIndex, powerW)); + host.samples.push({ ...emptySample(Math.round(seconds * 1000), gpuIndex), powerW }); const uuid = at(cells, 'gpu_uuid'); if (uuid && !(gpuIndex in host.gpuUuids)) host.gpuUuids[gpuIndex] = uuid; } diff --git a/packages/db/src/etl/power-publication.ts b/packages/db/src/etl/power-publication.ts index 0de76b89b..6f9af4a0c 100644 --- a/packages/db/src/etl/power-publication.ts +++ b/packages/db/src/etl/power-publication.ts @@ -85,11 +85,7 @@ export interface PowerPublicationManifest { points: PowerPublicationPoint[]; /** Fatal: verify-power-publication exits non-zero when this is non-empty. */ ingestErrors?: string[]; - /** - * Non-fatal: PowerX telemetry digest failures. Surfaced in the verification - * receipt so they stay visible, but they never fail the ingest — the benchmark - * rows landed, and the artifact can be re-digested by the backfill. - */ + /** Non-fatal PowerX telemetry digest failures; see `fatalPublicationErrors`. */ telemetryWarnings?: string[]; /** Attachment completeness, separate from benchmark/power publication validity. */ telemetry?: TelemetryReceipt; @@ -110,6 +106,7 @@ export interface PowerPublicationManifest { error?: string; }; } + /** * The errors that fail an ingest. `telemetryWarnings` is deliberately not among * them: a gpu_metrics digest failure costs one point's PowerX tab, while the diff --git a/packages/db/src/lib/backfill-benchmark-refresh.ts b/packages/db/src/lib/backfill-benchmark-refresh.ts index 15d2d60f6..46eedcdea 100644 --- a/packages/db/src/lib/backfill-benchmark-refresh.ts +++ b/packages/db/src/lib/backfill-benchmark-refresh.ts @@ -121,8 +121,7 @@ export async function refreshBackfillBenchmarks( `; if (rows.length !== ids.length) throw new Error('Refresh receipt IDs do not belong to its run/attempt'); - const changed = rows; - for (const row of changed) { + for (const row of rows) { const audit = row.power_audit; if (audit === null) throw new Error( @@ -159,7 +158,7 @@ export async function refreshBackfillBenchmarks( points[0]!.power_audit = update.powerAudit; } save(); - if (changed.length > 0) { + if (rows.length > 0) { phase = 'refresh latest_benchmarks'; await refreshLatestBenchmarks(sql); phase = 'invalidate cache'; @@ -178,7 +177,7 @@ export async function refreshBackfillBenchmarks( ) throw new Error('Invalid cache invalidation response'); phase = 'verify benchmark API'; - for (const model of new Set(changed.map((row) => row.model))) { + for (const model of new Set(rows.map((row) => row.model))) { const displayModel = DB_MODEL_TO_DISPLAY[model]; if (!displayModel) throw new Error(`Unmapped public model: ${model}`); const url = new URL('/api/v1/benchmarks', endpoint); @@ -197,7 +196,7 @@ export async function refreshBackfillBenchmarks( if (!api.ok) throw new Error(`HTTP ${api.status}`); const body: unknown = await api.json(); if (!Array.isArray(body)) throw new Error('Invalid benchmark API response'); - for (const row of changed.filter((entry) => entry.model === model)) { + for (const row of rows.filter((entry) => entry.model === model)) { const matches = body.filter((candidate) => { const raw = candidate?.id; const id = diff --git a/packages/db/src/lib/backfill-runner.ts b/packages/db/src/lib/backfill-runner.ts index 2051e493d..265cca382 100644 --- a/packages/db/src/lib/backfill-runner.ts +++ b/packages/db/src/lib/backfill-runner.ts @@ -166,11 +166,6 @@ export async function runCandidateIdBackfill( return true; } -/** - * jsonb parameter for a freshly computed value. `structuredClone` strips - * class instances/prototypes so postgres.js serializes plain data only — - * matches what the inline ingest path stores. - */ /** * List a candidate run's GitHub artifacts with transient-failure retry. * Returns `null` when GitHub no longer has the run at all, so a sweep over @@ -190,6 +185,11 @@ export async function listBackfillRunArtifacts( } } +/** + * jsonb parameter for a freshly computed value. `structuredClone` strips + * class instances/prototypes so postgres.js serializes plain data only — + * matches what the inline ingest path stores. + */ export function jsonbParam(sql: Sql, value: unknown): ReturnType { return sql.json(structuredClone(value) as unknown as Parameters[0]); } diff --git a/packages/db/src/lib/telemetry-purge.ts b/packages/db/src/lib/telemetry-purge.ts index 13b2b76dd..ef8b2d07a 100644 --- a/packages/db/src/lib/telemetry-purge.ts +++ b/packages/db/src/lib/telemetry-purge.ts @@ -28,8 +28,13 @@ export interface TelemetryPurgeCounts { export const NO_TELEMETRY: TelemetryPurgeCounts = { series: 0, samples: 0 }; +interface CountRow { + n: unknown; + samples: unknown; +} + /** `count`/`sum` come back as strings over the wire; `sum` is null on an empty set. */ -function toCounts(row: { n: unknown; samples: unknown } | undefined): TelemetryPurgeCounts { +function toCounts(row: CountRow | undefined): TelemetryPurgeCounts { if (!row) return NO_TELEMETRY; return { series: Number(row.n ?? 0), samples: Number(row.samples ?? 0) }; } @@ -44,12 +49,12 @@ export async function countRunTelemetry( workflowRunIds: readonly number[], ): Promise { if (workflowRunIds.length === 0) return NO_TELEMETRY; - const [row] = await sql` + const [row] = await sql` SELECT count(*)::int AS n, coalesce(sum(sample_count), 0)::bigint AS samples FROM gpu_metric_series WHERE workflow_run_id = ANY(${[...workflowRunIds]}) `; - return toCounts(row as { n: unknown; samples: unknown } | undefined); + return toCounts(row); } /** @@ -62,7 +67,7 @@ export async function deleteRunTelemetry( workflowRunIds: readonly number[], ): Promise { if (workflowRunIds.length === 0) return NO_TELEMETRY; - const [row] = await sql` + const [row] = await sql` WITH deleted AS ( DELETE FROM gpu_metric_series WHERE workflow_run_id = ANY(${[...workflowRunIds]}) @@ -70,7 +75,7 @@ export async function deleteRunTelemetry( ) SELECT count(*)::int AS n, coalesce(sum(sample_count), 0)::bigint AS samples FROM deleted `; - return toCounts(row as { n: unknown; samples: unknown } | undefined); + return toCounts(row); } /** @@ -83,7 +88,7 @@ export async function unlinkPointTelemetry( benchmarkResultIds: readonly number[], ): Promise { if (benchmarkResultIds.length === 0) return 0; - const [row] = await sql` + const [row] = await sql[]>` WITH deleted AS ( DELETE FROM benchmark_result_gpu_metrics WHERE benchmark_result_id = ANY(${[...benchmarkResultIds]}) @@ -91,7 +96,7 @@ export async function unlinkPointTelemetry( ) SELECT count(*)::int AS n FROM deleted `; - return Number((row as { n: unknown } | undefined)?.n ?? 0); + return Number(row?.n ?? 0); } /** One-line summary for preview and transcript output. */ diff --git a/packages/db/src/queries/gpu-metrics.ts b/packages/db/src/queries/gpu-metrics.ts index cc4767cb6..488fe0ca9 100644 --- a/packages/db/src/queries/gpu-metrics.ts +++ b/packages/db/src/queries/gpu-metrics.ts @@ -219,6 +219,20 @@ function toStatRow(raw: RawStatRow): GpuMetricStatRow { }; } +function groupBySeries( + rows: readonly Raw[], + toRow: (raw: Raw) => Row, +): Map { + const groups = new Map(); + for (const raw of rows) { + const key = Number(raw.series_id); + const bucket = groups.get(key); + if (bucket) bucket.push(toRow(raw)); + else groups.set(key, [toRow(raw)]); + } + return groups; +} + async function loadSeriesDetails( sql: DbClient, seriesRows: readonly RawSeriesRow[], @@ -284,26 +298,16 @@ async function loadSeriesDetails( throw new TelemetrySnapshotChangedError(ids); } - const statsBySeries = new Map(); - for (const raw of statRows) { - const key = Number(raw.series_id); - const bucket = statsBySeries.get(key); - if (bucket) bucket.push(toStatRow(raw)); - else statsBySeries.set(key, [toStatRow(raw)]); - } + const statsBySeries = groupBySeries(statRows, toStatRow); const staleIds = new Set( seriesRows .filter((row) => row.stats_version !== GPU_STATS_VERSION) .map((row) => Number(row.id)), ); - const staleSamples = new Map(); - for (const sample of sampleRows) { - const id = Number(sample.series_id); - if (!staleIds.has(id)) continue; - const rows = staleSamples.get(id) ?? []; - rows.push(sample); - staleSamples.set(id, rows); - } + const staleSamples = groupBySeries( + sampleRows.filter((sample) => staleIds.has(Number(sample.series_id))), + (sample) => sample, + ); for (const row of seriesRows) { const id = Number(row.id); if (!staleIds.has(id)) continue; @@ -322,13 +326,7 @@ async function loadSeriesDetails( })), ); } - const samplesBySeries = new Map(); - for (const raw of sampleRows) { - const key = Number(raw.series_id); - const bucket = samplesBySeries.get(key); - if (bucket) bucket.push(toSampleRow(raw)); - else samplesBySeries.set(key, [toSampleRow(raw)]); - } + const samplesBySeries = groupBySeries(sampleRows, toSampleRow); return seriesRows.map((row) => { const id = Number(row.id); diff --git a/packages/db/src/verify-power-publication.ts b/packages/db/src/verify-power-publication.ts index 6f7a68456..2b4e633e4 100644 --- a/packages/db/src/verify-power-publication.ts +++ b/packages/db/src/verify-power-publication.ts @@ -99,8 +99,7 @@ try { status: errors.length > 0 ? 'failed' : manifest.points.length > 0 ? 'matched' : 'no_power_points', errors, - // Kept out of `errors` on purpose: a telemetry digest failure costs one - // point's PowerX tab, not its benchmark data, so it must not fail the ingest. + // Non-fatal by design; see `fatalPublicationErrors`. telemetryWarnings: manifest.telemetryWarnings ?? [], ...(manifest.telemetry ? { telemetry: manifest.telemetry } : {}), }; diff --git a/packages/skills/skills/inferencex-api/integrity.json b/packages/skills/skills/inferencex-api/integrity.json index 014cf3568..3efb0807a 100644 --- a/packages/skills/skills/inferencex-api/integrity.json +++ b/packages/skills/skills/inferencex-api/integrity.json @@ -8,7 +8,7 @@ "references/cli-contract.md": "fa44faa38d889b4fbdee5ba42752b47758ee3706fab150513bd4e6db87a86cb0", "references/cli.md": "96b228f34cb3600f4548f4dc84df8506531747286763c56ca834c77bee05e1eb", "references/collectivex.md": "eb794f9c28d4a27bec4db80c42c4685ff3b1d204fd9b12a6788511fb258a789f", - "references/dashboard-views.md": "37cfe76b1e8cf825a589dd2e478f2a2238a9e5c423a48dd19b1b2ac77833f713", + "references/dashboard-views.md": "2061763bcd7bf784a3b212e9472f3a7c3be5a8816984dc8bd1599f28ea9caabe", "references/offline-exports.md": "95aa565dcd4c9e592159baa560a9a58217bc61e395de1c6daf0576795c7f4ec9", "references/pareto.md": "1b4d2d163f982e3f2d97addce310789eae9c1e5badd82dd7501708c9e0385027", "references/powerx.md": "ae01164107aec35247f729e425379b350d2e7ecdd8d4668fc68dbd879f3d965d", diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md index c72bf4584..2ab3eefd0 100644 --- a/packages/skills/skills/inferencex-api/references/dashboard-views.md +++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md @@ -50,30 +50,30 @@ https://inferencex.semianalysis.com/api/v1/views/profit-estimator?model=DeepSeek All paths below are relative to `/api/v1/views/`. The maintained exhaustive query-key list is in the app's OpenAPI contract. The following groups explain -which controls belong together. - -| View | Selection and calculation | -| ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `options` | Static registries and defaults, JSON only. Data-dependent run/config choices come from their own views. | -| `inference` | Model, sequence, precisions, date/exact run, GPU/vendor/framework/deployment/spec/power filters; metric, xmode, xmetric, percentile, optimal/best/allPoints; TCO, custom costs/powers, pricing, unofficial runs, comparison dates/endpoints. JSON or CSV. | -| `historical` | Model, sequence, target, metric, precisions, GPUs/vendors/frameworks/deployment, start/end, TCO and pricing. `extendToDate` labels synthetic extension, defaults to current UTC date. JSON or CSV. | -| `calculator` | Model/sequence/run/date, precisions/GPUs, percentile, target/mode, token type, owning/renting cost, TCO, power budget, cost cap, hide-above-limit and public unofficial runs. JSON or CSV. | -| `first-token` | Benchmark selection plus 1–8 positive TTFT caps in seconds, minimum interactivity, cost provider and token type. Shared dashboard winner selection. | -| `cache-reuse` | AgentX selection plus `config` from configurations and `recipe` from data.recipes. See recipe selection below. | -| `profit-estimator` | AgentX selection, comparison dates/runs, target, utilization, lab revenue share, list/OpenRouter/custom token prices, own/rent/custom chip costs, TCO and provisioned/modeled/compare power. USD/chip-hour. | -| `profit-estimator-per-gigawatt` | Same selectors; USD/GW-year basis and owning-cost default. | -| `fleet` | Model/sequence/precision/GPUs, target/percentile, TCO/token type/cost provider, MW, input/output prices, ramp, cache discount, MTBI, recovery, horizon and metric. JSON or CSV. | -| `evaluation` | Model, task, date, precision and GPU selection; public unofficial runs remain separately labeled and independent of the official date cutoff. JSON or CSV. | -| `reliability` | Rolling range, GPU selection and optional `asOf` date for reproducibility. JSON or CSV. | -| `gpu-specs` | Full hardware properties or selected metric ranking; table/radar/bar are renderings of these values. JSON or CSV. | -| `overview` | Models, model/hardware row limits, hardware tier, engine, comparison mode and reference. JSON or CSV. | -| `rankings` | Ranking kind (model or chip), selected model/scenario and format. Use the exact values in OpenAPI. | -| `compare` | GPU pairs, model/slug, scenario, tier selection and variant. JSON or CSV. | -| `collectivex` | Ordered run selection and suite; EP size/phase/modes/precision/operation/percentile/axis/SKU/backend/series; KV page size/x/y/pull/push/overlap ISL/series; swap direction/layout/metric/percentile/series. | -| `submissions` | Search, table sort/direction/offset/limit, weekly/cumulative chart, on-change cutoff and NVIDIA/AMD/total lines. Table search does not filter the independent submission-volume chart. | -| `current-inferencex-image` | Model, sequence, hardware, precision, speculation, node type, framework families and `asOf` for image age/release status. | -| `gpu-metrics` | Required run, file/host artifact, GPU indices, metric or correlation axes, statistics sorting, chart mode and interactive downsampling preference. DB-first telemetry; full-record statistics use stored digests. Responses are no-store. | -| `video` | CI run/artifact discovery; only already-published artifact reads; source, serving cell, media/fidelity slot, power phase and GPU denominator; same-workload comparisons, axes, deployment costs and selected tradeoff point. No-store. | +which API selectors belong together. + +| View | Selection and calculation | +| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `options` | Static registries and defaults, JSON only. Data-dependent run/config choices come from their own views. | +| `inference` | Model, sequence, precisions, date/exact run, GPU/vendor/framework/deployment/spec/power filters; metric, xmode, xstat, xmetric, percentile, optimal/best/allPoints; API-only equal-service sources/target; roleShare and powerFit panels; TCO, custom costs/powers, pricing, unofficial runs, comparison dates/endpoints. JSON or CSV. | +| `historical` | Model, sequence, target, metric, precisions, GPUs/vendors/frameworks/deployment, start/end, TCO and pricing. `extendToDate` labels synthetic extension, defaults to current UTC date. JSON or CSV. | +| `calculator` | Model/sequence/run/date, precisions/GPUs, percentile, target/mode, token type, owning/renting cost, TCO, power budget, cost cap, hide-above-limit and public unofficial runs. JSON or CSV. | +| `first-token` | Benchmark selection plus 1–8 positive TTFT caps in seconds, minimum interactivity, cost provider and token type. Shared dashboard winner selection. | +| `cache-reuse` | AgentX selection plus exact `config` from returned configurations and `recipe` from data.recipes. Cache-reuse curves retain official versus unofficial evidence. | +| `profit-estimator` | AgentX selection, comparison dates/runs, target, utilization, lab revenue share, list/OpenRouter/custom token prices, own/rent/custom chip costs, TCO, actual/theoretical cache-hit mode and provisioned/modeled/compare power. USD/chip-hour. | +| `profit-estimator-per-gigawatt` | Same selectors; USD/GW-year basis and owning-cost default. | +| `fleet` | Model/sequence/precision/GPUs, target/percentile, TCO/token type/cost provider, MW, input/output prices, ramp, cache discount, MTBI, recovery, horizon and metric. JSON or CSV. | +| `evaluation` | Model, task, date, precision and GPU selection; public unofficial runs remain separately labeled and independent of the official date cutoff. JSON or CSV. | +| `reliability` | Rolling range, GPU selection and optional `asOf` date for reproducibility. JSON or CSV. | +| `gpu-specs` | Full hardware properties or selected metric ranking; table/radar/bar are renderings of these values. JSON or CSV. | +| `overview` | Models, model/hardware row limits, hardware tier, engine, comparison mode and reference. JSON or CSV. | +| `rankings` | Ranking kind (model or chip), selected model/scenario and format. Use the exact values in OpenAPI. | +| `compare` | GPU pairs, model/slug, scenario, tier selection and variant. JSON or CSV. | +| `collectivex` | Ordered run selection and suite; EP size/phase/modes/precision/operation/percentile/axis/SKU/backend/series; KV page size/x/y/pull/push/overlap ISL/series; swap direction/layout/metric/percentile/series. | +| `submissions` | Search, table sort/direction/offset/limit, weekly/cumulative chart, on-change cutoff and NVIDIA/AMD/total lines. Table search does not filter the independent submission-volume chart. | +| `current-inferencex-image` | Model, sequence, hardware, precision, speculation, node type, framework families and `asOf` for image age/release status. | +| `gpu-metrics` | Required run, file/host artifact, GPU indices, metric or correlation axes, statistics sorting, chart mode and interactive downsampling preference. DB-first raw-explorer projection; full-record statistics use stored digests. Not a serving-window/role-pool projection. Responses are no-store. | +| `video` | CI run/artifact discovery; only already-published artifact reads; source, serving cell, media/fidelity slot, power phase and GPU denominator; same-workload comparisons, axes, deployment costs and selected tradeoff point. No-store. | ## Interpretation and maintenance @@ -93,11 +93,96 @@ day. Date-only comparisons select that day's exact logical snapshot; the primary source observations inclusively. Public unofficial overlays must not be relabeled as official results. +`i_rulers` belongs to browser share links, including run-specific comparison +curves. It is not a views API parameter. Tooltip scrolling and viewport limits +only affect access to existing actions; neither changes API data or calculations. + For measured-power gauges, `optimal=true` keeps the chart's higher-power outer envelope. `frontier.direction` describes that boundary; `metric.direction` retains the optimization direction used by `best=true`. Interpret the envelope as a load boundary, not evidence that those points are more energy efficient. +The dashboard names its four boundaries GPU Level Measured, GPU Level Provisioned +(TDP), All in Provisioned, and All in Measured. All in Measured combines measured +GPU power with modeled unmeasured components and PUE; do not describe it as measured +wall power. Metric IDs and API selector values are unchanged. Profit `powerBasis` +still accepts `provisioned`, `modeled`, or `compare`; `powerLabel` is display text. +Expanding assumptions or unavailable-estimate details does not change returned data. + +Prefer equal-service comparisons for article-facing hardware analysis. Use +`xstat=mean` only for fixed-sequence service axes when that statistic is intended: +streaming speed then means **1 / mean TPOT**, not arithmetic mean request speed. +TTFT/E2E use recorded means. Default is median; absent means are not replaced. +AgentX still uses `percentile`; concurrency uses no statistic. Check resolved +`params.xstat` and `xAxis.statistic`, not the requested parameter alone. + +Equal-service and matched-concurrency analysis are API-only. The dashboard keeps role and +power-fit panels, but has no service-comparison control, source selector, target input or +matched-concurrency table. Use `serviceCompare`, `serviceBaseline`, `serviceComparator` and +`serviceTarget` as API query parameters; they have no dashboard share-parameter equivalents. + +`serviceCompare=true` returns exact opaque `serviceSources` keys, an +`equalServiceCurve`, and an optional `equalServiceComparison` at `serviceTarget`. +Use returned keys verbatim for `serviceBaseline` and `serviceComparator` via +`URLSearchParams`; never substitute bare GPU names. Each `label` is display text +(hardware and date, plus precision, topology or run only where two sources would +otherwise look alike); select and join by `key`, never by label. Omitted keys select the first +two sources; unknown explicit keys stay unavailable. Omitted target means no +selected-target result. Targets are tok/s/user for streaming speed, seconds for +TTFT/E2E. Concurrency is unsupported for equal-service interpolation. + +A source groups one logical curve snapshot and configuration. Stitched points retain +their original producer provenance; differing telemetry producer/exporter hashes do not +split the snapshot. Recipe, image and topology remain distinct. Rows without a snapshot +retain their own run, measured date and telemetry producer/exporter hashes in the source key. + +The API uses scoped observed points before frontier/best +pruning, with no power-comparison clones. It interpolates raw quantities +linearly only inside each exact source range, then reports signed +`100 × (comparator / baseline − 1)`. Metrics are measured GPU W/GPU, +whole-deployment output tokens/s, and validated GPU J/output token. Negative energy +change means lower comparator energy. Preserve bracket endpoint identities, +`interpolated`, missing reasons and nulls. Never call interpolated points new +measurements or bridge different sources, recipes, topologies or missing endpoints. + +`serviceCompare=true` also returns `matchedConcurrency`: the two selected sources +paired at each concurrency either observed. Sides are `observed`, `missing`, or +`ambiguous` (disagreeing duplicates, none chosen); `changePercent` needs both +observed. It is a same-load diagnostic: speeds usually differ, so never present it +as an equal-service result. + +`roleShare=true` returns validated disaggregated prefill/decode J/output and their +shares of reconstructed total energy. It converts prefill J/input using the +same-window aggregate J/output-to-J/input ratio; it does not compare unlike token +denominators. Missing roles are omitted. `rolePoints` adds per-observation role +W/GPU, role-local J/input and J/output (different denominators, never add them) +and the output-token reconstruction, with nulls for missing figures. + +`powerFit=true` returns `powerFits`: per source, an OLS line of measured mean W/GPU +against output tok/s per allocated GPU, with `P₀`, slope `m` (J/output token), R², +n, x-range, registry `tdpWatts` and point identities. Fewer than three distinct +rates return `fit: null`, `reason: "too-few-points"`. Call `P₀` an extrapolated +intercept, not idle power, and do not read the line outside its x-range. + +These analytical results require JSON; enabling any with CSV returns 400. Existing CSV remains +a plotted-point export. + +同等服务对比与同并发诊断仅通过 API 提供,查询参数为 `serviceCompare`、`serviceBaseline`、 +`serviceComparator` 和 `serviceTarget`;仪表板没有对应的服务对比控件、来源选择、目标值输入、 +同并发表格或分享参数。角色分析和功耗拟合面板仍在仪表板中提供,分别对应 `roleShare` 和 +`powerFit`。来源标识、插值规则、缺失原因和仅 JSON 的响应约束以上文说明为准。 + +For exact-load comparisons, use `xmode=concurrency`. The response resolves +`optimal=false` and `best=false`, retains every eligible observed load, and sets +`frontier.direction=null`; concurrency is not a speed or quality preference. +Filter `topologies` with one or more exact `point.topologyKey` values from a first +response (encode with `URLSearchParams`). This selects the same GPU count, +parallelism, role-pool split and offload mode across hardware without removing +loads. Unknown metadata stays unknown, not equivalent to an explicit setting. +Keep hardware, precision, recipe, run/date and software provenance separate; +matching topology alone does not establish a controlled hardware comparison. +Missing concurrency points must remain absent, not interpolated or extrapolated. + Fleet lifecycle defaults (ramp, cached-input percent, MTBI, recovery) follow the dashboard's lifecycle panel; read the current values from `/api/v1/views/options` rather than hard-coding them.