diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md index c65916275..dd99cf859 100644 --- a/docs/dashboard-readonly-views.md +++ b/docs/dashboard-readonly-views.md @@ -66,8 +66,34 @@ state. GPU interactive downsampling does not alter returned raw data or statisti Power boundary labels are GPU Level Measured, GPU Level Provisioned (TDP), All in Provisioned, and All in Measured. The last combines measured GPU power with modeled unmeasured components and PUE; it is not a wall-meter measurement. These labels and -collapsed power-assumption/availability notes do not change metric IDs, API selectors, +power-assumption notes do not change metric IDs, API selectors, or calculations. Profit comparison `powerLabel` display text follows the same names. +All in Measured watts and energy accept validated 8K/1K and AgentX rows through the +shared chart/API transform, including historical and unofficial rows. AgentX reuses +the chassis or rack model without independent workload calibration. Telemetry and +topology gates still apply; NVL72 needs complete Grace or module power. The standalone +Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope. + +For All in Measured, `tableRows` retains every GPU-valid observation in the selected +scope and best-series selection, including axis-clipped and non-frontier points. +Missing system estimates use `y: null`, `status: "unavailable"`, and +`unavailableReason`; `measuredGpuWatts` remains available. CSV exports these rows +with blank missing values. Numeric `series` and `count` are unchanged. Each date +comparison and unofficial overlay has its own `tableRows`; latest does not pool history. + +For GW-year profit, `modeled` selects valid system-power points before building the +curve at the same target, without extrapolation or another snapshot. `compare` uses +the same valid-curve throughput for both power budgets, retaining the original +provisioned estimate when no valid curve covers the target. Provisioned-only keeps +the original performance curve. Official, comparison and unofficial scopes remain +independent; CPU/module telemetry and compatible sensor-basis requirements still apply. + +Dense profit charts reserve readable space per bar and scroll within the plot on narrow +screens; captions and controls stay fixed. This is presentation-only: API selectors, +calculations, source identities and CSV rows are unchanged. PNG export includes the full +plot regardless of its current scroll position, so no API or skills contract change is needed. +Profit charts omit the per-configuration unavailable-estimate list. This is presentation-only; +the API retains skipped rows and reasons, and calculations and CSV exports are unchanged. The GPU statistics table includes startup and warmup for all chips in the selected series, regardless of chip visibility. It is separate from serving-window power, J/token and selected-time-window calculations. Run telemetry is DB-first with an @@ -201,8 +227,7 @@ point identity. Fewer than three distinct output rates return `fit: null` with power, and R² is null when power did not vary. These analytical results are JSON-only: `format=csv` with any analysis enabled returns -400, rather than silently exporting only the primary chart. Ordinary CSV retains -its existing plotted-point contract. +400, rather than silently exporting only the primary chart. CSV exports plotted points except for All in Measured, which exports `tableRows`. | Surface | Dashboard control / share parameter | Read-only API coverage | | ---------------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------- | @@ -233,6 +258,17 @@ those properties. @semianalysisai/inferencex-skills 包。所有仪表板路由(含隐藏和功能开关控制的 视图)均在覆盖表中登记;上表列出各只读接口接受的全部查询参数名。 接口复用现有计算函数,公开运行与非官方叠加数据保留各自来源。 +整体实测指标的 `tableRows` 保留当前筛选范围和最优曲线选择内所有 GPU 遥测有效的观测点, +不按前沿或坐标轴显示范围裁剪。估算不可用时 `y` 为 null,`status` 为 `unavailable`, +`unavailableReason` 给出原因;`measuredGpuWatts` 保留实测 GPU 功耗。CSV 导出同一组行,缺失值留空。 +`series` 和 `count` 保持不变,仍只包含可绘制的数值点;各日期对比和非官方叠加分别返回自己的 `tableRows`, +Latest 不会合并历史数据。 + +按 GW 年估算利润时,modeled 先筛选满足系统功耗要求的数据点,再在原目标值上构建曲线, +不外推,也不借用其他快照。compare 的两种功耗方案使用同一条有效曲线的吞吐量; +没有有效曲线覆盖目标时,保留原曲线的预配估算。provisioned 单独使用时沿用原性能曲线。 +官方数据、日期对比和非官方叠加各自独立计算;CPU/模块遥测要求和传感器口径兼容性要求同样不变。 + 私有上传、密钥、提示词、反馈及管理操作不作为公开读取接口。 OperatorX 的入口受功能开关控制,页面使用专属的 `/api/v1/operatorx/*` 接口;目前没有发布 `/api/v1/views/operatorx` 契约。 diff --git a/docs/data-transforms.md b/docs/data-transforms.md index ebf7e541e..22325d3ca 100644 --- a/docs/data-transforms.md +++ b/docs/data-transforms.md @@ -79,8 +79,8 @@ Returns `{ chartData: InferenceData[][], hardwareConfig: HardwareConfig }`. | B4 utility modeled (measured → chassis → PUE) | `modeledSystemPower.deploymentFacilityWatts ÷ modeledSystemPower.gpuCount` | `B1 J/out × (B4 W ÷ B1 W)` | `utilityModeledWatts`, `utilityModeledJPerOutputToken` | - **N_alloc and total throughput** (`powerBasisNormalization`). Aggregate rows report output per allocated GPU, so `N_alloc` cancels and the builder uses `W ÷ output_tput_per_gpu` without trusting display counts (legacy ingest can encode TP × EP twice). Fixed-sequence disaggregated rows (`disagg && benchmark_type === 'single_turn'`) report output per decode GPU while the deployment also powers the prefill pool, so `total output tok/s = output_tput_per_gpu × num_decode_gpu` and `N_alloc = num_prefill_gpu + num_decode_gpu`; the article's 4P+4D counts eight GPUs, and B3 energy is `(P + D) / D` times `jOutput` on such rows. Other disaggregated benchmark types emit no provisioned energy because it is not verifiable in-app whether AgentX throughput already divides by all GPUs. -- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. `deploymentFacilityWatts` is chassis AC × PUE with PUE applied exactly once inside `estimateChassisPower`, divided by the physical measured GPU count (not `modeledGpuCount`, which over-counts partially allocated chassis: a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8, and the unit test pins that divisor). This is a different quantity from the existing `modeledChassisPowerPerGpu` (chassis AC ÷ modeled GPU count, no PUE). Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator). -- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware such as GB200/GB300 NVL72 (`reason: 'hardware'`), and telemetry/topology failures withhold it. The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here). +- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. B4 W/GPU is `deploymentFacilityWatts ÷ gpuCount`. PUE is applied exactly once to rounded AC power inside `estimateChassisPower` or `estimateRackPower`. `modelSystemPower` combines the resulting chassis or tray shares and attributes partially allocated units to their measured GPUs; NVL72 tray shares come from one rack evaluated at the measured trays’ mean input. Divide the deployment total by the physical measured GPU count, not `modeledGpuCount` (a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8). The existing `modeledChassisPowerPerGpu` instead uses `chassisAcWatts ÷ modeledGpuCount`, without PUE. Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator). +- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware, and telemetry/topology failures withhold it. GB200/GB300 NVL72 are supported when validated GPU telemetry is accompanied by complete, matching Grace or module telemetry (`reason: 'cpu-telemetry'` when that CPU-side evidence is missing or invalid). The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here). - **Ordering invariant** (unit-tested on real B200 telemetry): B3 ≥ B4 ≥ B1 and B2 ≥ B1 for both W/GPU and J/out on the same point. - **Reconstructed prefill energy** (`reconstructedPrefillJPerOutputToken`, `utils/role-energy.ts`). For validated (`power_valid === 1`, schema 2) disaggregated rows, `prefill_joules_per_input_token × (joules_per_output_token ÷ joules_per_input_token)` carries the prefill pool's energy onto the output-token axis: the ratio is the served input:output token count because schema-2 aggregate energy has one numerator. With `decode_joules_per_output_token` it sums back to the deployment's J/out. It feeds only the `i_pcompare=roles` comparison on the energy axis (PowerX Figure 7) and is never a y-axis of its own; aggregate rows and rows missing any of the four inputs omit it. - Historical Trends substitutes `output_tput_per_gpu := tput_per_gpu` for legacy rows lacking output throughput; provisioned energies in trends inherit that fallback. diff --git a/docs/index.md b/docs/index.md index 2e6c0224a..a6ce4bb46 100644 --- a/docs/index.md +++ b/docs/index.md @@ -8,7 +8,7 @@ Design rationale and non-obvious conventions. See [CLAUDE.md](../CLAUDE.md) for - [Pareto Boundary API](./pareto-api.md): Query frontier and hinterland observations, preserve provenance, and distinguish API scope from chart highlights. - [API Skill Examples](./inferencex-api-examples.md) — Install the public skill, query benchmarks, export measured PowerX data, and explain empty results -- [PowerX System Power](./powerx-system-power.md) — Pinned chassis model, measured-input guards, assumptions, and reproducible article exports +- [PowerX System Power](./powerx-system-power.md) / [简体中文](./powerx-system-power.zh.md) — Measured-curve steps, chassis/NVL72 requirements, a worked planning example, missing-data diagnosis, and reproducible exports - [PowerX Permanent View](./powerx-permanent-view.md) — Power boundaries as gated Measured Energy metrics, `i_metric`/`i_rulers` share links, missing-value states - [PowerX Persistence and Recovery](./powerx-persistence-recovery.md) — Telemetry receipts, migration prerequisites, and targeted repair - [API Skill Releases](./inferencex-skills-release.md) — Prepare an immutable package, verify clean installations and agent exports, and publish through the package-specific workflow diff --git a/docs/powerx-permanent-view.md b/docs/powerx-permanent-view.md index ba51b2e70..a1bc6b33b 100644 --- a/docs/powerx-permanent-view.md +++ b/docs/powerx-permanent-view.md @@ -23,17 +23,17 @@ group stays out of the selector otherwise. The boundary metrics are members of t ## Boundaries -| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source | -| --------------------- | ---------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------- | -| `gpu-measured` | GPU measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged | -| `gpu-provisioned` | GPU provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` | -| `utility-provisioned` | Utility provisioned (all-in) | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) | -| `utility-modeled` | Utility modeled (PUE) | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` chassis AC × PUE 1.3 (applied once), ÷ measured GPUs | +| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source | +| --------------------- | --------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------------------------- | +| `gpu-measured` | GPU Level Measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged | +| `gpu-provisioned` | GPU Level Provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` | +| `utility-provisioned` | All in Provisioned | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) | +| `utility-modeled` | All in Measured | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` deployment AC ÷ measured GPUs × PUE (1.3 air, 1.1 NVL72; applied once) | Formulas, the all-GPU normalization (`N_alloc` = prefill + decode GPUs for disaggregated rows) and the null rules are specified in -[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis model -itself in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its +[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis and NVL72 rack models +in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its per-decode-GPU normalization; its labels and the boundary metric's `all GPUs` label keep the two distinguishable in the selector, the availability list and CSV headers. diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md index 4775f66de..1b5be58cd 100644 --- a/docs/powerx-system-power.md +++ b/docs/powerx-system-power.md @@ -1,280 +1,245 @@ -# Modeled system power in PowerX - -PowerX can compare measured GPU-board watts with estimated chassis AC watts for -the non-agentic 8192-input/1024-output workload. The existing app transformation -and the offline article exporter both call `modelSystemPower`; the API and -benchmark producer continue returning their original measurements. - -`system-power-model.profiles.json` records the pinned power model revision, -component source hashes, hardware mapping, complete platform configuration, and -fixed inference assumptions. Its profiles come from executing the original -Python components. `system-power-model.ts` preserves their nonlinear fan curve, -PSU efficiency interpolation, intermediate rounding, and PUE ordering. The -Python-generated reference cases test this implementation against the source. - -The pinned source currently identifies itself as **DRAFT / pending human -verification**. Numerical parity establishes implementation equivalence, not -empirical chassis calibration. - -## Updating the model for historical results - -Modeled power is derived from retained measurements when the browser or a shared -views API transforms a benchmark row. Changing the model does not rewrite the -original GPU measurements or require a per-run database backfill. - -1. Update `REVISION` in `packages/app/scripts/generate-system-power-reference.py` - to the intended clean Python model commit, and update the recorded assumptions - when required. -2. Run that script with the path to the pinned model checkout to regenerate - `system-power-model.profiles.json` and `system-power-model.reference.json`. - If equations or load-dependent components changed, update the TypeScript - implementation too; regenerating constants alone is insufficient. -3. Run the system-power model parity and admission tests, then deploy the app. - Existing browser sessions need the updated bundle. Derived API responses need - the normal authenticated cache invalidation or cache expiry; deployment alone - does not establish that every cached response uses the new revision. -4. Regenerate frozen CSV/JSON exports separately. If the revised model needs - inputs that were never recorded, those rows stay unavailable until the input - gap is resolved. A new benchmark's power must not be attached to an older - benchmark's throughput. - -## Boundary and assumptions - -The input is measured mean GPU power during a validated serving window. The -modeled chassis AC output adds the source model's CPU, DRAM, networking, storage, -board, fans, and PSU conversion losses. Facility power is a separate estimate: -PUE is applied after chassis AC, including the source's rounding order. - -The fixed README inference sweep uses `u_cpu=0.20`, `u_ram=0.20`, `u_pcie=0.05`, -and `u_nvme=0.0`. The pinned Python model defaults to PUE `1.2`; PowerX uses -`1.3` for its supported air-cooled chassis profiles. Utility power = critical IT -power × PUE (`1.3` air, `1.1` DLC). -The factor applies after chassis AC; measured GPU power and chassis AC do not change. -Cooling describes the modeled chassis, not verified benchmark-site cooling. -The current profiles do not model DLC; `--pue` remains an explicit facility-factor -override and does not convert an air-cooled chassis model into a DLC model. -Platform-specific network assumptions, -fan control, component counts, and chassis defaults are preserved in the -generated profile; every JSON export includes that profile and every CSV row -includes its applicable assumptions and profile hash. These are model inputs, -not measured CPU/DRAM utilization. - -| Hardware identity | Source chassis implementation | -| ----------------- | ------------------------------------------------------------- | -| `h100` | `human_verified/hgx_h100_chassis/h100_chassis_power_model.py` | -| `h200` | `human_verified/hgx_h200_chassis/h200_chassis_power_model.py` | -| `b200` | `human_verified/hgx_b200_chassis/b200_chassis_power_model.py` | -| `b300` | `human_verified/hgx_b300_chassis/b300_chassis_power_model.py` | -| `mi300x` | `human_verified/mi300x_chassis/mi300x_chassis_power_model.py` | -| `mi325x` | `human_verified/mi325x_chassis/mi325x_chassis_power_model.py` | -| `mi355x` | `human_verified/mi355x_chassis/mi355x_chassis_power_model.py` | - -All listed profiles describe a complete eight-GPU chassis. GB200 and GB300 have -no matching model and are unsupported. Their rack topology is not substituted -with B200 or B300. - -A partially allocated chassis (one to seven measured GPUs on one host) is -modeled at measured per-GPU power × 8. That is the same `n_gpu × W/GPU` input -the source sweep scripts feed each chassis model, and it assumes the unmeasured -GPUs run the same workload. The estimate is labeled `chassisBasis: -'extrapolated'`: per-GPU values divide by the modeled chassis GPU count -(`modeledGpuCount`), while `deploymentAcWatts` / `deploymentFacilityWatts` keep -only the measured GPUs' share of each chassis. This is not a proportional share -of a chassis evaluated at partial load; fixed components, the fan curve, and PSU -efficiency are all evaluated at full-chassis load. Missing or invalid telemetry, -inconsistent counts, missing host placement, more than one chassis per host, and -model-domain overflow remain unavailable. - -For a single-node deployment, the producer's physical width is `TP * PP * PCP`. -EP partitions that width. Some existing API configuration aliases contain -`TP * EP`; the model cross-checks the physical width against measured total and -per-GPU watts instead of trusting or summing those aliases. Multi-node and -disaggregated inputs require one chassis (one to eight GPUs) per measured -worker, distinct worker hosts, and consistent total/role watts. A role average -alone cannot establish physical placement or evaluate each host's nonlinear -model. CPU-only frontend workers are excluded from GPU-chassis counting. Separate CPU-only -frontend/router hosts are outside this estimate; CPU power within GPU chassis -still uses the source's fixed 20% utilization assumption. - -The default measured contract is numeric `power_valid=1` and metric schema 2. -The original validated single-node producer predates the schema marker but -already defines both watts fields identically. This path retains the absent -schema and reports `validated-unversioned-single-node`; it does not upgrade the -source or admit unversioned disaggregated power. The article receipt additionally -pins the producer checkout and retains each original audit artifact. - -## Profit Estimator power basis - -The per-GW Profit Estimator offers provisioned power, measured + modeled power, -and a paired comparison in Benchmark Config. Provisioned remains the default. -The control uses the existing insider feature gate and is hidden while locked. -Unlock with ↑↑↓↓ (`inferencex-feature-gate=1` in local storage). While locked, -`c_power` cannot activate an alternative calculation or fetch full power rows; -relocking restores provisioned estimates immediately. -The alternative reuses the same hardware, P90 target, throughput frontier, -token mix, prices, utilization, and per-GPU-hour costs. It changes only the -facility kW/GPU used to calculate capacity per GW. Consequently, revenue, -compute expense, license fee, and profit scale together; profit margin does not -change. Electricity expense is not recomputed separately. - -This opt-in AgentX estimate requires validated schema-v2 telemetry and chassis -supported by the pinned model. Fully measured eight-GPU chassis are supported -on a single node, per measured worker host, or across an aggregate multinode -deployment without per-worker telemetry at the deployment-mean GPU power -(`topologyBasis: 'uniform-hosts'`; symmetric TP/PP/DP shards load each host alike). -Validated single-node 1/2/4-GPU allocations use full-chassis extrapolation: fill -an eight-GPU server with whole replicas at the measured per-GPU power and -throughput, then divide modeled facility power by eight. This assumes replica -co-location does not change performance or power; it is not a measurement of a -partly idle server. The chart, tooltip, and CSV label every extrapolated estimate, -including interpolation with one partial knot. Unsupported GB200/GB300 chassis, -partial multi-host allocations, disaggregated deployments without per-worker -telemetry, allocations that cannot tile eight GPUs, and missing/invalid -measurements stay unavailable with distinct reasons. -The ordinary 8K/1K transformation keeps its existing admission policy. - -At an exact frontier point, use that point's modeled power. Between points, -estimate power linearly using the same two knots as the existing throughput -interpolation; never select a different point to fill a power gap. The estimate -uses PUE 1.3 and an additional 10% planning margin. These assumptions, including -the fixed CPU/DRAM utilization above, are not validated peak-load provisioning or -AgentX system calibration. The UI and CSV label the estimate and its assumptions. -`c_power=modeled` and `c_power=compare` preserve the selection in share URLs. -Unavailable historical estimates use the hardware registry when a chip is absent -from today's results and include the source date/run label. - -每 GW 利润估算器在基准测试配置中提供预配功耗、实测加建模功耗,以及两种方式的同口径 -对比;默认仍采用预配功耗。两种方式使用同一硬件、P90 目标、吞吐量前沿、token 比例、 -价格、利用率和每 GPU 小时成本,仅改变换算每 GW 容量时采用的设施功率。因此收入、 -计算成本、模型许可费和利润按相同比例变化,利润率不变;不会另行重新计算电费。 - -该选项由现有内部功能开关控制,锁定时隐藏。按 ↑↑↓↓ 解锁(本地存储 -`inferencex-feature-gate=1`)。锁定时,`c_power` 不会启用其他估算方式或触发完整功耗 -数据请求;重新锁定后立即恢复预配功耗估算。 - -AgentX 估算仅接纳通过验证的 schema-v2 功耗,且要求单节点机箱及适用模型。 -实测单卡、双卡或四卡配置可复用现有整机外推:假设在八卡服务器上部署多个完整实例, -每卡功耗和吞吐量保持不变,再将建模设施功耗除以八。这要求实例共置不改变性能或功耗, -不代表部分 GPU 闲置时的整机实测功耗。图表、提示框和 CSV 均标注整机外推;若插值 -使用的任一数据点采用外推,也保留该标注。GB200/GB300 等无匹配模型的机箱、多节点 -配置、无法整除八卡的实例,以及缺失或无效功耗仍不可用,并分别说明原因。原有 -8K/1K 转换路径的接纳规则不变。精确前沿点使用自身的功耗;点间采用原吞吐量插值的 -同一对数据点线性估算功耗,不换用其他点填补缺失。PUE 取 1.3,另加 10% 功耗余量; -这些假设和上述固定 CPU/DRAM 利用率尚未通过 AgentX 系统校准,也不构成峰值供电容量 -验证。界面与 CSV 会注明估算及其假设,分享链接通过 `c_power` 保留所选方式。 -历史估算不可用时,若当天结果不含该芯片,则从硬件注册表获取名称;提示会附上来源 -日期或运行标签,避免与当前结果混淆。 - -## Offline comparison export - -The exporter reads a local cohort envelope and writes a **new** output directory: +# PowerX system power and smart provisioning + +[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md) + +Use this guide to inspect measured curves, estimate system power, and compare +capacity under a fixed facility power budget. Start with [the dashboard steps](#inspect-a-measured-curve), +then use [the hardware requirements](#hardware-and-telemetry-requirements) or +[troubleshooting](#when-a-curve-or-estimate-is-missing) when a result is unavailable. + +The system model is **DRAFT / pending human verification**. Its regression +fixtures establish numerical consistency, not empirical calibration. AgentX +estimates, including Kimi K3, are planning previews; they do not establish AgentX +calibration or safe peak-load provisioning. + +## Choose the power boundary + +| Dashboard boundary | What the value represents | +| --------------------------- | ---------------------------------------------------------------------------------------------------- | +| GPU Level Measured | Validated GPU-board power during the benchmark window. | +| GPU Level Provisioned (TDP) | Hardware TDP; a reference, not a reading from this run. | +| All in Provisioned | The hardware registry's fixed facility kW/GPU allowance. | +| All in Measured | Measured inputs plus modeled system components and facility overhead. It is not measured wall power. | + +The examples below use W per GPU and J per output token. GPU measurements remain +available independently of whether the system model can accept the row. Measured +P75/P90 values are time-weighted percentiles of synchronized fleet GPU power, +divided by GPU count; neither the average nor individual-device percentiles +substitute for them. + +## Inspect a measured curve + +**Prerequisites:** a benchmark selection with retained power data. Unlock the +experimental power controls with ↑↑↓↓ if they are hidden. + +1. Open `/inference`, select the model (for example, Kimi K3), workload, date/run, + engine, precision, and hardware. Keep those selections fixed when comparing + boundaries; different engines or historical runs are different curves. +2. Choose a measured-power metric and **GPU Level Measured**. Use **Table** to + inspect numeric rows and select a chart point to inspect its measurement + provenance. Start with average W/GPU: energy also requires a valid token + denominator, and P75/P90 require retained percentile measurements. +3. Switch to **All in Measured** to inspect facility estimates. This boundary + supports 8K/1K single-turn results and AgentX previews. The separate + **Modeled Chassis AC** metric remains limited to 8K/1K single-turn results. +4. For a capacity comparison, open `/profit-estimator-per-gigawatt`, choose the + model and a supported interactivity target, then select **Compare both** in + Benchmark Config. Match each result's workload, engine and precision to the + inference selection. Hover or select a bar for its power basis and read the + formula notes below. + +**Expected result:** the All in Measured table keeps every GPU-valid record in +the selected scope, including B200/H200 multi-node deployments. It shows measured +GPU power even when an all-in estimate is unavailable; the estimate displays +`—` with a reason and stays blank in CSV. The graph plots numeric estimates only. +The Profit Estimator adds target-range and financial requirements. All in +Measured builds its performance frontier from power-valid measurements in the +selected scope. Inference charts and tables also support unofficial-run overlays. + +## Hardware and telemetry requirements + +| Hardware / deployment | System estimate requires | +| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| H100, H200, B200, B300, MI300X, MI325X, MI355X | Valid GPU-board telemetry; an eight-GPU chassis profile models CPU, DRAM, networking, storage, board, fans and PSU losses. | +| B200/H200 and other supported chassis across multiple hosts | Consistent GPU counts and watts. Per-worker data must identify one chassis per distinct host. Aggregate non-disaggregated rows without workers may use complete eight-GPU hosts at the deployment mean (`uniform-hosts`). | +| Prefill/decode disaggregation | Per-worker host placement and role power consistent with deployment totals. A role average alone cannot establish each host's load. Separate CPU-only frontend/router hosts are outside the estimate. | +| GB200 / GB300 NVL72 | Valid GPU telemetry **and** complete, independently valid compute-module or Grace-socket telemetry from the same window: four GPUs and two sockets per compute tray. | + +The normal contract is numeric `power_valid=1`, +`power_metric_schema_version=2`, positive `avg_power_w` (W/GPU) and +`avg_total_gpu_power_w` (deployment W), and a consistent physical GPU count. +The model retains a legacy validated single-node exception with no schema +marker as `validated-unversioned-single-node`; this does not admit unversioned +multi-node, disaggregated or NVL72 rows. Profit planning always requires schema 2. + +**NVL72 sensor boundary:** `cpu_power_valid=1` and `power_audit.cpu` must record +matching expected/observed socket counts, two per tray. Either: + +- `avg_total_module_power_w` with `sensor_kind: module` covers GPU, HBM, Grace + and LPDDR5X together. Do not add GPU or Grace watts again. +- `avg_total_cpu_power_w` and `avg_cpu_socket_power_w` with + `sensor_kind: grace_socket` must agree with the socket count. The model adds + GPU-board power and a regulator allowance of GPU W × 0.15 / 0.85. + +CPU-rail-only, missing or unknown sensor provenance is insufficient. A present +but invalid module measurement remains unavailable; it does not silently fall +back to Grace readings. Model support does not imply that every producer version +collects these fields. + +**Partial allocations:** a one-to-seven-GPU chassis is modeled as eight GPUs at +the measured per-GPU load and labeled `extrapolated`. Deployment totals retain +only the measured GPUs' share. Profit planning accepts single-node 1/2/4-GPU +chassis allocations as whole-replica extrapolations, assuming co-location leaves +performance and power unchanged. Other partial layouts, including partial NVL72 +trays, are not accepted for profit planning. A module sensor already covers its +whole tray, so that reading is never scaled to fill unmeasured GPUs. + +## One worked NVL72 example + +**Input:** the [GB300 reference fixture](../packages/app/src/lib/system-power-model.reference.json) +with **3,000.75 W of module power per complete tray** and PUE 1.1. +This is a numerical fixture, not a measured Kimi K3 result. An actual benchmark +must separately satisfy both GPU and CPU/module validation. + +1. Scale the measured mean module input to 18 compute trays. Add modeled tray + components, nine NVSwitch trays, conversion losses and management switches. +2. Evaluate the power-shelf efficiency once at the combined rack load. Every + measured tray receives the same 1/18 rack share; this assumes the remaining + trays run at that mean load, rather than measuring actual rack occupancy. +3. Apply PUE once to rack AC power, then divide by 72 GPUs. +4. Apply the separate 10% planning reserve for the Profit Estimator. + +| Stage | Watts | +| ----------------------------------------------- | -------: | +| Module input × 18 trays | 54,013.5 | +| Modeled compute-tray components | 11,466.0 | +| Modeled NVSwitch trays | 4,107.6 | +| Tray conversion losses | 1,967.8 | +| Rack DC, including 200 W of management switches | 71,754.9 | +| Rack AC, after shelf losses | 74,904.6 | +| Facility power after PUE 1.1 | 82,395.1 | + +```text +Planning kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814 +GPU capacity per GW = 1,000,000 / 1.258814 ≈ 794,399 +GPU-hours per GW-year = GPU capacity × 8,760 +``` -```sh -bun packages/app/scripts/export-modeled-system-power.ts \ - --input /path/to/original-qwen-article.input.json \ - --output /path/to/new-original-qwen-comparison +An eight-GPU benchmark on two complete trays receives +82,395.1 × 8 / 72 ≈ 9,155.0 W of facility power, giving the same per-GPU result. +It has not measured all 72 GPUs. Intermediate values are rounded; summing the +displayed components can differ by 0.1 W. Planning uses the retained facility +total, not the fixture's rounded per-GPU display value. + +## Assumptions behind the estimate + +- **PUE:** the dashboard applies 1.3 to air-cooled chassis and 1.1 to NVL72, + after AC conversion losses. The historical profile default of 1.2 is not the + dashboard default. Cooling describes the model, not verified site cooling. + Changing PUE does not convert an air-cooled chassis model into a DLC model. +- **Chassis overhead:** coefficients reflect fixed assumptions of 20% CPU/DRAM + utilization, 5% PCIe utilization and idle NVMe. These are not live utilization + readings. Each host's nonlinear fan/PSU model is evaluated at its own load; + `uniform-hosts` explicitly substitutes the deployment mean for every host. +- **NVL72 overhead:** Grace and LPDDR5X are measured. Rack networking, switches, + fans, board residuals, conversion losses and power shelves are modeled. + [Profiles](../packages/app/src/lib/system-power-model.profiles.json) retain + component values, source status and ranges for unverified parameters. Rack DC + above the 264 kW installed shelf capacity is outside the model domain. +- **Planning reserve:** facility kW/GPU × 1.10 is a separate capacity buffer. + Average power plus this reserve is not a validated electrical peak limit. +- **Measured curves:** All in Measured selects power-valid points before building + the performance frontier. Throughput and power use that curve at the requested + target; the curve uses a compatible model, PUE, topology and sensor basis. + No target extrapolation or borrowing from unselected history occurs. +- **Matched comparison:** Compare both uses the same power-valid curve, target + and financial inputs for its paired bars. If no measured estimate is available, + the ordinary provisioned result remains. All in Provisioned keeps the ordinary + performance frontier. + +Lower planning power increases GPU capacity per GW. Revenue, compute cost and +license fees scale with that capacity under the fixed per-GPU assumptions; profit +margin and per-chip-hour economics do not improve. Electricity expense is not +recomputed separately. + +## When a curve or estimate is missing + +Compare the same model, workload, date/run, engine, precision and metric first. +Chart/table rows and target-based profit estimates answer different questions. + +| Symptom / reason | Check and next action | +| ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| GPU curve exists, All in Measured is absent | Check supported workload/hardware and the row's system-model status. Valid GPU power alone does not establish a system estimate. | +| B200/H200 multi-node system estimate is unavailable (`topology`, `role-power`, `gpu-count`) | Inspect physical GPU count, host placement and total/role watts. Use the original producer topology; do not infer chassis placement from a display label or sum TP/EP aliases. | +| NVL72 reports `cpu-telemetry` / `no-cpu-power` | Inspect the same-window CPU audit, sensor kind and complete socket coverage. GPU validity remains independent. | +| `telemetry` / `no-measured-power` | Check the original validation audit and raw samples. Reprocess only when the retained evidence supports the original window; otherwise collect replacement performance and power together. | +| `outside-measured-range` or power-invalid target bracket | Choose a target within the selected power-valid curve. Both bounding points need compatible valid power; points outside the selected scope cannot fill the gap. | +| `incompatible-power-basis` | Do not interpolate between module and GPU-plus-Grace readings, or different model/PUE bases. | +| No cost, token mix or provisioned power | Inspect the financial inputs. This can prevent both profit estimates even when power is valid. | +| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. | + +In **Compare both**, a valid provisioned result remains when its measured estimate +is unavailable. Measured-only mode never substitutes provisioned watts. See +[persistence and recovery](./powerx-persistence-recovery.md) for retained telemetry +and targeted repair; new power readings cannot be attached to old +throughput results. + +## Provenance and reproducible exports + +Inspect the point's measurement source and the estimate's model revision, PUE, +sensor basis and topology. Profit tooltips identify the power basis; formula +notes and CSV columns `Power basis`, `Power sensor` and `System power profile` +provide the assumptions and source details. +The model revision is an `app-sha256:` digest, not a benchmark run ID or Git commit. + +The [offline exporter](../packages/app/scripts/export-modeled-system-power.ts) +accepts a local `ComparisonInput` envelope with original `BenchmarkRow` entries, +metadata, and optional original artifacts/audits. Use the script's type for the +complete shape; preserve run, attempt, producer revision, capture time and hashes. +The exporter follows the ordinary 8K/1K model policy, not the AgentX preview opt-in. +Run from the repository root with installed dependencies and a new output path: +```sh bun packages/app/scripts/export-modeled-system-power.ts \ - --input /path/to/qwen35-current.input.json \ - --output /path/to/new-current-qwen-comparison --pue 1.3 + --input /path/to/cohort.input.json \ + --output /path/to/new-comparison ``` -The maintained input shape is `ComparisonInput` in the script: - -```ts -{ - cohort: string, - metadata: { /* source URLs, capture times, hashes and cohort selection */ }, - rows: [{ - id: string, // stable observation identity - cell?: string, // optional group of original replicates - benchmark: BenchmarkRow, // original API row or existing ETL output - rawInput?: unknown, // original artifact before ETL normalization - source?: object, // run, attempt, producer revision and artifact receipt - audit?: object // matching original power-validation sidecar - }] -} +Outputs are `comparison.json`, `comparison.csv` and optional `cells.csv`. JSON +retains original rows, audits, validity, model outputs and provenance; unavailable +CSV values stay blank. Optional `--pue 1.3` overrides the facility factor for +**every** row and is recorded in metadata. Without it, per-hardware defaults apply. +See [API examples](./inferencex-api-examples.md) for measured-data extraction. + +Modeled energy requires a matching audit with exact telemetry duration, physical +GPU count and successful request/token denominators. It is modeled deployment +power × duration, not integrated measured wall energy. No kernel-level +prefill/decode energy is inferred. Each replicate is modeled before aggregation; +a cell mean remains unavailable if any scoped replicate is unavailable. + +## Maintain the model + +| Responsibility | Source | +| --------------------------------------------------- | ------------------------------------------------------------------------------- | +| Equations, nonlinear curves and rounding | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) | +| Component parameters, assumptions and source status | [profiles](../packages/app/src/lib/system-power-model.profiles.json) | +| Workload, telemetry, topology and PUE admission | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) | +| Matched frontier and planning reserve | [profit-power.ts](../packages/app/src/components/calculator/profit-power.ts) | +| Frozen numerical baseline | [reference fixtures](../packages/app/src/lib/system-power-model.reference.json) | +| Model identity and source hashes | [provenance](../packages/app/src/lib/system-power-model.provenance.json) | + +Edit equations or active coefficients together with justified expected values and +assumption metadata. Changing labels such as `u_cpu` alone does not change watts. +Refresh and check the manifest: + +```sh +bun packages/app/scripts/update-system-power-provenance.ts +bun packages/app/scripts/update-system-power-provenance.ts --check ``` -Use the existing `normalizeArtifactRows` / `mapBenchmarkRow` for raw producer -aggregates. Keep the original aggregate as `rawInput`, retain original schema -markers, and check that its measured metrics survive normalization unchanged. -Use complete raw API responses for current snapshots, then select the exact -`single_turn`, `isl=8192`, `osl=1024` workload locally. Retain every scoped row, -including unsupported hardware and missing/invalid power; never mix a current -snapshot into the frozen article campaign. - -Outputs are `comparison.json`, `comparison.csv`, and optional `cells.csv`. -The JSON includes original input rows, measured validity, modeled outputs, -assumptions, audit windows, and source provenance. CSV includes separate measured -and modeled columns; unavailable numbers are blank. Invalid/unverified raw -values remain in the raw-input record and are not labeled valid measurements. -Metadata records input and implementation SHA-256 hashes, model and application -revisions, worktree state, generation time, and full profile provenance. Generate -the final release export from the intended application commit; file hashes also -identify any local changes during development. - -Modeled energy is available only when a matching valid audit sidecar supplies an -exact telemetry duration, physical GPU count, and successful request/token -denominators. It is modeled deployment power (the measured GPUs' share of each -chassis) multiplied by that duration, not a time integral of measured wall power. Actual output-token counts are used; -nominal `1024` tokens per query never replace recorded counts. No kernel-level -prefill/decode energy is inferred. API snapshots without these sidecars receive -power estimates only. - -Each replicate is modeled before aggregation. A cell mean averages its modeled -replicate outputs; it does not evaluate the model at mean watts. If any replicate -is unavailable, the corresponding mean remains unavailable rather than silently -dropping that replicate. - -## 中文说明 - -模型结果在浏览器或共享 views API 转换 benchmark 数据时计算,不写回原始 GPU -测量值。更新模型时,先修改生成脚本中的固定版本及相关假设,再生成 profiles 和 -reference JSON;如果公式或随负载变化的组件有改动,还需同步 TypeScript 实现。 -通过一致性及准入测试后部署,刷新浏览器,并使派生 API 缓存失效或等待其过期。 -冻结的 CSV/JSON 需另行导出。通常无需逐 run 回填数据库;若新模型需要历史记录中 -没有的输入,应保留不可用状态,也不能把新一轮测得的功耗配到旧吞吐结果上。 - -PowerX 的系统功耗结果以实测 GPU 功率为输入,使用固定版本的功耗模型估算 -8-GPU 机箱的 AC 输入功率,再单独应用 PUE 得到设施功率估计。CPU 和 DRAM 利用率 -均假设为 20%;这些是模型参数,不是实测利用率。完整平台配置、源码版本和校验和 -随导出结果保留。模型源码仍标记为待人工核验,数值一致性不代表完成了实机校准。 - -固定版本的 Python 模型默认 PUE 为 1.2;PowerX 对当前风冷机箱模型 -采用 1.3。市电侧功率 = IT 负载功率 × PUE,风冷取 1.3,直接液冷(DLC) -取 1.1。PUE 仅作用于机箱交流功率,不改变 GPU 实测功率或机箱交流功率。这里的 -冷却方式指建模机箱,并非已核实的测试站点配置。当前模型不支持 DLC;`--pue` 仅 -覆盖设施功率系数,不会把风冷机箱模型转换为液冷模型。 - -仅使用部分 GPU 的机箱(单台主机上实测 1–7 张 GPU)按实测每卡功率 × 8 建模, -与模型源码 sweep 脚本喂给各机箱模型的 `n_gpu × W/GPU` 输入一致,并假设未实测的 -GPU 运行相同负载。结果标记为 `chassisBasis: 'extrapolated'`:每卡数值按建模机箱 -的 GPU 总数分摊,`deploymentAcWatts` 只保留实测 GPU 在各机箱中的份额。这不是把 -部分分配的机箱按比例分摊:固定组件、风扇曲线和 PSU 效率都在满机箱负载点求值。 -GB200、GB300 没有匹配模型,也不能套用 B200、B300 模型。缺失、无效和不支持的 -情况保持不可用。纯 CPU frontend worker 不计入 GPU 机箱数;独立的纯 CPU -frontend/router 主机不在估算范围内,GPU 机箱内的 CPU 功率仍按 20% 利用率计算。 - -导出时每次测量先独立计算,再对同一 cell 的重复测量取平均。能耗使用审计记录中的 -实际窗口和成功 token 数,按实测 GPU 的份额计算,明确标记为估计值,不改写原有 -GPU 实测指标。当前 API 快照与原文章冻结数据分别导出,避免混用不同时间和配置的 -结果。 - -## Measured P75 and P90 GPU power - -`y_measuredP75Power` and `y_measuredP90Power` show the time-weighted P75 and P90 of -synchronized fleet GPU-board power over the validated load window, divided by GPU -count. They share the regular measured-power chart path for official points and -unofficial overlays. Missing or unvalidated percentile data remains unavailable; -average power is never used as a substitute. These metrics are separate from modeled -chassis AC power and individual-device percentiles. - -P75 and P90 backfills use the same 34 original validated traces and exact windows -recorded in `docs/data/power-p90-backfill.json`. - -`y_measuredP75Power` 和 `y_measuredP90Power` 分别显示已验证负载窗口内整组 GPU -功耗按时间加权的 P75 和 P90,再按参与测量的 GPU 数量均摊。正式数据与非正式 -运行叠加层使用同一计算和绘图路径。缺少测量值或未通过验证时保持不可用, -不会用平均功耗替代。该指标与机箱交流功耗估算、单个设备的功耗分位数不同。 -两个分位数均由审计记录中的同一批 34 份原始遥测及其测量窗口重新计算。 +Run affected model, admission, planning, views API and export checks. Keep the +496 historical reference cases frozen; baseline parity does not establish +calibration. The app owns the model and parameters; no private Python repository +is required. Historical measurements are modeled on read, so a model-only change +needs an updated app bundle and API cache refresh/expiry, not a raw-data backfill. +Regenerate frozen exports separately; missing source measurements remain missing. diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md new file mode 100644 index 000000000..84e7cb700 --- /dev/null +++ b/docs/powerx-system-power.zh.md @@ -0,0 +1,211 @@ +# PowerX 系统功耗与智能容量规划 + +[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md) + +本指南介绍如何查看实测曲线、估算系统功耗,以及比较固定设施功率预算下的 GPU 容量。 +先按[仪表板操作步骤](#查看实测曲线)选择数据;结果缺失时,再查阅 +[硬件要求](#硬件与遥测要求)和[排查说明](#曲线或估算结果缺失时如何排查)。 + +系统模型仍为 **DRAFT / pending human verification(草稿,待人工核验)**。 +回归用例只能证明数值实现一致,不能证明已完成实机校准。包括 Kimi K3 在内的 +AgentX 估算属于容量规划预览,不代表已完成 AgentX 校准,也不能作为安全峰值供电规划的依据。 + +## 选择功耗边界 + +| 仪表板边界 | 数值含义 | +| --------------------------- | ---------------------------------------------------------------- | +| GPU Level Measured | 基准测试窗口内通过验证的 GPU 板卡实测功率。 | +| GPU Level Provisioned (TDP) | 硬件 TDP 参考值,不是本次运行的测量值。 | +| All in Provisioned | 硬件注册表中固定的设施 kW/GPU 配额。 | +| All in Measured | 实测输入加上系统组件和设施开销的模型估算,不是墙上功率计的读数。 | + +下文示例采用 W/GPU 和 J/输出 token。系统模型不支持某条记录,并不影响 +该记录有效 GPU 测量值的展示。实测 P75/P90 是全部 GPU 同步汇总功率的时间加权 +分位数,再除以 GPU 数量;平均值和单卡分位数都不能替代。 + +## 查看实测曲线 + +**前提:**所选基准测试保留了功耗数据。若实验性功耗控件未显示,用 ↑↑↓↓ 解锁。 + +1. 打开 `/inference`,选择模型(如 Kimi K3)、工作负载、日期/运行、引擎、精度和 + 硬件。比较不同功耗边界时保持这些选项一致;不同引擎或历史运行属于不同曲线。 +2. 选择实测功耗指标和 **GPU Level Measured**,通过 **Table** 查看数值记录,点击 + 图表上的数据点查看测量来源。先看平均 W/GPU:能耗还需要有效的 token 分母, + P75/P90 则需要保留相应分位数测量值。 +3. 切换到 **All in Measured** 查看设施功率估算。该边界支持 8K/1K 单轮结果和 + AgentX 预览;独立的 **Modeled Chassis AC** 指标仍仅支持 8K/1K 单轮结果。 +4. 要比较容量,打开 `/profit-estimator-per-gigawatt`,选择模型和曲线支持的 + 交互性目标,再在 Benchmark Config 中选择 **Compare both**。逐项核对结果的 + 工作负载、引擎和精度是否与推理页面所选一致。悬停或选中柱形查看功耗依据, + 并阅读图表下方的公式说明。 + +**预期结果:**All in Measured 表格保留当前筛选范围内所有 GPU 功耗有效的记录, +包括 B200/H200 多节点部署。即使无法计算整体功耗估算,仍会显示实测 GPU 功率; +缺失的估算显示为 `—` 并注明原因,CSV 中对应数值留空。图表只绘制有数值的估算。 +利润估算器另有目标范围和财务输入要求。All in Measured 在当前筛选范围内, +先选出功耗有效的实测点,再构建性能前沿。推理图表和表格也支持非官方运行叠加。 + +## 硬件与遥测要求 + +| 硬件 / 部署方式 | 系统估算所需输入 | +| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| H100、H200、B200、B300、MI300X、MI325X、MI355X | 有效 GPU 板卡遥测;八卡机箱模型估算 CPU、DRAM、网络、存储、主板、风扇和 PSU 损耗。 | +| B200/H200 等受支持机箱的多主机部署 | GPU 数量与功率须一致。逐 worker 数据须明确每台独立主机对应一个机箱。没有 worker 数据的非分离式多节点记录,可按完整八卡主机使用部署平均负载(`uniform-hosts`)。 | +| Prefill/decode 分离式部署 | 逐 worker 的主机分布和角色功率须与部署总量一致。仅有角色平均值无法确定各主机负载。独立的纯 CPU frontend/router 主机不在估算范围内。 | +| GB200 / GB300 NVL72 | 有效 GPU 遥测,**以及**同一窗口内完整且独立通过验证的计算模块或 Grace socket 遥测:每个计算 tray 为四张 GPU、两个 socket。 | + +常规输入要求数值型 `power_valid=1`、`power_metric_schema_version=2`, +`avg_power_w`(W/GPU)和 `avg_total_gpu_power_w`(部署总 W)均为正值,且物理 +GPU 数量一致。模型对已验证但没有 schema 标记的历史单节点记录保留例外,并标为 +`validated-unversioned-single-node`;该例外不适用于未标版本的多节点、分离式或 +NVL72 记录。利润规划始终要求 schema 2。 + +**NVL72 传感器边界:**要求 `cpu_power_valid=1`,且 `power_audit.cpu` 记录的预期 +和实际 socket 数量一致,每个 tray 两个 socket。支持两种输入: + +- `avg_total_module_power_w` 配合 `sensor_kind: module`,已覆盖 GPU、HBM、Grace + 和 LPDDR5X,不能再次加上 GPU 或 Grace 功率。 +- `avg_total_cpu_power_w` 和 `avg_cpu_socket_power_w` 配合 + `sensor_kind: grace_socket`,总功率须与单 socket 平均功率及 socket 数量相符。 + 模型另加 GPU 板卡功率, + 以及 GPU W × 0.15 / 0.85 的稳压损耗估算。 + +仅有 CPU rail 读数、缺少来源信息或传感器类型未知,都不满足要求。module 字段 +存在但无效时,结果保持不可用,不会自动改用 Grace 读数。模型支持某字段,并不 +代表每个生产端版本都已采集该字段。 + +**部分 GPU 分配:**单机箱仅测量一至七张 GPU 时,按实测单卡负载外推到八卡机箱, +标为 `extrapolated`;部署总量只计入已测 GPU 的份额。利润规划接受单节点 1/2/4 卡 +机箱配置的整副本外推,前提是假设副本共置不改变性能和功耗。其他部分分配方式, +包括未测满的 NVL72 tray,不用于利润规划。module 传感器已覆盖整个 tray,因此 +不会放大其读数来补足未测 GPU。 + +## 一个 NVL72 计算示例 + +**输入:**[GB300 参考用例](../packages/app/src/lib/system-power-model.reference.json) +中,每个完整 tray 的 module 功率为 **3,000.75 W**,PUE 为 1.1。这是数值回归 +用例,不是 Kimi K3 实测结果。实际基准测试还须分别通过 GPU 和 CPU/module 验证。 + +1. 将实测平均 module 输入扩展到 18 个计算 tray,加上模型中的 tray 组件、九个 + NVSwitch tray、转换损耗及管理交换机。 +2. 在整个机架的合计负载上计算一次电源架效率。每个已测 tray 分得 1/18 的机架 + 功率;这是假设其余 tray 也处于相同平均负载,并非测量了实际机架占用情况。 +3. 对机架 AC 功率应用一次 PUE,再除以 72 张 GPU。 +4. 利润估算器另加 10% 的规划余量。 + +| 阶段 | 功率(W) | +| ------------------------------ | --------: | +| Module 输入 × 18 个 tray | 54,013.5 | +| 建模的计算 tray 组件 | 11,466.0 | +| 建模的 NVSwitch tray | 4,107.6 | +| Tray 转换损耗 | 1,967.8 | +| 机架 DC,包含 200 W 管理交换机 | 71,754.9 | +| 计入电源架损耗后的机架 AC | 74,904.6 | +| 应用 PUE 1.1 后的设施功率 | 82,395.1 | + +```text +规划 kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814 +每 GW 的 GPU 容量 = 1,000,000 / 1.258814 ≈ 794,399 +每 GW 每年的 GPU 小时 = GPU 容量 × 8,760 +``` + +若基准测试使用两个完整 tray、共八张 GPU,其设施功率份额为 +82,395.1 × 8 / 72 ≈ 9,155.0 W,归一到每张 GPU 后结果相同。这不代表测量了全部 +72 张 GPU。中间值经过舍入,展示值相加可能相差 0.1 W。规划计算使用保留的设施 +总功率,不使用参考用例中已舍入的单卡展示值。 + +## 估算采用的假设 + +- **PUE:**仪表板对风冷机箱使用 1.3,对 NVL72 使用 1.1,均在 AC 转换损耗之后 + 应用。历史 profile 的默认值 1.2 不是仪表板默认值。冷却类型描述的是模型,不是 + 经核实的现场实际冷却方式;改变 PUE 不会把风冷机箱模型变成 DLC 模型。 +- **机箱开销:**系数采用固定假设:CPU/DRAM 利用率 20%、PCIe 利用率 5%、NVMe + 空闲。这些不是实时利用率读数。各主机的非线性风扇/PSU 模型按本机负载计算; + `uniform-hosts` 则明确假设每台主机都采用部署平均负载。 +- **NVL72 开销:**Grace 和 LPDDR5X 使用实测值;机架网络、交换机、风扇、主板 + 其余开销、转换损耗和电源架由模型估算。[Profiles](../packages/app/src/lib/system-power-model.profiles.json) + 保留组件参数、来源状态和未核实参数的范围。机架 DC 超过已安装电源架容量 + 264 kW 时,超出模型定义域。 +- **规划余量:**设施 kW/GPU × 1.10 是独立的容量缓冲。平均功率加此余量不等于 + 经过验证的供电峰值上限。 +- **实测曲线:**All in Measured 先选出功耗有效的数据点,再构建性能前沿。 + 吞吐量和功耗都按这条曲线在所选目标处计算;整条曲线的模型版本、PUE、拓扑 + 和传感器口径须兼容。不做目标范围外推,也不借用未选中历史曲线上的数据点。 +- **匹配比较:**Compare both 的实测与预配柱形使用同一条功耗有效曲线、同一目标 + 和相同财务输入。没有可用实测估算时,仍保留常规预配结果。All in Provisioned + 继续使用常规性能前沿。 + +较低的规划功率会提高每 GW 可容纳的 GPU 数量。在固定单卡假设下,收入、计算成本 +和授权费用都随容量变化;利润率和每芯片小时的经济指标不会改善,也不会另行重算电费。 + +## 曲线或估算结果缺失时如何排查 + +先确认比较的是同一模型、工作负载、日期/运行、引擎、精度和指标。图表/表格记录与 +按性能目标计算的利润估算,回答的问题不同。 + +| 现象 / 原因 | 检查项与下一步 | +| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | +| 有 GPU 曲线,但没有 All in Measured | 检查工作负载、硬件是否受支持,以及该记录的系统模型状态。GPU 功耗有效不等于系统估算可用。 | +| B200/H200 多节点系统估算不可用(`topology`、`role-power`、`gpu-count`) | 检查物理 GPU 数量、主机分布和总功率/角色功率。使用原生产端拓扑,不从展示名称推断机箱分布,也不累加 TP/EP 别名。 | +| NVL72 返回 `cpu-telemetry` / `no-cpu-power` | 检查同一窗口的 CPU 审计、传感器类型和完整 socket 覆盖;GPU 有效性独立判断。 | +| `telemetry` / `no-measured-power` | 检查原验证审计及原始样本。只有保留证据足以支持原窗口时才重处理,否则须重新采集匹配的性能和功耗。 | +| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前功耗有效曲线范围内的目标。插值两端的功耗须有效且口径兼容,不能用筛选范围外的数据点补齐缺口。 | +| `incompatible-power-basis` | 不在 module 读数与 GPU 加 Grace 读数之间插值,也不在不同模型版本或 PUE 取值之间插值。 | +| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 | +| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 | + +**Compare both** 在实测估算不可用时仍保留有效预配结果。仅实测模式不会用预配 +功率代替。保留遥测和定向修复方式见[持久化与恢复](./powerx-persistence-recovery.md);新采集的功率不能附到旧吞吐量上。 + +## 来源与可复现导出 + +核对数据点的测量来源,以及估算的模型版本、PUE、传感器边界和拓扑。利润图表提示框 +标明功耗依据;公式说明及 CSV 中的 `Power basis`、`Power sensor`、 +`System power profile` 列提供假设和来源细节。模型版本是 `app-sha256:` 摘要,不是基准 +测试运行 ID,也不是 Git commit。 + +[离线导出器](../packages/app/scripts/export-modeled-system-power.ts)读取本地 +`ComparisonInput`,其中包含原始 `BenchmarkRow`、元数据,以及可选的原始产物/ +审计。完整结构以脚本中的类型为准;保留运行、attempt、生产端版本、获取时间和哈希。 +导出器采用常规 8K/1K 模型策略,不启用 AgentX 预览。安装项目依赖后,从仓库根目录 +运行,并指定一个新输出路径: + +```sh +bun packages/app/scripts/export-modeled-system-power.ts \ + --input /path/to/cohort.input.json \ + --output /path/to/new-comparison +``` + +输出包括 `comparison.json`、`comparison.csv` 和可选的 `cells.csv`。JSON 保留原始 +记录、审计、有效性、模型输出和来源;CSV 中不可用数值留空。可选参数 `--pue 1.3` +会覆盖**所有**记录的设施系数,并写入元数据;不传时采用各硬件默认值。实测数据提取 +方法见 [API 示例](./inferencex-api-examples.md)。 + +建模能耗要求匹配的审计提供精确遥测时长、物理 GPU 数量,以及成功请求/token 分母。 +计算为建模部署功率 × 时长,不是实测墙上功率的积分,也不会推断 kernel 级 +prefill/decode 能耗。各次重复测试先独立建模再聚合;只要范围内任一次不可用,该单元 +的均值就保持不可用。 + +## 维护模型 + +| 职责 | 源码 | +| ----------------------------------- | ---------------------------------------------------------------------------- | +| 公式、非线性曲线和舍入 | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) | +| 组件参数、假设和来源状态 | [profiles](../packages/app/src/lib/system-power-model.profiles.json) | +| 工作负载、遥测、拓扑和 PUE 接纳规则 | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) | +| 匹配前沿和规划余量 | [profit-power.ts](../packages/app/src/components/calculator/profit-power.ts) | +| 冻结的数值基线 | [参考用例](../packages/app/src/lib/system-power-model.reference.json) | +| 模型标识和源码哈希 | [provenance](../packages/app/src/lib/system-power-model.provenance.json) | + +修改公式或生效系数时,同时提供有依据的期望值和假设元数据。仅修改 `u_cpu` 等标签 +不会改变功率。刷新并检查 manifest: + +```sh +bun packages/app/scripts/update-system-power-provenance.ts +bun packages/app/scripts/update-system-power-provenance.ts --check +``` + +运行受影响的模型、接纳规则、规划、views API 和导出检查。保留 496 个历史参考用例, +不要重写其基线;数值一致不代表完成校准。模型和参数由应用维护,不依赖私有 Python +仓库。历史测量数据在读取时建模,因此仅修改模型需要更新应用包并刷新 API 缓存或 +等待其过期,不需要回填原始数据。冻结的导出文件需另行生成,缺失的源测量仍然缺失。 diff --git a/docs/tco-calculator.md b/docs/tco-calculator.md index 756ccf4d8..4b03e3393 100644 --- a/docs/tco-calculator.md +++ b/docs/tco-calculator.md @@ -1135,7 +1135,7 @@ and the two can be collapsed into one once both are on master. input / $0.06 cached / $1.20 output per M tok, the permanent 50%-off rate on its pay-as-you-go page); the OpenRouter aggregate also sits below it. At 83 tok/s/user the B200, B300, GB200, and MI355X agentic curves are priced, and the - H100, H200, MI300X, and MI325X curves top out below it and list as not priced. + H100, H200, MI300X, and MI325X curves top out below it and are omitted at that target. GLM 5.2/5.3 opens on a 10% model license fee and MiniMax M3 on 20%; Kimi K3 opens on the 30% `DEFAULT_LAB_CUT_PCT`. DeepSeek V4 Pro opens on 24 tok/s/user, the speed DeepSeek's own API serves at, DeepSeek's peak-hour list price for @@ -1144,14 +1144,14 @@ and the two can be collapsed into one once both are on master. and a 0% model license fee, since the weights ship under the MIT license. At 24 tok/s/user the B200, B300, and MI355X agentic curves are priced; the GB200, GB300, and H200 curves bottom out above it (their lowest measured - points sit at roughly 40, 30, and 27 tok/s/user) and list as not priced until + points sit at roughly 40, 30, and 27 tok/s/user) and are omitted until a lower-interactivity run lands. DeepSeek V4.1 Flash opens on 125 tok/s/user, the speed DeepSeek's own API serves the Flash tier at, DeepSeek's peak-hour list price for `deepseek-flash` ($0.30 input / $0.006 cached / $1.20 output per M tok; off-peak is half that), and a 0% model license fee, since the weights ship under the MIT license. It entered the fleet on AgentX only (InferenceX#2961), so the page is wired ahead of the first published rows; SKUs whose agentic curves stop short - of 125 tok/s/user list as not priced rather than extrapolated. A + of 125 tok/s/user are omitted rather than extrapolated. A model with a list price gets a third Token Price option, ` list price`, next to OpenRouter and Custom; the caption names the source in force and links the lab's pricing page when the list price is used. Switching to Custom diff --git a/packages/app/cypress/component/power-compare.cy.tsx b/packages/app/cypress/component/power-compare.cy.tsx index f2a51b485..da9ec59d0 100644 --- a/packages/app/cypress/component/power-compare.cy.tsx +++ b/packages/app/cypress/component/power-compare.cy.tsx @@ -53,10 +53,14 @@ function measuredCurve(hwKey: string, run_url?: string): InferenceData[] { ); } -function mountCompare(data: InferenceData[], overlay: InferenceData[]) { +function mountCompare( + data: InferenceData[], + overlay: InferenceData[], + { locale = 'en', width = 1000 }: { locale?: 'en' | 'zh'; width?: number } = {}, +) { mountWithProviders( - -
+ +
{ ); }); }); + +describe('Modeled power source links', () => { + for (const locale of ['en', 'zh'] as const) { + for (const width of [1280, 390]) { + const overlay = width === 390; + it(`links to app-owned source and ${locale} assumptions from a ${overlay ? 'mobile overlay' : 'desktop official'} tooltip`, () => { + cy.viewport(width, 720); + const modeledCurve = (hwKey: string, runUrl?: string) => + measuredCurve(hwKey, runUrl).map((point) => + createMockInferenceData({ + ...point, + disagg: false, + modeledSystemPower: { + status: 'supported', + hardware: hwKey, + modelRevision: 'model-content-digest-for-tooltip-fixture', + modelPath: 'packages/app/src/lib/system-power-model.ts', + gpuCount: 8, + chassisCount: 1, + modeledGpuCount: 8, + measuredGpuWattsPerGpu: 600, + chassisAcWatts: 6400, + chassisAcWattsPerGpu: 800, + facilityWatts: 8320, + deploymentAcWatts: 6400, + deploymentFacilityWatts: 8320, + pue: 1.3, + topologyBasis: 'single-node', + chassisBasis: 'full', + telemetryBasis: 'validated-v2', + }, + }), + ); + mountCompare(modeledCurve('b200'), modeledCurve('h100', OVERLAY_RUN_URL), { + locale, + width: Math.min(1000, width - 32), + }); + cy.get(`${svg} ${overlay ? '.unofficial-overlay-pt' : '.dot-group'}`) + .eq(1) + .click({ force: true }); + cy.get('[data-chart-tooltip]:visible').within(() => { + if (overlay) cy.contains('powerx-compare').should('exist'); + cy.get('[data-testid="tooltip-modeled-system-power"]').within(() => { + const base = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${process.env.NEXT_PUBLIC_APP_SOURCE_REF ?? 'master'}`; + cy.get(`a[href="${base}/packages/app/src/lib/system-power-model.ts"]`) + .should('have.attr', 'title', 'model-content-digest-for-tooltip-fixture') + .and('have.attr', 'target', '_blank') + .and('have.attr', 'rel', 'noopener noreferrer') + .scrollIntoView() + .should('be.visible'); + cy.contains('a', locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions') + .should( + 'have.attr', + 'href', + `${base}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`, + ) + .and('have.attr', 'target', '_blank') + .and('have.attr', 'rel', 'noopener noreferrer') + .then(($link) => { + $link[0].scrollIntoView({ block: 'center' }); + }) + .should('be.visible') + .then(($link) => { + const bounds = $link[0].getBoundingClientRect(); + expect(bounds.left).to.be.at.least(0); + expect(bounds.right).to.be.at.most(width); + const shell = $link.closest('[data-chart-tooltip]')[0].firstElementChild!; + const frame = shell.getBoundingClientRect(); + expect(bounds.top).to.be.at.least(frame.top); + expect(bounds.bottom).to.be.at.most(frame.bottom); + }); + }); + }); + cy.screenshot(`power-model-links-${locale}-${width}`, { capture: 'viewport' }); + }); + } + } +}); diff --git a/packages/app/cypress/component/profit-estimator-chart.cy.tsx b/packages/app/cypress/component/profit-estimator-chart.cy.tsx index be46f8df6..16da8ef3e 100644 --- a/packages/app/cypress/component/profit-estimator-chart.cy.tsx +++ b/packages/app/cypress/component/profit-estimator-chart.cy.tsx @@ -58,7 +58,7 @@ function colorForRow(row: ProfitEstimatorRow): string { function mountChart(widthPx: number, rows: ProfitEstimatorRow[] = ROWS) { cy.mount( -
+
{ }); } -/** Left-to-right boxes of every revenue figure and margin line above the bars. */ +/** Combined revenue and margin bounds for each bar, in left-to-right order. */ function labelBoxes(): Cypress.Chainable { return cy - .get('[data-testid="profit-estimator-chart"] .revenue-label tspan') + .get('[data-testid="profit-estimator-chart"] .revenue-label') .then(($tspans) => [...$tspans] .filter((el) => (el.textContent ?? '') !== '') @@ -102,8 +102,31 @@ function overlaps(a: DOMRect, b: DOMRect): boolean { } describe('ProfitEstimatorChart revenue labels', () => { + it('preserves the rendered SVG at the exact scrolling threshold', () => { + cy.viewport(1280, 900); + mountChart(390, ROWS.slice(0, 3)); + cy.get('[data-testid="d3-chart-svg"]').then(($svg) => { + const svg = $svg[0]; + const width = svg.getBoundingClientRect().width; + cy.get('[data-testid="profit-chart-container"]').invoke('css', 'width', `${width + 32}px`); + cy.get('[data-chart-scroll][tabindex]').should('not.exist'); + cy.get('[data-testid="d3-chart-svg"]').should(($current) => { + expect($current[0], 'same SVG at the threshold').to.equal(svg); + expect($current[0].getBoundingClientRect().width).to.equal(width); + expect($current.find('.revenue-label')).to.have.length(3); + }); + cy.get('[data-chart-scroll]').should('not.have.attr', 'tabindex'); + cy.get('[data-testid="profit-chart-container"]').invoke('css', 'width', `${width + 31}px`); + cy.get('[data-chart-scroll]').should('have.attr', 'tabindex', '0'); + cy.get('[data-testid="d3-chart-svg"]').should(($current) => { + expect($current[0], 'same SVG below the threshold').to.equal(svg); + expect($current.find('.revenue-label')).to.have.length(3); + }); + }); + }); + it('never lets neighbouring revenue figures overlap on a phone', () => { - // iPhone 15/16 CSS viewport; the card padding leaves the chart ~361px. + // iPhone 15/16 viewport: the 361px card viewport scrolls across the wider plot. cy.viewport(393, 852); mountChart(393); cy.screenshot('profit-estimator-chart-phone', { overwrite: true }); @@ -125,9 +148,12 @@ describe('ProfitEstimatorChart revenue labels', () => { mountChart(393, LOSING_ROWS); cy.get('[data-testid="profit-estimator-chart"] .loss-label').should('have.length', 2); cy.screenshot('profit-estimator-chart-phone-loss', { overwrite: true }); - // The wide figure loses the word and its decimal; the narrow one keeps its decimal. + // The scrollable plot has room to retain the loss label as well as the signed amount. cy.get('[data-testid="profit-estimator-chart"] .loss-label').then(($labels) => { - expect([...$labels].map((el) => el.textContent)).to.deep.equal(['-$561M', '-$4.9B']); + expect([...$labels].map((el) => el.textContent)).to.deep.equal([ + 'Loss -$561M', + 'Loss -$4.9B', + ]); }); cy.get('[data-testid="profit-estimator-chart"] rect.bar') .first() diff --git a/packages/app/cypress/component/scatter-graph.cy.tsx b/packages/app/cypress/component/scatter-graph.cy.tsx index 7ab5903c7..2598d8632 100644 --- a/packages/app/cypress/component/scatter-graph.cy.tsx +++ b/packages/app/cypress/component/scatter-graph.cy.tsx @@ -8,6 +8,7 @@ import { import ScatterGraph from '@/components/inference/ui/ScatterGraph'; import { useParetoHighlightToggle } from '@/components/inference/hooks/useParetoHighlightToggle'; import ChartDisplay from '@/components/inference/ui/ChartDisplay'; +import chartDefinitions from '@/components/inference/metric-registry'; import { mountWithProviders } from '../support/test-utils'; import { expandLegendAdvanced } from '../support/legend-advanced'; import { @@ -466,6 +467,7 @@ describe('ScatterGraph', () => { }); it('explains why All in Measured has no points', () => { + cy.viewport(1280, 720); mountWithProviders(
{ 'be.visible', ); cy.contains('No measurements to plot for this selection.').should('not.exist'); + cy.contains('NVL72 also needs complete Grace or module telemetry').should('be.visible'); + cy.contains('not NVL72 systems').should('not.exist'); + cy.screenshot('nvl72-empty-en-desktop', { overwrite: true }); }); it('localizes the All in Measured explanation', () => { + cy.viewport(390, 720); mountWithProviders(
@@ -521,6 +527,9 @@ describe('ScatterGraph', () => { cy.contains('当前选择没有可用的整体实测功耗数值。').should('be.visible'); cy.contains('当前选择没有可绘制的测量数据。').should('not.exist'); + cy.contains('NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测').should('be.visible'); + cy.contains('不含 NVL72 系统').should('not.exist'); + cy.screenshot('nvl72-empty-zh-mobile', { overwrite: true }); }); for (const selectedYAxisMetric of ['y_tpPerGpu', 'y_measuredPrefillJPerInputToken'] as const) { @@ -1786,6 +1795,66 @@ describe('ScatterGraph', () => { }); }); +describe('ChartDisplay modeled power disclosures', () => { + for (const locale of ['en', 'zh'] as const) { + it(`explains NVL72 telemetry and cooling PUE without hiding its modeled point (${locale})`, () => { + const point = createMockInferenceData({ + hwKey: 'gb200', + hw: 'NVIDIA GB200', + model: Model.Qwen3_5, + y: 1000, + utilityModeledWatts: { y: 1000, roof: false }, + }); + mountWithProviders( + +
+ +
+
, + { + inference: { + selectedModel: Model.Qwen3_5, + selectedYAxisMetric: 'y_utilityModeledWatts', + activeHwTypes: new Set(['gb200']), + hwTypesWithData: new Set(['gb200']), + hardwareConfig: { gb200: { name: 'gb200', label: 'GB200', suffix: '', gpu: 'GB200' } }, + graphs: [ + { + model: Model.Qwen3_5, + sequence: Sequence.EightK_OneK, + chartDefinition: chartDefinitions[0], + data: [point], + }, + ], + }, + globalFilters: { selectedModel: Model.Qwen3_5 }, + unofficial: {}, + }, + ); + for (const width of [1280, 390]) { + cy.viewport(width, 900); + cy.get('[data-testid="power-basis-assumptions"]') + .should('be.visible') + .and('contain.text', 'PUE 1.3') + .and('contain.text', 'PUE 1.1') + .and('contain.text', 'Grace') + .and('not.contain.text', 'app-sha') + .and( + 'not.contain.text', + locale === 'en' + ? 'NVL72 systems (GB200, GB300) and points without values are omitted' + : 'NVL72 系统(GB200、GB300)及缺少数值的数据点不绘制', + ) + .should(($note) => { + expect($note[0].scrollWidth).to.be.at.most($note[0].clientWidth); + }); + cy.get('.dot-group .visible-shape').should('exist'); + cy.screenshot(`nvl72-boundary-${locale}-${width}`, { overwrite: true }); + } + }); + } +}); + describe('ChartDisplay responsive status notes', () => { for (const locale of ['en', 'zh'] as const) { for (const width of [390, 1440]) { diff --git a/packages/app/cypress/e2e/powerx-compare.cy.ts b/packages/app/cypress/e2e/powerx-compare.cy.ts index 159b01683..648aaa2b6 100644 --- a/packages/app/cypress/e2e/powerx-compare.cy.ts +++ b/packages/app/cypress/e2e/powerx-compare.cy.ts @@ -216,3 +216,207 @@ describe('PowerX article panels', () => { }); }); }); + +// Aggregate role counts describe shared devices: B200 TP8 × PP2, H200 TP16 × 2 workers. +const agenticRows = [ + ['b200', 'dynamo-vllm', 8, 2, 1, 16, 760], + ['h200', 'vllm', 16, 1, 2, 32, 168], + ['gb200', 'trt', 4, 1, 1, 4, 450], + ['b300', 'vllm', 8, 1, 1, 8, 0], +].map(([hardware, framework, tp, pp, replicas, chips, watts], index) => ({ + ...rows(null, 'b200')[0], + id: 990100 + index, + model: 'kimik3', + hardware, + framework, + benchmark_type: 'agentic_traces', + disagg: false, + isl: null, + osl: null, + prefill_tp: tp, + decode_tp: tp, + prefill_num_workers: replicas, + decode_num_workers: replicas, + num_prefill_gpu: chips, + num_decode_gpu: chips, + metrics: { + power_valid: watts ? 1 : 0, + power_metric_schema_version: 2, + avg_power_w: watts, + avg_total_gpu_power_w: Number(watts) * Number(chips), + prefill_pp: pp, + decode_pp: pp, + p90_itl: 0.02 + index * 0.01, + median_itl: 0.01 + index * 0.01, + median_intvty: 100 / (index + 1), + tput_per_gpu: 200 + index * 100, + output_tput_per_gpu: 100 + index * 50, + joules_per_output_token: 2 + index, + }, +})); + +describe('AgentX All in Measured chart and table', () => { + for (const [locale, width] of [ + ['en', 1280], + ['zh', 390], + ] as const) { + it(`retains B200/H200 multinode rows and export at ${locale} ${width}px`, () => { + cy.viewport(width, 900); + cy.intercept('GET', '/api/v1/availability', { body: agenticRows }).as('agenticAvailability'); + cy.intercept('GET', '/api/v1/benchmarks*', { body: agenticRows }).as('agenticBenchmarks'); + cy.intercept('GET', '/api/v1/workflow-info*', { + body: { runs: [], changelogs: [], configs: [] }, + }); + cy.intercept('GET', '/api/v1/trace-availability*', { body: {} }); + cy.intercept('GET', '/api/v1/log-availability*', { body: {} }); + cy.intercept('GET', '/api/v1/resident-sequence-lengths*', { body: {} }); + const overlayBody = { + runInfos: [ + { + id: OVERLAY_RUN_ID, + name: 'agentic-power', + branch: 'agentic-power', + sha: 'abc000', + createdAt: `${DATE}T00:00:00Z`, + url: OVERLAY_RUN_URL, + conclusion: 'success', + status: 'completed', + isNonMainBranch: true, + }, + ], + benchmarks: [agenticRows[1], agenticRows[2]].map((row) => ({ + ...row, + id: 0, + run_url: OVERLAY_RUN_URL, + })), + evaluations: [], + }; + cy.intercept('GET', '/api/unofficial-run*', { body: overlayBody }).as('agenticOverlay'); + let csvBlob: Blob | undefined; + cy.visit( + `${locale === 'zh' ? '/zh' : ''}/inference?g_model=Kimi-K3&i_seq=agentic-traces&i_prec=fp4&i_pctl=p90&i_metric=y_utilityModeledWatts&i_optimal=0&i_best=0&unofficialrun=${OVERLAY_RUN_ID}`, + { + onBeforeLoad(win) { + win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); + win.localStorage.setItem('inferencex-feature-gate', '1'); + win.URL.createObjectURL = (object) => { + if (object instanceof win.Blob) csvBlob = object; + return 'blob:agentic-power'; + }; + win.HTMLAnchorElement.prototype.click = () => {}; + }, + }, + ); + cy.wait(['@agenticAvailability', '@agenticBenchmarks', '@agenticOverlay']); + cy.get('[data-testid="chart-figure"]').first().find('.dot-group').should('have.length', 2); + cy.get('[data-testid="chart-figure"]') + .first() + .find('.unofficial-overlay-pt') + .should('have.length', 1); + cy.get('[data-testid="power-agentic-model-note"]') + .first() + .should( + 'contain.text', + locale === 'en' ? 'not been independently calibrated' : '尚未针对 AgentX', + ); + // A fixed page header otherwise repeats over the stitched element capture. + cy.get('header').invoke('css', 'visibility', 'hidden'); + cy.get('[data-testid="chart-figure"]') + .first() + .scrollIntoView() + .screenshot(`agentic-all-in-${locale}-chart`, { overwrite: true }); + cy.get('[data-testid="inference-table-view-btn"]').first().click(); + cy.get('[data-testid="chart-figure"]') + .first() + .find('tbody tr') + .should('have.length', 5) + .then(($rows) => { + expect($rows.text()).to.contain('B200').and.contain('H200'); + expect($rows.text()).to.contain('GB200').and.not.to.contain('B300'); + const missing = [...$rows].filter((row) => row.textContent!.includes('GB200')); + expect(missing).to.have.length(2); + for (const row of missing) { + expect(row.textContent).to.contain('450').and.contain('—'); + expect(row.textContent).to.contain( + locale === 'en' + ? 'Grace or module telemetry missing or invalid' + : 'Grace 或 module 遥测缺失或无效', + ); + } + }); + cy.get('[data-testid="chart-figure"]') + .first() + .screenshot(`agentic-all-in-${locale}-table`, { overwrite: true }); + if (locale === 'zh') { + cy.get('[data-testid="inference-results-table"]') + .first() + .contains('th', '整体估算状态') + .then(($status) => { + const scroll = $status[0].closest('table')!.parentElement!; + const pinnedWidth = + $status[0].parentElement!.firstElementChild!.getBoundingClientRect().width; + cy.wrap(scroll).scrollTo($status[0].offsetLeft - pinnedWidth, 0); + cy.wrap($status).should(($cell) => { + expect($cell[0].getBoundingClientRect().right).to.be.at.most( + scroll.getBoundingClientRect().right + 1, + ); + }); + }); + cy.get('[data-testid="chart-figure"]') + .first() + .screenshot('agentic-all-in-zh-table-status', { overwrite: true }); + } + cy.get('header').invoke('css', 'visibility', ''); + cy.get('[data-testid="export-button"]').first().click(); + cy.get('[data-testid="export-csv-button"]').click(); + cy.then(() => csvBlob!.text()).then((csv) => { + const [header, ...data] = csv.split('\n').filter((line) => !line.startsWith('#')); + const columns = header.split(','); + const values = data.map((line) => line.split(',')); + expect(values.map((row) => row[columns.indexOf('Hardware')]).sort()).to.deep.equal([ + 'b200', + 'gb200', + 'gb200', + 'h200', + 'h200', + ]); + expect( + values.map((row) => Number(row[columns.indexOf('Physical Chips')])).sort((a, b) => a - b), + ).to.deep.equal([4, 4, 16, 32, 32]); + const missing = values.filter((row) => row[columns.indexOf('Hardware')] === 'gb200'); + for (const row of missing) { + expect(row[10]).to.equal(''); + expect(row[columns.indexOf('Measured GPU Power (W/chip)')]).to.equal('450'); + expect(row[columns.indexOf('All-in Estimate Status')]).to.equal( + 'Grace or module telemetry missing or invalid', + ); + } + }); + cy.document().then((doc) => expect(doc.documentElement.scrollWidth).to.be.at.most(width)); + // A selection containing only GPU measurements must still have a usable table. + cy.intercept('GET', '/api/v1/benchmarks*', { body: [agenticRows[2]] }).as('missingOnly'); + cy.intercept('GET', '/api/unofficial-run*', { + body: { + ...overlayBody, + benchmarks: [{ ...agenticRows[2], id: 0, run_url: OVERLAY_RUN_URL }], + }, + }).as('missingOverlay'); + cy.reload(); + cy.wait(['@missingOnly', '@missingOverlay']); + cy.get('[data-testid="inference-table-view-btn"]').first().click(); + cy.get('[data-testid="chart-figure"]') + .first() + .find('tbody tr') + .should('have.length', 2) + .and('contain.text', 'GB200'); + cy.location('href').then((href) => { + const url = new URL(href); + url.searchParams.set('i_best', '1'); + url.searchParams.set('i_xmode', 'concurrency'); + cy.visit(url.toString()); + }); + cy.get('[data-testid="inference-table-view-btn"]').first().click(); + cy.get('[data-testid="chart-figure"]').first().find('tbody tr').should('have.length', 2); + }); + } +}); diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts index e8a59f7fa..dfb6648ed 100644 --- a/packages/app/cypress/e2e/profit-estimator.cy.ts +++ b/packages/app/cypress/e2e/profit-estimator.cy.ts @@ -38,6 +38,7 @@ import { interceptVrPublicationData } from '../support/vr-publication-fixtures'; import { interceptProfitData, profitBenchmarkRows, + profitNvl72Rows, PROFIT_CHANGELOG_NOTES, PROFIT_DATE, PROFIT_HISTORY_DATE, @@ -94,21 +95,42 @@ const chart = () => cy.get('[data-testid="profit-estimator-chart"]'); const chartSvg = () => chart().find('svg').filter(':has(.chart-root)').first(); const bars = () => chart().find('rect.bar'); -function assertDisclosureOpen(testId: string, open: boolean) { - cy.get(`[data-testid="${testId}"]`).should(($details) => { - expect($details[0].open, `${testId} native disclosure state`).to.equal(open); - const content = $details[0].querySelector('p'); - expect(content, `${testId} content`).not.to.equal(null); - // Cypress visibility omits native closed-details rendering in some browsers. - if (content && typeof content.checkVisibility === 'function') { - expect(content.checkVisibility(), `${testId} browser visibility`).to.equal(open); - } - }); -} - // Clear the preceding chart before each case changes the viewport. describe('Profit estimator power option', { testIsolation: true }, () => { for (const locale of ['en', 'zh'] as const) { + it(`keeps available estimates without a per-configuration warning list (${locale})`, () => { + stubOpenRouter(); + const width = locale === 'en' ? 1280 : 390; + cy.viewport(width, 900); + cy.intercept('GET', '/api/v1/benchmarks*', { + body: profitBenchmarkRows().map((row) => ({ + ...row, + metrics: { + ...row.metrics, + power_valid: row.hardware === 'b300' ? 0 : 1, + power_metric_schema_version: 2, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + }, + })), + }); + cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=compare`, { + onBeforeLoad: unlockPowerGate, + }); + chart().find('text.revenue-label').should('have.length', 6); + chart() + .find('.x-axis') + .should('not.contain', 'H200') + .and('contain', 'B300') + .and('contain', 'GB300'); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + chart().scrollIntoView(); + cy.screenshot(`profit-no-unavailable-list-${locale}`, { + capture: 'viewport', + overwrite: true, + }); + }); + it(`prices DeepSeek Flash partial chassis with a one-line power note and CSV labels (${locale})`, () => { stubOpenRouter(); cy.viewport(locale === 'en' ? 1280 : 393, 900); @@ -146,10 +168,6 @@ describe('Profit estimator power option', { testIsolation: true }, () => { }, ); const label = locale === 'en' ? 'Full-chassis extrapolation' : '整机外推'; - const hardwareReason = - locale === 'en' - ? 'no system power model for this hardware' - : '该硬件暂无适用的系统功耗模型'; cy.get('#profit-target').should('have.value', '125'); chart().find('text.revenue-label').should('have.length', 3); chart() @@ -157,19 +175,11 @@ describe('Profit estimator power option', { testIsolation: true }, () => { .and('contain', 'B300') .and('contain', 'MI355X') .and('contain', label); - cy.get('[data-testid="profit-power-unavailable"]') - .should('contain', 'GB300') - .and('contain', hardwareReason); - assertDisclosureOpen('profit-power-unavailable', false); - cy.get('[data-testid="profit-power-unavailable"] > summary').click(); - assertDisclosureOpen('profit-power-unavailable', true); - cy.get('[data-testid="profit-power-unavailable"] > p').should('be.visible'); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); cy.get('[data-testid="profit-power-note"]') .should('contain', locale === 'en' ? 'All in Measured' : '整体实测功耗') .and('not.contain', locale === 'en' ? 'unmeasured components' : '未实测的组件'); cy.get('[data-testid="profit-power-assumptions"]').should('not.exist'); - cy.get('[data-testid="profit-power-unavailable"] > summary').click(); - assertDisclosureOpen('profit-power-unavailable', false); cy.get('[data-testid="profit-power-note"]').then(($note) => { const box = $note[0].getBoundingClientRect(); expect(box.left).to.be.at.least(0); @@ -226,7 +236,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.then(() => expect(rawRequests).to.equal(0)); cy.get('body').type('{uparrow}{uparrow}{downarrow}{downarrow}'); cy.get('#profit-power').should('contain', 'Compare both'); - chart().find('text.revenue-label').should('have.length', 6); + chart().find('text.revenue-label').should('have.length', 7); cy.window().then((win) => { win.localStorage.removeItem('inferencex-feature-gate'); win.dispatchEvent(new Event('inferencex:feature-gate:locked')); @@ -237,7 +247,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => { }); for (const currentValid of [true, false]) { - it(`dates historical power skips without current hardware metadata (${currentValid ? 'with current bars' : 'empty chart'})`, () => { + it(`keeps historical provisioned bars when measured power is unavailable (${currentValid ? 'with current bars' : 'provisioned bars only'})`, () => { stubOpenRouter(); cy.intercept('GET', '/api/v1/benchmarks*', (req) => { const historical = req.query['date'] === PROFIT_HISTORY_DATE; @@ -262,16 +272,9 @@ describe('Profit estimator power option', { testIsolation: true }, () => { `/profit-estimator-per-gigawatt?c_power=compare&i_gpus=b200_sglang,b300_vllm&i_dstart=${PROFIT_HISTORY_DATE}&i_dend=${PROFIT_HISTORY_DATE}`, { onBeforeLoad: unlockPowerGate }, ); - cy.get('[data-testid="profit-power-unavailable"]').should( - 'contain', - `B300 (vLLM) (FP4) • ${PROFIT_HISTORY_DATE}`, - ); - cy.get('[data-testid="profit-power-unavailable"]').should( - 'contain', - `B200 (SGLang) (FP4) • ${PROFIT_HISTORY_DATE}`, - ); - if (currentValid) chart().find('text.revenue-label').should('have.length', 2); - else cy.get('[data-testid="profit-estimator-chart"]').should('not.exist'); + chart() + .find('text.revenue-label') + .should('have.length', currentValid ? 4 : 3); }); } @@ -301,12 +304,246 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.wait('@power-rows').its('request.query').should('not.have.property', 'view'); cy.get('#profit-power').should('contain', 'Compare both'); cy.get('#profit-target').should('have.value', '45'); - chart().find('text.revenue-label').should('have.length', 6); + chart().find('text.revenue-label').should('have.length', 7); chart().should('contain', 'B200').and('contain', 'B300').and('contain', 'MI355X'); chart().should('contain', 'All in Measured').and('contain', 'All in Provisioned'); - cy.get('[data-testid="profit-power-unavailable"]').should('contain', 'GB300'); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + }); + + it('prices a GB200 NVL72 tray on its measured compute module and names the basis', () => { + stubOpenRouter(); + cy.viewport(1280, 900); + cy.intercept('GET', '/api/v1/benchmarks*', (req) => { + req.reply({ + body: [ + ...profitBenchmarkRows().map((row) => ({ + ...row, + metrics: { + ...row.metrics, + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + }, + })), + ...profitNvl72Rows(), + ], + }); + }); + let csv: Blob | undefined; + // Locked: the control stays hidden and the tray prices on provisioned power like every SKU. + cy.visit('/profit-estimator-per-gigawatt?c_power=compare', { + onBeforeLoad: (win) => { + suppressNudges(win); + win.localStorage.removeItem('inferencex-feature-gate'); + win.URL.createObjectURL = (blob) => { + if (blob instanceof win.Blob) csv = blob; + return 'blob:profit-nvl72-csv-test'; + }; + win.HTMLAnchorElement.prototype.click = () => {}; + }, + }); + chart().find('text.revenue-label').should('have.length', 5); + cy.get('#profit-power').should('not.exist'); + cy.get('[data-testid="profit-power-note"]').should('not.exist'); + cy.get('body').type('{uparrow}{uparrow}{downarrow}{downarrow}'); + cy.get('#profit-power').should('contain', 'Compare both'); + // Four supported SKUs get pairs; GB300 keeps its provisioned bar without CPU power. + chart().find('text.revenue-label').should('have.length', 9); + chart().should('contain', 'GB200').and('contain', 'All in Measured'); + // The header keeps its one line; the NVL72 basis and modeled components + // live in the Power Estimation help and the CSV caption. + cy.get('[data-testid="profit-power-note"]') + .should('contain', 'Compare both') + .and('not.contain', 'NVSwitch trays'); + cy.get('[data-testid="profit-power-assumptions"]').should('not.exist'); + cy.get('[data-testid="option-help-profit-power"]').click(); + cy.get('[data-testid="option-help-content-profit-power"]') + .should('be.visible') + .and('contain', 'PUE 1.3 for air-cooled chassis or 1.1 for NVL72') + .and('contain', 'GB200 NVL72') + .and('contain', 'measured module (GPU + HBM + Grace + LPDDR5X; module sensor)') + .and('contain', 'NVSwitch trays') + .and('contain', 'DLC PUE 1.1'); + cy.get('body').type('{esc}'); + cy.get('[data-testid="option-help-content-profit-power"]').should('not.exist'); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + chart().scrollIntoView(); + cy.screenshot('profit-nvl72-compare-desktop', { capture: 'viewport', overwrite: true }); + cy.viewport(393, 900); + cy.get('[data-testid="profit-power-note"]').then(($note) => { + const bounds = $note[0].getBoundingClientRect(); + expect(bounds.left).to.be.at.least(0); + expect(bounds.right).to.be.at.most(393); + }); + chart().find('text.revenue-label').should('have.length', 9); + chart().scrollIntoView(); + cy.screenshot('profit-nvl72-compare-mobile', { capture: 'viewport', overwrite: true }); + chart().find('[data-chart-scroll]').scrollIntoView().scrollTo('left').should('be.visible'); + chart() + .find('[data-chart-scroll]') + .then(($scroll) => { + const bounds = $scroll[0].getBoundingClientRect(); + expect(bounds.left).to.be.at.least(0); + expect(bounds.right).to.be.at.most(393); + }); + cy.screenshot('profit-nvl72-chart-mobile', { capture: 'viewport', overwrite: true }); + cy.get('[data-testid="export-button"]').first().click(); + cy.get('[data-testid="export-csv-button"]').click(); + cy.then(() => csv!.text()).then((text) => { + expect(text).to.contain('Power basis,Power sensor,System power profile'); + expect(text).to.contain('PUE 1.3 for air-cooled chassis or 1.1 for NVL72'); + expect(text).to.contain('GB200 NVL72'); + expect(text).to.contain('NVSwitch trays'); + expect(text).to.contain('DLC PUE 1.1'); + const rows = text.split('\n').filter((line) => line.startsWith('GB200')); + expect(rows).to.have.length(2); + expect(rows.some((row) => row.includes('All in Provisioned'))).to.equal(true); + const measured = rows.find((row) => row.includes('measured module')); + expect(measured).to.contain('GPU + HBM + Grace + LPDDR5X; module sensor'); + expect(measured).to.contain(',module,'); + expect(measured).to.contain('system-power-model.ts @ app-sha256:'); + expect(measured).to.contain(' sha256:'); + }); }); + for (const locale of ['en', 'zh'] as const) { + it(`keeps all nine power comparison bars readable and reachable on mobile (${locale})`, () => { + stubOpenRouter(); + cy.viewport(390, 900); + cy.intercept('GET', '/api/v1/benchmarks*', { + body: [ + ...profitBenchmarkRows().map((row) => ({ + ...row, + metrics: { + ...row.metrics, + power_valid: 1, + power_metric_schema_version: 2, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + }, + })), + ...profitNvl72Rows(), + ], + }); + cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=compare`, { + onBeforeLoad: unlockPowerGate, + }); + chart() + .find('text.revenue-label') + .should('have.length', 9) + .should(($labels) => { + const boxes = [...$labels].map((label) => label.getBoundingClientRect()); + for (let i = 1; i < boxes.length; i++) { + expect( + boxes[i].left - boxes[i - 1].right, + 'space between revenue and margin labels', + ).to.be.at.least(4); + } + }); + chart() + .find('image.bar-vendor-mark') + .should(($marks) => { + const boxes = [...$marks].map((mark) => mark.getBoundingClientRect()); + for (let i = 1; i < boxes.length; i++) { + expect(boxes[i].left - boxes[i - 1].right, 'space between vendor marks').to.be.at.least( + 4, + ); + } + }); + const scroller = () => chart().find('[data-chart-scroll]'); + scroller().should('have.attr', 'tabindex', '0').and('have.attr', 'role', 'region'); + scroller() + .scrollIntoView({ offset: { top: -70, left: 0 } }) + .focus() + .should('have.focus'); + scroller().then(($scroll) => { + const el = $scroll[0]; + const event = new el.ownerDocument.defaultView!.KeyboardEvent('keydown', { + key: 'ArrowRight', + bubbles: true, + cancelable: true, + }); + el.dispatchEvent(event); + expect( + event.defaultPrevented, + `handled key on ${el.clientWidth}/${el.scrollWidth}`, + ).to.equal(true); + expect(el.scrollLeft, 'synchronous scroll').to.be.greaterThan(0); + }); + scroller().should(($scroll) => expect($scroll[0].scrollLeft).to.be.greaterThan(0)); + scroller().scrollTo('right'); + chart().find('rect.bar-profit').last().should('be.visible'); + chart().find('rect.bar-profit').last().click({ scrollBehavior: false }); + cy.get('[data-chart-tooltip="profit-estimator-chart"]') + .should('be.visible') + .and('contain', 'MI355X'); + // Clear the pinned tooltip for screenshots; the SVG center is outside the scroll viewport. + chartSvg().trigger('click', { force: true, scrollBehavior: false }); + cy.get('[data-chart-tooltip="profit-estimator-chart"]').should('not.be.visible'); + cy.document().should((doc) => { + expect(doc.documentElement.scrollWidth).to.be.at.most(390); + }); + cy.get('[data-testid="profit-caption"]').should(($caption) => { + const bounds = $caption[0].getBoundingClientRect(); + expect(bounds.left).to.be.at.least(0); + expect(bounds.right).to.be.at.most(390); + }); + scroller() + .scrollIntoView({ offset: { top: -70, left: 0 } }) + .scrollTo('left'); + chart().find('rect.bar-tco').first().should('be.visible'); + cy.screenshot(`profit-dense-${locale}-mobile-left`, { capture: 'viewport', overwrite: true }); + scroller().scrollTo('right'); + cy.screenshot(`profit-dense-${locale}-mobile-right`, { + capture: 'viewport', + overwrite: true, + }); + let exportedPng = ''; + let exportedLabels: string[] = []; + cy.window().then((win) => { + win.HTMLAnchorElement.prototype.click = function () { + if (this.download.endsWith('.png')) { + exportedPng = this.href; + exportedLabels = [ + ...win.document.querySelectorAll('#profit-estimator-chart-export text.revenue-label'), + ].map((label) => label.textContent ?? ''); + } + }; + }); + cy.get('[data-testid="export-button"]').first().click(); + cy.get('[data-testid="export-png-button"]').click(); + cy.window() + .should(() => expect(exportedPng).to.match(/^data:image\/png;base64,/u)) + .then((win) => { + expect(exportedLabels).to.have.length(9); + cy.writeFile( + `cypress/downloads/profit-dense-${locale}.png`, + exportedPng.split(',')[1], + 'base64', + ); + return new Cypress.Promise((resolve, reject) => { + const png = new win.Image(); + png.addEventListener('load', () => { + expect(png.naturalWidth).to.be.greaterThan(1500); + resolve(); + }); + png.addEventListener('error', () => reject(new Error('Profit PNG did not decode'))); + png.src = exportedPng; + }); + }); + scroller().should(($scroll) => expect($scroll[0].scrollLeft).to.be.greaterThan(0)); + cy.viewport(1280, 900); + scroller().should('not.have.attr', 'tabindex'); + scroller().should(($scroll) => { + expect($scroll[0].scrollWidth).to.equal($scroll[0].clientWidth); + }); + chart().find('text.revenue-label').should('have.length', 9); + chartSvg().scrollIntoView({ offset: { top: -70, left: 0 } }); + cy.screenshot(`profit-dense-${locale}-desktop`, { capture: 'viewport', overwrite: true }); + }); + } + it('keeps the benchmark settings and restores the original chart after unavailable power', () => { stubOpenRouter(); cy.visit('/profit-estimator-per-gigawatt', { onBeforeLoad: unlockPowerGate }); @@ -317,10 +554,8 @@ describe('Profit estimator power option', { testIsolation: true }, () => { cy.get('#profit-power').click(); cy.get('[role="option"]').contains('All in Measured').click(); // These existing fixtures intentionally have throughput but no validated power. - cy.get('[data-testid="profit-power-unavailable"]').should( - 'contain', - 'no usable measured power', - ); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + cy.contains('No SKU can be priced for the current selection.').should('be.visible'); cy.get('#profit-target').should('have.value', '45'); cy.get('[data-testid="profit-model-selector"]').should('contain', 'Kimi K3'); cy.get('[data-testid="profit-price-source-selector"]').should('contain', 'Moonshot'); diff --git a/packages/app/cypress/e2e/profit-power-curves.cy.ts b/packages/app/cypress/e2e/profit-power-curves.cy.ts new file mode 100644 index 000000000..d06d53559 --- /dev/null +++ b/packages/app/cypress/e2e/profit-power-curves.cy.ts @@ -0,0 +1,126 @@ +import { + interceptProfitData, + profitBenchmarkRows, + PROFIT_DATE, + PROFIT_HISTORY_DATE, +} from '../support/profit-fixtures'; + +// Invalid MI355X knots dominate the valid curve in the performance frontier. +// Power selection must recover the remaining curve before interpolation. +function rowsFor(date = PROFIT_DATE) { + const rows = profitBenchmarkRows('kimik3', date).map((row) => ({ + ...row, + metrics: { + ...row.metrics, + power_valid: Number( + row.hardware !== 'b300' && !(row.hardware === 'mi355x' && row.conc === 16), + ), + power_metric_schema_version: 2, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + }, + })); + return [ + ...rows, + ...rows + .filter((row) => row.hardware === 'mi355x') + .map((row) => ({ + ...row, + id: row.id + 100_000, + framework: 'atom', + metrics: { ...row.metrics, power_valid: 1 }, + })), + ]; +} + +function setup() { + interceptProfitData(); + cy.intercept('GET', 'https://openrouter.ai/api/v1/models', { data: [] }); + cy.intercept('GET', '/api/v1/benchmarks*', (req) => { + req.reply({ + body: rowsFor(req.query['date'] === PROFIT_HISTORY_DATE ? PROFIT_HISTORY_DATE : PROFIT_DATE), + }); + }); +} + +function unlock(win: Cypress.AUTWindow) { + win.localStorage.setItem('inferencex-feature-gate', '1'); + win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now())); + win.sessionStorage.setItem('inferencex-reproducibility-nudge-shown', '1'); +} + +const chart = () => cy.get('[data-testid="profit-estimator-chart"]'); +const barCount = (count: number) => chart().find('text.revenue-label').should('have.length', count); + +describe('Profit power-valid curves', { testIsolation: true }, () => { + for (const locale of ['en', 'zh'] as const) { + it(`keeps the target and pairs the selected valid curve (${locale})`, () => { + setup(); + const width = locale === 'en' ? 1280 : 390; + cy.viewport(width, 900); + let csv: Blob | undefined; + cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=modeled`, { + onBeforeLoad: (win) => { + unlock(win); + win.URL.createObjectURL = (blob) => { + if (blob instanceof win.Blob) csv = blob; + return 'blob:profit-test'; + }; + win.HTMLAnchorElement.prototype.click = () => {}; + }, + }); + cy.get('#profit-target').should('have.value', '45'); + barCount(3); + chart() + .find('.x-axis') + .should('contain', 'MI355X') + .and('not.contain', 'H200') + .and('not.contain', 'GB300') + .and('not.contain', 'B300'); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + cy.get('body').then(($body) => expect($body[0].scrollWidth).to.be.at.most(width)); + chart().scrollIntoView(); + cy.screenshot(`profit-valid-curves-${locale}`, { capture: 'viewport', overwrite: true }); + cy.get('#profit-power').click(); + cy.get('[role="option"]') + .contains(locale === 'en' ? 'Compare both' : '对比两种估算方式') + .click(); + barCount(8); + cy.get('#profit-target').should('have.value', '45'); + cy.get('[data-testid="export-button"]').first().click(); + cy.get('[data-testid="export-csv-button"]').click(); + cy.then(() => csv!.text()).then((text) => { + const rows = text + .split('\n') + .filter((line) => /^(?:B200|B300|GB300|MI355X)/.test(line)) + .map((line) => line.split(',')); + expect(rows).to.have.length(8); + const paired = rows.filter((row) => row[0].includes('MI355X')); + expect(paired).to.have.length(4); + const revenue = paired.map((row) => row[9]); + expect(new Set(revenue).size).to.equal(2); + for (const value of new Set(revenue)) + expect(revenue.filter((r) => r === value)).to.have.length(2); + expect(text).not.to.contain('NaN'); + }); + cy.get('#profit-power').click(); + cy.get('[role="option"]') + .contains(locale === 'en' ? 'All in Provisioned' : '整体预配功耗') + .click(); + barCount(5); + }); + } + + it('uses valid curves independently for historical comparisons', () => { + setup(); + cy.viewport(1280, 900); + cy.visit( + `/profit-estimator-per-gigawatt?c_power=compare&i_gpus=mi355x_vllm&i_dstart=${PROFIT_HISTORY_DATE}&i_dend=${PROFIT_HISTORY_DATE}`, + { onBeforeLoad: unlock }, + ); + cy.get('#profit-target').should('have.value', '45'); + barCount(4); + chart().should('contain', PROFIT_HISTORY_DATE); + cy.get('[data-testid="profit-power-unavailable"]').should('not.exist'); + }); +}); diff --git a/packages/app/cypress/support/profit-fixtures.ts b/packages/app/cypress/support/profit-fixtures.ts index 030042997..c4c98bf23 100644 --- a/packages/app/cypress/support/profit-fixtures.ts +++ b/packages/app/cypress/support/profit-fixtures.ts @@ -91,6 +91,10 @@ interface ProfitSku { precision: string; curve: Curve; tputScale: number; + /** Physical GPUs per row (TP); eight when unset. */ + gpus?: number; + /** Extra metric keys every row of this SKU carries, e.g. validated power telemetry. */ + metrics?: Record; } export const PROFIT_SKUS: ProfitSku[] = [ @@ -101,14 +105,41 @@ export const PROFIT_SKUS: ProfitSku[] = [ { hardware: 'mi355x', framework: 'vllm', precision: 'fp4', curve: WIDE_CURVE, tputScale: 0.9 }, ]; +/** + * One GB200 NVL72 compute tray (four GPUs, two Grace sockets) with the CPU-side + * keys the srt-slurm CPU power leg publishes, module sensor included, so the + * smart-provisioning basis can price NVL72. Kept out of `PROFIT_SKUS` so the + * default bar counts the other specs lock down do not move. Watts are + * controlled inputs, not published constants. + */ +const NVL72_SKU: ProfitSku = { + hardware: 'gb200', + framework: 'sglang', + precision: 'fp4', + curve: WIDE_CURVE, + tputScale: 1.3, + gpus: 4, + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 1, + avg_power_w: 900.25, + avg_total_gpu_power_w: 3601, + avg_cpu_socket_power_w: 250.5, + avg_total_cpu_power_w: 501, + avg_total_module_power_w: 4300.75, + }, +}; + let idCursor = 800_000; export const profitBenchmarkRows = ( dbKey: string = PROFIT_MODEL_DB_KEY, date = PROFIT_DATE, runId?: number, + skus: readonly ProfitSku[] = PROFIT_SKUS, ) => - PROFIT_SKUS.flatMap((sku) => + skus.flatMap((sku) => sku.curve.map(([conc, intvty, tput, e2el]) => ({ id: idCursor++, hardware: sku.hardware, @@ -118,21 +149,20 @@ export const profitBenchmarkRows = ( spec_method: 'none', disagg: false, is_multinode: false, - prefill_tp: 8, - decode_tp: 8, - num_prefill_gpu: 8, - num_decode_gpu: 8, + prefill_tp: sku.gpus ?? 8, + decode_tp: sku.gpus ?? 8, + num_prefill_gpu: sku.gpus ?? 8, + num_decode_gpu: sku.gpus ?? 8, isl: null, osl: null, conc, offload_mode: 'on', benchmark_type: 'agentic_traces', image: `${sku.framework}:test`, - metrics: metricsFor( - intvty, - Math.round(tput * sku.tputScale * tputScaleFor(date, runId)), - e2el, - ), + metrics: { + ...metricsFor(intvty, Math.round(tput * sku.tputScale * tputScaleFor(date, runId)), e2el), + ...sku.metrics, + }, workers: null, date, workflow_run_id: runId ?? PROFIT_SINGLE_RUN_ID[date], @@ -141,6 +171,15 @@ export const profitBenchmarkRows = ( })), ); +/** GB200 NVL72 rows with measured compute-module power, for the smart-provisioning basis. */ +export const profitNvl72Rows = (dbKey: string = PROFIT_MODEL_DB_KEY, date = PROFIT_DATE) => + profitBenchmarkRows(dbKey, date, undefined, [NVL72_SKU]).map((row) => ({ + ...row, + power_audit: { + cpu: { sensor_kind: 'module', expected_sockets: 2, observed_sockets: 2 }, + }, + })); + export const profitAvailabilityRows = (dbKeys: readonly string[] = PROFIT_DB_KEYS) => dbKeys.flatMap((dbKey) => PROFIT_DATES.flatMap((date) => diff --git a/packages/app/next.config.ts b/packages/app/next.config.ts index 0ac25747e..2fceaf35f 100644 --- a/packages/app/next.config.ts +++ b/packages/app/next.config.ts @@ -12,6 +12,11 @@ const nextConfig: NextConfig = { // distDir, so distinct dirs let the two coexist. distDir: process.env.NEXT_DIST_DIR || '.next', allowedDevOrigins: allowedDevOriginsFromEnv(), + env: { + // Preview source links must follow the deployed commit, not the model content digest. + NEXT_PUBLIC_APP_SOURCE_REF: + process.env.VERCEL_GIT_COMMIT_SHA || process.env.GITHUB_SHA || 'master', + }, transpilePackages: ['@semianalysisai/inferencex-constants'], serverExternalPackages: ['shiki'], redirects() { diff --git a/packages/app/scripts/export-modeled-system-power.ts b/packages/app/scripts/export-modeled-system-power.ts index 6e3a59a40..94242ffa7 100644 --- a/packages/app/scripts/export-modeled-system-power.ts +++ b/packages/app/scripts/export-modeled-system-power.ts @@ -6,8 +6,13 @@ import { pathToFileURL } from 'node:url'; import { parseArgs } from 'node:util'; import type { BenchmarkRow } from '../src/lib/api'; -import { AIR_COOLED_SYSTEM_PUE, modelSystemPower } from '../src/lib/modeled-system-power'; -import profileData from '../src/lib/system-power-model.profiles.json'; +import { + AIR_COOLED_SYSTEM_PUE, + DLC_SYSTEM_PUE, + defaultSystemPue, + modelSystemPower, +} from '../src/lib/modeled-system-power'; +import { SYSTEM_POWER_MODEL_METADATA as profileData } from '../src/lib/system-power-model'; interface PowerAudit { power_valid: boolean; @@ -48,6 +53,34 @@ const mean = (values: (number | null)[]) => ? values.reduce((sum, value) => sum + value, 0) / values.length : null; +interface SystemPowerProfile { + assumptions: Record; + modelPath: string; +} +const CHASSIS_PROFILES: Record = profileData.profiles; +const RACK_PROFILES: Record = profileData.rackProfiles; +const isRackHardware = (hardware: string) => Object.hasOwn(RACK_PROFILES, hardware); +const profileFor = (hardware: string): SystemPowerProfile | null => + isRackHardware(hardware) + ? RACK_PROFILES[hardware] + : Object.hasOwn(CHASSIS_PROFILES, hardware) + ? CHASSIS_PROFILES[hardware] + : null; + +// The chassis notes are unchanged for x86 rows; NVL72 rows carry their own. +const CHASSIS_NOTES = { + boundary: + 'Measured GPU-board inputs; modeled GPU-chassis AC includes their CPU/DRAM, other model components, and PSU loss. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after GPU-chassis AC.', + extrapolation: + 'A partially allocated chassis is modeled at measured per-GPU power × 8 (the source sweep input), assuming the unmeasured GPUs run the same workload. Deployment values are the measured GPUs’ share of that chassis; per-GPU values divide by the modeled chassis GPU count.', +}; +const RACK_NOTES = { + boundary: + 'Measured compute-module input per NVL72 tray (module sensor, or GPU board + Grace socket with the source’s regulator-loss allowance); the Grace CPU and LPDDR5X are never modeled. Modeled rack residual: NVSwitch trays, NICs/DPUs, NVMe, tray fans and board, tray 50 V → 12 V conversion, power shelves and management switches, evaluated once for a rack of 18 trays at the measured trays’ mean input and amortised over 72 GPUs. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after rack AC.', + extrapolation: + 'A partially allocated tray is modeled at measured per-GPU power × 4 on the GPU-board share only, assuming the unmeasured GPUs run the same workload; a module reading already covers the whole tray and is never scaled. Deployment values are the measured GPUs’ share of that tray; per-GPU values divide by the modeled tray GPU count.', +}; + function estimatedEnergy( row: BenchmarkRow, modeled: ReturnType, @@ -119,13 +152,19 @@ function estimatedEnergy( }; } -export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_PUE) { +/** + * `pue` overrides every row; without it each row takes the dashboard's default for + * its hardware (1.3 air-cooled chassis, 1.1 DLC NVL72 rack) so article figures match + * chart hovers. + */ +export function buildComparison(input: ComparisonInput, pue?: number) { if (!input || typeof input.cohort !== 'string' || !Array.isArray(input.rows)) { throw new Error( 'Expected a cohort envelope with a rows array. See docs/powerx-system-power.md.', ); } - if (!finite(pue) || pue < 1) throw new Error('PUE must be a finite number >= 1.'); + if (pue !== undefined && (!finite(pue) || pue < 1)) + throw new Error('PUE must be a finite number >= 1.'); const ids = new Set(); const rows = input.rows.map((entry) => { const row = entry.benchmark; @@ -141,16 +180,22 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_ throw new Error(`Invalid benchmark input or duplicate id: ${entry.id}`); } ids.add(entry.id); + const hardware = row.hardware.toLowerCase(); + const rack = isRackHardware(hardware); + // Same estimate path and PUE selection as the dashboard (`modelSystemPower(row)`). const modeled = modelSystemPower(row, pue); - const profile = Object.entries(profileData.profiles).find( - ([key]) => key === row.hardware.toLowerCase(), - )?.[1]; + const rowPue = pue ?? defaultSystemPue(hardware); + const profile = profileFor(hardware); const measurementStatus = row.metrics.power_valid === 1 ? 'producer-valid' : row.metrics.power_valid === 0 ? 'invalid' : 'unverified'; + // The CPU-side keys carry their own verdict; only NVL72 rows report them. + const cpuValid = row.metrics.cpu_power_valid === 1; + const trayEstimate = + modeled.status === 'supported' && modeled.topologyBasis === 'nvl72-trays' ? modeled : null; return { id: entry.id, cell: entry.cell ?? null, @@ -165,10 +210,28 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_ total_gpu_w: measurement(row.metrics.avg_total_gpu_power_w), total_gpu_j: measurement(row.metrics.total_gpu_energy_j), gpu_j_per_output_token: measurement(row.metrics.joules_per_output_token), + ...(rack + ? { + cpu_power_valid: row.metrics.cpu_power_valid ?? null, + total_grace_w: cpuValid ? measurement(row.metrics.avg_total_cpu_power_w) : null, + total_grace_j: cpuValid ? measurement(row.metrics.total_cpu_energy_j) : null, + total_module_w: cpuValid + ? measurement(row.metrics.avg_total_module_power_w) + : null, + total_module_j: cpuValid + ? measurement(row.metrics.total_module_energy_j) + : null, + } + : {}), } : null, - assumptions: profile ? { ...profile.assumptions, pue } : null, + pue: rowPue, + measured_basis: trayEstimate?.measuredBasis ?? null, + sensor_kind: trayEstimate?.sensorKind ?? null, + assumptions: profile ? { ...profile.assumptions, pue: rowPue } : null, model_path: profile?.modelPath ?? null, + calculation_boundary: rack ? RACK_NOTES.boundary : CHASSIS_NOTES.boundary, + extrapolation_note: rack ? RACK_NOTES.extrapolation : CHASSIS_NOTES.extrapolation, modeled, estimated_energy: estimatedEnergy(row, modeled, entry.audit), audit: entry.audit ?? null, @@ -214,7 +277,11 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_ model_revision: profileData.modelRevision, model_path: replicates[0].model_path, assumptions: replicates[0].assumptions, - pue, + // Replicates share hardware, so they share the PUE selection. + pue: replicates[0].pue, + measured_bases: [ + ...new Set(replicates.flatMap((row) => (row.measured_basis ? [row.measured_basis] : []))), + ], status: complete ? 'supported' : 'unsupported', unsupported_reasons: [ ...new Set( @@ -276,16 +343,18 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_ metadata: { cohort: input.cohort, source: input.metadata, - pue, + // Explicit --pue applies to every row; otherwise each row records its own default. + pue_override: pue ?? null, + pue_defaults: { air_cooled_chassis: AIR_COOLED_SYSTEM_PUE, dlc_nvl72_rack: DLC_SYSTEM_PUE }, scope: { benchmark_type: 'single_turn', isl: 8192, osl: 1024 }, selection: 'Every supplied row is retained, including unsupported, invalid, and missing-input cases.', aggregation: 'Each replicate is modeled first. Cell means include every replicate; any unavailable value leaves its cell mean unavailable.', - boundary: - 'Measured GPU-board inputs; modeled GPU-chassis AC includes their CPU/DRAM, other model components, and PSU loss. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after GPU-chassis AC.', - extrapolation: - 'A partially allocated chassis is modeled at measured per-GPU power × 8 (the source sweep input), assuming the unmeasured GPUs run the same workload. Deployment values are the measured GPUs’ share of that chassis; per-GPU values divide by the modeled chassis GPU count.', + boundary: CHASSIS_NOTES.boundary, + extrapolation: CHASSIS_NOTES.extrapolation, + rack_boundary: RACK_NOTES.boundary, + rack_extrapolation: RACK_NOTES.extrapolation, energy_caveat: 'Energy from modeled average power is an estimate. Nonlinear fan/PSU behavior is not integrated over time. Energy requires an exact matching audit window and successful token counts.', model: profileData, @@ -321,7 +390,7 @@ async function main() { }); if (!values.input || !values.output) throw new Error( - 'Usage: bun packages/app/scripts/export-modeled-system-power.ts --input cohort.json --output NEW_DIRECTORY [--pue 1.3]', + 'Usage: bun packages/app/scripts/export-modeled-system-power.ts --input cohort.json --output NEW_DIRECTORY [--pue 1.3]. Without --pue each row uses the dashboard default for its hardware (1.3 air-cooled chassis, 1.1 DLC NVL72 rack).', ); const inputBytes = await readFile(values.input); const result = buildComparison( @@ -334,6 +403,7 @@ async function main() { 'packages/app/src/lib/modeled-system-power.ts', 'packages/app/src/lib/system-power-model.ts', 'packages/app/src/lib/system-power-model.profiles.json', + 'packages/app/src/lib/system-power-model.provenance.json', ]; const hashes: Record = {}; for (const path of codePaths) @@ -377,6 +447,11 @@ async function main() { measured_total_gpu_w: row.measured_inputs?.total_gpu_w, measured_total_gpu_j: row.measured_inputs?.total_gpu_j, measured_gpu_j_per_output_token: row.measured_inputs?.gpu_j_per_output_token, + cpu_power_valid: row.measured_inputs?.cpu_power_valid, + measured_total_grace_w: row.measured_inputs?.total_grace_w, + measured_total_module_w: row.measured_inputs?.total_module_w, + measured_basis: row.measured_basis, + sensor_kind: row.sensor_kind, modeled_status: row.modeled.status, unsupported_reason: row.modeled.status === 'unsupported' ? row.modeled.reason : null, modeled_chassis_ac_w: row.modeled.status === 'supported' ? row.modeled.chassisAcWatts : null, @@ -395,11 +470,11 @@ async function main() { topology_basis: row.modeled.status === 'supported' ? row.modeled.topologyBasis : null, model_revision: row.modeled.modelRevision, model_status: profileData.status, - calculation_boundary: metadata.boundary, - extrapolation_note: metadata.extrapolation, + calculation_boundary: row.calculation_boundary, + extrapolation_note: row.extrapolation_note, energy_caveat: metadata.energy_caveat, model_path: row.model_path, - pue: metadata.pue, + pue: row.pue, assumptions: row.assumptions, estimated_energy: row.estimated_energy, estimated_energy_status: row.estimated_energy.status, diff --git a/packages/app/scripts/generate-system-power-reference.py b/packages/app/scripts/generate-system-power-reference.py deleted file mode 100644 index 1c303cfe2..000000000 --- a/packages/app/scripts/generate-system-power-reference.py +++ /dev/null @@ -1,155 +0,0 @@ -#!/usr/bin/env python3 -"""Regenerate fixed 8k1k profiles and parity cases from Oren's pinned Python models. - -Usage: python3 packages/app/scripts/generate-system-power-reference.py /path/to/inferencex_power_model -Only Python's standard library is required. No telemetry, dependencies, or GPUs are fetched. -Apply the repository formatter to generated JSON before committing. -""" - -import argparse -from dataclasses import asdict -import hashlib -import importlib -import json -from pathlib import Path -import subprocess -import sys - -sys.dont_write_bytecode = True -REVISION = "ca4403aa527069857351ad8047dbb726844b3382" -SOURCE = "https://github.com/SemiAnalysisAI/inferencex_power_model" -MODELS = { - "h100": ("hgx_h100_chassis/h100_chassis_power_model.py", "h100_chassis_power", "make_h100_config"), - "h200": ("hgx_h200_chassis/h200_chassis_power_model.py", "h200_chassis_power", "make_h200_config"), - "b200": ("hgx_b200_chassis/b200_chassis_power_model.py", "b200_chassis_power", "B200ChassisMasterConfig"), - "b300": ("hgx_b300_chassis/b300_chassis_power_model.py", "b300_chassis_power", "B300ChassisConfig"), - "mi300x": ("mi300x_chassis/mi300x_chassis_power_model.py", "mi300x_chassis_power", "MI300XChassisConfig"), - "mi325x": ("mi325x_chassis/mi325x_chassis_power_model.py", "mi325x_chassis_power", "MI325XChassisConfig"), - "mi355x": ("mi355x_chassis/mi355x_chassis_power_model.py", "mi355x_chassis_power", "MI355XChassisConfig"), -} -ASSUMPTIONS = {"u_pcie": 0.05, "u_cpu": 0.20, "u_ram": 0.20, "u_nvme": 0.0, "pue": 1.20} - - -def git(repo, *args): - return subprocess.check_output(["git", "-C", str(repo), *args], text=True).strip() - - -def main(): - parser = argparse.ArgumentParser(description=__doc__) - parser.add_argument("model_repo", type=Path) - parser.add_argument("--output-dir", type=Path, default=Path(__file__).resolve().parents[1] / "src/lib") - args = parser.parse_args() - repo = args.model_repo.resolve() - if git(repo, "rev-parse", "HEAD") != REVISION: - raise SystemExit(f"Model checkout must be pinned to {REVISION}") - if git(repo, "status", "--porcelain", "--untracked-files=all", "--", "*.py", "README.md", "AGENTS.md"): - raise SystemExit("Model Python sources or methodology files are dirty; use the clean pinned revision") - - profiles, cases = {}, [] - for hardware, (model_path, function_name, config_factory) in MODELS.items(): - path = repo / "human_verified" / model_path - sys.path.insert(0, str(path.parent)) - module = importlib.import_module(path.stem) - function, cfg = getattr(module, function_name), getattr(module, config_factory)() - assumptions = dict(ASSUMPTIONS) - assumptions.update({"u_eth": 0.0} if hardware.startswith("mi") else {"u_nvlink": 0.50, "u_ib": 0.0}) - if hardware in ("b200", "b300"): - assumptions["u_dpu"] = 0.0 - baseline = function(0.0, cfg=cfg, **assumptions) - fixed_components = {key: value for key, value in baseline["components_dc_w"].items() - if not key.endswith("_measured") and key != "chassis_fans"} - fans, psu = cfg.fans, cfg.psu - capacity = psu.active_capacity_w if hardware == "b200" else psu.load_sharing_capacity_w - limit = psu.active_capacity_w if hardware == "b200" else ( - psu.redundant_capacity_w if hardware in ("h100", "h200") else psu.modeled_capacity_w) - profiles[hardware] = { - "modelPath": "human_verified/" + model_path, - "functionName": function_name, - "configFactory": config_factory, - "gpuCount": 8, - "assumptions": {**assumptions, "fan_pwm": None}, - "defaultConfig": asdict(cfg), - "fixedComponentsDcWatts": fixed_components, - "fan": { - "electricalNameplateWatts": fans.electrical_nameplate_w, - "electricalGroupsWatts": [fans.n_80mm * fans.rated_80mm_w, fans.n_60mm * fans.rated_60mm_w] - if hardware == "b200" else [fans.electrical_nameplate_w], - "minPwm": fans.min_pwm_frac, - "maxPwm": fans.normal_max_pwm_frac, - "fullCoolingLoadWatts": fans.full_cooling_load_w, - "exponent": fans.fan_curve_exponent, - }, - "psu": { - "loadSharingCapacityWatts": capacity, - "maxDcWatts": limit, - "efficiencyCurve": sorted(psu.efficiency_curve.items()), - }, - } - - def evaluate(gpu, pue=1.2): - return function(gpu, cfg=cfg, **{**assumptions, "pue": pue}) - - fixed = sum(fixed_components.values()) - samples = {0.05, 1.25, 1000.25, 1000.75, 2400.0, 4000.0, 5600.0} - # Both sides of fan saturation and all reachable PSU interpolation knots. - for boundary in (fans.full_cooling_load_w - fixed,): - samples.update(round(boundary + offset, 3) for offset in (-0.2, 0.0, 0.2) if boundary + offset > 0) - for fraction in sorted(psu.efficiency_curve): - target = fraction * capacity - if not baseline["dc_total_w"] < target <= limit: - continue - lo, hi = 0.0, limit - for _ in range(50): - mid = (lo + hi) / 2 - try: - below = evaluate(mid)["dc_total_w"] < target - except ValueError: - below = False - if below: - lo = mid - else: - hi = mid - samples.update(round(hi + offset, 3) for offset in (-0.2, 0.0, 0.2)) - samples.add(limit) # Capacity overflow must be unavailable, never clamped. - for gpu in sorted(samples): - for pue in (1.0, 1.2): - case = {"hardware": hardware, "measuredGpuWatts": gpu, "pue": pue} - try: - result = evaluate(gpu, pue) - case["expected"] = { - "preFanDcWatts": result["pre_fan_dc_w"], - "fanWatts": result["components_dc_w"]["chassis_fans"], - "dcWatts": result["dc_total_w"], - "psuEfficiency": result["psu_efficiency"], - "psuLossWatts": result["psu_conversion_loss_w"], - "chassisAcWatts": result["ac_wall_w"], - "facilityWatts": result["utility_power_w"], - } - except ValueError as error: - case["expected"] = None - case["referenceError"] = str(error) - cases.append(case) - - source_hashes = {} - # Include the complete pinned Python implementation and plot entry points. - for relative in git(repo, "ls-files", "*.py").splitlines(): - source_hashes[relative] = hashlib.sha256((repo / relative).read_bytes()).hexdigest() - provenance = {"modelRevision": REVISION, "source": SOURCE, "status": "DRAFT / pending human verification"} - outputs = { - "system-power-model.profiles.json": { - **provenance, - "assumptionsSource": f"{SOURCE}/blob/{REVISION}/README.md#chassis-models", - "assumptions": ASSUMPTIONS, - "sourceSha256": dict(sorted(source_hashes.items())), - "profiles": profiles, - }, - "system-power-model.reference.json": {**provenance, "cases": cases}, - } - args.output_dir.mkdir(parents=True, exist_ok=True) - for filename, payload in outputs.items(): - (args.output_dir / filename).write_text(json.dumps(payload, indent=2, allow_nan=False) + "\n") - print(f"Generated {len(profiles)} profiles and {len(cases)} Python reference cases at {REVISION}") - - -if __name__ == "__main__": - main() diff --git a/packages/app/scripts/update-system-power-provenance.ts b/packages/app/scripts/update-system-power-provenance.ts new file mode 100644 index 000000000..38c085d7a --- /dev/null +++ b/packages/app/scripts/update-system-power-provenance.ts @@ -0,0 +1,51 @@ +import { createHash } from 'node:crypto'; +import { readFile, writeFile } from 'node:fs/promises'; +import { resolve } from 'node:path'; +import { pathToFileURL } from 'node:url'; + +const ROOT = resolve(import.meta.dirname, '../../..'); +const PROFILE_PATH = 'packages/app/src/lib/system-power-model.profiles.json'; +const MANIFEST_PATH = 'packages/app/src/lib/system-power-model.provenance.json'; +const SOURCE_PATHS = [ + 'packages/app/src/lib/modeled-system-power.ts', + PROFILE_PATH, + 'packages/app/src/lib/system-power-model.ts', +]; +const sha256 = (value: string | Uint8Array) => createHash('sha256').update(value).digest('hex'); + +export async function buildSystemPowerProvenance(root = ROOT) { + const sourceSha256 = Object.fromEntries( + await Promise.all( + SOURCE_PATHS.map(async (path) => [path, sha256(await readFile(resolve(root, path)))]), + ), + ); + return { + modelRevision: `app-sha256:${sha256(JSON.stringify(sourceSha256))}`, + modelRevisionStatus: 'App-owned TypeScript equations, parameters and admission/PUE policy', + source: 'https://github.com/SemiAnalysisAI/InferenceX-app', + status: 'DRAFT / pending human verification', + assumptionsSource: `${PROFILE_PATH}#/assumptions`, + rackAssumptionsSource: `${PROFILE_PATH}#/rackAssumptions`, + sourceSha256, + }; +} + +async function main() { + const args = process.argv.slice(2); + if (args.length > 1 || (args.length === 1 && args[0] !== '--check')) + throw new Error('Usage: bun packages/app/scripts/update-system-power-provenance.ts [--check]'); + const current = await buildSystemPowerProvenance(); + const target = resolve(ROOT, MANIFEST_PATH); + if (args[0] === '--check') { + const stored = JSON.parse(await readFile(target, 'utf8')); + if (JSON.stringify(stored) !== JSON.stringify(current)) + throw new Error('Model provenance is stale; run this script without --check.'); + } else { + await writeFile(target, `${JSON.stringify(current, null, 2)}\n`); + } + console.log(current.modelRevision); +} + +if (process.argv[1] && pathToFileURL(resolve(process.argv[1])).href === import.meta.url) { + await main(); +} diff --git a/packages/app/src/app/api/v1/views/extensions.test.ts b/packages/app/src/app/api/v1/views/extensions.test.ts index 6e3722cf4..344ea8e68 100644 --- a/packages/app/src/app/api/v1/views/extensions.test.ts +++ b/packages/app/src/app/api/v1/views/extensions.test.ts @@ -349,7 +349,7 @@ describe('new dashboard projections', () => { expect([firstFetchCount, mocks.unofficial.mock.calls.length]).toEqual([1, 2]); }, ); - it.each(['modeled', 'compare'])( + it.each(['modeled'])( 'uses exact comparison snapshots with %s power while keeping the primary date cutoff', async (powerBasis) => { const rows = [ @@ -391,6 +391,164 @@ describe('new dashboard projections', () => { }, ); it.each(['modeled', 'compare'])( + 'uses valid power curves at the same target in official, historical and overlay %s estimates', + async (powerBasis) => { + const curve = (scale: number, date: string) => + [ + [20, 9000, 1], + [40, 8000, 0], + [60, 3000, 1], + ].map(([interactivity, throughput, powerValid], index) => + agenticRow({ + id: scale * 100 + index, + conc: 3 - index, + // The exact logical snapshot includes a retained endpoint from an + // older producer run; power selection must not split that curve. + date: scale === 2 && index === 0 ? '2026-09-08' : date, + run_url: + scale === 2 && index === 0 + ? 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/122' + : agenticRow().run_url, + curve_workflow_run_id: scale * 1000, + curve_date: date, + metrics: { + ...agenticRow().metrics, + p90_itl: 1 / interactivity, + tput_per_gpu: throughput * scale, + input_tput_per_gpu: throughput * scale * 0.9, + output_tput_per_gpu: throughput * scale * 0.1, + power_valid: powerValid, + }, + }), + ); + mocks.benchmarks.mockImplementation((request: NextRequest) => + Response.json( + request.nextUrl.searchParams.get('exact') === 'true' + ? curve(2, '2026-09-09') + : curve(1, '2026-09-10'), + ), + ); + mocks.unofficial.mockImplementation(() => + Response.json({ benchmarks: curve(3, '2026-09-10'), evaluations: [] }), + ); + const query = `model=DeepSeek-V4-Pro&precisions=fp4&target=45&priceSource=custom&inputPrice=1&cachedInputPrice=1&outputPrice=1&powerBasis=${powerBasis}&dates=2026-09-09&unofficialrun=456`; + const response = await gw(req('profit-estimator-per-gigawatt', query)); + expect(response.status).toBe(200); + const body = await response.json(); + for (const [output, scale] of [ + [body.data, 1], + [body.comparisons[0].data, 2], + [body.overlays, 3], + ]) { + expect(output.skipped).toEqual([]); + expect(output.rows).toHaveLength(powerBasis === 'compare' ? 2 : 1); + // The existing Steffen curve over valid endpoints at 20/60 yields + // 4766.6015625 tok/s/GPU at 45. + // Both power budgets use that throughput, even though the full performance + // frontier includes the faster, power-invalid knot at 40. + for (const row of output.rows) + expect(row.revenuePerGpuHour).toBeCloseTo(17.159765625 * scale); + } + const provisioned = await gw( + req( + 'profit-estimator-per-gigawatt', + query.replace(`powerBasis=${powerBasis}`, 'powerBasis=provisioned'), + ), + ); + const baseline = await provisioned.json(); + expect(baseline.data.rows).toHaveLength(1); + expect(baseline.data.rows[0].revenuePerGpuHour).toBeGreaterThan(17.159765625); + }, + ); + it('returns NVL72 measured basis and matching capacity for official and overlay estimates', async () => { + const tray = agenticRow({ + hardware: 'gb200', + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + power_audit: { + cpu: { sensor_kind: 'module', expected_sockets: 2, observed_sockets: 2 }, + }, + metrics: { + ...agenticRow().metrics, + avg_power_w: 900.25, + avg_total_gpu_power_w: 3601, + cpu_power_valid: 1, + avg_total_module_power_w: 4300.75, + }, + }); + mocks.benchmarks.mockImplementation(() => Response.json([tray])); + mocks.unofficial.mockImplementation(() => + Response.json({ benchmarks: [{ ...tray, id: 456 }], evaluations: [] }), + ); + const response = await gw( + req( + 'profit-estimator-per-gigawatt', + 'model=DeepSeek-V4-Pro&target=45&priceSource=custom&powerBasis=compare&unofficialrun=456', + ), + ); + expect(response.status).toBe(200); + const body = await response.json(); + for (const output of [body.data, body.overlays]) { + expect(output.skipped).toEqual([]); + expect(output.rows).toHaveLength(2); + const [provisioned, modeled] = output.rows; + expect(provisioned.powerLabel).toBe('All in Provisioned'); + expect(provisioned.powerSource).toBeUndefined(); + expect(modeled.powerLabel).toBe('All in Measured'); + expect(modeled.powerSource).toMatchObject({ + topology: 'nvl72-trays', + measuredBasis: 'module', + sensorKind: 'module', + pue: 1.1, + modelPath: 'packages/app/src/lib/system-power-model.ts', + modelRevision: expect.stringMatching(/^app-sha256:[0-9a-f]{64}$/u), + profileSha256: expect.stringMatching(/^[0-9a-f]{64}$/u), + }); + // Pinned rack reference: 1.6077325 kW/GPU, including PUE and planning margin. + expect(modeled.gpuHours).toBeCloseTo((1_000_000 / 1.6077325) * 8760, 2); + expect(modeled.revenuePerGpuHour).toBe(provisioned.revenuePerGpuHour); + } + }); + it.each([ + [1, 'no-cpu-power'], + [0, 'no-measured-power'], + ] as const)( + 'retains provisioned NVL72 estimates with GPU verdict %i and explains missing measured power', + async (powerValid, reason) => { + const missing = agenticRow({ + hardware: 'gb300', + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + metrics: { + ...agenticRow().metrics, + avg_total_gpu_power_w: 2400, + power_valid: powerValid, + }, + }); + mocks.benchmarks.mockImplementation(() => Response.json([missing])); + for (const powerBasis of ['compare', 'modeled']) { + const response = await gw( + req( + 'profit-estimator-per-gigawatt', + `model=DeepSeek-V4-Pro&target=45&priceSource=custom&powerBasis=${powerBasis}`, + ), + ); + expect(response.status).toBe(200); + const body = await response.json(); + expect(body.data.rows).toHaveLength(powerBasis === 'compare' ? 1 : 0); + if (powerBasis === 'compare') { + expect(body.data.rows[0].powerLabel).toBe('All in Provisioned'); + expect(body.data.rows[0].powerSource).toBeUndefined(); + } + expect(body.data.skipped).toMatchObject([{ reason }]); + } + }, + ); + it.each(['modeled'])( 'labels full-chassis extrapolation for official and overlay %s estimates', async (powerBasis) => { const partial = agenticRow({ diff --git a/packages/app/src/app/api/v1/views/inference/route.test.ts b/packages/app/src/app/api/v1/views/inference/route.test.ts index 662b9717a..a7d3a8d29 100644 --- a/packages/app/src/app/api/v1/views/inference/route.test.ts +++ b/packages/app/src/app/api/v1/views/inference/route.test.ts @@ -102,6 +102,68 @@ beforeEach(() => { }); describe('GET /api/v1/views/inference', () => { + it('keeps unavailable all-in table rows in official, compared and unofficial scopes and CSV', async () => { + const gpuOnly = makeRow({ + hardware: 'gb200', + metrics: { + ...makeRow().metrics, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + power_valid: 1, + power_metric_schema_version: 2, + }, + }); + const historical = { ...gpuOnly, id: 998, date: '2026-02-28' }; + const unofficial = { + ...gpuOnly, + id: 999, + run_url: 'https://github.com/org/repo/actions/runs/999', + }; + mockGetLatestBenchmarks.mockImplementation((_db, _models, date) => + Promise.resolve(date === historical.date ? [historical] : [gpuOnly]), + ); + mockUnofficialRun.mockImplementation(() => + Response.json({ benchmarks: [unofficial], evaluations: [] }), + ); + const query = + '/api/v1/views/inference?model=DeepSeek-R1-0528&metric=utilityModeledWatts&best=false&optimal=true&dates=2026-02-28&unofficialrun=999'; + const response = await GET(request(query)); + const body = await response.json(); + expect(response.status).toBe(200); + for (const [scope, id] of [ + [body, gpuOnly.id], + [body.comparisons[0], historical.id], + [body.overlays[0], unofficial.id], + ]) { + expect(scope.series).toEqual([]); + expect(scope.count).toBe(0); + expect(scope.tableRows).toMatchObject([ + { + id, + y: null, + measuredGpuWatts: 500, + status: 'unavailable', + unavailableReason: 'cpu-telemetry', + }, + ]); + expect(scope).not.toHaveProperty('observedPoints'); + } + expect(body.comparisons[0].tableRows[0].date).toBe(historical.date); + expect(body.overlays[0].tableRows[0].runId).toBe(999); + + const csvResponse = await GET(request(`${query}&format=csv`)); + const csv = await csvResponse.text(); + const [header, ...lines] = csv.trim().split('\r\n'); + const columns = header.split(','); + expect(lines).toHaveLength(3); + for (const line of lines) { + const values = line.split(','); + expect(values[columns.indexOf('y')]).toBe(''); + expect(values[columns.indexOf('measuredGpuWatts')]).toBe('500'); + expect(values[columns.indexOf('unavailableReason')]).toBe('cpu-telemetry'); + } + }); + it('compares stitched observations before frontier pruning and preserves each producer endpoint', async () => { const rows = ['h200', 'mi300x'].flatMap((hardware, index) => [20, 60].map((x, position) => diff --git a/packages/app/src/app/api/v1/views/inference/route.ts b/packages/app/src/app/api/v1/views/inference/route.ts index 8ae6f0d7d..7fc666470 100644 --- a/packages/app/src/app/api/v1/views/inference/route.ts +++ b/packages/app/src/app/api/v1/views/inference/route.ts @@ -197,7 +197,14 @@ function buildView( return { resolvedPrecisions, result }; } -function csvRows(data: { result: Pick }) { +function csvRows(data: { result: Pick }) { + if (data.result.tableRows) + return data.result.tableRows.map(({ metrics, ...row }) => ({ + ...row, + ...Object.fromEntries( + Object.entries(metrics).map(([key, value]) => [`metric_${key}`, value]), + ), + })); return data.result.series.flatMap((entry) => entry.points.map((point) => ({ hwKey: entry.hwKey, @@ -583,6 +590,7 @@ export function GET(request: NextRequest) { frontier: data.result.frontier, hardware: data.result.hardware, series: data.result.series, + ...(data.result.tableRows === undefined ? {} : { tableRows: data.result.tableRows }), count: data.result.count, comparisons, overlays, diff --git a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx index afab4e98b..5d0d0d06b 100644 --- a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx +++ b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx @@ -28,6 +28,7 @@ import { type ProfitEstimatorAssumptions, type ProfitEstimatorRow, } from './profit-estimator'; +import type { ProfitPowerSource } from './profit-power'; export type ProfitSegmentKind = 'tco' | 'labCut' | 'profit' | 'loss'; @@ -96,6 +97,8 @@ const GLYPH_WIDTH_EM = 0.55; const LABEL_SIDE_PAD_PX = 4; /** The labels above a bar may borrow this much of the gap to each neighbour, in px. */ const X_GAP_ALLOWANCE = 12; +/** Keep comparison labels and vendor marks readable when bars outgrow the viewport. */ +const MIN_BAR_STEP_PX = 96; /** * Horizontal room the labels above a bar may use, in px. A label may overhang @@ -265,7 +268,15 @@ const STRINGS = { ofRevenue: 'of revenue', dismiss: 'Click anywhere to dismiss', runDate: 'Run date', + powerBasis: 'Power basis', + powerBasisLabels: { + chassis: 'measured GPU board; CPU, DRAM, networking, storage, board, fans and PSU modeled', + module: 'measured module (GPU + HBM + Grace + LPDDR5X; module sensor)', + 'grace-socket': + 'measured GPU board + Grace socket (Grace socket sensor), regulator loss modeled', + }, noData: 'No SKU can be priced for the current selection.', + scrollHint: 'Scroll horizontally to view the full chart.', }, zh: { yAxisModeled: '每吉瓦设施总功耗对应的年收入(美元)', @@ -291,7 +302,15 @@ const STRINGS = { ofRevenue: '(占收入)', dismiss: '点击任意位置关闭', runDate: '运行日期', + powerBasis: '功耗口径', + powerBasisLabels: { + chassis: '实测 GPU 板卡功耗;CPU、DRAM、网络、存储、主板、风扇和 PSU 由模型估算', + module: '实测模块功耗(GPU + HBM + Grace + LPDDR5X;模块传感器)', + 'grace-socket': + '实测 GPU 板卡 + Grace socket 功耗(Grace socket 传感器),稳压损耗由模型估算', + }, noData: '当前选择下没有可定价的 SKU。', + scrollHint: '横向滚动查看完整图表。', }, } as const; @@ -299,6 +318,12 @@ export function profitEstimatorChartStrings(locale: Locale) { return STRINGS[locale]; } +/** Name what a measured + modeled row's power was measured on. */ +export function powerBasisLabel(source: ProfitPowerSource, locale: Locale): string { + const labels = STRINGS[locale].powerBasisLabels; + return labels[source.topology === 'chassis' ? 'chassis' : source.sensorKind]; +} + /** Segment label lines that fit a given pixel height: name and amount, amount only, or none. */ export function segmentLabelLines( kind: ProfitSegmentKind, @@ -630,6 +655,7 @@ export function generateProfitTooltipHTML( ${isPinned ? `
${t.dismiss}
` : ''}
${label}
${row.date ? line(t.runDate, escapeHtml(row.dateLabel ?? row.date)) : ''} + ${row.powerSource ? line(t.powerBasis, escapeHtml(powerBasisLabel(row.powerSource, locale))) : ''} ${line(`${t.revenue} (${t.utilization} ${assumptions.utilizationPct}%)`, usd(row.revenue))} ${line(t.tco, usd(row.tco), skuColor, TCO_OPACITY)} ${line(t.grossMargin, usd(row.grossMargin))} @@ -1010,12 +1036,22 @@ export default function ProfitEstimatorChart({ ); const baseMargin = compact ? CHART_MARGIN_COMPACT : CHART_MARGIN; + const minimumMargin = slantedMargins( + [...labelMap.values()], + MIN_BAR_STEP_PX, + CHART_TYPE.axisLabelSub, + baseMargin, + ); + const chartWidth = Math.max( + dimensions.width, + rows.length * MIN_BAR_STEP_PX + minimumMargin.left + minimumMargin.right, + ); // Upright two-line labels when each SKU has room for them; slanted otherwise. const labelLayout = useMemo(() => { - const plotWidth = dimensions.width - baseMargin.left - baseMargin.right; + const plotWidth = chartWidth - baseMargin.left - baseMargin.right; const slot = rows.length > 0 ? plotWidth / rows.length : 0; return xLabelLayout([...labelMap.values()], slot, CHART_TYPE.axisLabelSub); - }, [dimensions.width, baseMargin, rows.length, labelMap]); + }, [chartWidth, baseMargin, rows.length, labelMap]); const margin = useMemo(() => { if (labelLayout === 'stacked') { const dated = [...labelMap.values()].some((label) => splitHistoryLabel(label)[1] !== ''); @@ -1024,15 +1060,15 @@ export default function ProfitEstimatorChart({ bottom: X_LABEL_STACKED_BOTTOM + (dated ? X_LABEL_HISTORY_LINE_PX : 0), }; } - const plotWidth = dimensions.width - baseMargin.left - baseMargin.right; + const plotWidth = chartWidth - baseMargin.left - baseMargin.right; const slot = rows.length > 0 ? plotWidth / rows.length : 0; return slantedMargins([...labelMap.values()], slot, CHART_TYPE.axisLabelSub, baseMargin); - }, [baseMargin, labelLayout, dimensions.width, rows.length, labelMap]); + }, [baseMargin, labelLayout, chartWidth, rows.length, labelMap]); const plotHeight = chartHeight - margin.top - margin.bottom; // The vendor mark grows with the bar, so the headroom above the tallest stack // has to be sized from the same band width the renderer will see. const bandwidth = useMemo(() => { - const plotWidth = dimensions.width - margin.left - margin.right; + const plotWidth = chartWidth - margin.left - margin.right; if (plotWidth <= 0 || rows.length === 0) return 0; return d3 .scaleBand() @@ -1040,7 +1076,7 @@ export default function ProfitEstimatorChart({ .range([0, plotWidth]) .padding(BAND_PADDING) .bandwidth(); - }, [dimensions.width, margin, rows]); + }, [chartWidth, margin, rows]); const yDomain = useMemo( () => profitYDomain(rows, plotHeight, stackHeadroomPx(barMarkHeight(bandwidth))), [rows, plotHeight, bandwidth], @@ -1167,6 +1203,11 @@ export default function ProfitEstimatorChart({ instructions="" legendElement={legendElement} caption={caption} + scrollablePlot={{ + minWidth: chartWidth, + label: t.scrollHint, + enabled: chartWidth > dimensions.width, + }} />
); diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx index 6c1be5d32..e737a393d 100644 --- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx +++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx @@ -91,10 +91,9 @@ import { profitModelDefaults, type ProfitBasis, type ProfitEstimatorRow, - type ProfitEstimatorSkipReason, } from './profit-estimator'; -import { profitEstimatorChartStrings, rowLabel } from './ProfitEstimatorChart'; -import { estimateProfitByPower, type ProfitPowerBasis } from './profit-power'; +import { powerBasisLabel, profitEstimatorChartStrings, rowLabel } from './ProfitEstimatorChart'; +import { estimateProfitByPower, powerSourceKey, type ProfitPowerBasis } from './profit-power'; import { buildProfitHistoryResults, historyFadeShare, @@ -206,7 +205,7 @@ const STRINGS = { benchmarkGroup: 'Benchmark Config', powerLabel: 'Power Estimation', powerTooltip: - 'Change only the power budget used to scale the same benchmark result to one GW. Pricing, throughput, utilization and unit costs stay the same.', + 'All in Measured uses the best power-valid curve at the selected target. Compare both uses that same curve for both bars when available. Pricing, utilization and unit costs stay the same.', powerOptions: { provisioned: POWER_BASIS_LABELS['utility-provisioned'].en, modeled: POWER_BASIS_LABELS['utility-modeled'].en, @@ -219,8 +218,10 @@ const STRINGS = { }, powerPreview: `${ALL_IN_MEASURED_NOTE.en} AgentX system power is not yet qualified.`, powerDetails: - 'GPU power is interpolated between the same throughput points. Includes PUE 1.3 and 10% headroom. Full-chassis extrapolation fills an eight-GPU server with replicas of the measured 1/2/4-GPU workload at the same per-GPU power and throughput; it does not measure a partly idle server.', - unavailableEstimates: (count: number) => `Unavailable estimates (${count})`, + 'GPU power is interpolated between the same throughput points. Includes PUE 1.3 for air-cooled chassis or 1.1 for NVL72, and 10% headroom. Aggregate multinode hosts use the measured deployment mean. Full-chassis extrapolation fills an eight-GPU server with replicas of the measured 1/2/4-GPU workload at the same per-GPU power and throughput; it does not measure a partly idle server.', + powerNvl72Note: (hardware: string, basis: string, pue: number) => + `${hardware}: ${basis}. Modeled: NVSwitch trays, NICs/DPUs, NVMe, power shelves, DLC PUE ${pue}.`, + csvPowerHeaders: ['Power basis', 'Power sensor', 'System power profile'], pricingGroup: 'Pricing Config', costProviderLabel: 'Cost Provider', costProviderTooltip: @@ -297,18 +298,6 @@ const STRINGS = { 'Revenue ($/GPU/hr, 100% util)', ], }, - skipped: (entries: string) => `Not priced: ${entries}.`, - skipReason: { - 'outside-measured-range': 'no measured point at the target interactivity', - 'no-power': 'no all-in power figure', - 'no-measured-power': 'no usable measured power for these benchmark points', - 'unsupported-power-hardware': 'no system power model for this hardware', - 'unsupported-power-topology': - 'this topology cannot be modeled as whole replicas on one eight-GPU server', - 'outside-power-model': 'these benchmark points are outside the supported power model', - 'no-cost': 'no TCO for this tier', - 'no-token-mix': 'no input/output token mix recorded', - } satisfies Record, compareHistory: 'Compare history', gpuConfig: 'Chip Config', gpuConfigTooltip: `Select up to ${PROFIT_HISTORY_MAX_GPUS} chip configurations to compare how their estimated revenue and profit have moved over time. Each config is priced again on every compared date (the ends of the date range, plus any date or run added from the Config Changelog below) using the run measured then, so software updates show up as a change in the bar.`, @@ -330,7 +319,7 @@ const STRINGS = { benchmarkGroup: '基准测试配置', powerLabel: '功耗估算方式', powerTooltip: - '仅更改将同一基准测试结果换算为每 GW 收益时采用的功耗预算。价格、吞吐量、利用率和单位成本保持不变。', + '整体实测功耗在选定目标下采用功耗有效的最优曲线。对比两种估算方式时,若实测估算可用,两根柱子采用同一条曲线。价格、利用率和单位成本保持不变。', powerOptions: { provisioned: POWER_BASIS_LABELS['utility-provisioned'].zh, modeled: POWER_BASIS_LABELS['utility-modeled'].zh, @@ -343,8 +332,10 @@ const STRINGS = { }, powerPreview: `${ALL_IN_MEASURED_NOTE.zh} AgentX 系统功耗模型尚未完成验证。`, powerDetails: - 'GPU 功耗在相同的吞吐量数据点间插值,计入 PUE 1.3 和 10% 功耗余量。整机外推假设八卡服务器部署多个相同的实测单卡、双卡或四卡实例,每卡功耗和吞吐量保持不变;它不代表部分 GPU 闲置时的整机实测功耗。', - unavailableEstimates: (count: number) => `无法估算(${count} 项)`, + 'GPU 功耗在相同的吞吐量数据点间插值,风冷机箱 PUE 为 1.3,NVL72 为 1.1,另加 10% 功耗余量。聚合多节点按部署平均功耗估算各台服务器。整机外推假设八卡服务器部署多个相同的实测单卡、双卡或四卡实例,每卡功耗和吞吐量保持不变;它不代表部分 GPU 闲置时的整机实测功耗。', + powerNvl72Note: (hardware: string, basis: string, pue: number) => + `${hardware}:${basis}。建模部分:NVSwitch tray、网卡/DPU、NVMe、电源架,液冷 PUE ${pue}。`, + csvPowerHeaders: ['功耗口径', '功耗传感器', '系统功耗 profile'], pricingGroup: '定价配置', costProviderLabel: '成本供应商', costProviderTooltip: @@ -421,17 +412,6 @@ const STRINGS = { '收入($/GPU/hr,100% 利用率)', ], }, - skipped: (entries: string) => `未定价:${entries}。`, - skipReason: { - 'outside-measured-range': '未在该交互性下实测', - 'no-power': '缺少全电源配置功率数据', - 'no-measured-power': '同一组基准测试数据点缺少有效功耗', - 'unsupported-power-hardware': '该硬件暂无适用的系统功耗模型', - 'unsupported-power-topology': '该拓扑无法按完整实例部署在单台八卡服务器上建模', - 'outside-power-model': '这些基准测试数据点超出功耗模型的适用范围', - 'no-cost': '该层级无 TCO 数据', - 'no-token-mix': '未记录输入/输出 token 比例', - } satisfies Record, compareHistory: '对比历史趋势', gpuConfig: '芯片配置', gpuConfigTooltip: `最多选择 ${PROFIT_HISTORY_MAX_GPUS} 个芯片配置,对比其收入与利润估算随时间的变化。每个配置都会用当日实测的运行结果,在每个对比日期(日期范围的起止两端,以及从下方配置变更日志中添加的日期或运行)重新估价,软件更新带来的差异会直接体现在柱形上。`, @@ -1003,7 +983,15 @@ function ProfitEstimatorInner({ // from that date's run with the same target, prices, and TCO tier. const fullEstimate = useMemo(() => { if (!hasData || !pricing) return { rows: [], skipped: [] }; - const current = getResults(targetValue, mode, interpolationCostProvider); + const curvePowerBasis = basis === 'gw-year' ? powerBasis : 'provisioned'; + const current = getResults( + targetValue, + mode, + interpolationCostProvider, + undefined, + false, + curvePowerBasis, + ); const results = historyActive ? [ ...current.filter((r) => selectedGPUs.includes(r.hwKey)), @@ -1014,6 +1002,7 @@ function ProfitEstimatorInner({ targetValue, mode, costProvider: interpolationCostProvider, + powerBasis: curvePowerBasis, currentRunIds: historyCurrentRunIds, }), ] @@ -1041,6 +1030,7 @@ function ProfitEstimatorInner({ hasData, pricing, getResults, + basis, powerBasis, t.powerBarLabels, targetValue, @@ -1359,21 +1349,24 @@ function ProfitEstimatorInner({ historyCurrentRunIds, ]); - const powerUnavailable = useMemo( - () => - t.skipped( - fullEstimate.skipped - .map((row) => { - const label = rowLabel( - { ...row, dateLabel: row.date ? historyEntryLabel(row.date) : undefined }, - hardwareConfig, - ); - return `${label}: ${t.skipReason[row.reason]}`; - }) - .join('; '), - ), - [fullEstimate.skipped, hardwareConfig, historyEntryLabel, t], - ); + const powerBasisNotes = useMemo(() => { + const notes = new Map(); + for (const row of estimate.rows) { + const source = row.powerSource; + if (source?.topology !== 'nvl72-trays') continue; + const key = `${row.hwKey}|${powerSourceKey(source)}`; + if (notes.has(key)) continue; + notes.set( + key, + t.powerNvl72Note( + rowLabel({ hwKey: row.hwKey }, hardwareConfig), + powerBasisLabel(source, locale), + source.pue, + ), + ); + } + return [...notes.values()]; + }, [estimate.rows, hardwareConfig, locale, t]); // Rendered as the chart's figcaption so it is part of the PNG export. const caption = useMemo(() => { @@ -1396,20 +1389,6 @@ function ProfitEstimatorInner({ {t.powerLabel}: {t.powerOptions[powerBasis]}

)} - {basis === 'gw-year' && powerBasis !== 'provisioned' && fullEstimate.skipped.length > 0 && ( -
- track('profit_estimator_power_unavailable_toggled')} - > - {t.unavailableEstimates(fullEstimate.skipped.length)} - -

{powerUnavailable}

-
- )} { // Whole dollars are plenty per GW-year; per chip-hour the cents are the figure. const usd = (value: number) => (basis === 'gw-year' ? Math.round(value) : value.toFixed(4)); + // Measured + modeled rows name their basis, sensor, and pinned profile so a + // spreadsheet can tell a measured module from a modeled chassis per row. + const includeBasis = powerControlsEnabled && powerBasis !== 'provisioned'; + const basisColumns = (row: ProfitEstimatorRow) => { + const source = row.powerSource; + if (!source) return [t.powerBarLabels.provisioned, '', '']; + return [ + powerBasisLabel(source, locale), + source.topology === 'nvl72-trays' ? source.sensorKind : '', + `${source.modelPath} @ ${source.modelRevision}${source.profileSha256 ? ` sha256:${source.profileSha256}` : ''}`, + ]; + }; const rows = estimate.rows.map((row) => [ rowLabel({ ...row, date: undefined }, hardwareConfig), row.precision?.toUpperCase() ?? '', @@ -1533,14 +1522,24 @@ function ProfitEstimatorInner({ row.revenuePerGpuHour.toFixed(4), // GPU-hours is 1 per chip-hour, so that basis has no column for it. ...(basis === 'gw-year' ? [Math.round(row.gpuHours)] : []), + ...(includeBasis ? basisColumns(row) : []), ]); const [sku, precision, ...rest] = t.csvHeaders[basis]; - exportToCsv(exportFileName, [sku, precision, t.csvDateHeader, ...rest], rows, [ + const headers = [ + sku, + precision, + t.csvDateHeader, + ...rest, + ...(includeBasis ? t.csvPowerHeaders : []), + ]; + exportToCsv(exportFileName, headers, rows, [ t.captionFormula[basis](assumptions.utilizationPct, assumptions.labCutPct), ...(powerControlsEnabled ? [ `${t.powerLabel}: ${t.powerOptions[powerBasis]}`, - ...(powerBasis === 'provisioned' ? [] : [t.powerPreview, t.powerDetails]), + ...(powerBasis === 'provisioned' + ? [] + : [t.powerPreview, t.powerDetails, ...powerBasisNotes]), ] : []), ]); @@ -1549,10 +1548,12 @@ function ProfitEstimatorInner({ hardwareConfig, exportFileName, t, + locale, assumptions, basis, selectedRunDate, powerBasis, + powerBasisNotes, powerControlsEnabled, ]); @@ -1626,7 +1627,12 @@ function ProfitEstimatorInner({
- {basis === 'gw-year' && - estimate.rows.length === 0 && - powerBasis !== 'provisioned' && ( -

- {t.powerPreview} {powerUnavailable} -

- )} { costProvider: 'costh' as const, }; + it('selects valid knots within each historical run without borrowing from older runs', () => { + const row = (x: number, throughput: number, valid: boolean, run = 200) => { + const point = agenticRow('2026-06-14', 'b200', x, throughput, { + workflow_run_id: run, + run_started_at: `2026-06-14T${run === 200 ? '12' : '06'}:00:00Z`, + }); + return { + ...point, + metrics: { + ...point.metrics, + power_valid: Number(valid), + power_metric_schema_version: 2, + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + }, + }; + }; + const rows = [row(20, 8000, true), row(40, 9000, false), row(80, 3000, true)]; + const selected = buildProfitHistoryResults([{ date: '2026-06-14', rows }], { + ...options, + powerBasis: 'modeled', + }); + expect(selected).toHaveLength(1); + expect(selected[0].nearestPoints.map((p) => p.interactivity)).toEqual([20, 80]); + expect(modeledPowerAtTarget(selected[0], 60)).toHaveProperty('kwPerGpu'); + const noLatestPower = buildProfitHistoryResults( + [ + { + date: '2026-06-14', + rows: [ + row(20, 8000, true, 100), + row(80, 3000, true, 100), + row(20, 8000, false), + row(80, 3000, false), + ], + }, + ], + { ...options, powerBasis: 'compare' }, + ); + expect(noLatestPower[0].nearestPoints.every((p) => p.sourceRow?.workflow_run_id === 200)).toBe( + true, + ); + expect(modeledPowerAtTarget(noLatestPower[0], 60)).toEqual({ reason: 'no-measured-power' }); + }); + it('interpolates the selected chip at the target on each date and stamps the date', () => { const rowsByDate = [ { diff --git a/packages/app/src/components/calculator/profit-history.ts b/packages/app/src/components/calculator/profit-history.ts index cf1c821a7..dc4513aa9 100644 --- a/packages/app/src/components/calculator/profit-history.ts +++ b/packages/app/src/components/calculator/profit-history.ts @@ -29,7 +29,7 @@ import { getHardwareConfig, getModelSortIndex, isKnownGpu } from '@/lib/constant import { type Percentile, Sequence } from '@/lib/data-mappings'; import { getDisplayLabel } from '@/lib/utils'; -import { interpolateForGPU } from './interpolation'; +import { interpolateProfitForGPU, type ProfitPowerBasis } from './profit-power'; import type { ProfitEstimatorRow } from './profit-estimator'; import { buildGpuGroups, type GroupMeta } from './throughput-data'; import type { CalculatorMode, CostProvider, InterpolatedResult } from './types'; @@ -168,6 +168,7 @@ export function buildProfitHistoryResults( targetValue: number; mode: CalculatorMode; costProvider: CostProvider; + powerBasis?: ProfitPowerBasis; /** * Per chip, the run its current bar is built from * (`profitHistoryCurrentRunIds`); a pinned run entry for that same run is @@ -183,6 +184,7 @@ export function buildProfitHistoryResults( targetValue, mode, costProvider, + powerBasis = 'provisioned', currentRunIds = {}, } = options; if (selectedGPUs.length === 0) return []; @@ -211,7 +213,7 @@ export function buildProfitHistoryResults( for (const [groupKey, points] of Object.entries(grouped)) { const meta = groupMeta[groupKey]; if (!meta) continue; - const result = interpolateForGPU(points, targetValue, mode, costProvider); + const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis); if (!result || !(result.value > 0)) continue; results.push({ ...result, diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts index 69ca900ae..f297608f3 100644 --- a/packages/app/src/components/calculator/profit-power.test.ts +++ b/packages/app/src/components/calculator/profit-power.test.ts @@ -2,10 +2,16 @@ import { describe, expect, it } from 'vitest'; import type { BenchmarkRow } from '@/lib/api'; import { modelSystemPower } from '@/lib/modeled-system-power'; +import { estimateRackPower } from '@/lib/system-power-model'; import { Percentile, Sequence } from '@/lib/data-mappings'; import { buildGpuGroups, interpolateForGPU } from './useThroughputData'; import { estimateProfitRows } from './profit-estimator'; -import { estimateProfitByPower, modeledPowerAtTarget } from './profit-power'; +import { + estimateProfitByPower, + interpolateProfitForGPU, + modeledPowerAtTarget, + type ProfitPowerSource, +} from './profit-power'; import type { GPUDataPoint, InterpolatedResult } from './types'; // Power telemetry from MI355X Kimi K3 source row 441385; the target/rates below @@ -51,6 +57,8 @@ const point: GPUDataPoint = { throughput: 6000, inputThroughput: 5940, outputThroughput: 60, + inputTokenShare: 0.99, + cacheHitRate: 0.9, concurrency: 16, tp: 8, precision: 'fp4', @@ -95,8 +103,249 @@ const labels = { extrapolated: 'Full-chassis extrapolation', }; +// One GB200 NVL72 compute tray on the AgentX workload: four GPUs on one host, two +// Grace sockets, module sensor present. Watts are controlled inputs, not published +// constants; the CPU-side keys follow the producer contract (InferenceX docs/results-and-ingestion.md). +const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 }; +const traySource: BenchmarkRow = { + ...source, + hardware: 'gb200', + framework: 'sglang', + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + power_audit: { cpu: { sensor_kind: 'module', expected_sockets: 2, observed_sockets: 2 } }, + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 1, + avg_power_w: 900.25, + avg_total_gpu_power_w: 3601, + ...GRACE, + avg_total_module_power_w: 4300.75, + }, +}; +const trayPoint: GPUDataPoint = { ...point, sourceRow: traySource, hwKey: 'gb200_sglang', tp: 4 }; +const trayResult: InterpolatedResult = { + ...result, + hwKey: trayPoint.hwKey, + resultKey: trayPoint.hwKey, + nearestPoints: [trayPoint], +}; +const withPoints = (base: InterpolatedResult, points: GPUDataPoint[]): InterpolatedResult => ({ + ...base, + nearestPoints: points, +}); + describe('profit power basis preview', () => { - it.each([8, 4, 2])( + it('retains the existing Steffen result for the Kimi K3 MI355X valid knots at 45', () => { + const knots = [ + { ...point, interactivity: 14.39677512237259, throughput: 12090.08106 }, + { ...point, interactivity: 53.85029617662897, throughput: 7227.86065 }, + ]; + expect( + interpolateProfitForGPU(knots, 45, 'interactivity_to_throughput', 'costh', 'modeled')?.value, + ).toBeCloseTo(8059.609379677125, 8); + }); + + it('selects power-valid raw knots before the frontier at the unchanged target', () => { + const low = { + ...point, + interactivity: 14.396775, + throughput: 8000, + sourceRow: { ...source, id: 443687, conc: 48 }, + }; + const high = { + ...point, + interactivity: 53.850296, + throughput: 4000, + sourceRow: { ...source, id: 443686, conc: 14 }, + }; + const invalid = { + ...point, + interactivity: 15.951507, + throughput: 9000, + sourceRow: { + ...source, + id: 443692, + conc: 44, + metrics: { ...source.metrics, power_valid: 0 }, + }, + }; + const points = [low, invalid, high]; + const original = interpolateForGPU(points, 45, 'interactivity_to_throughput', 'costh')!; + expect(original.nearestPoints.map((p) => p.sourceRow?.id)).toEqual([443692, 443686]); + expect(modeledPowerAtTarget(original, 45)).toEqual({ reason: 'no-measured-power' }); + const eligible = interpolateForGPU([low, high], 45, 'interactivity_to_throughput', 'costh')!; + for (const basis of ['modeled', 'compare'] as const) { + const selected = interpolateProfitForGPU( + points, + 45, + 'interactivity_to_throughput', + 'costh', + basis, + )!; + expect(selected).toEqual(eligible); + expect(modeledPowerAtTarget(selected, 45)).toHaveProperty('kwPerGpu'); + const output = estimateProfitByPower( + [selected], + specs, + pricing, + assumptions, + basis, + 45, + labels, + ); + expect(output.skipped).toEqual([]); + expect(output.rows).toHaveLength(basis === 'compare' ? 2 : 1); + if (basis === 'compare') { + expect(output.rows[0].revenue / output.rows[0].gpuHours).toBeCloseTo( + output.rows[1].revenue / output.rows[1].gpuHours, + 10, + ); + } + } + expect( + interpolateProfitForGPU(points, 45, 'interactivity_to_throughput', 'costh', 'provisioned'), + ).toEqual(original); + }); + + it('keeps inherited producer rows in the selected logical curve', () => { + const points = [20, 60].map((interactivity, i) => ({ + ...point, + interactivity, + throughput: 8000 - i * 4000, + sourceRow: { + ...source, + date: `2026-09-${10 + i}`, + workflow_run_id: 100 + i, + run_url: `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${100 + i}`, + curve_date: '2026-09-12', + curve_workflow_run_id: 102, + }, + })); + const selected = interpolateProfitForGPU( + points, + 45, + 'interactivity_to_throughput', + 'costh', + 'modeled', + )!; + expect(selected.clamped).toBe(false); + expect(selected.nearestPoints).toEqual(points); + expect(modeledPowerAtTarget(selected, 45)).toHaveProperty('kwPerGpu'); + }); + + it('accepts an exact valid point and preserves provisioned fallback outside the valid range', () => { + const exact = interpolateProfitForGPU( + [point], + 45, + 'interactivity_to_throughput', + 'costh', + 'modeled', + )!; + expect(modeledPowerAtTarget(exact, 45)).toHaveProperty('kwPerGpu'); + const invalid = { + ...point, + interactivity: 60, + throughput: 3000, + sourceRow: { ...source, metrics: { ...source.metrics, power_valid: 0 } }, + }; + const points = [{ ...point, interactivity: 20 }, invalid]; + const original = interpolateForGPU(points, 45, 'interactivity_to_throughput', 'costh')!; + const fallback = interpolateProfitForGPU( + points, + 45, + 'interactivity_to_throughput', + 'costh', + 'compare', + )!; + expect(fallback).toEqual(original); + const estimate = estimateProfitByPower( + [fallback], + specs, + pricing, + assumptions, + 'compare', + 45, + labels, + ); + expect(estimate.rows).toHaveLength(1); + expect(estimate.rows[0].powerLabel).toBe(labels.provisioned); + expect(estimate.skipped[0].reason).toBe('no-measured-power'); + const clamped = interpolateProfitForGPU( + [{ ...point, interactivity: 38 }], + 45, + 'interactivity_to_throughput', + 'costh', + 'modeled', + )!; + expect(modeledPowerAtTarget(clamped, 45)).toEqual({ reason: 'outside-measured-range' }); + const missingCpu = { + ...trayPoint, + sourceRow: { ...traySource, metrics: { ...traySource.metrics, cpu_power_valid: 0 } }, + }; + const excluded = interpolateProfitForGPU( + [missingCpu], + 45, + 'interactivity_to_throughput', + 'costh', + 'modeled', + )!; + expect(modeledPowerAtTarget(excluded, 45)).toEqual({ reason: 'no-cpu-power' }); + }); + + it('partitions the whole frontier by sensor basis before interpolation', () => { + const graceSource: BenchmarkRow = { + ...traySource, + power_audit: { + cpu: { sensor_kind: 'grace_socket', expected_sockets: 2, observed_sockets: 2 }, + }, + metrics: { ...traySource.metrics }, + }; + delete graceSource.metrics.avg_total_module_power_w; + const grace = [30, 60].map((interactivity, i) => ({ + ...trayPoint, + sourceRow: graceSource, + interactivity, + throughput: 9000 - i * 5000, + cacheHitRate: 0.5 + i * 0.2, + inputTokenShare: 0.8 + i * 0.1, + })); + const modules = [10, 20].map((interactivity, i) => ({ + ...trayPoint, + interactivity, + throughput: 12000 - i * 1000, + cacheHitRate: 0.99, + inputTokenShare: 0.99, + })); + const expected = interpolateForGPU(grace, 45, 'interactivity_to_throughput', 'costh'); + expect( + interpolateForGPU([...modules, ...grace], 45, 'interactivity_to_throughput', 'costh')?.value, + ).not.toBe(expected?.value); + for (const points of [[...modules, ...grace], [...grace, ...modules].toReversed()]) { + expect( + interpolateProfitForGPU(points, 45, 'interactivity_to_throughput', 'costh', 'compare'), + ).toEqual(expected); + } + const fasterModules = modules.map((p, i) => ({ + ...p, + interactivity: 30 + i * 30, + throughput: 11000 - i * 5000, + })); + expect( + interpolateProfitForGPU( + [...fasterModules, ...grace], + 45, + 'interactivity_to_throughput', + 'costh', + 'modeled', + ), + ).toEqual(interpolateForGPU(fasterModules, 45, 'interactivity_to_throughput', 'costh')); + }); + + it.each([8])( 'keeps raw power attached through official and run-keyed %i-GPU frontiers', (gpus) => { const row = { @@ -133,28 +382,167 @@ describe('profit power basis preview', () => { }, ); - it('distinguishes partial chassis from unsupported rack hardware with valid telemetry', () => { - for (const hardware of ['gb200', 'gb300']) { - expect( - modeledPowerAtTarget( - { ...result, nearestPoints: [{ ...point, sourceRow: { ...source, hardware } }] }, - 45, - ), - ).toEqual({ reason: 'unsupported-power-hardware' }); - } - const partial = { - ...source, + it('accepts fully measured NVL72 trays and records the measured basis behind the estimate', () => { + // 1.1 × the tray's amortised facility watts per GPU from the pinned GB200 rack profile. + const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!; + const power = modeledPowerAtTarget(trayResult, 45); + if (!('kwPerGpu' in power)) throw new Error(power.reason); + const kw = power.kwPerGpu; + expect(kw).toBeCloseTo((rack.facilityWatts / rack.gpuCount / 1000) * 1.1, 8); + expect(kw).toBeCloseTo(1.6077325, 7); + const output = estimateProfitByPower( + [trayResult], + specs, + pricing, + assumptions, + 'compare', + 45, + labels, + ); + expect(output.skipped).toEqual([]); + const [provisioned, modeled] = output.rows; + expect(provisioned.powerSource).toBeUndefined(); + expect(modeled.powerSource).toEqual({ + topology: 'nvl72-trays', + measuredBasis: 'module', + sensorKind: 'module', + pue: 1.1, + modelPath: 'packages/app/src/lib/system-power-model.ts', + modelRevision: rack.modelRevision, + profileSha256: expect.stringMatching(/^[0-9a-f]{64}$/u), + } satisfies ProfitPowerSource); + expect(modeled.gpuHours / provisioned.gpuHours).toBeCloseTo(2.09 / kw, 10); + + // Two fully measured GB300 trays on distinct hosts, Grace-socket basis. + const twoTrays: BenchmarkRow = { + ...traySource, + hardware: 'gb300', + power_audit: { + cpu: { sensor_kind: 'grace_socket', expected_sockets: 4, observed_sockets: 4 }, + }, + disagg: true, + is_multinode: true, prefill_tp: 4, decode_tp: 4, - metrics: { ...source.metrics, avg_power_w: 500, avg_total_gpu_power_w: 2000 }, + metrics: { + ...traySource.metrics, + avg_power_w: 900, + avg_total_gpu_power_w: 7200, + prefill_avg_power_w: 950, + decode_avg_power_w: 850, + avg_cpu_socket_power_w: 260, + avg_total_cpu_power_w: 1040, + }, + workers: [ + { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 950 }, + { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 850 }, + ], }; - expect(modelSystemPower(partial, undefined, true)).toMatchObject({ + delete (twoTrays.metrics as Record).avg_total_module_power_w; + const estimate = modelSystemPower(twoTrays, undefined, true); + expect(estimate).toMatchObject({ status: 'supported', chassisBasis: 'full', gpuCount: 8 }); + const rows = estimateProfitByPower( + [withPoints(trayResult, [{ ...trayPoint, sourceRow: twoTrays }])], + specs, + pricing, + assumptions, + 'modeled', + 45, + labels, + ).rows; + expect(rows).toHaveLength(1); + expect(rows[0].powerSource).toMatchObject({ + topology: 'nvl72-trays', + measuredBasis: 'gpu-plus-grace', + sensorKind: 'grace-socket', + pue: 1.1, + }); + }); + + it('rejects partially measured trays and never mixes measured bases between knots', () => { + const partialTray: BenchmarkRow = { + ...traySource, + prefill_tp: 2, + decode_tp: 2, + metrics: { + ...traySource.metrics, + avg_total_gpu_power_w: 1800.5, + avg_total_module_power_w: 4300.75, + }, + }; + expect(modelSystemPower(partialTray, undefined, true)).toMatchObject({ status: 'supported', chassisBasis: 'extrapolated', }); expect( - modeledPowerAtTarget({ ...result, nearestPoints: [{ ...point, sourceRow: partial }] }, 45), - ).toMatchObject({ extrapolated: true }); + modeledPowerAtTarget(withPoints(trayResult, [{ ...trayPoint, sourceRow: partialTray }]), 45), + ).toEqual({ reason: 'unsupported-power-topology' }); + + const graceOnly: BenchmarkRow = { + ...traySource, + metrics: { ...traySource.metrics }, + power_audit: { + cpu: { sensor_kind: 'grace_socket', expected_sockets: 2, observed_sockets: 2 }, + }, + }; + delete (graceOnly.metrics as Record).avg_total_module_power_w; + const mixed = withPoints(trayResult, [ + { ...trayPoint, interactivity: 30 }, + { ...trayPoint, interactivity: 60, sourceRow: graceOnly }, + ]); + expect(modeledPowerAtTarget(mixed, 30)).toHaveProperty('kwPerGpu'); + expect(modeledPowerAtTarget(mixed, 60)).toHaveProperty('kwPerGpu'); + expect(modeledPowerAtTarget(mixed, 45)).toEqual({ reason: 'incompatible-power-basis' }); + const same = withPoints(trayResult, [ + { ...trayPoint, interactivity: 30 }, + { ...trayPoint, interactivity: 60 }, + ]); + expect(modeledPowerAtTarget(same, 45)).toEqual(modeledPowerAtTarget(trayResult, 45)); + }); + + it('plans aggregate multinode NVL72 rows without a worker array as inferred trays', () => { + // Kimi K3 GB200 dynamo-vLLM TP16: sixteen GPUs on four trays, no per-worker array. + const inferred: BenchmarkRow = { + ...traySource, + power_audit: { cpu: { sensor_kind: 'module', expected_sockets: 8, observed_sockets: 8 } }, + is_multinode: true, + prefill_tp: 16, + decode_tp: 0, + num_prefill_gpu: 16, + num_decode_gpu: 16, + metrics: { + ...traySource.metrics, + avg_power_w: 441.741, + avg_total_gpu_power_w: 7067.859, + avg_total_cpu_power_w: 2004, + avg_total_module_power_w: 9071.859, + }, + }; + const rack = estimateRackPower( + 'gb200', + { basis: 'module', moduleWattsPerTray: 9071.859 / 4 }, + 1.1, + )!; + const inferredResult = withPoints(trayResult, [{ ...trayPoint, sourceRow: inferred }]); + expect(modeledPowerAtTarget(inferredResult, 45)).toMatchObject({ + kwPerGpu: expect.closeTo((((rack.facilityWatts / 18) * 4) / 16 / 1000) * 1.1, 8), + }); + const rows = estimateProfitByPower( + [inferredResult], + specs, + pricing, + assumptions, + 'modeled', + 45, + labels, + ).rows; + expect(rows).toHaveLength(1); + expect(rows[0].powerSource).toMatchObject({ + topology: 'nvl72-trays', + measuredBasis: 'module', + sensorKind: 'module', + pue: 1.1, + }); }); it('leaves the default estimator and default AgentX model gate unchanged', () => { @@ -164,7 +552,7 @@ describe('profit power basis preview', () => { expect(modelSystemPower(source)).toMatchObject({ status: 'unsupported', reason: 'workload' }); }); - it.each([2, 4])( + it.each([2])( 'prices a validated %i-GPU allocation as a labeled full-chassis extrapolation', (gpus) => { const partial = { @@ -240,6 +628,12 @@ describe('profit power basis preview', () => { ); expect(output.skipped).toEqual([]); const [baseline, modeled] = output.rows; + expect(baseline.powerSource).toBeUndefined(); + expect(modeled.powerSource).toMatchObject({ + topology: 'chassis', + pue: 1.3, + modelPath: 'packages/app/src/lib/system-power-model.ts', + }); expect(modeled.revenuePerGpuHour).toBe(baseline.revenuePerGpuHour); const ratio = 2.09 / 1.5976675; for (const field of ['gpuHours', 'revenue', 'tco', 'labCut', 'profit'] as const) { @@ -260,6 +654,12 @@ describe('profit power basis preview', () => { expect(modeledPowerAtTarget(bracket, 45)).toEqual({ reason: 'no-measured-power' }); expect( estimateProfitByPower([bracket], specs, pricing, assumptions, 'compare', 45, labels), + ).toMatchObject({ + rows: [{ powerLabel: labels.provisioned }], + skipped: [{ reason: 'no-measured-power' }], + }); + expect( + estimateProfitByPower([bracket], specs, pricing, assumptions, 'modeled', 45, labels), ).toMatchObject({ rows: [], skipped: [{ reason: 'no-measured-power' }] }); expect(modeledPowerAtTarget({ ...result, nearestPoints: [point, missing] }, 45)).toMatchObject({ kwPerGpu: expect.closeTo(1.5976675, 8), @@ -314,7 +714,7 @@ describe('profit power basis preview', () => { const high = modeledPowerAtTarget(bracket, 60); if (!('kwPerGpu' in low) || !('kwPerGpu' in high)) throw new Error('Expected valid power knots'); - expect(modeledPowerAtTarget(bracket, 40)).toEqual({ + expect(modeledPowerAtTarget(bracket, 40)).toMatchObject({ kwPerGpu: low.kwPerGpu + (high.kwPerGpu - low.kwPerGpu) / 3, extrapolated: true, }); @@ -331,10 +731,6 @@ describe('profit power basis preview', () => { }, 'unsupported-power-topology', ], - [ - { metrics: { ...source.metrics, avg_power_w: 500, avg_total_gpu_power_w: 2000 } }, - 'no-measured-power', - ], [ { metrics: { power_valid: 1, avg_power_w: 796.131, avg_total_gpu_power_w: 6369.045 } }, 'no-measured-power', diff --git a/packages/app/src/components/calculator/profit-power.ts b/packages/app/src/components/calculator/profit-power.ts index 1c6bc2a1b..bbaa1e01c 100644 --- a/packages/app/src/components/calculator/profit-power.ts +++ b/packages/app/src/components/calculator/profit-power.ts @@ -1,4 +1,5 @@ -import { modelSystemPower } from '@/lib/modeled-system-power'; +import { modelSystemPower, type SystemPowerSensorKind } from '@/lib/modeled-system-power'; +import { type RackMeasuredBasis, systemPowerSourceSha256 } from '@/lib/system-power-model'; import type { TokenRevenuePricing } from '@/components/inference/types'; import { @@ -10,13 +11,51 @@ import { type ProfitEstimatorSkipReason, type ProfitEstimatorSpecs, } from './profit-estimator'; -import type { GPUDataPoint, InterpolatedResult } from './types'; +import { interpolateForGPU } from './interpolation'; +import type { CalculatorMode, CostProvider, GPUDataPoint, InterpolatedResult } from './types'; export type ProfitPowerBasis = 'provisioned' | 'modeled' | 'compare'; -type PlanningPower = - | { kwPerGpu: number; extrapolated: boolean } - | { reason: ProfitEstimatorSkipReason }; +interface ProfitPowerProfile { + /** Facility PUE the estimate applied once at the chassis or rack AC boundary. */ + pue: number; + modelPath: string; + modelRevision: string; + /** Equation-file hash; modelRevision also covers parameters and admission/PUE policy. */ + profileSha256: string | null; +} + +/** + * What the measured + modeled budget was measured on. An eight-GPU x86 chassis + * measures the GPU boards and models the rest; an NVL72 tray measures the compute + * module (or GPU board + Grace socket) and models only the rack residual. + */ +export type ProfitPowerSource = + | (ProfitPowerProfile & { topology: 'chassis' }) + | (ProfitPowerProfile & { + topology: 'nvl72-trays'; + measuredBasis: RackMeasuredBasis; + sensorKind: SystemPowerSensorKind; + }); + +export interface ProfitPlanningPower { + kwPerGpu: number; + source: ProfitPowerSource; + extrapolated: boolean; +} + +/** Identity of a source for deduplicating notes and refusing to mix bases between knots. */ +export function powerSourceKey(source: ProfitPowerSource): string { + return [ + source.topology, + source.pue, + source.modelPath, + source.modelRevision, + source.topology === 'nvl72-trays' ? `${source.measuredBasis}/${source.sensorKind}` : '', + ].join('|'); +} + +type PlanningPower = ProfitPlanningPower | { reason: ProfitEstimatorSkipReason }; function planningPower(point: GPUDataPoint): PlanningPower { const row = point.sourceRow; @@ -31,6 +70,9 @@ function planningPower(point: GPUDataPoint): PlanningPower { case 'role-power': { return { reason: 'unsupported-power-topology' }; } + case 'cpu-telemetry': { + return { reason: 'no-cpu-power' }; + } case 'telemetry': case 'gpu-count': { return { reason: 'no-measured-power' }; @@ -47,13 +89,59 @@ function planningPower(point: GPUDataPoint): PlanningPower { (estimate.topologyBasis !== 'single-node' || 8 % estimate.gpuCount !== 0) ) return { reason: 'unsupported-power-topology' }; + const profile: ProfitPowerProfile = { + pue: estimate.pue, + modelPath: estimate.modelPath, + modelRevision: estimate.modelRevision, + profileSha256: systemPowerSourceSha256(estimate.modelPath), + }; + const source: ProfitPowerSource = + estimate.topologyBasis === 'nvl72-trays' + ? { + ...profile, + topology: 'nvl72-trays', + measuredBasis: estimate.measuredBasis, + sensorKind: estimate.sensorKind, + } + : { ...profile, topology: 'chassis' }; return { kwPerGpu: (estimate.deploymentFacilityWatts / estimate.gpuCount / 1000) * 1.1, extrapolated: estimate.chassisBasis === 'extrapolated', + source, }; } -/** Reusing the original frontier prevents the power choice from changing throughput. */ +/** Select a compatible power-valid frontier within the caller's selected curve. */ +export function interpolateProfitForGPU( + points: GPUDataPoint[], + target: number, + mode: CalculatorMode, + costProvider: CostProvider, + powerBasis: ProfitPowerBasis, +): InterpolatedResult | null { + const original = interpolateForGPU(points, target, mode, costProvider); + if (powerBasis === 'provisioned' || mode !== 'interactivity_to_throughput') return original; + const cohorts = new Map(); + for (const point of points) { + const power = planningPower(point); + if ('reason' in power) continue; + const key = powerSourceKey(power.source); + const cohort = cohorts.get(key) ?? []; + cohort.push(point); + cohorts.set(key, cohort); + } + let best: InterpolatedResult | null = null; + // Source ordering makes equal-throughput selection independent of input order. + for (const key of [...cohorts.keys()].toSorted()) { + const candidate = interpolateForGPU(cohorts.get(key)!, target, mode, costProvider); + if (!candidate || 'reason' in modeledPowerAtTarget(candidate, target)) continue; + if (!best || candidate.value > best.value) best = candidate; + } + // Preserve provisioned-only fallback and the existing unavailability reason. + return best ?? original; +} + +/** Read power from the same selected knots as throughput, without extrapolation. */ export function modeledPowerAtTarget(result: InterpolatedResult, target: number): PlanningPower { if (result.clamped) return { reason: 'outside-measured-range' }; const exact = result.nearestPoints.find((p) => Math.abs(p.interactivity - target) < 1e-9); @@ -71,12 +159,15 @@ export function modeledPowerAtTarget(result: InterpolatedResult, target: number) upper = planningPower(right); if ('reason' in lower) return lower; if ('reason' in upper) return upper; + if (powerSourceKey(lower.source) !== powerSourceKey(upper.source)) + return { reason: 'incompatible-power-basis' }; return { kwPerGpu: lower.kwPerGpu + ((upper.kwPerGpu - lower.kwPerGpu) * (target - left.interactivity)) / (right.interactivity - left.interactivity), extrapolated: lower.extrapolated || upper.extrapolated, + source: lower.source, }; } @@ -100,6 +191,13 @@ export function estimateProfitByPower( output.skipped.push(baseline); continue; } + if (powerBasis === 'compare') { + output.rows.push({ + ...baseline, + resultKey: `${baseline.resultKey}__provisioned`, + powerLabel: labels.provisioned, + }); + } const power = modeledPowerAtTarget(result, target); if ('reason' in power) { output.skipped.push({ @@ -111,31 +209,25 @@ export function estimateProfitByPower( }); continue; } - const modeled = estimateSkuProfit( + const estimated = estimateSkuProfit( result, { ...specs, powerKwPerGpu: power.kwPerGpu }, pricing, assumptions, ); - if (!isProfitEstimatorRow(modeled)) { - output.skipped.push(modeled); + if (!isProfitEstimatorRow(estimated)) { + output.skipped.push(estimated); continue; } + const modeled = { ...estimated, powerSource: power.source }; if (powerBasis === 'compare') { - output.rows.push( - { - ...baseline, - resultKey: `${baseline.resultKey}__provisioned`, - powerLabel: labels.provisioned, - }, - { - ...modeled, - resultKey: `${modeled.resultKey}__modeled`, - powerLabel: power.extrapolated - ? `${labels.modeled} · ${labels.extrapolated}` - : labels.modeled, - }, - ); + output.rows.push({ + ...modeled, + resultKey: `${modeled.resultKey}__modeled`, + powerLabel: power.extrapolated + ? `${labels.modeled} · ${labels.extrapolated}` + : labels.modeled, + }); } else output.rows.push( power.extrapolated ? { ...modeled, powerLabel: labels.extrapolated } : modeled, diff --git a/packages/app/src/components/calculator/useThroughputData.ts b/packages/app/src/components/calculator/useThroughputData.ts index 0f87a2afe..41dfaeab5 100644 --- a/packages/app/src/components/calculator/useThroughputData.ts +++ b/packages/app/src/components/calculator/useThroughputData.ts @@ -24,6 +24,7 @@ import { sign, } from './interpolation'; import type { CostProvider, CostType, GPUDataPoint, InterpolatedResult } from './types'; +import { interpolateProfitForGPU, type ProfitPowerBasis } from './profit-power'; // Re-export pure functions so existing imports from this module keep working. export { @@ -261,6 +262,7 @@ export function useThroughputData( costProvider: CostProvider, visibleHwKeys?: Set, hideSkuAboveConfigLimit = false, + powerBasis: ProfitPowerBasis = 'provisioned', ): InterpolatedResult[] => { const results: InterpolatedResult[] = []; @@ -270,7 +272,7 @@ export function useThroughputData( // Skip GPUs that are not visible (legend filters by hwKey) if (visibleHwKeys && !visibleHwKeys.has(hwKey)) continue; - const result = interpolateForGPU(points, targetValue, mode, costProvider); + const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis); if (result && result.value > 0 && !(hideSkuAboveConfigLimit && result.clampedAbove)) { results.push({ ...result, @@ -305,6 +307,7 @@ export function useThroughputData( visibleHwKeys?: Set, runInfoByIndex?: Record, hideSkuAboveConfigLimit = false, + powerBasis: ProfitPowerBasis = 'provisioned', ): InterpolatedResult[] => { const results: InterpolatedResult[] = []; @@ -313,7 +316,7 @@ export function useThroughputData( if (!meta) continue; if (visibleHwKeys && !visibleHwKeys.has(meta.hwKey)) continue; - const result = interpolateForGPU(points, targetValue, mode, costProvider); + const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis); if (result && result.value > 0 && !(hideSkuAboveConfigLimit && result.clampedAbove)) { results.push({ ...result, diff --git a/packages/app/src/components/inference/InferenceContext.tsx b/packages/app/src/components/inference/InferenceContext.tsx index d29bc5759..38759dbea 100644 --- a/packages/app/src/components/inference/InferenceContext.tsx +++ b/packages/app/src/components/inference/InferenceContext.tsx @@ -1245,6 +1245,14 @@ export function InferenceProvider({ const wantedType = selectedXAxisMode === 'interactivity' ? 'interactivity' : 'e2e'; const graph = graphs.find((candidate) => candidate.chartDefinition.chartType === wantedType); if (!graph) return hwTypesWithData; + // All in Measured can have GPU-valid table rows without a numeric estimate to rank. + const unrankedHwTypes = graph.tableData + ? new Set( + graph.tableData + .filter((point) => effectivePrecisions.includes(point.precision)) + .map(extractHwKey), + ) + : hwTypesWithData; const direction = graph.chartDefinition[ `${selectedYAxisMetric}_roofline` as keyof typeof graph.chartDefinition @@ -1255,11 +1263,19 @@ export function InferenceProvider({ direction !== 'lower_left' && direction !== 'lower_right' ) { - return hwTypesWithData; + return unrankedHwTypes; } const best = bestSeriesPerSku(graph.data, direction); - return best.size > 0 ? best : hwTypesWithData; - }, [graphs, hwTypesWithData, selectedXAxisMode, selectedYAxisMetric]); + if (best.size > 0) return best; + return unrankedHwTypes; + }, [ + graphs, + hwTypesWithData, + selectedXAxisMode, + selectedYAxisMetric, + effectivePrecisions, + extractHwKey, + ]); const setBestPerSkuAndApply = useCallback( (enabled: boolean, options?: { applySelection?: boolean }) => { diff --git a/packages/app/src/components/inference/hooks/useChartData.ts b/packages/app/src/components/inference/hooks/useChartData.ts index 3401a6002..281c50b75 100644 --- a/packages/app/src/components/inference/hooks/useChartData.ts +++ b/packages/app/src/components/inference/hooks/useChartData.ts @@ -16,8 +16,10 @@ import { rowToSequence } from '@semianalysisai/inferencex-constants'; import { useQueries, useQuery } from '@tanstack/react-query'; import chartDefinitions, { + isAllInMeasuredConfigKey, tokenMetricTypeForConfigKey, } from '@/components/inference/metric-registry'; +import { allInMeasuredTableData } from '@/components/inference/utils/inference-table-data'; import { applyTokenRevenuePricing, usesTokenSalePricing, @@ -597,6 +599,15 @@ export function useChartData( chartDefinition, data: processedData, clippedData, + ...(isAllInMeasuredConfigKey(selectedYAxisMetric) + ? { + tableData: expandPowerCompareSeries( + allInMeasuredTableData(filteredData, metricKey, xAxisField), + selectedYAxisMetric, + powerCompare, + ), + } + : {}), }; }, ); diff --git a/packages/app/src/components/inference/types.ts b/packages/app/src/components/inference/types.ts index 33a491b16..296c5d140 100644 --- a/packages/app/src/components/inference/types.ts +++ b/packages/app/src/components/inference/types.ts @@ -494,6 +494,8 @@ export interface RenderableGraph { sequence: string; chartDefinition: ChartDefinition; data: InferenceData[]; + /** All GPU-valid rows for the All in Measured table, including unavailable estimates. */ + tableData?: InferenceData[]; clippedData?: ClippedInferenceData[]; } /** @@ -511,6 +513,7 @@ export interface RenderableGraph { export interface OverlayData { /** The data points to overlay */ data: InferenceData[]; + tableData?: InferenceData[]; /** Overlay points hidden by the same display limits as official data. */ clippedData?: ClippedInferenceData[]; /** Hardware configuration for the overlay data (may have different hardware types) */ diff --git a/packages/app/src/components/inference/ui/ChartDisplay.tsx b/packages/app/src/components/inference/ui/ChartDisplay.tsx index 768d87c48..544c188dc 100644 --- a/packages/app/src/components/inference/ui/ChartDisplay.tsx +++ b/packages/app/src/components/inference/ui/ChartDisplay.tsx @@ -14,9 +14,13 @@ import chartDefinitions, { } from '@/components/inference/metric-registry'; import { metricRowLabel } from '@/components/inference/axis-metric-explanations'; import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config'; -import { AIR_COOLED_SYSTEM_PUE } from '@/lib/modeled-system-power'; +import { AIR_COOLED_SYSTEM_PUE, DLC_SYSTEM_PUE } from '@/lib/modeled-system-power'; import { SYSTEM_POWER_MODEL_REVISION } from '@/lib/system-power-model'; -import { ALL_IN_MEASURED_EMPTY, ALL_IN_MEASURED_NOTE } from '@/lib/power-basis'; +import { + ALL_IN_MEASURED_AGENTIC_NOTE, + ALL_IN_MEASURED_EMPTY, + ALL_IN_MEASURED_NOTE, +} from '@/lib/power-basis'; import { applyTokenRevenuePricing, cachedInputPricePerMillion, @@ -120,6 +124,8 @@ import WorkflowInfoDisplay from './WorkflowInfoDisplay'; type InferenceViewMode = 'chart' | 'table'; +const modelRevisionLabel = SYSTEM_POWER_MODEL_REVISION.replace(/^app-sha256:/u, '').slice(0, 12); + const STRINGS = { en: { inferencePerformance: 'Inference Performance', @@ -137,14 +143,14 @@ const STRINGS = { e2eNormIntvtyDisclaimer: 'E2E Normalized Interactivity requires persisted per-request traces, so unofficial-run overlays are unavailable for this experimental view.', systemPowerAssumptions: - '8k1k estimate from validated GPU telemetry · CPU/DRAM utilization 20% · Eight-GPU chassis models; a partially allocated chassis is extrapolated to a full chassis at the measured per-GPU power. Chassis AC includes platform overheads; PUE is applied separately for facility power. Click a point for measured GPU power, topology, and power model provenance. Unsupported inputs are omitted.', + '8K / 1K estimates from validated telemetry. Eight-GPU chassis assume 20% CPU/DRAM utilization; partial allocations are extrapolated to a full chassis at the measured per-GPU power. NVL72 uses measured Grace or module power plus modeled rack overhead. Facility power applies PUE once: 1.3 for air-cooled chassis, 1.1 for NVL72. Click a point for measurement and model provenance. Unsupported inputs are omitted.', completedSequenceLengths: (count: string) => `Completed requests across all resident points (n=${count})`, viewMode: 'View mode', noChartData: 'No benchmark data matches the current model, scenario, and filter selection. Adjust the filters above to see results.', noSystemPowerData: - 'No system-power estimates are available for this selection. Choose 8K / 1K with validated GPU telemetry, supported hardware, and known eight-GPU chassis placement. Measured GPU power remains available separately where telemetry exists.', + 'No system-power estimates are available for this selection. Choose 8K / 1K with validated GPU telemetry and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Measured GPU power remains available separately where telemetry exists.', // Boundary disclosures for the derived power axes (lib/power-basis.ts). // Formulas in words; constants named so a screenshot records its method. powerBasisAssumptions: { @@ -152,7 +158,7 @@ const STRINGS = { 'GPU Level Provisioned (TDP) · Watts are the rated TDP per GPU from the hardware registry, so the power curve is flat per hardware. Joules per output token = TDP × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together. Hardware without a published TDP is omitted.', 'utility-provisioned': 'All in Provisioned · Watts are the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model), so the power curve is flat per hardware. Joules per output token = all-in W × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together, unlike the ungated All-in Provisioned J per Output Token, which divides per decode GPU.', - 'utility-modeled': `All in Measured · Measured GPU power carried through the modeled chassis (CPU, DRAM, platform, PSU losses) to the utility meter: modeled chassis AC × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled, applied once), divided by the measured GPUs; joules per output token scale measured joules by the same ratio. Chassis power model revision ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}. Available for 8K / 1K with validated telemetry on supported hardware only; NVL72 systems (GB200, GB300) and points without values are omitted.`, + 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K and AgentX on supported hardware; unavailable estimates stay in the table with their measured GPU power and reason.`, }, vsTtft: (word: string) => `vs. ${word} Time To First Token`, vsE2eLatency: (pctl?: string) => @@ -174,18 +180,18 @@ const STRINGS = { e2eNormIntvtyDisclaimer: '端到端归一化交互性需要持久化的逐请求 trace 数据,因此该实验性视图不支持非官方运行覆盖。', systemPowerAssumptions: - '基于已验证 GPU 遥测的 8k1k 估算 · CPU/DRAM 利用率 20% · 采用八卡机箱模型;仅使用部分 GPU 的机箱按实测每卡功耗外推至满机箱。机箱交流功耗包含平台开销;数据中心功耗另行应用 PUE。点击数据点可查看 GPU 实测功耗、拓扑和功耗模型来源。不支持的输入不绘制。', + '基于已验证遥测的 8K / 1K 估算。八卡机箱假设 CPU/DRAM 利用率为 20%;仅使用部分 GPU 时,按实测每卡功耗外推至满机箱。NVL72 使用实测 Grace 或 module 功耗,加上模型估算的机架开销。数据中心功耗只应用一次 PUE:风冷机箱为 1.3,NVL72 为 1.1。点击数据点可查看测量与模型来源。不支持的输入不绘制。', completedSequenceLengths: (count: string) => `当前所有数据点的已完成请求(n=${count})`, viewMode: '视图模式', noChartData: '当前模型、场景与筛选条件下没有匹配的基准测试数据。请调整上方筛选条件查看结果。', noSystemPowerData: - '当前选择没有可用的系统功耗估算。请选择 8K / 1K 场景;估算仅覆盖 GPU 遥测已验证、硬件受支持、八卡机箱位置已知的运行。存在遥测数据时,仍可单独查看 GPU 实测功耗。', + '当前选择没有可用的系统功耗估算。请选择 8K / 1K 场景,并确保 GPU 遥测已验证、机箱或机架功耗模型受支持。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。存在遥测数据时,仍可单独查看 GPU 实测功耗。', powerBasisAssumptions: { 'gpu-provisioned': 'GPU 额定功耗(TDP)· 功率取硬件注册表中每 GPU 的额定 TDP,因此每种硬件的功率曲线为水平线。每输出 token 能耗 = TDP × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入。未公布 TDP 的硬件不绘制。', 'utility-provisioned': '整体预配功耗 · 功率取硬件注册表中每 GPU 的全电源配置(all-in)市电功率(来源:SemiAnalysis Datacenter Industry Model),因此每种硬件的功率曲线为水平线。每输出 token 能耗 = all-in 功率 × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入,这与未加门控的“每输出 token 全电源配置能耗”按 decode GPU 计算不同。', - 'utility-modeled': `整体实测功耗 · 将 GPU 实测功耗经机箱功耗模型(CPU、DRAM、平台开销、PSU 损耗)推算至市电侧:机箱交流功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷,仅应用一次),再除以实测 GPU 数;每输出 token 能耗按同一比例放大实测能耗。机箱功耗模型版本 ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}。仅适用于 8K / 1K、遥测已验证且硬件受支持的运行;NVL72 系统(GB200、GB300)及缺少数值的数据点不绘制。`, + 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。适用于受支持硬件的 8K / 1K 和 AgentX 场景。估算不可用的数据点仍保留在表格中,并显示实测 GPU 功耗和不可用原因。`, }, vsTtft: (word: string) => `vs. ${word === 'Median' ? '中位' : word} 首 token 延迟(TTFT)`, vsE2eLatency: (pctl?: string) => (pctl ? `vs. ${pctl} 端到端延迟` : 'vs. 端到端延迟'), @@ -497,7 +503,11 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean let overlayPoints = processed.data; let clippedOverlayPoints = processed.clippedData; + let tableOverlayPoints = processed.tableData; if (compareGpuPair?.length === 2) { + tableOverlayPoints = tableOverlayPoints?.filter((p) => + hardwareKeyMatchesAnyBase(String(p.hwKey), compareGpuPair), + ); overlayPoints = overlayPoints.filter((p) => hardwareKeyMatchesAnyBase(String(p.hwKey), compareGpuPair), ); @@ -506,10 +516,15 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean ); } - if (overlayPoints.length === 0 && clippedOverlayPoints.length === 0) return null; + if ( + overlayPoints.length === 0 && + clippedOverlayPoints.length === 0 && + !tableOverlayPoints?.length + ) + return null; const keySet = new Set([ - ...overlayPoints.map((p) => String(p.hwKey)), + ...(tableOverlayPoints ?? overlayPoints).map((p) => String(p.hwKey)), ...clippedOverlayPoints.map(({ point }) => String(point.hwKey)), ]); const hardwareConfigFiltered = Object.fromEntries( @@ -519,6 +534,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean return { data: overlayPoints, clippedData: clippedOverlayPoints, + tableData: tableOverlayPoints, hardwareConfig: hardwareConfigFiltered, label: unofficialRunInfo.branch, runUrl: unofficialRunInfo.url, @@ -553,7 +569,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const eligibleKeys = new Set(); for (const overlay of [overlayDataByChartType.e2e, overlayDataByChartType.interactivity]) { const points = [ - ...(overlay?.data ?? []), + ...(overlay?.tableData ?? overlay?.data ?? []), ...(overlay?.clippedData ?? []).map((entry) => entry.point), ]; for (const point of points) { @@ -571,7 +587,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const officialScope = useMemo(() => { const eligibleKeys = new Set(); for (const graph of graphs) { - const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)]; + const points = graph.tableData ?? [ + ...graph.data, + ...(graph.clippedData ?? []).map((entry) => entry.point), + ]; for (const point of points) { if ( selectedPrecisions.includes(point.precision) && @@ -730,6 +749,8 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean const effectiveGraphs = useMemo(() => { if (graphs.length > 0) return graphs; const hasOverlay = + (overlayDataByChartType.e2e?.tableData?.length ?? 0) > 0 || + (overlayDataByChartType.interactivity?.tableData?.length ?? 0) > 0 || (overlayDataByChartType.e2e?.data.length ?? 0) > 0 || (overlayDataByChartType.e2e?.clippedData?.length ?? 0) > 0 || (overlayDataByChartType.interactivity?.data.length ?? 0) > 0 || @@ -741,6 +762,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean chartDefinition, data: [] as InferenceData[], clippedData: [], + tableData: undefined, })); }, [graphs, overlayDataByChartType, selectedModel, selectedSequence]); @@ -755,7 +777,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean if (!isAgenticSequence) return [] as number[]; const ids = new Set(); for (const graph of visibleGraphs) { - const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)]; + const points = graph.tableData ?? [ + ...graph.data, + ...(graph.clippedData ?? []).map((entry) => entry.point), + ]; for (const point of points) { if ( selectedPrecisions.includes(point.precision) && @@ -785,7 +810,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean if (!useDerivedXAxis) return [] as number[]; const ids = new Set(); for (const graph of visibleGraphs) { - const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)]; + const points = graph.tableData ?? [ + ...graph.data, + ...(graph.clippedData ?? []).map((entry) => entry.point), + ]; for (const point of points) { if (point.benchmark_type === 'agentic_traces' && isPersistedBenchmarkId(point.id)) { ids.add(point.id); @@ -809,7 +837,12 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean // Legacy AgentX axes can still render transient/non-persisted rows, which // have no ids to request. if (!derivedSpec && derivedTargetIds.length === 0) return visibleGraphs; - return visibleGraphs.map((graph) => ({ ...graph, data: [], clippedData: [] })); + return visibleGraphs.map((graph) => ({ + ...graph, + data: [], + clippedData: [], + tableData: graph.tableData ? [] : undefined, + })); } return visibleGraphs.map((graph) => { const rooflineKey = `${selectedYAxisMetric}_roofline` as keyof typeof graph.chartDefinition; @@ -838,7 +871,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean }) .filter((entry): entry is NonNullable => entry !== null); - if (!derivedSpec) return { ...graph, data, clippedData }; + const tableData = graph.tableData?.map( + (point) => preparePoint(point) ?? { ...point, x: NaN }, + ); + if (!derivedSpec) return { ...graph, data, clippedData, tableData }; const chartDefinition = { ...graph.chartDefinition, @@ -848,7 +884,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean y_latency_limit: undefined, ...(derivedCorner ? { [rooflineKey]: derivedCorner } : {}), }; - return { ...graph, chartDefinition, data, clippedData }; + return { ...graph, chartDefinition, data, clippedData, tableData }; }); }, [ isAgenticSequence, @@ -1007,12 +1043,29 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean graph.chartDefinition.chartType, overlayDataByChartType, ); + const tableMode = getViewMode(graphIndex) === 'table'; + const exportData = tableMode + ? (graph.tableData ?? [ + ...graph.data, + ...(graph.clippedData ?? []).map((entry) => entry.point), + ]) + : graph.data; + const exportOverlay = + tableMode && overlay + ? { + ...overlay, + data: overlay.tableData ?? [ + ...overlay.data, + ...(overlay.clippedData ?? []).map((entry) => entry.point), + ], + } + : overlay; const { officialRows: visibleData, overlayRows: visibleOverlayRowsForExport, } = isGpuComparison - ? visibleDateComparisonRows(graph.data, overlay) - : visibleComparisonRows(graph.data, overlay); + ? visibleDateComparisonRows(exportData, exportOverlay) + : visibleComparisonRows(exportData, exportOverlay); const { headers, rows } = inferenceChartToCsv( visibleData, graph.model, @@ -1218,6 +1271,17 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean {ALL_IN_MEASURED_NOTE[locale]}

)} + {isAgenticSequence && + selectedPowerBasis && + (selectedPowerBasis === 'utility-modeled' || + powerCompare === 'boundaries') && ( +

+ {ALL_IN_MEASURED_AGENTIC_NOTE[locale]} +

+ )} {isUnofficialRun && selectedXAxisMode === 'e2e-normalized-interactivity' && (

@@ -1238,14 +1302,14 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean // must not silently remove measured rows from the table. // Restore both official and unofficial clipped points before // applying the shared precision, quick-filter, and legend gates. - const tableOfficialData = [ + const tableOfficialData = graph.tableData ?? [ ...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point), ]; const tableOverlay = overlay ? { ...overlay, - data: [ + data: overlay.tableData ?? [ ...overlay.data, ...(overlay.clippedData ?? []).map((entry) => entry.point), ], diff --git a/packages/app/src/components/inference/ui/InferenceTable.tsx b/packages/app/src/components/inference/ui/InferenceTable.tsx index 0d4e8ab74..3eb31a667 100644 --- a/packages/app/src/components/inference/ui/InferenceTable.tsx +++ b/packages/app/src/components/inference/ui/InferenceTable.tsx @@ -5,8 +5,15 @@ import { useMemo } from 'react'; import type { ChartDefinition, InferenceData } from '@/components/inference/types'; import { type DataTableColumn, DataTable } from '@/components/ui/data-table'; import { chipCounts } from '@/lib/chip-counts'; -import { getNestedYValue, metricLabel, xAxisLabel } from '@/lib/chart-utils'; -import { isModeledSystemPowerConfigKey } from '@/components/inference/metric-registry'; +import { metricLabel, xAxisLabel } from '@/lib/chart-utils'; +import { + isAllInMeasuredConfigKey, + isModeledSystemPowerConfigKey, +} from '@/components/inference/metric-registry'; +import { + allInMeasuredStatusLabel, + inferenceTableYValue, +} from '@/components/inference/utils/inference-table-data'; import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare'; import { sortRowsByYMetric } from '@/components/inference/ui/inference-table-sort'; import { type Precision, getPrecisionLabel } from '@/lib/data-mappings'; @@ -22,7 +29,8 @@ interface InferenceTableProps { } /** Format a number for table display — picks sensible precision and groups thousands. */ -export function formatInferenceTableNumber(value: number, decimals?: number): string { +export function formatInferenceTableNumber(value: number | null, decimals?: number): string { + if (value === null || !Number.isFinite(value)) return '—'; const fixedDecimals = decimals ?? (Math.abs(value) >= 100 ? 0 : Math.abs(value) >= 1 ? 1 : Math.abs(value) >= 0.01 ? 3 : 4); @@ -48,6 +56,8 @@ export function inferenceTableHeaderLabels( yMetric: metricLabel(chartDefinition, selectedYAxisMetric, locale), xMetric: xAxisLabel(chartDefinition, locale), throughput: locale === 'zh' ? '单芯片吞吐量 (tok/s)' : 'Throughput/Chip (tok/s)', + measuredGpuPower: locale === 'zh' ? 'GPU 实测功耗 (W/芯片)' : 'Measured GPU Power (W/chip)', + estimateStatus: locale === 'zh' ? '整体估算状态' : 'All-in Estimate Status', }; } @@ -59,6 +69,7 @@ export default function InferenceTable({ const locale = useLocale(); const yPath = chartDefinition[selectedYAxisMetric as keyof ChartDefinition] as string | undefined; const showModeledPower = isModeledSystemPowerConfigKey(selectedYAxisMetric); + const showAllInMeasured = isAllInMeasuredConfigKey(selectedYAxisMetric); const headers = useMemo( () => inferenceTableHeaderLabels(chartDefinition, selectedYAxisMetric, locale), [chartDefinition, selectedYAxisMetric, locale], @@ -148,14 +159,33 @@ export default function InferenceTable({ header: headers.yMetric, align: 'right', // Comparison clones keep the source metrics; y holds the plotted role/boundary. - cell: (row) => - formatInferenceTableNumber( - row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath), - ), - sortValue: (row) => (row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath)), - className: 'tabular-nums', + cell: (row) => formatInferenceTableNumber(inferenceTableYValue(row, yPath)), + sortValue: (row) => inferenceTableYValue(row, yPath) ?? '', + className: showAllInMeasured ? 'tabular-nums min-w-36' : 'tabular-nums', importance: 'key', }, + ...(showAllInMeasured + ? [ + { + header: headers.measuredGpuPower, + align: 'right' as const, + cell: (row: InferenceData) => + formatInferenceTableNumber(row.measuredAvgPower?.y ?? null), + sortValue: (row: InferenceData) => row.measuredAvgPower?.y ?? '', + className: 'tabular-nums min-w-32', + importance: 'key' as const, + }, + { + header: headers.estimateStatus, + cell: (row: InferenceData) => + allInMeasuredStatusLabel(row, selectedYAxisMetric.slice(2), locale), + sortValue: (row: InferenceData) => + allInMeasuredStatusLabel(row, selectedYAxisMetric.slice(2), locale), + className: 'w-32 min-w-32 sm:w-auto sm:min-w-48', + importance: 'key' as const, + }, + ] + : []), { header: headers.xMetric, align: 'right', @@ -173,7 +203,15 @@ export default function InferenceTable({ importance: 'key', }, ], - [yPath, headers, showModeledPower, powerCompare, selectedYAxisMetric, locale], + [ + yPath, + headers, + showModeledPower, + showAllInMeasured, + powerCompare, + selectedYAxisMetric, + locale, + ], ); return ( diff --git a/packages/app/src/components/inference/ui/inference-table-sort.ts b/packages/app/src/components/inference/ui/inference-table-sort.ts index 112ec5062..6a384a768 100644 --- a/packages/app/src/components/inference/ui/inference-table-sort.ts +++ b/packages/app/src/components/inference/ui/inference-table-sort.ts @@ -1,5 +1,5 @@ import type { ChartDefinition, InferenceData } from '@/components/inference/types'; -import { getNestedYValue } from '@/lib/chart-utils'; +import { inferenceTableYValue } from '@/components/inference/utils/inference-table-data'; /** * Default row order for the inference table. @@ -25,8 +25,10 @@ export function sortRowsByYMetric( const yAscending = rooflineDir?.startsWith('lower'); return [...data].toSorted((a, b) => { - const ay = a.powerVariant ? a.y : getNestedYValue(a, yPath); - const by = b.powerVariant ? b.y : getNestedYValue(b, yPath); + const ay = inferenceTableYValue(a, yPath); + const by = inferenceTableYValue(b, yPath); + if (ay === null) return by === null ? 0 : 1; + if (by === null) return -1; return yAscending ? ay - by : by - ay; }); } diff --git a/packages/app/src/components/inference/utils.test.ts b/packages/app/src/components/inference/utils.test.ts index 2b07210fb..4f9168c56 100644 --- a/packages/app/src/components/inference/utils.test.ts +++ b/packages/app/src/components/inference/utils.test.ts @@ -25,6 +25,38 @@ describe('selectUnofficialOverlayForMode', () => { ); }); }); + +describe('All in Measured table coverage', () => { + it('retains GPU-valid overlay rows with missing all-in power without plotting a fallback', () => { + const result = processOverlayChartDataWithClipping( + [ + pt({ + id: 1, + measuredAvgPower: { y: 450, roof: false }, + utilityModeledWatts: { y: 700, roof: false }, + }), + pt({ + id: 2, + measuredAvgPower: { y: 500, roof: false }, + modeledSystemPower: { + status: 'unsupported', + reason: 'cpu-telemetry', + modelRevision: 'test', + }, + }), + pt({ id: 3 }), + ], + 'interactivity', + 'y_utilityModeledWatts', + null, + ); + expect(result.data.map((point) => [point.id, point.y])).toEqual([[1, 700]]); + expect(result.tableData?.map((point) => [point.id, point.y])).toEqual([ + [1, 700], + [2, NaN], + ]); + }); +}); // --------------------------------------------------------------------------- // fixture factories // --------------------------------------------------------------------------- diff --git a/packages/app/src/components/inference/utils.ts b/packages/app/src/components/inference/utils.ts index f459b5639..4dadf08e3 100644 --- a/packages/app/src/components/inference/utils.ts +++ b/packages/app/src/components/inference/utils.ts @@ -5,7 +5,8 @@ import { getGpuSpecs, type TcoBasis } from '@/lib/constants'; * For Pareto front calculations, see @/lib/chart-utils */ -import chartDefinitions from '@/components/inference/metric-registry'; +import chartDefinitions, { isAllInMeasuredConfigKey } from '@/components/inference/metric-registry'; +import { allInMeasuredTableData } from '@/components/inference/utils/inference-table-data'; import { resolveXAxisField, type FixedSequenceStatistic, @@ -40,6 +41,7 @@ export function selectUnofficialOverlayForMode( export interface ProcessedChartData { data: InferenceData[]; clippedData: ClippedInferenceData[]; + tableData?: InferenceData[]; } /** @@ -260,7 +262,7 @@ export function processOverlayChartDataWithClipping( } } - return partitionChartDataByLimits( + const partition = partitionChartDataByLimits( processedData, { ...chartDef, x_scale_field: xAxisField }, selectedYAxisMetric, @@ -269,4 +271,16 @@ export function processOverlayChartDataWithClipping( isAgentic, }, ); + return { + ...partition, + ...(isAllInMeasuredConfigKey(selectedYAxisMetric) + ? { + tableData: expandPowerCompareSeries( + allInMeasuredTableData(sourceData, metricKey, xAxisField), + selectedYAxisMetric, + options?.powerCompare ?? 'none', + ), + } + : {}), + }; } diff --git a/packages/app/src/components/inference/utils/inference-table-data.test.ts b/packages/app/src/components/inference/utils/inference-table-data.test.ts new file mode 100644 index 000000000..bcf85b42e --- /dev/null +++ b/packages/app/src/components/inference/utils/inference-table-data.test.ts @@ -0,0 +1,107 @@ +import { describe, expect, it } from 'vitest'; +import { createElement } from 'react'; +import { renderToStaticMarkup } from 'react-dom/server'; +import type { InferenceData } from '@/components/inference/types'; +import { chartDefinitions } from '@/components/inference/metric-registry'; +import InferenceTable from '@/components/inference/ui/InferenceTable'; +import { sortRowsByYMetric } from '@/components/inference/ui/inference-table-sort'; +import { inferenceChartToCsv } from '@/lib/csv-export-helpers'; +import { + allInMeasuredTableData, + allInMeasuredUnavailableReason, + inferenceTableYValue, +} from './inference-table-data'; + +const point = (overrides: Partial = {}): InferenceData => + ({ + id: 1, + x: 20, + y: 999, + date: '2026-09-01', + tp: 4, + conc: 8, + hwKey: 'gb200_trt', + hw: 'gb200', + precision: 'fp4', + measuredAvgPower: { y: 450, roof: false }, + modeledSystemPower: { status: 'unsupported', reason: 'cpu-telemetry', modelRevision: 'test' }, + ...overrides, + }) as InferenceData; + +describe('All in Measured table values', () => { + it('retains only positive GPU measurements and never substitutes throughput for a missing estimate', () => { + const rows = allInMeasuredTableData( + [ + point(), + point({ measuredAvgPower: undefined }), + point({ measuredAvgPower: { y: 0, roof: false } }), + ], + 'utilityModeledWatts', + 'median_intvty', + ); + expect(rows).toHaveLength(1); + expect(rows[0].y).toBeNaN(); + expect(inferenceTableYValue(rows[0], 'utilityModeledWatts.y')).toBeNull(); + expect(allInMeasuredUnavailableReason(rows[0], 'utilityModeledWatts')).toBe('cpu-telemetry'); + expect( + allInMeasuredUnavailableReason( + { ...rows[0], y: 1400, powerVariant: { kind: 'basis', id: 'gpu-provisioned' } }, + 'utilityModeledWatts', + ), + ).toBe('cpu-telemetry'); + }); + + it('sorts unavailable estimates last and renders their measured watts and reason', () => { + const rows = allInMeasuredTableData( + [point(), point({ id: 2, utilityModeledWatts: { y: 700, roof: false } })], + 'utilityModeledWatts', + 'median_intvty', + ); + expect( + sortRowsByYMetric(rows, chartDefinitions[0], 'y_utilityModeledWatts').map((row) => row.id), + ).toEqual([2, 1]); + const html = renderToStaticMarkup( + createElement(InferenceTable, { + data: rows, + chartDefinition: chartDefinitions[0], + selectedYAxisMetric: 'y_utilityModeledWatts', + }), + ); + expect(html).toContain('Grace or module telemetry missing or invalid'); + expect(html).toContain('450'); + expect(html).toContain('—'); + expect(html).not.toContain('NaN'); + expect(html).not.toContain('999'); + }); + + it('exports unavailable estimates as blanks alongside GPU watts and reason', () => { + const rows = allInMeasuredTableData( + [point({ x: NaN })], + 'utilityModeledWatts', + 'median_intvty', + ); + const csv = inferenceChartToCsv(rows, 'Kimi-K3', 'agentic-traces', [], { + yHeader: 'All in Measured', + yPath: 'utilityModeledWatts.y', + xHeader: 'Interactivity', + }); + expect(csv.rows[0][csv.headers.indexOf('All in Measured')]).toBe(''); + expect(csv.rows[0][csv.headers.indexOf('Measured GPU Power (W/chip)')]).toBe(450); + expect(csv.rows[0][csv.headers.indexOf('All-in Estimate Status')]).toBe( + 'Grace or module telemetry missing or invalid', + ); + expect(csv.rows[0][csv.headers.indexOf('Interactivity')]).toBe(''); + const cloneCsv = inferenceChartToCsv( + [{ ...rows[0], powerVariant: { kind: 'basis', id: 'utility-modeled' } }], + 'Kimi-K3', + 'agentic-traces', + [], + { + yHeader: 'All in Measured', + yPath: 'utilityModeledWatts.y', + xHeader: 'Interactivity', + }, + ); + expect(cloneCsv.rows[0][cloneCsv.headers.indexOf('All in Measured')]).toBe(''); + }); +}); diff --git a/packages/app/src/components/inference/utils/inference-table-data.ts b/packages/app/src/components/inference/utils/inference-table-data.ts new file mode 100644 index 000000000..f1228483d --- /dev/null +++ b/packages/app/src/components/inference/utils/inference-table-data.ts @@ -0,0 +1,66 @@ +import type { AggDataEntry, InferenceData, YAxisMetricKey } from '@/components/inference/types'; +import { remapInferencePoint } from '@/lib/chart-utils'; +import type { Locale } from '@/lib/i18n'; +import type { SystemPowerUnsupportedReason } from '@/lib/modeled-system-power'; + +/** Table candidates keep GPU measurements even when the all-in model cannot run. */ +export function allInMeasuredTableData( + data: readonly InferenceData[], + metricKey: YAxisMetricKey, + xAxisField: keyof AggDataEntry, +): InferenceData[] { + return data + .filter((point) => Number.isFinite(point.measuredAvgPower?.y) && point.measuredAvgPower!.y > 0) + .map((point) => ({ + ...remapInferencePoint(point, metricKey, xAxisField), + // InferenceData requires a number. This sentinel is table-only; exports use null/blank. + y: point[metricKey]?.y ?? NaN, + })); +} + +export function inferenceTableYValue(point: InferenceData, yPath?: string): number | null { + const key = yPath?.split('.')[0] as YAxisMetricKey | undefined; + const value = point.powerVariant || !key ? point.y : point[key]?.y; + return typeof value === 'number' && Number.isFinite(value) ? value : null; +} + +type UnavailableReason = SystemPowerUnsupportedReason | 'energy' | 'model-unavailable'; + +export function allInMeasuredUnavailableReason( + point: InferenceData, + metricKey: string, +): UnavailableReason | null { + if (Number.isFinite(point[metricKey as YAxisMetricKey]?.y)) return null; + if (point.modeledSystemPower?.status === 'unsupported') return point.modeledSystemPower.reason; + if ( + point.modeledSystemPower?.status === 'supported' && + metricKey === 'utilityModeledJPerOutputToken' + ) + return 'energy'; + return 'model-unavailable'; +} + +const REASONS: Record> = { + workload: { en: 'Unsupported workload', zh: '不支持此工作负载' }, + hardware: { en: 'Unsupported hardware', zh: '不支持此硬件' }, + telemetry: { en: 'GPU telemetry missing or invalid', zh: 'GPU 遥测缺失或无效' }, + 'cpu-telemetry': { + en: 'Grace or module telemetry missing or invalid', + zh: 'Grace 或 module 遥测缺失或无效', + }, + 'gpu-count': { en: 'GPU count missing or inconsistent', zh: 'GPU 数量缺失或不一致' }, + topology: { en: 'Unsupported topology', zh: '不支持此拓扑' }, + 'role-power': { en: 'Role power incomplete', zh: 'Prefill 或 decode 功耗不完整' }, + 'model-domain': { en: 'Outside model range', zh: '超出模型适用范围' }, + energy: { en: 'Measured energy unavailable', zh: '缺少实测能耗' }, + 'model-unavailable': { en: 'All-in estimate unavailable', zh: '整体功耗估算不可用' }, +}; + +export function allInMeasuredStatusLabel( + point: InferenceData, + metricKey: string, + locale: Locale, +): string { + const reason = allInMeasuredUnavailableReason(point, metricKey); + return reason === null ? (locale === 'zh' ? '可用' : 'Available') : REASONS[reason][locale]; +} diff --git a/packages/app/src/components/inference/utils/tooltip-utils.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.test.ts index d90342a63..6a3709271 100644 --- a/packages/app/src/components/inference/utils/tooltip-utils.test.ts +++ b/packages/app/src/components/inference/utils/tooltip-utils.test.ts @@ -1,4 +1,4 @@ -import { describe, it, expect } from 'vitest'; +import { describe, it, expect, vi } from 'vitest'; import type { HardwareConfig, InferenceData } from '@/components/inference/types'; import type { SystemPowerEstimate } from '@/lib/modeled-system-power'; @@ -72,8 +72,8 @@ function tooltipConfig(overrides: Partial = {}): TooltipConfig { const systemPower = { status: 'supported', hardware: 'h100', - modelRevision: 'ca4403aa527069857351ad8047dbb726844b3382', - modelPath: 'chassis/H100.py', + modelRevision: `app-sha256:${'a'.repeat(64)}`, + modelPath: 'packages/app/src/lib/system-power-model.ts', gpuCount: 16, chassisCount: 2, chassisAcWatts: 12000, @@ -134,6 +134,26 @@ describe('modeled system-power tooltip', () => { ...overrides, }); + it.each(['en', 'zh'] as const)( + 'discloses the AgentX estimate in the %s All in Measured tooltip', + (locale) => { + const html = generateTooltipContent( + config({ + locale, + selectedYAxisMetric: 'y_utilityModeledWatts', + data: pt({ modeledSystemPower: systemPower, benchmark_type: 'agentic_traces' }), + }), + ); + expect(html).toContain('AgentX'); + expect(html).toContain( + locale === 'en' + ? 'not been independently calibrated' + : '尚未针对 AgentX 工作负载进行独立校准', + ); + expect(html).not.toContain('8k1k'); + }, + ); + it('separates measured input, normalized chassis AC, and whole-deployment facility power', () => { const html = generateTooltipContent(config()); expect(html).toContain('500 W/GPU'); @@ -146,11 +166,29 @@ describe('modeled system-power tooltip', () => { expect(html).toContain( 'Includes GPU chassis CPUs; excludes separate CPU-only frontend/router hosts.', ); - expect(html).toContain(`/blob/${systemPower.modelRevision}/${systemPower.modelPath}`); + expect(html).toContain(`/blob/master/${systemPower.modelPath}`); expect(html).not.toContain('12,000 W/GPU'); expect(html).not.toContain('Unmeasured chassis GPUs'); }); + it.each(['zh'] as const)('links %s model provenance to the deployed app source', (locale) => { + const buildRef = 'b'.repeat(40); + vi.stubEnv('NEXT_PUBLIC_APP_SOURCE_REF', buildRef); + try { + const html = generateTooltipContent(config({ locale })); + const app = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${buildRef}`; + expect(html).toContain(`${app}/${systemPower.modelPath}`); + expect(html).toContain(`${app}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`); + expect(html).toContain(locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions'); + expect(html).toContain(`title="${systemPower.modelRevision}"`); + expect(html).toContain('h100 · aaaaaaaaaaaa'); + expect(html).not.toContain('inferencex_power_model'); + expect(html).not.toContain(`/blob/${systemPower.modelRevision}/`); + } finally { + vi.unstubAllEnvs(); + } + }); + it('labels an extrapolated partial chassis and reports the measured GPUs’ share', () => { const data = pt({ physicalChips: 4, @@ -185,6 +223,56 @@ describe('modeled system-power tooltip', () => { expect(zh).not.toContain('6000 W'); }); + it('names NVL72 compute trays and the measured basis instead of eight-GPU chassis', () => { + const trays = { + ...systemPower, + hardware: 'gb200', + modelPath: 'packages/app/src/lib/system-power-model.ts', + gpuCount: 8, + chassisCount: 2, + modeledGpuCount: 8, + pue: 1.1, + topologyBasis: 'nvl72-trays', + measuredBasis: 'module', + sensorKind: 'module', + } satisfies SystemPowerEstimate; + const html = generateTooltipContent(config({ data: pt({ modeledSystemPower: trays }) })); + expect(html).toContain('2 full NVL72 compute trays · 8 GPUs'); + expect(html).toContain('Measured: module sensor (GPU + HBM + Grace + LPDDR5X)'); + expect(html).toContain('Rack AC is divided by all 72 GPUs'); + expect(html).toContain('Grace CPU and LPDDR5X are measured'); + expect(html).toContain('PUE 1.1'); + expect(html).not.toContain('eight-GPU chassis'); + expect(html).not.toContain('CPU/DRAM utilization'); + expect(html).not.toContain('Includes GPU chassis CPUs'); + + const partial = pt({ + physicalChips: 3, + modeledSystemPower: { + ...trays, + gpuCount: 3, + chassisCount: 1, + modeledGpuCount: 4, + chassisBasis: 'extrapolated', + measuredBasis: 'gpu-plus-grace', + sensorKind: 'grace-socket', + }, + }); + const en = generateTooltipContent(config({ data: partial })); + expect(en).toContain('1 NVL72 compute tray · 3 of 4 GPUs measured, extrapolated to full tray'); + expect(en).toContain('Unmeasured tray GPUs are assumed to run the same workload'); + expect(en).toContain('Measured: GPU board + Grace socket. Modeled: regulator loss'); + expect(en).not.toContain('Unmeasured chassis GPUs'); + + const zh = generateTooltipContent(config({ data: partial, locale: 'zh' })); + expect(zh).toContain('1 个 NVL72 计算 tray · 实测 3/4 张 GPU,按满 tray 外推'); + expect(zh).toContain('假设 tray 内未实测的 GPU 运行相同负载'); + expect(zh).toContain('实测:GPU 板卡 + Grace socket'); + expect(zh).toContain('Grace CPU 与 LPDDR5X 为实测值'); + expect(zh).not.toContain('八卡机箱'); + expect(zh).not.toContain('CPU/DRAM 利用率'); + }); + it('preserves the same model provenance in unofficial and date-comparison tooltips', () => { const official = config(); const overlay = generateOverlayTooltipContent({ @@ -228,17 +316,6 @@ describe('modeled system-power tooltip', () => { ); }); - it('localizes the measurement boundary and occupancy assumptions', () => { - const html = generateTooltipContent(config({ locale: 'zh' })); - expect(html).toContain('GPU 实测功耗'); - expect(html).toContain('整个部署的机箱交流功耗估算'); - expect(html).toContain('数据中心功耗估算'); - expect(html).toContain('2 个完整八卡机箱 · 16 张 GPU'); - expect(html).toContain('CPU/DRAM 利用率:20%'); - expect(html).toContain('计入 GPU 机箱内的 CPU'); - expect(html).toContain('不计入独立的纯 CPU 前端或路由主机。'); - }); - it('breaks normalization and host scope into two compact lines in pinned tooltips', () => { for (const locale of ['en', 'zh'] as const) { const html = generateTooltipContent(config({ locale })); @@ -1032,15 +1109,6 @@ describe('generateGPUGraphTooltipContent', () => { describe('measured-power withheld tooltip line', () => { const reasons = ['sampling_gap_exceeded', 'expected_gpu_count_mismatch']; - it('renders the withheld line with humanized codes (en)', () => { - const html = generateTooltipContent( - tooltipConfig({ data: pt({ power_valid: 0, power_invalid_reasons: reasons }) }), - ); - expect(html).toContain('Measured power withheld'); - expect(html).toContain('sampling gap exceeded'); - expect(html).toContain('expected gpu count mismatch'); - }); - it('renders the withheld line in Chinese on /zh surfaces', () => { const html = generateTooltipContent( tooltipConfig({ @@ -1073,11 +1141,7 @@ describe('measured-power withheld tooltip line', () => { expect(html).not.toContain('Measured power withheld'); }); - it.each([ - ['absent reasons', pt({ power_valid: 0 })], - ['empty reasons', pt({ power_valid: 0, power_invalid_reasons: [] })], - ['valid row', pt({ power_valid: 1 })], - ])('omits the line for %s', (_name, data) => { + it.each([['absent reasons', pt({ power_valid: 0 })]])('omits the line for %s', (_name, data) => { const html = generateTooltipContent(tooltipConfig({ data })); expect(html).not.toContain('Measured power withheld'); }); @@ -1169,15 +1233,6 @@ describe('worker power drilldown', () => { expect(generateGPUGraphTooltipContent(config)).not.toContain('tooltip-worker-power'); }); - it('renders nothing when workers is absent or empty', () => { - expect(generateTooltipContent(tooltipConfig({ isPinned: true }))).not.toContain( - 'tooltip-worker-power', - ); - expect( - generateTooltipContent(tooltipConfig({ data: pt({ workers: [] }), isPinned: true })), - ).not.toContain('tooltip-worker-power'); - }); - it('caps the table at 8 rows with a "+N more workers" line', () => { const many = Array.from({ length: 10 }, (_, i) => ({ role: 'decode', @@ -1232,18 +1287,6 @@ describe('worker power drilldown', () => { }); describe('power tier tooltip line', () => { - it('states the tier for a legacy point on a measured axis', () => { - const html = generateTooltipContent( - tooltipConfig({ - selectedYAxisMetric: 'y_measuredJPerOutputToken', - data: pt({ power_tier: 'legacy' }), - }), - ); - expect(html).toContain( - 'Power Measurement: Historical (not validated under the current method)', - ); - }); - it('states the certified tier on a measured axis', () => { const html = generateTooltipContent( tooltipConfig({ @@ -1254,16 +1297,6 @@ describe('power tier tooltip line', () => { expect(html).toContain('Power Measurement: Validated (current PowerX method)'); }); - it('omits the tier line on non-measured axes', () => { - const html = generateTooltipContent( - tooltipConfig({ - selectedYAxisMetric: 'y_tpPerGpu', - data: pt({ power_tier: 'legacy' }), - }), - ); - expect(html).not.toContain('Power Measurement'); - }); - it('omits the tier line when the point carries no tier', () => { const html = generateTooltipContent( tooltipConfig({ selectedYAxisMetric: 'y_measuredAvgPower', data: pt() }), diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts index a4877ffe4..3590e5dda 100644 --- a/packages/app/src/components/inference/utils/tooltipUtils.ts +++ b/packages/app/src/components/inference/utils/tooltipUtils.ts @@ -8,10 +8,15 @@ import type { Locale } from '@/lib/i18n'; import { isKvOffloadEnabled } from '@/lib/kv-offload'; import { chartStateHref } from '@/lib/url-state'; import { chipCounts } from '@/lib/chip-counts'; -import type { SystemPowerUnsupportedReason } from '@/lib/modeled-system-power'; +import { ALL_IN_MEASURED_AGENTIC_NOTE } from '@/lib/power-basis'; +import type { + SystemPowerSensorKind, + SystemPowerUnsupportedReason, +} from '@/lib/modeled-system-power'; import type { HardwareConfig, InferenceData, OverlayData } from '@/components/inference/types'; import { + isAllInMeasuredConfigKey, isMeasuredEnergyConfigKey, isModeledSystemPowerConfigKey, } from '@/components/inference/metric-registry'; @@ -257,13 +262,14 @@ const escapeHtml = (s: string): string => const SYSTEM_POWER_STRINGS = { en: { heading: 'Draft System-Power Model · 8k1k', + agenticHeading: 'Draft System-Power Model · AgentX', measuredGpu: 'Measured GPU power', normalizedAc: 'Modeled chassis AC per GPU', deploymentAc: 'Modeled deployment chassis AC', facility: 'Modeled facility power', assumptions: 'CPU/DRAM utilization: 20%; PCIe: 5%; NVMe: 0%; fans: auto.', platformAssumptions: 'NVIDIA NVLink: 50%, IB: 0%; AMD Ethernet: 0%.', - sweep: 'Fixed README inference sweep', + guide: 'Power model assumptions', topology: (chassis: number, measured: number, modeled: number) => measured === modeled ? `${chassis} full eight-GPU chassis · ${measured} GPUs` @@ -274,6 +280,24 @@ const SYSTEM_POWER_STRINGS = { 'No per-host telemetry for this multinode deployment; every chassis is modeled at the deployment-mean GPU power.', normalization: 'AC power is divided by all modeled chassis GPUs, including prefill and decode.', boundary: 'Includes GPU chassis CPUs; excludes separate CPU-only frontend/router hosts.', + // NVL72 compute trays: the compute module is measured, the rack residual modeled. + trayTopology: (trays: number, measured: number, modeled: number) => { + const unit = trays === 1 ? 'tray' : 'trays'; + return measured === modeled + ? `${trays} full NVL72 compute ${unit} · ${measured} GPUs` + : `${trays} NVL72 compute ${unit} · ${measured} of ${modeled} GPUs measured, extrapolated to full ${unit}`; + }, + trayExtrapolation: + 'Unmeasured tray GPUs are assumed to run the same workload at the measured per-GPU power; a module reading already covers the whole tray. Deployment values are the measured GPUs’ share.', + trayAssumptions: { + module: + 'Measured: module sensor (GPU + HBM + Grace + LPDDR5X). Modeled: NVSwitch trays, NICs/DPUs, NVMe, power shelves.', + 'grace-socket': + 'Measured: GPU board + Grace socket. Modeled: regulator loss, NVSwitch trays, NICs/DPUs, NVMe, power shelves.', + } satisfies Record, + trayPlatformAssumptions: 'NVIDIA NVLink: 50%, IB: 0%; PCIe: 5%.', + trayNormalization: 'Rack AC is divided by all 72 GPUs of a rack of matching trays.', + trayBoundary: 'Grace CPU and LPDDR5X are measured; excludes CPU-only frontend/router hosts.', model: 'Power model source', unavailable: 'System-power estimate unavailable', reasons: { @@ -284,17 +308,20 @@ const SYSTEM_POWER_STRINGS = { topology: 'The available topology does not establish chassis placement.', 'role-power': 'Valid measured power and topology are required for every GPU worker role.', 'model-domain': 'The measured input is outside the source model’s supported range.', + 'cpu-telemetry': + 'Validated measured Grace-side (CPU) power is required for NVL72 compute trays.', } satisfies Record, }, zh: { heading: '系统功耗模型(草案)· 8k1k', + agenticHeading: '系统功耗模型(草案)· AgentX', measuredGpu: 'GPU 实测功耗', normalizedAc: '每 GPU 分摊的机箱交流功耗估算', deploymentAc: '整个部署的机箱交流功耗估算', facility: '数据中心功耗估算', assumptions: 'CPU/DRAM 利用率:20%;PCIe:5%;NVMe:0%;风扇:自动。', platformAssumptions: 'NVIDIA NVLink:50%,IB:0%;AMD Ethernet:0%。', - sweep: 'README 中的固定推理参数扫描', + guide: '功耗模型与假设', topology: (chassis: number, measured: number, modeled: number) => measured === modeled ? `${chassis} 个完整八卡机箱 · ${measured} 张 GPU` @@ -304,6 +331,21 @@ const SYSTEM_POWER_STRINGS = { uniformHosts: '该多节点部署没有逐主机功耗数据;每个机箱按部署平均每卡功耗建模。', normalization: '交流功耗按所有建模机箱的 GPU 总数分摊,包括 Prefill 与 Decode。', boundary: '计入 GPU 机箱内的 CPU;不计入独立的纯 CPU 前端或路由主机。', + trayTopology: (trays: number, measured: number, modeled: number) => + measured === modeled + ? `${trays} 个完整 NVL72 计算 tray · ${measured} 张 GPU` + : `${trays} 个 NVL72 计算 tray · 实测 ${measured}/${modeled} 张 GPU,按满 tray 外推`, + trayExtrapolation: + '假设 tray 内未实测的 GPU 运行相同负载、功耗与实测每卡功耗相同;模块读数本身已覆盖整个 tray。部署数值为实测 GPU 所占份额。', + trayAssumptions: { + module: + '实测:模块传感器(GPU + HBM + Grace + LPDDR5X)。建模:NVSwitch tray、网卡/DPU、NVMe、电源架。', + 'grace-socket': + '实测:GPU 板卡 + Grace socket。建模:稳压损耗、NVSwitch tray、网卡/DPU、NVMe、电源架。', + } satisfies Record, + trayPlatformAssumptions: 'NVIDIA NVLink:50%,IB:0%;PCIe:5%。', + trayNormalization: '机架交流功耗按由相同 tray 组成的整机架的 72 张 GPU 分摊。', + trayBoundary: 'Grace CPU 与 LPDDR5X 为实测值;不计入独立的纯 CPU 前端或路由主机。', model: '功耗模型来源', unavailable: '无法估算系统功耗', reasons: { @@ -314,6 +356,7 @@ const SYSTEM_POWER_STRINGS = { topology: '现有拓扑信息无法确认 GPU 所在的机箱。', 'role-power': '每个 GPU worker 角色都需要有效的实测功耗和拓扑信息。', 'model-domain': '实测输入超出功耗模型的支持范围。', + 'cpu-telemetry': 'NVL72 计算 tray 需要通过验证的 Grace 侧(CPU)实测功耗。', } satisfies Record, }, } as const; @@ -328,7 +371,8 @@ const modeledSystemPowerHTML = ( if ( !estimate || (!isMeasuredEnergyConfigKey(selectedYAxisMetric) && - !isModeledSystemPowerConfigKey(selectedYAxisMetric)) + !isModeledSystemPowerConfigKey(selectedYAxisMetric) && + !isAllInMeasuredConfigKey(selectedYAxisMetric)) ) { return ''; } @@ -337,10 +381,31 @@ const modeledSystemPowerHTML = ( if (!isPinned || estimate.reason === 'workload') return ''; return tooltipLine(t.unavailable, t.reasons[estimate.reason]); } - const sourceUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/${estimate.modelPath}`; - const readmeUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/README.md`; + const sourceRef = encodeURIComponent(process.env.NEXT_PUBLIC_APP_SOURCE_REF || 'master'); + const appSource = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${sourceRef}`; + const sourceUrl = `${appSource}/${estimate.modelPath}`; + const guideUrl = `${appSource}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`; + // Tray estimates measure the compute module; chassis estimates model the CPU/DRAM. + const tray = estimate.topologyBasis === 'nvl72-trays' ? estimate : null; + const topology = tray + ? t.trayTopology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount) + : t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount); + const extrapolation = + estimate.chassisBasis === 'extrapolated' + ? `
${tray ? t.trayExtrapolation : t.extrapolation}` + : ''; + const uniformHosts = estimate.topologyBasis === 'uniform-hosts' ? `
${t.uniformHosts}` : ''; + const notes = tray + ? [ + t.trayAssumptions[tray.sensorKind], + t.trayPlatformAssumptions, + t.trayNormalization, + t.trayBoundary, + ] + : [t.assumptions, t.platformAssumptions, t.normalization, t.boundary]; return `

- ${t.heading} + ${d.benchmark_type === 'agentic_traces' ? t.agenticHeading : t.heading} + ${d.benchmark_type === 'agentic_traces' ? `
${ALL_IN_MEASURED_AGENTIC_NOTE[locale]}
` : ''} ${tooltipLine(t.measuredGpu, `${fmt(estimate.measuredGpuWattsPerGpu)} W/GPU`)} ${tooltipLine(t.normalizedAc, `${fmt(estimate.chassisAcWattsPerGpu)} W/GPU`)} ${ @@ -348,9 +413,9 @@ const modeledSystemPowerHTML = ( ? ` ${tooltipLine(t.deploymentAc, `${fmt(estimate.deploymentAcWatts)} W`)} ${tooltipLine(`${t.facility} (PUE ${fmt(estimate.pue)})`, `${fmt(estimate.deploymentFacilityWatts)} W`)} -
${t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount)}${estimate.chassisBasis === 'extrapolated' ? `
${t.extrapolation}` : ''}${estimate.topologyBasis === 'uniform-hosts' ? `
${t.uniformHosts}` : ''}
${t.assumptions}
${t.platformAssumptions}
${t.normalization}
${t.boundary}
- ${tooltipLine(t.model, `${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.slice(0, 12))}`)} - ${t.sweep} +
${topology}${extrapolation}${uniformHosts}
${notes.join('
')}
+ ${tooltipLine(t.model, `${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.replace(/^app-sha256:/u, '').slice(0, 12))}`)} + ${t.guide} ` : '' } diff --git a/packages/app/src/components/ui/d3-chart-wrapper.tsx b/packages/app/src/components/ui/d3-chart-wrapper.tsx index 1d91d5103..5880cb96e 100644 --- a/packages/app/src/components/ui/d3-chart-wrapper.tsx +++ b/packages/app/src/components/ui/d3-chart-wrapper.tsx @@ -65,6 +65,7 @@ export interface D3ChartWrapperProps { instructions?: string; testId?: string; grabCursor?: boolean; + scrollablePlot?: { minWidth: number; label: string; enabled: boolean }; } export function D3ChartWrapper({ @@ -83,83 +84,117 @@ export function D3ChartWrapper({ instructions, testId, grabCursor = true, + scrollablePlot, }: D3ChartWrapperProps) { const locale = useLocale(); const resolvedInstructions = instructions ?? DEFAULT_CHART_INSTRUCTIONS[locale]; - return ( -
- {caption &&
{caption}
} -
-
-
- {/* Stable hook for tests. `[data-testid="scatter-graph"] svg` also + const plot = ( +
+
+
+ {/* Stable hook for tests. `[data-testid="scatter-graph"] svg` also matches every Lucide icon inside the card — dozens of them — so picking "the first svg" silently grabs an icon whenever the selected metric renders one above the chart. */} - { - (e.currentTarget as SVGSVGElement).style.cursor = 'grabbing'; - } - : undefined - } - onMouseUp={ - grabCursor - ? (e) => { - (e.currentTarget as SVGSVGElement).style.cursor = 'grab'; - } - : undefined + { + (e.currentTarget as SVGSVGElement).style.cursor = 'grabbing'; + } + : undefined + } + onMouseUp={ + grabCursor + ? (e) => { + (e.currentTarget as SVGSVGElement).style.cursor = 'grab'; + } + : undefined + } + onClick={() => { + if (isPinned()) { + dismissTooltip(); + hideTooltipElements(tooltipRef, svgRef); } - onClick={() => { - if (isPinned()) { - dismissTooltip(); - hideTooltipElements(tooltipRef, svgRef); - } - }} - /> - {/* Tooltip is portalled to with position:fixed so it can + }} + /> + {/* Tooltip is portalled to with position:fixed so it can rise above sibling chart cards' stacking contexts. The d3 layer writes viewport-coords into style.left/top — see computeTooltipPosition. */} - - {noDataOverlay} -
- {resolvedInstructions && ( -

- {resolvedInstructions} -

- )} -
-
-
+ + {noDataOverlay}
- {legendElement && ( - /* Sizes to the legend content: when the sidebar legend panel is open + {resolvedInstructions && ( +

+ {resolvedInstructions} +

+ )} +
+
+
+
+ {legendElement && ( + /* Sizes to the legend content: when the sidebar legend panel is open (.sidebar-legend present) the column grows to fit the widest legend label (capped) so full names display without truncation, while still sitting next to the plot without overlapping it; when closed the legend renders only a small reopen button and the chart reclaims the width. Height belongs to the legend itself: short lists should not reserve an empty chart-height column. */ +
+ {legendElement} +
+ )} +
+ ); + + return ( +
+ {caption &&
{caption}
} + {scrollablePlot ? ( + <> + {scrollablePlot.enabled && ( +

{scrollablePlot.label}

+ )}
{ + if (!scrollablePlot.enabled) return; + if (event.target !== event.currentTarget) return; + if (event.key !== 'ArrowLeft' && event.key !== 'ArrowRight') return; + event.preventDefault(); + event.currentTarget.scrollBy({ + left: (event.key === 'ArrowRight' ? 1 : -1) * event.currentTarget.clientWidth * 0.8, + }); + }} > - {legendElement} + {plot}
- )} -
+ + ) : ( + plot + )}
); } diff --git a/packages/app/src/hooks/useChartExport.test.ts b/packages/app/src/hooks/useChartExport.test.ts index 0c28af0cf..2fb740f01 100644 --- a/packages/app/src/hooks/useChartExport.test.ts +++ b/packages/app/src/hooks/useChartExport.test.ts @@ -131,6 +131,37 @@ describe('useChartExport failure messages', () => { }, ); + it.each([false, true])( + 'exports the whole plot with scrolling=%s without moving the live chart', + async (scrolling) => { + const plot = + '
First barLast bar
'; + chart.innerHTML = `
Power comparison
${scrolling ? `
${plot}
` : plot}`; + const liveScroller = chart.querySelector('[data-chart-scroll]'); + if (liveScroller) liveScroller.scrollLeft = 250; + const original = chart.innerHTML; + let snapshot: HTMLElement; + exportMocks.toPng.mockImplementationOnce((element: HTMLElement) => { + snapshot = element.cloneNode(true) as HTMLElement; + throw new Error('stop after capture'); + }); + vi.spyOn(window, 'alert').mockImplementation(() => {}); + vi.spyOn(console, 'error').mockImplementation(() => {}); + await act(() => current.exportToImage()); + expect(exportMocks.toPng).toHaveBeenCalledOnce(); + expect(snapshot!.textContent).toContain('Power comparison'); + expect(snapshot!.textContent).toContain('First bar'); + expect(snapshot!.textContent).toContain('Last bar'); + const exportedScroller = snapshot!.querySelector('[data-chart-scroll]'); + if (scrolling) { + expect(exportedScroller?.style.overflow).toBe('visible'); + expect(exportedScroller?.scrollLeft).toBe(0); + } else expect(exportedScroller).toBeNull(); + expect(chart.innerHTML).toBe(original); + if (liveScroller) expect(liveScroller.scrollLeft).toBe(250); + }, + ); + it('includes the legend again after line labels are toggled off', async () => { chart.innerHTML = '
B200
'; diff --git a/packages/app/src/hooks/useChartExport.ts b/packages/app/src/hooks/useChartExport.ts index 514d7e82a..a105c20a6 100644 --- a/packages/app/src/hooks/useChartExport.ts +++ b/packages/app/src/hooks/useChartExport.ts @@ -356,7 +356,13 @@ export function useChartExport({ // Layout: force side-by-side flex row for export applyStyles(exportElement, { width: 'fit-content', overflow: 'visible', padding: '16px' }); - const flexContainer = clone.querySelector(':scope > .flex') as HTMLElement | null; + for (const scroller of clone.querySelectorAll('[data-chart-scroll]')) { + applyStyles(scroller, { width: 'fit-content', overflow: 'visible' }); + scroller.scrollLeft = 0; + } + const flexContainer = clone.querySelector( + ':scope > .flex, [data-chart-scroll] > .flex', + ) as HTMLElement | null; applyStyles(flexContainer, { flexDirection: 'row', width: 'fit-content', diff --git a/packages/app/src/lib/api-route-catalog.ts b/packages/app/src/lib/api-route-catalog.ts index b989cadd6..5f51eaa46 100644 --- a/packages/app/src/lib/api-route-catalog.ts +++ b/packages/app/src/lib/api-route-catalog.ts @@ -139,7 +139,7 @@ export const apiRouteCatalog = [ method: 'GET', classification: 'published-read', operationId: 'get-inference-view', - sourceSha256: 'bafac08dbe6e6e4b75dd15987e79f84b514f3e93351149ad44b296f6dd2a9662', + sourceSha256: '428fab797ad0d304ee7ee193fd2e115219e986e87294d4f561953ab64e0fd5ef', }, { source: 'src/app/api/v1/views/options/route.ts', @@ -843,6 +843,22 @@ export interface ApiContractSourceDigest { * touching a route module. Digest changes require an explicit documentation review. */ export const apiContractSourceDigests = [ + { + source: 'src/components/calculator/profit-power.ts', + sourceSha256: 'ff92b0a954afb78769ab8a29ca6ca7bdc52f83bb49379e2f2f274ef4c179330a', + reviewArea: { + en: 'Power-valid curve selection at fixed targets, compatible power bases, paired throughput and provisioned fallback shared by Profit UI and API.', + zh: '利润界面与 API 共用的有效功耗曲线选择、固定目标值、功耗口径兼容性、配对吞吐量与预配估算回退。', + }, + }, + { + source: 'src/components/inference/utils/inference-table-data.ts', + sourceSha256: '816eb0dbc416f6c322fd3955af13fc03f8c6988704d82f36a9bfd8ec9abc9669', + reviewArea: { + en: 'All in Measured table eligibility, nullable values and unavailable reasons shared by UI, CSV and public views.', + zh: '界面、CSV 和公开视图共用的整体实测表格行筛选、可空数值与不可用原因。', + }, + }, { source: 'src/components/inference/utils/resolveXAxisField.ts', sourceSha256: '4783579c7b3c1a21b91968cb03e9c35a57a85251f4992cb5667a855ba7c77497', @@ -885,7 +901,7 @@ export const apiContractSourceDigests = [ }, { source: 'src/lib/benchmark-transform.ts', - sourceSha256: '53fb9102b1ddd0c597b3b2a414894564d2deda3da5ea2cf0059f997b2db852c8', + sourceSha256: 'e92e215ead748eafa2a6b496201adcd0f9387110d0843cb8fe7ebfdbcd8a59ef', reviewArea: { en: 'Raw benchmark means and derived reciprocal mean-TPOT interactivity used by Dashboard and read-only views.', zh: '仪表板和只读视图共用的原始 benchmark 均值与 mean TPOT 倒数形式的 interactivity。', @@ -1015,7 +1031,7 @@ export const apiContractSourceDigests = [ { source: 'src/lib/views-api/calculator-extensions.ts', - sourceSha256: '2bcd27b5f5fbea9dcd75b32f7ee88e1c92343bed3ff65a293fc9a8ff5ecf04c3', + sourceSha256: '0c571acc23940e65290a5421c0872eb63f2531b289699c520209910b50dee784', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', @@ -1033,7 +1049,7 @@ export const apiContractSourceDigests = [ { source: 'src/lib/views-api/series.ts', - sourceSha256: '9974fe5166ac847f4b9284eda8a901ee0f9c8b432567ef184890453c715d2f4f', + sourceSha256: '9bf521e64e111965b597cb0f55d2f2b4b765b09f8168e8399e9c4316921cc643', reviewArea: { en: 'Dashboard read-only selector and calculation parity.', zh: '仪表板只读接口的选择项与计算一致性。', diff --git a/packages/app/src/lib/benchmark-transform.ts b/packages/app/src/lib/benchmark-transform.ts index 3c7263a56..788f5a382 100644 --- a/packages/app/src/lib/benchmark-transform.ts +++ b/packages/app/src/lib/benchmark-transform.ts @@ -236,7 +236,7 @@ export function rowToAggDataEntry(row: BenchmarkRow): AggDataEntry { ? row.power_invalid_reasons : undefined, power_metric_schema_version: m.power_metric_schema_version, - modeledSystemPower: modelSystemPower(row), + modeledSystemPower: modelSystemPower(row, undefined, true), power_tier: resolvePowerTier({ powerValid: m.power_valid, wholeDeploymentSemantics: hasWholeDeploymentEnergySemantics, diff --git a/packages/app/src/lib/chart-utils.ts b/packages/app/src/lib/chart-utils.ts index a1ae03ac2..e9aae745f 100644 --- a/packages/app/src/lib/chart-utils.ts +++ b/packages/app/src/lib/chart-utils.ts @@ -442,7 +442,13 @@ export function buildDerivedChartFields( if (wants(key)) fields[key] = value; } - if (wants('modeledChassisPowerPerGpu') && entry.modeledSystemPower?.status === 'supported') { + if ( + wants('modeledChassisPowerPerGpu') && + entry.benchmark_type === 'single_turn' && + entry.isl === 8192 && + entry.osl === 1024 && + entry.modeledSystemPower?.status === 'supported' + ) { fields.modeledChassisPowerPerGpu = chartMetric(entry.modeledSystemPower.chassisAcWattsPerGpu); } diff --git a/packages/app/src/lib/csv-export-helpers.ts b/packages/app/src/lib/csv-export-helpers.ts index a58ba948f..b70bfa4ec 100644 --- a/packages/app/src/lib/csv-export-helpers.ts +++ b/packages/app/src/lib/csv-export-helpers.ts @@ -7,7 +7,11 @@ * plotted x/y axes. */ -import { METRIC_REGISTRY } from '@/components/inference/metric-registry'; +import { METRIC_REGISTRY, isAllInMeasuredConfigKey } from '@/components/inference/metric-registry'; +import { + allInMeasuredStatusLabel, + inferenceTableYValue, +} from '@/components/inference/utils/inference-table-data'; import type { InferenceData, TrendDataPoint } from '@/components/inference/types'; import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare'; import { chipCounts } from '@/lib/chip-counts'; @@ -31,9 +35,9 @@ function nestedMetric(point: InferenceData, path: string): number | '' { const value = point[key as keyof InferenceData]; if (nestedKey && typeof value === 'object' && value !== null && nestedKey in value) { const nestedValue = (value as Record)[nestedKey]; - return typeof nestedValue === 'number' ? nestedValue : ''; + return typeof nestedValue === 'number' && Number.isFinite(nestedValue) ? nestedValue : ''; } - return typeof value === 'number' ? value : ''; + return typeof value === 'number' && Number.isFinite(value) ? value : ''; } /** Preserve a real zero while leaving source metrics that were not measured blank. */ @@ -64,6 +68,7 @@ export function inferenceChartToCsv( const powerCompare = inferPowerCompare(allPoints); const showPowerSeries = powerCompare !== 'none'; const plottedMetric = displayedMetrics ? `y_${displayedMetrics.yPath.split('.')[0]}` : ''; + const showAllInMeasured = isAllInMeasuredConfigKey(plottedMetric); const headers = [ 'Model', 'ISL', @@ -119,6 +124,7 @@ export function inferenceChartToCsv( 'DP', ...(showModeledPower ? ['Configured Chip Count'] : []), ...(showPowerSeries ? ['Power Series'] : []), + ...(showAllInMeasured ? ['Measured GPU Power (W/chip)', 'All-in Estimate Status'] : []), ]; const displayedColumns = displayedMetrics @@ -126,9 +132,14 @@ export function inferenceChartToCsv( { header: displayedMetrics.yHeader, value: (point: InferenceData) => - point.powerVariant ? point.y : nestedMetric(point, displayedMetrics.yPath), + point.powerVariant + ? (inferenceTableYValue(point) ?? '') + : nestedMetric(point, displayedMetrics.yPath), + }, + { + header: displayedMetrics.xHeader, + value: (point: InferenceData) => (Number.isFinite(point.x) ? point.x : ''), }, - { header: displayedMetrics.xHeader, value: (point: InferenceData) => point.x }, ].filter( (column, index, columns) => !headers.includes(column.header) && @@ -187,6 +198,9 @@ export function inferenceChartToCsv( d.dp ?? '', ...(showModeledPower ? [chips.configured] : []), ...(showPowerSeries ? [powerSeriesLabel(d, plottedMetric, powerCompare, 'en')] : []), + ...(showAllInMeasured + ? [d.measuredAvgPower?.y ?? '', allInMeasuredStatusLabel(d, plottedMetric.slice(2), 'en')] + : []), ]; row.splice(10, 0, ...displayedColumns.map((column) => column.value(d))); return row; diff --git a/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx b/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx index 6553b1611..21bf95998 100644 --- a/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx +++ b/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx @@ -39,6 +39,7 @@ function D3ChartInner( legendElement, noDataOverlay, caption, + scrollablePlot, onRender, onDisplayUpdate, }: D3ChartProps, @@ -159,6 +160,7 @@ function D3ChartInner( legendElement={legendElement} noDataOverlay={noDataOverlay} caption={caption} + scrollablePlot={scrollablePlot} /> ); } diff --git a/packages/app/src/lib/d3-chart/D3Chart/types.ts b/packages/app/src/lib/d3-chart/D3Chart/types.ts index e60543d0d..e8de3b7b9 100644 --- a/packages/app/src/lib/d3-chart/D3Chart/types.ts +++ b/packages/app/src/lib/d3-chart/D3Chart/types.ts @@ -294,6 +294,8 @@ export interface D3ChartProps { legendElement?: React.ReactNode; noDataOverlay?: React.ReactNode; caption?: React.ReactNode; + /** Keep the SVG mounted so D3 retains its layout when scrolling toggles. */ + scrollablePlot?: { minWidth: number; label: string; enabled: boolean }; /** Called after all layers render. Useful for one-off DOM manipulations. */ onRender?: (ctx: RenderContext) => void; diff --git a/packages/app/src/lib/modeled-system-power-export.test.ts b/packages/app/src/lib/modeled-system-power-export.test.ts index 59f6c57fa..e0fc5af57 100644 --- a/packages/app/src/lib/modeled-system-power-export.test.ts +++ b/packages/app/src/lib/modeled-system-power-export.test.ts @@ -5,7 +5,7 @@ import { csv, type ComparisonInput, } from '../../scripts/export-modeled-system-power'; -import { estimateChassisPower } from '@/lib/system-power-model'; +import { estimateChassisPower, estimateRackPower } from '@/lib/system-power-model'; // Original H200 c1, run 31672765610, artifact 9171086754; schema marker was absent. function input(): ComparisonInput { @@ -86,11 +86,20 @@ describe('offline modeled PowerX comparisons', () => { const source = input(); const before = structuredClone(source); const result = buildComparison(source); - expect(result.metadata.pue).toBe(1.3); + expect(result.metadata.model.source).toBe('https://github.com/SemiAnalysisAI/InferenceX-app'); + expect(result.metadata.model.modelRevision).toMatch(/^app-sha256:[0-9a-f]{64}$/u); + expect(result.rows[0].model_path).toBe('packages/app/src/lib/system-power-model.ts'); + expect(result.rows[0].modeled.modelRevision).toBe(result.metadata.model.modelRevision); + expect(result.metadata.pue_override).toBeNull(); + expect(result.metadata.pue_defaults).toEqual({ air_cooled_chassis: 1.3, dlc_nvl72_rack: 1.1 }); expect(result.metadata.model.assumptions.pue).toBe(1.2); + expect(result.rows[0].pue).toBe(1.3); expect(result.rows[0].modeled).toMatchObject({ pue: 1.3 }); expect(result.rows[0].assumptions).toMatchObject({ pue: 1.3 }); - expect(buildComparison(source, 1.1).rows[0].modeled).toMatchObject({ pue: 1.1 }); + expect(result.cells[0].pue).toBe(1.3); + const overridden = buildComparison(source, 1.1); + expect(overridden.metadata.pue_override).toBe(1.1); + expect(overridden.rows[0]).toMatchObject({ pue: 1.1, modeled: { pue: 1.1 } }); expect(result.rows[0].estimated_energy).toMatchObject({ status: 'estimated', output_tokens: 9303, @@ -153,17 +162,14 @@ describe('offline modeled PowerX comparisons', () => { }); }); - it.each([0, -1, Infinity, NaN, 1.5])( - 'withholds energy for invalid output denominator %s', - (tokens) => { - const source = input(); - source.rows[0].audit!.benchmark_window.total_output_tokens = tokens; - expect(buildComparison(source).rows[0].estimated_energy).toEqual({ - status: 'unavailable', - reason: 'audit-does-not-match-measured-input', - }); - }, - ); + it.each([0, 1.5])('withholds energy for invalid output denominator %s', (tokens) => { + const source = input(); + source.rows[0].audit!.benchmark_window.total_output_tokens = tokens; + expect(buildComparison(source).rows[0].estimated_energy).toEqual({ + status: 'unavailable', + reason: 'audit-does-not-match-measured-input', + }); + }); it('withholds energy for missing, mismatched, and invalid audit receipts', () => { for (const mutate of [ @@ -216,17 +222,99 @@ describe('offline modeled PowerX comparisons', () => { expect(() => buildComparison(source)).toThrow('different benchmark configurations'); }); + it('routes GB200 NVL72 rows through the tray estimate with the DLC PUE and leaves x86 rows unchanged', () => { + const source = input(); + const baseline = buildComparison(structuredClone(source)); + // One GB200 compute tray on the module basis, as the dashboard would receive it. + const gb200 = structuredClone(source.rows[0]); + gb200.id = 'gb200:tray'; + gb200.cell = 'gb200:c1'; + gb200.audit = undefined; + Object.assign(gb200.benchmark, { + hardware: 'gb200', + power_audit: { cpu: { sensor_kind: 'module', expected_sockets: 2, observed_sockets: 2 } }, + framework: 'dynamo-trt', + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + metrics: { + pp: 1, + pcp_size: 1, + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 1, + avg_power_w: 900.25, + avg_total_gpu_power_w: 3601, + avg_cpu_socket_power_w: 250.5, + avg_total_cpu_power_w: 501, + avg_total_module_power_w: 4300.75, + total_module_energy_j: 258045, + }, + }); + source.rows.push(gb200); + const result = buildComparison(source); + const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!; + expect(result.rows[1]).toMatchObject({ + pue: 1.1, + measured_basis: 'module', + sensor_kind: 'module', + model_path: 'packages/app/src/lib/system-power-model.ts', + assumptions: { u_nvlink: 0.5, pue: 1.1 }, + measured_inputs: { + avg_gpu_w: 900.25, + cpu_power_valid: 1, + total_grace_w: 501, + total_module_w: 4300.75, + total_module_j: 258045, + }, + modeled: { + status: 'supported', + topologyBasis: 'nvl72-trays', + pue: 1.1, + chassisAcWatts: rack.rackAcWatts / 18, + facilityWatts: rack.facilityWatts / 18, + }, + }); + expect(result.rows[1].calculation_boundary).toContain('NVL72'); + expect(result.rows[1].extrapolation_note).toContain('tray'); + expect(result.cells[1]).toMatchObject({ + cell: 'gb200:c1', + pue: 1.1, + measured_bases: ['module'], + modeled_chassis_ac_w_mean: rack.rackAcWatts / 18, + }); + // The x86 row and its cell are byte-identical to an export without the NVL72 row. + expect(result.rows[0]).toEqual(baseline.rows[0]); + expect(result.cells[0]).toEqual(baseline.cells[0]); + expect(result.rows[0].measured_inputs).not.toHaveProperty('total_module_w'); + expect(result.rows[0]).toMatchObject({ measured_basis: null, sensor_kind: null }); + // An explicit --pue still overrides every row, rack and chassis alike. + const overridden = buildComparison(source, 1.3); + expect(overridden.rows.map((row) => row.pue)).toEqual([1.3, 1.3]); + expect(overridden.rows[1].modeled).toMatchObject({ pue: 1.3, topologyBasis: 'nvl72-trays' }); + // Without cpu_power_valid the row stays unavailable and reports no CPU-side inputs. + delete gb200.benchmark.metrics.cpu_power_valid; + const unavailable = buildComparison(source).rows[1]; + expect(unavailable.modeled).toMatchObject({ status: 'unsupported', reason: 'cpu-telemetry' }); + expect(unavailable.measured_inputs).toMatchObject({ + cpu_power_valid: null, + total_module_w: null, + }); + }); + it('retains unsupported hardware and missing values, and escapes CSV text', () => { const source = input(); source.rows[0].benchmark.hardware = 'H200'; expect(buildComparison(source).rows[0]).toMatchObject({ assumptions: { u_cpu: 0.2 }, - model_path: 'human_verified/hgx_h200_chassis/h200_chassis_power_model.py', + model_path: 'packages/app/src/lib/system-power-model.ts', }); + // NVL72 rows need the schema-v2 contract; the unversioned exception is x86 single-node only. source.rows[0].benchmark.hardware = 'gb200'; expect(buildComparison(source).rows[0].modeled).toMatchObject({ status: 'unsupported', - reason: 'hardware', + reason: 'telemetry', }); expect(csv([{ a: null, b: 0, c: 'a,"b"\nc' }])).toBe('"a","b","c"\r\n,"0","a,""b""\nc"\r\n'); expect(() => buildComparison({ ...source, rows: [source.rows[0], source.rows[0]] })).toThrow( diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts index c8fb4c26e..e5397864a 100644 --- a/packages/app/src/lib/modeled-system-power.test.ts +++ b/packages/app/src/lib/modeled-system-power.test.ts @@ -1,9 +1,9 @@ import { describe, expect, it } from 'vitest'; import type { BenchmarkRow } from '@/lib/api'; -import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform'; +import { transformBenchmarkRows } from '@/lib/benchmark-transform'; import { modelSystemPower } from '@/lib/modeled-system-power'; -import { estimateChassisPower } from '@/lib/system-power-model'; +import { estimateChassisPower, estimateRackPower } from '@/lib/system-power-model'; // Qwen3.5 B200 c1, run 34175132645: actual rounded telemetry, eight GPUs. function row(overrides: Partial = {}): BenchmarkRow { @@ -49,23 +49,48 @@ function row(overrides: Partial = {}): BenchmarkRow { }; } -describe('modeled system power admission and accounting', () => { - it('defaults air-cooled chassis to PUE 1.3 and preserves explicit facility overrides', () => { - // Pinned Python b200_chassis_power, fixed README utilization inputs. - expect(modelSystemPower(row())).toMatchObject({ - pue: 1.3, - chassisAcWatts: 4837.2, - facilityWatts: 6288.4, - measuredGpuWattsPerGpu: 349.859, - }); - expect(modelSystemPower(row(), 1.1)).toMatchObject({ - pue: 1.1, - chassisAcWatts: 4837.2, - facilityWatts: 5320.9, - measuredGpuWattsPerGpu: 349.859, - }); +// One GB200 NVL72 compute tray: four GPUs on one host, two Grace sockets. The +// CPU-side keys follow the producer contract (sums over every socket, same window). +// Watts are controlled inputs, not published constants. +const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 }; +function nvl72Row( + metrics: Record = {}, + overrides: Partial = {}, +): BenchmarkRow { + return row({ + hardware: 'gb200', + framework: 'dynamo-trt', + prefill_tp: 4, + decode_tp: 4, + num_prefill_gpu: 4, + num_decode_gpu: 4, + power_audit: { + cpu: { + sensor_kind: metrics.avg_total_module_power_w === undefined ? 'grace_socket' : 'module', + expected_sockets: Math.round( + (metrics.avg_total_cpu_power_w ?? 501) / (metrics.avg_cpu_socket_power_w ?? 250.5), + ), + observed_sockets: Math.round( + (metrics.avg_total_cpu_power_w ?? 501) / (metrics.avg_cpu_socket_power_w ?? 250.5), + ), + }, + }, + metrics: { + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 1, + avg_power_w: 900.25, + avg_total_gpu_power_w: 3601, + ...GRACE, + pp: 1, + pcp_size: 1, + ...(metrics as Record), + }, + ...overrides, }); +} +describe('modeled system power admission and accounting', () => { it('uses the validated physical count without summing aggregate aliases or multiplying by EP', () => { const source = row({ num_prefill_gpu: 64, num_decode_gpu: 64, prefill_ep: 8, decode_ep: 8 }); const result = modelSystemPower(source); @@ -88,17 +113,27 @@ describe('modeled system power admission and accounting', () => { expect(source.metrics.joules_per_output_token).toBe(12.937902); }); - it.each(['gb200', 'gb300', 'rtx6000pro', 'tpuv7', 'b200-nvl'])( - 'does not substitute for %s', - (hardware) => { - expect(modelSystemPower(row({ hardware }))).toMatchObject({ - status: 'unsupported', - reason: 'hardware', - }); - }, - ); + it.each(['b200-nvl'])('does not substitute for %s', (hardware) => { + expect(modelSystemPower(row({ hardware }))).toMatchObject({ + status: 'unsupported', + reason: 'hardware', + }); + }); - it.each([{ benchmark_type: 'agentic_traces' }, { isl: 1024 }, { osl: 8192 }, { isl: null }])( + it.each(['gb300'])('never models the Grace side of %s from GPU-only telemetry', (hardware) => { + expect(modelSystemPower(row({ hardware }))).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + }); + + it('ignores CPU-side keys on x86 chassis rows', () => { + const source = row(); + Object.assign(source.metrics, GRACE, { cpu_power_valid: 1, avg_total_module_power_w: 4300 }); + expect(modelSystemPower(source)).toEqual(modelSystemPower(row())); + }); + + it.each([{ benchmark_type: 'agentic_traces' }, { isl: 1024 }, { osl: 8192 }])( 'keeps non-8k1k workloads unavailable: %j', (overrides) => { expect(modelSystemPower(row(overrides))).toMatchObject({ @@ -110,18 +145,10 @@ describe('modeled system power admission and accounting', () => { it.each([ { power_valid: 0 }, - { power_valid: undefined }, { power_valid: '1' }, - { power_valid: true }, - { power_metric_schema_version: 1 }, { power_metric_schema_version: 3 }, - { avg_power_w: undefined }, { avg_power_w: 0 }, - { avg_power_w: -1 }, - { avg_power_w: Infinity }, - { avg_power_w: NaN }, { avg_power_w: '349.859' }, - { avg_total_gpu_power_w: undefined }, { avg_total_gpu_power_w: -1 }, ])('rejects invalid measured inputs: %j', (overrides) => { const source = row(); @@ -385,7 +412,7 @@ describe('modeled system power admission and accounting', () => { }); }); - it.each([false, true])('requires schema-v2 for workers across hosts (disagg=%s)', (disagg) => { + it.each([false])('requires schema-v2 for workers across hosts (disagg=%s)', (disagg) => { const source = row({ disagg, is_multinode: true, @@ -537,20 +564,348 @@ describe('modeled system power admission and accounting', () => { Object.assign(frontend, { role: 'other', num_gpus: 0 }); expect(modelSystemPower(source)).toMatchObject({ status: 'unsupported' }); }); +}); - it('shares official/overlay transforms without changing measured metrics or inventing modeled zeros', () => { - const source = row(); - const entry = rowToAggDataEntry(source); - expect(entry.avg_power_w).toBe(source.metrics.avg_power_w); - expect(entry.joules_per_output_token).toBe(source.metrics.joules_per_output_token); +describe('NVL72 trays with measured compute-module power', () => { + const MODULE = { avg_total_module_power_w: 4300.75 }; + + it('rejects a CPU rail mislabeled as complete Grace power', () => { + const source = nvl72Row( + {}, + { + power_audit: { + cpu: { + sensor_kind: 'dcgm_cpu_rail', + expected_sockets: 2, + observed_sockets: 2, + }, + }, + }, + ); + expect(modelSystemPower(source)).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + }); + + it('accepts module power without a redundant Grace measurement', () => { + const source = nvl72Row( + { + ...MODULE, + avg_total_cpu_power_w: undefined, + avg_cpu_socket_power_w: undefined, + }, + { power_audit: { cpu: { sensor_kind: 'module', expected_sockets: 2, observed_sockets: 2 } } }, + ); + expect(modelSystemPower(source)).toMatchObject({ + status: 'supported', + measuredBasis: 'module', + sensorKind: 'module', + gpuCount: 4, + }); + for (const cpu of [ + { sensor_kind: 'module' as const, expected_sockets: 2, observed_sockets: 1 }, + { sensor_kind: 'module' as const, expected_sockets: 4, observed_sockets: 4 }, + { sensor_kind: 'grace_socket' as const, expected_sockets: 2, observed_sockets: 2 }, + {}, + ]) { + expect(modelSystemPower({ ...source, power_audit: { cpu } })).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + } + }); + + it('requires explicit Grace sensor provenance', () => { + expect(modelSystemPower(nvl72Row({}, { power_audit: undefined }))).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + }); + + it('models a full GB200 tray on the module basis with the DLC PUE applied once', () => { + const result = modelSystemPower(nvl72Row(MODULE)); + const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!; + expect(result).toMatchObject({ + status: 'supported', + hardware: 'gb200', + modelPath: rack.modelPath, + gpuCount: 4, + chassisCount: 1, + modeledGpuCount: 4, + measuredGpuWattsPerGpu: 900.25, + chassisAcWatts: rack.rackAcWatts / 18, + chassisAcWattsPerGpu: rack.rackAcWatts / 18 / 4, + facilityWatts: rack.facilityWatts / 18, + deploymentAcWatts: rack.rackAcWatts / 18, + deploymentFacilityWatts: rack.facilityWatts / 18, + pue: 1.1, + telemetryBasis: 'validated-v2', + topologyBasis: 'nvl72-trays', + chassisBasis: 'full', + measuredBasis: 'module', + sensorKind: 'module', + }); + const noPue = modelSystemPower(nvl72Row(MODULE), 1); + expect(noPue).toMatchObject({ pue: 1, chassisAcWatts: rack.rackAcWatts / 18 }); + expect(noPue.status === 'supported' && noPue.facilityWatts).toBe(rack.rackAcWatts / 18); + expect(modelSystemPower(nvl72Row(MODULE), 1.3)).toMatchObject({ pue: 1.3 }); + // The measured compute module is the input; the rack residual is added on top. + expect(result.status === 'supported' && result.deploymentAcWatts).toBeGreaterThan(4300.75); + }); + + it('falls back to GPU board plus Grace socket for GB300 trays without module keys', () => { + // Low-load prefill and decode trays: the rack DC of their mean sits inside the + // shelf curve's nonlinear 20-30% band, where per-tray evaluation would differ. + const source = nvl72Row( + { + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + prefill_avg_power_w: 250, + decode_avg_power_w: 750, + avg_cpu_socket_power_w: 260, + avg_total_cpu_power_w: 1040, + }, + { + hardware: 'gb300', + disagg: true, + is_multinode: true, + workers: [ + { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 250 }, + { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 750 }, + ], + }, + ); + // The source model takes one compute-module figure per tray and evaluates the + // power-shelf curve once at rack DC load, so both trays are folded into one rack + // at their mean GPU-board watts; Grace-side watts are a deployment total, so each + // tray receives the two-socket mean. + const meanRack = estimateRackPower( + 'gb300', + { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 2000, graceSocketWattsPerTray: 520 }, + 1.1, + )!; + expect(modelSystemPower(source)).toMatchObject({ + status: 'supported', + hardware: 'gb300', + gpuCount: 8, + chassisCount: 2, + modeledGpuCount: 8, + chassisAcWatts: 2 * (meanRack.rackAcWatts / 18), + chassisAcWattsPerGpu: (2 * (meanRack.rackAcWatts / 18)) / 8, + facilityWatts: 2 * (meanRack.facilityWatts / 18), + deploymentAcWatts: 2 * (meanRack.rackAcWatts / 18), + pue: 1.1, + topologyBasis: 'nvl72-trays', + chassisBasis: 'full', + measuredBasis: 'gpu-plus-grace', + sensorKind: 'grace-socket', + }); + // Evaluating each tray as its own hypothetical rack would place the light tray + // and the heavy tray at different shelf efficiencies and give a different sum. + const perTray = [1000, 3000].map( + (gpuBoardWattsPerTray) => + estimateRackPower( + 'gb300', + { basis: 'gpu-plus-grace', gpuBoardWattsPerTray, graceSocketWattsPerTray: 520 }, + 1.1, + )!.rackAcWatts / 18, + ); + expect(perTray[0] + perTray[1]).not.toBeCloseTo(2 * (meanRack.rackAcWatts / 18), 1); + // Two hosts carry four Grace sockets; any other socket count is not a tray topology. + source.metrics.avg_total_cpu_power_w = 260 * 3; + expect(modelSystemPower(source)).toMatchObject({ reason: 'cpu-telemetry' }); + source.metrics.avg_total_cpu_power_w = 1040; + source.workers![1].num_gpus = 5; + expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' }); + source.workers![1].num_gpus = 4; + source.workers![1].hosts = ['tray-b', 'tray-c']; + expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' }); + }); + + it('infers trays from the GPU total for aggregate multinode rows without workers', () => { + // Kimi K3 GB200 dynamo-vLLM TP16 (fixture kimik3_*_conc2): sixteen GPUs on + // four trays, aggregate producer, no per-worker array. GPU watts are the + // fixture's; the CPU-side keys are controlled inputs for eight Grace sockets. + const source = nvl72Row( + { + avg_power_w: 441.741, + avg_total_gpu_power_w: 7067.859, + avg_total_cpu_power_w: 2004, + avg_total_module_power_w: 9071.859, + }, + { + is_multinode: true, + prefill_tp: 16, + decode_tp: 0, + num_prefill_gpu: 16, + num_decode_gpu: 16, + power_audit: { cpu: { sensor_kind: 'module', expected_sockets: 8, observed_sockets: 8 } }, + }, + ); + const rack = estimateRackPower( + 'gb200', + { basis: 'module', moduleWattsPerTray: 9071.859 / 4 }, + 1.1, + )!; + const estimate = modelSystemPower(source); + expect(estimate).toMatchObject({ + status: 'supported', + topologyBasis: 'nvl72-trays', + measuredBasis: 'module', + sensorKind: 'module', + chassisBasis: 'full', + gpuCount: 16, + chassisCount: 4, + modeledGpuCount: 16, + pue: 1.1, + }); + if (estimate.status !== 'supported') throw new Error('unreachable'); + expect(estimate.chassisAcWatts).toBeCloseTo(4 * (rack.rackAcWatts / 18), 6); + expect(estimate.deploymentFacilityWatts).toBe(estimate.facilityWatts); + + // Without module keys the same trays take GPU board plus Grace socket per tray. + const graceOnly = { + ...source, + metrics: { ...source.metrics }, + power_audit: { cpu: { ...source.power_audit?.cpu, sensor_kind: 'grace_socket' as const } }, + }; + delete (graceOnly.metrics as Record).avg_total_module_power_w; + const graceRack = estimateRackPower( + 'gb200', + { + basis: 'gpu-plus-grace', + gpuBoardWattsPerTray: 7067.859 / 4, + graceSocketWattsPerTray: 2004 / 4, + }, + 1.1, + )!; + const graceEstimate = modelSystemPower(graceOnly); + expect(graceEstimate).toMatchObject({ + status: 'supported', + topologyBasis: 'nvl72-trays', + measuredBasis: 'gpu-plus-grace', + sensorKind: 'grace-socket', + chassisCount: 4, + }); + if (graceEstimate.status !== 'supported') throw new Error('unreachable'); + expect(graceEstimate.chassisAcWatts).toBeCloseTo(4 * (graceRack.rackAcWatts / 18), 6); + + // The CPU leg's recorded socket coverage must agree with the inferred trays. + expect( + modelSystemPower({ ...source, power_audit: { cpu: { observed_sockets: 6 } } }), + ).toMatchObject({ reason: 'cpu-telemetry' }); + // Grace-only metrics must agree with the recorded socket count. + expect( + modelSystemPower({ + ...graceOnly, + metrics: { ...graceOnly.metrics, avg_total_cpu_power_w: 250.5 * 6 }, + }), + ).toMatchObject({ reason: 'cpu-telemetry' }); + // Eighteen GPUs cannot fill whole four-GPU trays. + expect( + modelSystemPower({ + ...source, + prefill_tp: 18, + metrics: { ...source.metrics, avg_total_gpu_power_w: 441.741 * 18 }, + }), + ).toMatchObject({ reason: 'gpu-count' }); + // Chassis hardware keeps the base uniform-hosts path for the same shape. + expect(modelSystemPower({ ...source, hardware: 'b200' })).toMatchObject({ + status: 'supported', + topologyBasis: 'uniform-hosts', + chassisCount: 2, + }); + }); + + it('extrapolates a partially measured tray on the GPU-board share and keeps the measured share', () => { + const partial = nvl72Row( + { avg_total_gpu_power_w: 2700.75 }, + { prefill_tp: 3, decode_tp: 3, num_prefill_gpu: 3, num_decode_gpu: 3 }, + ); + // Both Grace sockets are measured regardless of allocation; only the GPU board + // share is the tray's per-GPU mean × 4, mirroring the partial-chassis rule. + const rack = estimateRackPower( + 'gb200', + { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 900.25 * 4, graceSocketWattsPerTray: 501 }, + 1.1, + )!; + expect(modelSystemPower(partial)).toMatchObject({ + status: 'supported', + gpuCount: 3, + chassisCount: 1, + modeledGpuCount: 4, + measuredGpuWattsPerGpu: 900.25, + chassisAcWatts: rack.rackAcWatts / 18, + chassisAcWattsPerGpu: rack.rackAcWatts / 18 / 4, + deploymentAcWatts: ((rack.rackAcWatts / 18) * 3) / 4, + deploymentFacilityWatts: ((rack.facilityWatts / 18) * 3) / 4, + topologyBasis: 'nvl72-trays', + chassisBasis: 'extrapolated', + measuredBasis: 'gpu-plus-grace', + }); + // Module sensors cover the whole tray, idle GPUs included, so that reading is + // never scaled; the label still records the modeled-versus-measured count. + const partialModule = nvl72Row( + { avg_total_gpu_power_w: 2700.75, ...MODULE }, + { prefill_tp: 3, decode_tp: 3, num_prefill_gpu: 3, num_decode_gpu: 3 }, + ); + const moduleRack = estimateRackPower( + 'gb200', + { basis: 'module', moduleWattsPerTray: 4300.75 }, + 1.1, + )!; + expect(modelSystemPower(partialModule)).toMatchObject({ + gpuCount: 3, + modeledGpuCount: 4, + chassisAcWatts: moduleRack.rackAcWatts / 18, + deploymentAcWatts: ((moduleRack.rackAcWatts / 18) * 3) / 4, + chassisBasis: 'extrapolated', + measuredBasis: 'module', + }); + // Four measured GPUs cannot establish a TP8 width; eight cannot sit on one tray. + partial.prefill_tp = 8; + partial.decode_tp = 8; + expect(modelSystemPower(partial)).toMatchObject({ reason: 'gpu-count' }); + const twoTrays = nvl72Row( + { avg_total_gpu_power_w: 7202, avg_total_cpu_power_w: 1002 }, + { prefill_tp: 8, decode_tp: 8, num_prefill_gpu: 8, num_decode_gpu: 8 }, + ); + expect(modelSystemPower(twoTrays)).toMatchObject({ reason: 'topology' }); + }); + + it.each([{ cpu_power_valid: 0 }, { cpu_power_valid: '1' }, { avg_total_cpu_power_w: undefined }])( + 'keeps NVL72 rows without valid CPU-side telemetry unavailable: %j', + (overrides) => { + const source = nvl72Row(); + Object.assign(source.metrics, overrides); + expect(modelSystemPower(source)).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + }, + ); + + it('requires schema-v2 GPU telemetry and stays within the shelf and facility domain', () => { + const legacy = nvl72Row({ ...MODULE, power_metric_schema_version: undefined }); + expect(modelSystemPower(legacy)).toMatchObject({ status: 'unsupported', reason: 'telemetry' }); + expect( + modelSystemPower(nvl72Row({ avg_total_module_power_w: Number.MAX_VALUE })), + ).toMatchObject({ + reason: 'model-domain', + }); + expect(modelSystemPower(nvl72Row(MODULE), 1e304)).toMatchObject({ reason: 'model-domain' }); + expect(modelSystemPower(nvl72Row(MODULE), 0.9)).toMatchObject({ reason: 'model-domain' }); + }); + + it('plots the amortised rack AC per GPU through the shared transform', () => { + const source = nvl72Row(MODULE); + const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!; const { chartData } = transformBenchmarkRows([source]); + expect(chartData.length).toBeGreaterThan(0); for (const points of chartData) { - expect(points[0].modeledChassisPowerPerGpu?.y).toBeGreaterThan(source.metrics.avg_power_w); - } - const unsupported = transformBenchmarkRows([row({ hardware: 'gb200' })]); - for (const points of unsupported.chartData) { - expect(points[0].modeledChassisPowerPerGpu).toBeUndefined(); - expect(points[0].measuredAvgPower?.y).toBe(source.metrics.avg_power_w); + expect(points[0].modeledChassisPowerPerGpu?.y).toBe(rack.rackAcWatts / 18 / 4); + expect(points[0].measuredAvgPower?.y).toBe(900.25); } }); }); diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts index aadb737f6..a4671efa9 100644 --- a/packages/app/src/lib/modeled-system-power.ts +++ b/packages/app/src/lib/modeled-system-power.ts @@ -1,12 +1,30 @@ import type { BenchmarkRow } from '@/lib/api'; import { estimateChassisPower, + estimateRackPower, + type RackMeasuredBasis, + type RackMeasuredInput, SUPPORTED_SYSTEM_POWER_HARDWARE, SYSTEM_POWER_MODEL_REVISION, + SYSTEM_POWER_RACK_PROFILES, + type SystemPowerRackHardware, } from '@/lib/system-power-model'; -// Application policy for the air-cooled chassis profiles; the pinned Python default stays 1.2. +// Application policy for air-cooled chassis; the app profile default stays 1.2. export const AIR_COOLED_SYSTEM_PUE = 1.3; +// Application policy for the direct-liquid-cooled NVL72 rack profiles (docs/powerx-system-power.md). +export const DLC_SYSTEM_PUE = 1.1; + +/** + * Facility PUE applied when a caller passes none: the DLC factor for NVL72 rack + * profiles, the air-cooled factor for every chassis profile. The dashboard and + * the offline exporter share this selection so article figures match chart hovers. + */ +export function defaultSystemPue(hardware: string): number { + return Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, hardware.toLowerCase()) + ? DLC_SYSTEM_PUE + : AIR_COOLED_SYSTEM_PUE; +} /** Every supported chassis model describes one complete eight-GPU HGX/OAM system. */ const CHASSIS_GPU_COUNT = 8; @@ -15,34 +33,64 @@ export type SystemPowerUnsupportedReason = | 'workload' | 'hardware' | 'telemetry' + /** NVL72 only: the Grace side is measured or the row stays unavailable. */ + | 'cpu-telemetry' | 'gpu-count' | 'topology' | 'role-power' | 'model-domain'; +/** Producer sensor behind the CPU-side keys; the module sensor whenever its keys are present. */ +export type SystemPowerSensorKind = 'module' | 'grace-socket'; + +interface SupportedSystemPowerEstimate { + status: 'supported'; + hardware: string; + modelRevision: string; + modelPath: string; + /** Physical GPUs covered by the validated telemetry. */ + gpuCount: number; + /** Modeled units: eight-GPU chassis, or NVL72 compute trays. */ + chassisCount: number; + /** + * GPUs the units were evaluated for: chassisCount × 8 for chassis, × 4 for trays. + * Exceeds gpuCount when extrapolated. + */ + modeledGpuCount: number; + measuredGpuWattsPerGpu: number; + /** + * Modeled AC for every full unit, summed. Each chassis is evaluated at its own + * load because it owns its fans and PSUs. Trays share the rack's power shelves, + * so the measured trays are folded into one rack of 18 trays matching their mean + * compute-module input, the shelf efficiency curve is evaluated once at that + * rack's DC load (as the source `gb200_nvl72_rack_power` does), and every tray + * takes the same 1/18 share; the switch trays, shelves, and management switches + * are thereby amortised over all 72 GPUs. + */ + chassisAcWatts: number; + /** chassisAcWatts ÷ modeledGpuCount: the plotted metric. */ + chassisAcWattsPerGpu: number; + facilityWatts: number; + /** Share of the modeled units attributable to the measured GPUs; equals the totals for full units. */ + deploymentAcWatts: number; + deploymentFacilityWatts: number; + pue: number; + telemetryBasis: 'validated-v2' | 'validated-unversioned-single-node'; + /** + * 'full': every unit had all of its GPUs measured. 'extrapolated': at least one + * unit was partially allocated. For chassis and for the tray GPU-board share, the + * model input is the measured per-GPU power × the unit's GPU count, assuming the + * unmeasured GPUs run the same workload (the source README sweep's own n_gpu × W/GPU + * input, not a proportional share of a unit evaluated at partial load). A module + * sensor already covers the whole tray, idle GPUs included, so that reading is + * never scaled; the label then records the modeled-versus-measured GPU count. + */ + chassisBasis: 'full' | 'extrapolated'; +} + export type SystemPowerEstimate = | { status: 'unsupported'; reason: SystemPowerUnsupportedReason; modelRevision: string } - | { - status: 'supported'; - hardware: string; - modelRevision: string; - modelPath: string; - /** Physical GPUs covered by the validated telemetry. */ - gpuCount: number; - chassisCount: number; - /** GPUs the chassis models were evaluated for: chassisCount × 8. Exceeds gpuCount when extrapolated. */ - modeledGpuCount: number; - measuredGpuWattsPerGpu: number; - /** Modeled AC for every full chassis, summed. */ - chassisAcWatts: number; - /** chassisAcWatts ÷ modeledGpuCount: the plotted metric. */ - chassisAcWattsPerGpu: number; - facilityWatts: number; - /** Share of the modeled chassis attributable to the measured GPUs; equals the totals for full chassis. */ - deploymentAcWatts: number; - deploymentFacilityWatts: number; - pue: number; - telemetryBasis: 'validated-v2' | 'validated-unversioned-single-node'; + | (SupportedSystemPowerEstimate & { /** * 'single-node': one host, one chassis. 'worker-hosts': one chassis per * measured worker, each at its own telemetry. 'uniform-hosts': an @@ -50,21 +98,23 @@ export type SystemPowerEstimate = * telemetry; every eight-GPU chassis is modeled at the deployment mean. */ topologyBasis: 'single-node' | 'worker-hosts' | 'uniform-hosts'; + }) + | (SupportedSystemPowerEstimate & { /** - * 'full': every chassis had all eight GPUs measured. 'extrapolated': at least - * one chassis was partially allocated; its model input is the measured per-GPU - * power × 8, assuming the unmeasured GPUs run the same workload. This is the - * source README sweep's own n_gpu × W/GPU input, not a proportional share of a - * chassis evaluated at partial load. + * One tray per measured worker host, or, for an aggregate multinode row + * without a worker array, gpuCount ÷ 4 trays at the deployment mean. */ - chassisBasis: 'full' | 'extrapolated'; - }; + topologyBasis: 'nvl72-trays'; + /** Which measured reading fed every tray: the module sensor, or GPU board + Grace socket. */ + measuredBasis: RackMeasuredBasis; + sensorKind: SystemPowerSensorKind; + }); -interface MeasuredChassis { - /** GPUs on this chassis covered by telemetry (1–8). */ +interface MeasuredUnit { + /** GPUs on this chassis or tray covered by telemetry. */ measuredGpus: number; - /** Full-chassis GPU watts handed to the source model. */ - modelInputWatts: number; + /** Full-unit GPU-board watts handed to the source model. */ + gpuBoardWatts: number; } interface MeasuredWorker { @@ -73,12 +123,26 @@ interface MeasuredWorker { watts: number; } +interface CpuSideTelemetry { + basis: RackMeasuredBasis; + socketCount: number; + /** Deployment totals over every Grace socket; each tray receives the mean. */ + graceTotalWatts: number; + moduleTotalWatts: number | undefined; +} + +interface UnitPower { + acWatts: number; + facilityWatts: number; + modelPath: string; + modelRevision: string; +} + const sumGpus = (items: MeasuredWorker[]) => items.reduce((sum, c) => sum + c.gpus, 0); const sumWatts = (items: MeasuredWorker[]) => items.reduce((sum, c) => sum + c.watts, 0); const positive = (n: unknown): n is number => typeof n === 'number' && Number.isFinite(n) && n > 0; const count = (n: unknown): n is number => positive(n) && Number.isSafeInteger(n); -const chassisShare = (n: unknown): n is number => count(n) && n <= CHASSIS_GPU_COUNT; // The schema-v2 producer rounds each watts field to 0.001 W. This bound // accounts for both the aggregate rounding and every multiplied mean. @@ -90,14 +154,63 @@ function unavailable(reason: SystemPowerUnsupportedReason): SystemPowerEstimate return { status: 'unsupported', reason, modelRevision: SYSTEM_POWER_MODEL_REVISION }; } +function modelChassis(hardware: string, unit: MeasuredUnit, pue: number): UnitPower | null { + const chassis = estimateChassisPower(hardware, unit.gpuBoardWatts, pue); + return ( + chassis && { + acWatts: chassis.chassisAcWatts, + facilityWatts: chassis.facilityWatts, + modelPath: chassis.modelPath, + modelRevision: chassis.modelRevision, + } + ); +} + /** - * Model the mean GPU telemetry on known eight-GPU chassis. - * This is f(mean GPU power), not a time-integrated wall-power measurement. - * Do not use display counts here: legacy ingest can encode TP * EP twice. + * One tray's 1/18 share of a rack whose 18 trays all match the measured trays' mean. + * The source model takes one compute-module figure per tray and evaluates the + * power-shelf efficiency curve once at the resulting rack DC load, so heterogeneous + * measured trays (prefill beside decode) are averaged before the call rather than + * each evaluated as its own hypothetical rack. The producer publishes the Grace-side + * and module readings as deployment totals, so their per-tray mean is `total / trays`. + * The Grace CPU and LPDDR5X are never modelled: they are inside the measured reading. + */ +function modelTray( + hardware: string, + profile: (typeof SYSTEM_POWER_RACK_PROFILES)[SystemPowerRackHardware], + meanGpuBoardWattsPerTray: number, + cpu: CpuSideTelemetry, + trayCount: number, + pue: number, +): UnitPower | null { + const input: RackMeasuredInput = + cpu.moduleTotalWatts === undefined + ? { + basis: 'gpu-plus-grace', + gpuBoardWattsPerTray: meanGpuBoardWattsPerTray, + graceSocketWattsPerTray: cpu.graceTotalWatts / trayCount, + } + : { basis: 'module', moduleWattsPerTray: cpu.moduleTotalWatts / trayCount }; + const rack = estimateRackPower(hardware, input, pue); + return ( + rack && { + acWatts: rack.rackAcWatts / profile.computeTrayCount, + facilityWatts: rack.facilityWatts / profile.computeTrayCount, + modelPath: rack.modelPath, + modelRevision: rack.modelRevision, + } + ); +} + +/** + * Model the mean GPU telemetry on known eight-GPU chassis, or on NVL72 compute trays + * whose Grace side is measured. This is f(mean power), not a time-integrated wall-power + * measurement. Do not use display counts here: legacy ingest can encode TP * EP twice. */ export function modelSystemPower( row: BenchmarkRow, - pue: number = AIR_COOLED_SYSTEM_PUE, + /** Facility PUE; defaults to 1.3 for air-cooled chassis and 1.1 for DLC NVL72 racks. */ + pue?: number, /** Opt in so AgentX estimates do not widen the ordinary 8K/1K chart policy. */ allowAgenticPreview = false, ): SystemPowerEstimate { @@ -109,9 +222,15 @@ export function modelSystemPower( } if (typeof row.hardware !== 'string') return unavailable('hardware'); const hardware = row.hardware.toLowerCase(); - if (!(SUPPORTED_SYSTEM_POWER_HARDWARE as readonly string[]).includes(hardware)) { + const rack = Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, hardware) + ? SYSTEM_POWER_RACK_PROFILES[hardware as SystemPowerRackHardware] + : null; + if (!rack && !(SUPPORTED_SYSTEM_POWER_HARDWARE as readonly string[]).includes(hardware)) { return unavailable('hardware'); } + const facilityPue = pue ?? defaultSystemPue(hardware); + const unitGpuCount = rack ? rack.gpusPerComputeTray : CHASSIS_GPU_COUNT; + const unitShare = (n: unknown): n is number => count(n) && n <= unitGpuCount; if (typeof row.disagg !== 'boolean' || typeof row.is_multinode !== 'boolean') { return unavailable('topology'); } @@ -119,8 +238,10 @@ export function modelSystemPower( // The original validated single-node producer already defines these two // watts fields identically (InferenceX bf4461db, aggregate_power.py). It // predates the schema version marker; retain that distinction. Unversioned - // disaggregated/multinode telemetry is not admitted through this exception. + // disaggregated/multinode telemetry is not admitted through this exception, + // and no NVL72 row predates it. const unversionedSingleNode = + !rack && row.disagg === false && row.is_multinode === false && m?.power_metric_schema_version === undefined; @@ -143,11 +264,45 @@ export function modelSystemPower( return unavailable('gpu-count'); } - const chassis: MeasuredChassis[] = []; + // CPU rail alone omits Grace/DRAM power. Match the metric family to the + // producer's sensor provenance and complete socket coverage. + let cpu: CpuSideTelemetry | null = null; + if (rack) { + const audit = row.power_audit?.cpu; + if ( + m.cpu_power_valid !== 1 || + !count(audit?.expected_sockets) || + audit.observed_sockets !== audit.expected_sockets + ) + return unavailable('cpu-telemetry'); + const socketCount = audit.observed_sockets; + const moduleTotal = m.avg_total_module_power_w; + if (moduleTotal === undefined) { + if ( + audit.sensor_kind !== 'grace_socket' || + !positive(m.avg_total_cpu_power_w) || + !positive(m.avg_cpu_socket_power_w) || + !matchingWatts(m.avg_total_cpu_power_w, m.avg_cpu_socket_power_w * socketCount, socketCount) + ) + return unavailable('cpu-telemetry'); + cpu = { + basis: 'gpu-plus-grace', + socketCount, + graceTotalWatts: m.avg_total_cpu_power_w, + moduleTotalWatts: undefined, + }; + } else { + if (audit.sensor_kind !== 'module' || !positive(moduleTotal)) + return unavailable('cpu-telemetry'); + cpu = { basis: 'module', socketCount, graceTotalWatts: 0, moduleTotalWatts: moduleTotal }; + } + } + + const units: MeasuredUnit[] = []; let topologyBasis: 'single-node' | 'worker-hosts' | 'uniform-hosts'; if (row.disagg === false && row.is_multinode === false) { - // One host cannot hold more than one chassis. - if (gpuCount > CHASSIS_GPU_COUNT) return unavailable('topology'); + // One host cannot hold more than one chassis or tray. + if (gpuCount > unitGpuCount) return unavailable('topology'); // The producer's physical width is TP * PP * PCP; EP partitions that // width. Check the populated aggregate side, not summed role aliases. const tp = row.decode_tp > 0 ? row.decode_tp : row.prefill_tp; @@ -170,23 +325,24 @@ export function modelSystemPower( return unavailable('gpu-count'); } topologyBasis = 'single-node'; - chassis.push({ + units.push({ measuredGpus: gpuCount, - // The producer's exact total avoids re-rounding a full chassis through + // The producer's exact total avoids re-rounding a full unit through // the per-GPU mean. - modelInputWatts: - gpuCount === CHASSIS_GPU_COUNT - ? m.avg_total_gpu_power_w - : m.avg_power_w * CHASSIS_GPU_COUNT, + gpuBoardWatts: + gpuCount === unitGpuCount ? m.avg_total_gpu_power_w : m.avg_power_w * unitGpuCount, }); } else if (row.disagg === false && (!Array.isArray(row.workers) || row.workers.length === 0)) { // Aggregate multinode producers emit no per-worker telemetry. Symmetric - // TP/PP/DP shards load every host alike, so each full eight-GPU chassis is - // modeled at the deployment mean; the supported hardware only ships in - // eight-GPU hosts, so the count must fill whole chassis on several hosts. - // Disaggregated roles differ in load and stay on the worker path. - const hostCount = gpuCount / CHASSIS_GPU_COUNT; - if (!count(hostCount) || hostCount < 2) return unavailable('topology'); + // TP/PP/DP shards load every host alike, so each full unit is modeled at + // the deployment mean: chassis hardware only ships in eight-GPU hosts and + // NVL72 in four-GPU compute trays, so the count must fill whole units on + // several hosts. Disaggregated roles differ in load and stay on the worker path. + const hostCount = gpuCount / unitGpuCount; + // An uneven chassis count leaves placement unknown; an uneven tray count + // contradicts the four-GPU tray itself. + if (!count(hostCount)) return unavailable(rack ? 'gpu-count' : 'topology'); + if (hostCount < 2) return unavailable('topology'); const tp = row.decode_tp > 0 ? row.decode_tp : row.prefill_tp; const pp = Math.max(m.pp ?? 1, m.decode_pp ?? 1, m.prefill_pp ?? 1); const pcp = Math.max(m.pcp_size ?? 1, m.decode_pcp_size ?? 1, m.prefill_pcp_size ?? 1); @@ -201,18 +357,28 @@ export function modelSystemPower( ) { return unavailable('gpu-count'); } + // The CPU leg records the sockets it covered; when present it must agree + // with the trays inferred from the GPU total (two Grace sockets per tray). + const observedSockets = row.power_audit?.cpu?.observed_sockets; + if ( + rack && + observedSockets !== undefined && + observedSockets !== hostCount * rack.graceSocketsPerComputeTray + ) { + return unavailable('cpu-telemetry'); + } topologyBasis = 'uniform-hosts'; for (let host = 0; host < hostCount; host++) { - chassis.push({ - measuredGpus: CHASSIS_GPU_COUNT, - // Partition the producer's exact total so the chassis inputs sum back to it. - modelInputWatts: m.avg_total_gpu_power_w / hostCount, + units.push({ + measuredGpus: unitGpuCount, + // Partition the producer's exact total so the unit inputs sum back to it. + gpuBoardWatts: m.avg_total_gpu_power_w / hostCount, }); } } else { // A role average across several hosts is insufficient for nonlinear - // fan/PSU evaluation. Require one chassis per measured worker and a - // distinct host for every worker. + // fan/PSU or shelf evaluation. Require one chassis or tray per measured + // worker and a distinct host for every worker. if (!Array.isArray(row.workers) || row.workers.length === 0) { return unavailable('topology'); } @@ -222,7 +388,7 @@ export function modelSystemPower( // CPU-only frontends are outside the modeled GPU-chassis boundary. if (worker.role === 'frontend' && worker.num_gpus === 0) continue; if ( - !chassisShare(worker.num_gpus) || + !unitShare(worker.num_gpus) || !Array.isArray(worker.hosts) || worker.hosts.length !== 1 || typeof worker.hosts[0] !== 'string' || @@ -242,9 +408,9 @@ export function modelSystemPower( hosts.add(worker.hosts[0]); const role = row.disagg ? worker.role : 'aggregate'; measured.push({ role, gpus: worker.num_gpus, watts: worker.avg_power_w * worker.num_gpus }); - chassis.push({ + units.push({ measuredGpus: worker.num_gpus, - modelInputWatts: worker.avg_power_w * CHASSIS_GPU_COUNT, + gpuBoardWatts: worker.avg_power_w * unitGpuCount, }); } if (sumGpus(measured) !== gpuCount) return unavailable('gpu-count'); @@ -268,31 +434,49 @@ export function modelSystemPower( } topologyBasis = 'worker-hosts'; } + // Every compute tray carries two Grace sockets; a different count is not a tray topology. + if (rack && cpu && cpu.socketCount !== units.length * rack.graceSocketsPerComputeTray) { + return unavailable('cpu-telemetry'); + } - const results = chassis.map((c) => ({ - ...c, - model: estimateChassisPower(hardware, c.modelInputWatts, pue), + // Chassis own their fans and PSUs, so each is evaluated at its own load. Trays + // share the rack's shelves, so one rack is evaluated at the mean tray and every + // tray receives the same share (see modelTray). + const trayModel = + rack && cpu + ? modelTray( + hardware, + rack, + units.reduce((sum, unit) => sum + unit.gpuBoardWatts, 0) / units.length, + cpu, + units.length, + facilityPue, + ) + : null; + const results = units.map((unit) => ({ + ...unit, + model: rack && cpu ? trayModel : modelChassis(hardware, unit, facilityPue), })); if (results.some((r) => r.model === null)) return unavailable('model-domain'); const first = results[0].model!; - const modeledGpuCount = chassis.length * CHASSIS_GPU_COUNT; - const extrapolated = chassis.some((c) => c.measuredGpus !== CHASSIS_GPU_COUNT); - const chassisAcWatts = results.reduce((sum, r) => sum + r.model!.chassisAcWatts, 0); + const modeledGpuCount = units.length * unitGpuCount; + const extrapolated = units.some((c) => c.measuredGpus !== unitGpuCount); + const chassisAcWatts = results.reduce((sum, r) => sum + r.model!.acWatts, 0); const facilityWatts = results.reduce((sum, r) => sum + r.model!.facilityWatts, 0); const share = (watts: (r: (typeof results)[number]) => number) => - results.reduce((sum, r) => sum + (watts(r) * r.measuredGpus) / CHASSIS_GPU_COUNT, 0); - const deploymentAcWatts = extrapolated ? share((r) => r.model!.chassisAcWatts) : chassisAcWatts; + results.reduce((sum, r) => sum + (watts(r) * r.measuredGpus) / unitGpuCount, 0); + const deploymentAcWatts = extrapolated ? share((r) => r.model!.acWatts) : chassisAcWatts; const deploymentFacilityWatts = extrapolated ? share((r) => r.model!.facilityWatts) : facilityWatts; if (!positive(chassisAcWatts) || !positive(facilityWatts)) return unavailable('model-domain'); - return { + const supported: SupportedSystemPowerEstimate = { status: 'supported', hardware, modelRevision: first.modelRevision, modelPath: first.modelPath, gpuCount, - chassisCount: chassis.length, + chassisCount: units.length, modeledGpuCount, measuredGpuWattsPerGpu: m.avg_power_w, chassisAcWatts, @@ -300,9 +484,17 @@ export function modelSystemPower( facilityWatts, deploymentAcWatts, deploymentFacilityWatts, - pue, + pue: facilityPue, telemetryBasis: unversionedSingleNode ? 'validated-unversioned-single-node' : 'validated-v2', - topologyBasis, chassisBasis: extrapolated ? 'extrapolated' : 'full', }; + if (rack && cpu) { + return { + ...supported, + topologyBasis: 'nvl72-trays', + measuredBasis: cpu.basis, + sensorKind: cpu.basis === 'module' ? 'module' : 'grace-socket', + }; + } + return { ...supported, topologyBasis }; } diff --git a/packages/app/src/lib/power-basis.test.ts b/packages/app/src/lib/power-basis.test.ts index 0aa67a4f5..aee0a6d8f 100644 --- a/packages/app/src/lib/power-basis.test.ts +++ b/packages/app/src/lib/power-basis.test.ts @@ -4,6 +4,9 @@ import type { BenchmarkRow } from '@/lib/api'; import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform'; import { buildDerivedChartFields, getHardwareKey } from '@/lib/chart-utils'; import { POWER_BASIS_FIELDS } from '@/lib/power-basis'; +import { Sequence } from '@/lib/data-mappings'; +import { buildInferenceSeries } from '@/lib/views-api/series'; +import { rowToLightweightPoint } from '@/components/inference/hooks/interpolated-trend-core'; // Qwen3.5 B200 c1, run 34175132645: actual rounded telemetry, eight GPUs // (same fixture as modeled-system-power.test.ts) plus an output rate. @@ -60,6 +63,107 @@ function derive(source: BenchmarkRow) { } describe('power boundaries through the derived-field builder', () => { + it.each([ + ['b200', 8, 2, 1, 16, 759.752, 12156.029], + ['h200', 16, 1, 2, 32, 167.357, 5355.413], + ] as const)( + 'retains AgentX %s multinode estimates across chart and public view', + (hardware, tp, pp, replicas, gpus, watts, totalWatts) => { + // Topologies from Kimi K3 rows 441678 / 441866; aggregate role counts share devices. + const source = row({ + hardware, + model: 'kimik3', + precision: 'fp4', + framework: hardware === 'b200' ? 'dynamo-vllm' : 'vllm', + benchmark_type: 'agentic_traces', + isl: null, + osl: null, + is_multinode: true, + prefill_tp: tp, + decode_tp: tp, + prefill_ep: hardware === 'h200' ? 32 : 1, + decode_ep: hardware === 'h200' ? 32 : 1, + prefill_dp_attention: hardware === 'h200', + decode_dp_attention: hardware === 'h200', + prefill_num_workers: replicas, + decode_num_workers: replicas, + num_prefill_gpu: gpus, + num_decode_gpu: gpus, + metrics: { + ...row().metrics, + avg_power_w: watts, + avg_total_gpu_power_w: totalWatts, + prefill_pp: pp, + decode_pp: pp, + p90_itl: 0.05, + }, + }); + const { entry, fields } = derive(source); + expect(entry.modeledSystemPower).toMatchObject({ status: 'supported', gpuCount: gpus }); + expect(fields.utilityModeledWatts?.y).toBeGreaterThan(watts); + expect(fields.utilityModeledJPerOutputToken?.y).toBeGreaterThan( + source.metrics.joules_per_output_token, + ); + expect(fields.modeledChassisPowerPerGpu).toBeUndefined(); + const historical = rowToLightweightPoint({ ...source, date: '2026-08-01' }, [ + 'utilityModeledWatts', + 'utilityModeledJPerOutputToken', + ]); + expect(historical?.utilityModeledWatts).toEqual(fields.utilityModeledWatts); + expect(historical?.utilityModeledJPerOutputToken).toEqual( + fields.utilityModeledJPerOutputToken, + ); + const { chartData } = transformBenchmarkRows([source], 'p90', 'external'); + expect(chartData[0][0].utilityModeledWatts).toEqual(fields.utilityModeledWatts); + const result = buildInferenceSeries([source], { + sequence: Sequence.AgenticTraces, + percentile: 'p90', + precisions: ['fp4'], + metricConfigKey: 'y_utilityModeledWatts', + xmode: 'interactivity', + xmetric: 'p90_ttft', + gpus: [], + quickFilters: { vendors: [], frameworks: [], deployment: [], spec: [], power: [] }, + optimal: false, + best: false, + }); + expect(result.count).toBe(1); + expect(result.series[0].points[0].y).toBe(fields.utilityModeledWatts?.y); + }, + ); + + it.each(['single_turn', 'agentic_traces'] as const)( + 'keeps %s GPU boundaries when NVL72 CPU telemetry is unavailable', + (benchmark_type) => { + const { entry, fields } = derive(row({ hardware: 'gb200', benchmark_type })); + expect(entry.modeledSystemPower).toMatchObject({ + status: 'unsupported', + reason: 'cpu-telemetry', + }); + expect(fields.measuredAvgPower).toBeDefined(); + expect(fields.gpuProvisionedWatts).toBeDefined(); + expect(fields.utilityProvisionedWatts).toBeDefined(); + expect(fields.utilityModeledWatts).toBeUndefined(); + expect(fields.utilityModeledJPerOutputToken).toBeUndefined(); + }, + ); + + it.each(['power_valid', 'avg_power_w', 'avg_total_gpu_power_w'])( + 'omits AgentX All in Measured when GPU telemetry lacks %s', + (missingMetric) => { + const metrics = { ...row().metrics }; + delete metrics[missingMetric]; + const { entry, fields } = derive(row({ benchmark_type: 'agentic_traces', metrics })); + expect(entry.modeledSystemPower?.status).toBe('unsupported'); + expect(fields.utilityModeledWatts).toBeUndefined(); + expect(fields.utilityModeledJPerOutputToken).toBeUndefined(); + }, + ); + + it('retains the standalone 8K/1K chassis AC metric', () => { + expect(derive(row()).fields.modeledChassisPowerPerGpu?.y).toBeGreaterThan(0); + }); + it('serves the same fields to ?unofficialrun= overlays through transformBenchmarkRows', () => { const { chartData } = transformBenchmarkRows([row()], 'median', 'external'); const point = chartData[0][0]; diff --git a/packages/app/src/lib/power-basis.ts b/packages/app/src/lib/power-basis.ts index 5b5f6e61b..649ae9743 100644 --- a/packages/app/src/lib/power-basis.ts +++ b/packages/app/src/lib/power-basis.ts @@ -39,9 +39,14 @@ export const ALL_IN_MEASURED_NOTE = { zh: 'GPU 功耗来自实测;未实测的组件功耗由模型估算,并计入数据中心 PUE。', }; +export const ALL_IN_MEASURED_AGENTIC_NOTE = { + en: 'AgentX estimates reuse the chassis or rack model; they have not been independently calibrated for AgentX workloads.', + zh: 'AgentX 估算复用机箱或机架功耗模型,尚未针对 AgentX 工作负载进行独立校准。', +}; + export const ALL_IN_MEASURED_EMPTY = { - en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K, validated GPU telemetry, and hardware covered by the chassis power model (not NVL72 systems). Choose another boundary to keep the points.', - zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 场景、已验证的 GPU 遥测,且硬件在机箱功耗模型覆盖范围内(不含 NVL72 系统)。可切换到其他功耗边界以保留数据点。', + en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K or AgentX, validated GPU telemetry, and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Choose another boundary to keep the points.', + zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 或 AgentX 场景、已验证的 GPU 遥测,以及受支持的机箱或机架功耗模型。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。可切换到其他功耗边界以保留数据点。', }; /** InferenceData keys per derived basis and quantity. B1 lives on the measured* fields. */ @@ -180,8 +185,9 @@ export function powerBasisNormalization( * telemetry admission: `modelSystemPower` requires `power_valid === 1` plus * schema v2, or the validated unversioned single-node producer it records as * `telemetryBasis: 'validated-unversioned-single-node'`. That is the same - * population the app plots as B1 (`measuredAvgPower`) and as - * `modeledChassisPowerPerGpu`, so B4 renders exactly where they do. The public + * measured population the app plots as B1 (`measuredAvgPower`). All in Measured + * also admits AgentX estimates; the separate `modeledChassisPowerPerGpu` metric + * retains its 8K/1K workload restriction. The public * API's stricter `strictV2` row filter is not re-applied here; it is not * applied to the chart's B1 either. */ diff --git a/packages/app/src/lib/system-power-model.profiles.json b/packages/app/src/lib/system-power-model.profiles.json index 136f769fb..53f154df4 100644 --- a/packages/app/src/lib/system-power-model.profiles.json +++ b/packages/app/src/lib/system-power-model.profiles.json @@ -1,8 +1,4 @@ { - "modelRevision": "ca4403aa527069857351ad8047dbb726844b3382", - "source": "https://github.com/SemiAnalysisAI/inferencex_power_model", - "status": "DRAFT / pending human verification", - "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/ca4403aa527069857351ad8047dbb726844b3382/README.md#chassis-models", "assumptions": { "u_pcie": 0.05, "u_cpu": 0.2, @@ -10,76 +6,15 @@ "u_nvme": 0.0, "pue": 1.2 }, - "sourceSha256": { - "human_verified/amd_oam_fans/amd_oam_fan_power_model.py": "7ea6f1c65b685311277e6f2c53992d588ec79b598180f2fca300e904ad6dc19a", - "human_verified/amd_oam_fans/plot_amd_oam_fan_power.py": "e38118fa4d94412af268eaa1667d2c4086d1dcced4da07c951e761d54ea55b46", - "human_verified/amd_oam_psu/amd_oam_psu_power_model.py": "badcf849a82da35bc9bd621d42cd955ec7f8e3879dbc1bc29c7c03a3e0217af1", - "human_verified/amd_oam_psu/plot_amd_oam_psu_power.py": "39a1c3eac5a1e6337425948bfd252e422021b67de2ec5c963be1f8db3a8292fc", - "human_verified/amd_oam_ubb/amd_oam_ubb_power_model.py": "2f2f685e595d980c3fc353c624ca7183ee4b7f34c52d649066ef266664bccab4", - "human_verified/b200_fans/b200_fan_power_model.py": "163c6dd95f460cf6bde1b60cec061cce9be11e69cac1c6fe000b01a7ace81975", - "human_verified/b200_fans/plot_b200_fan_power.py": "8172b706c4a882ebb1fed220b3bcd602ea9db49019c4b1ac5a862d83be07bedd", - "human_verified/b200_psu/b200_psu_power_model.py": "58e3f81c7fbf722848184ddd73ed94edb065cfa96d6382d74bf42e8f7e15bbe8", - "human_verified/b200_psu/plot_b200_psu_power.py": "c96127308dab7105243ca21dcef07dce4fad8f674d9273006e3f4b1c29656b10", - "human_verified/b200_ubb/b200_ubb_power_model.py": "e3d3acf16a5ef9ca318ba49dc8b58d40823f8b76144d2999dffc9973db96648a", - "human_verified/b200_ubb/b200_ubb_residual_power_model.py": "85ca624b8d85d504f4e8a245f05aac07ab804b3fb262c2f00eb50ed35e22d273", - "human_verified/b200_ubb/plot_b200_ubb_residual_power.py": "a0d7a7e5f44948c59262bf62f4f088cb64fb17c29926f7b3327a18fb695903ed", - "human_verified/b300_fans/b300_fan_power_model.py": "9c5f85d0b7c9d4ee9fb8d233bb3d150a69643902970a68b475bb266918a495c7", - "human_verified/b300_fans/plot_b300_fan_power.py": "761a8c1c80f8efa0fe49bb749897da20764c1a207efe4d77c22a43b001b319c4", - "human_verified/b300_psu/b300_psu_power_model.py": "930f30f3a720b964623b35a7ee2d510a03f39ec6751dda5b480c7a8cfb0bd99e", - "human_verified/b300_psu/plot_b300_psu_power.py": "241880ab01f8d4012a6d2b358f85648bebb229b4966b1af081ecd331997cefc8", - "human_verified/b300_ubb/b300_ubb_power_model.py": "8cd694f1306ca09de53dcc9590dcc47a1b9bd5c10af9837bb875dc417fe74837", - "human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7", - "human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db", - "human_verified/chassis_plot_utils.py": "66f5d76f5ae5342433963595e58f665bdd0955eabdb31041a34c2f8180839af5", - "human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f", - "human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c", - "human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb", - "human_verified/generic/connectx8/plot_connectx8_power.py": "97921625373a479f03ad3930c8542e86da6c4ee4521654f77bcb965d93759433", - "human_verified/generic/cpu/cpu_power_model.py": "ab4315df415d70474ae4fb5c700fdf5b6de2a8a89f33823fbc8664efa01bf72f", - "human_verified/generic/cpu/plot_cpu_power.py": "bc5866234e7af605c8b43664cca1c1e96116d4fefb88c791cbe5e35f6234e821", - "human_verified/generic/dpu/dpu_power_model.py": "677cce41b990213c96f73e016c8bf24da9841ef995a7522e1038c832d502e398", - "human_verified/generic/dpu/plot_dpu_power.py": "d52357acdabb6a69b2ac87037cc339208e70b1bf7bdd31483b98dbedad06f05b", - "human_verified/generic/dram/dram_power_model.py": "b8ace82de182715a003c747710fc62275f2392d883aadb57c05f19beb1367b67", - "human_verified/generic/dram/plot_dram_power.py": "8a3d59fbf31492f9de04d8f3f80ddc33719018b6f47e8b0e671d3f63e29acead", - "human_verified/generic/nvme/nvme_power_model.py": "53fa1f59edf4427a62c946758d4f3df55b1591f8d43f824e07fd61e6d4196743", - "human_verified/generic/nvme/plot_nvme_power.py": "276d013eb26ff13c68e60cbf48bddc88fd6f5ff6c0a59a9854aec11da6ebf8b7", - "human_verified/generic/pcie_switches/pcie5_144lane_switch_power_model.py": "23d23322126c4251cb2ad5bee4448b0f9ad71b531607e8353cd96a90410cc28d", - "human_verified/generic/pcie_switches/plot_pcie_switch_power.py": "3cf3139b18495efc320c1f3d2754d8f732b16f1775959f66ff222ef97b64de90", - "human_verified/generic/pollara400/plot_pollara400_power.py": "22196c3037330b07903109e0f9a6917fd2e9e8e56ee942df552a84efd06a470e", - "human_verified/generic/pollara400/pollara400_power_model.py": "8dd5e674d9bfa5e19cb68c5684eb717df61063762afa555cd1a9c43f1d24c5d8", - "human_verified/generic/retimers/pcie5_x16_retimer_power_model.py": "393381bb40ef12cb81b681eff742129453136c6ba6c4b6950682e6cdc63a5335", - "human_verified/generic/retimers/plot_pcie5_x16_retimer_power.py": "c8f3a9aede8876cebab9f0674926b59f591f71ec1e58e86bdebe3ec63f61a208", - "human_verified/generic/thor2/plot_thor2_power.py": "7f03cd9af3dfd4e787e0e9d429a5758e24f4b62990af6d7d7e22a1a5224a78c7", - "human_verified/generic/thor2/thor2_power_model.py": "67b8e91a01dccd3abd7bdd37dbef86d1194d398c9326b8774eede1488962c54c", - "human_verified/hgx_b200_chassis/b200_chassis_power_model.py": "89d94969ce1acfeee67784ad269c431415995f9b18996d4b28d213d34393864a", - "human_verified/hgx_b200_chassis/plot_chassis_inference_gpu_sweep.py": "ed9378b471bf502adda4ac8f2467d803e48a5347ae329e118bfeebc9400ecc79", - "human_verified/hgx_b200_chassis_residual/b200_chassis_residual_power_model.py": "d23fe72c6039545e51d6271ddef28bee5d69f9ee7e4f8032a5def60021f973eb", - "human_verified/hgx_b200_chassis_residual/plot_chassis_residual_power.py": "a46d81768e0af6982bdb9d146b7e42fdbb86f1afcdb581d808657b5508a65414", - "human_verified/hgx_b300_chassis/b300_chassis_power_model.py": "68af8917ead2472cb0f6473784a5a75a24233cb766df8fa2684421a07d660fbc", - "human_verified/hgx_b300_chassis/plot_b300_chassis_inference_gpu_sweep.py": "bfe74771dd61b6dbb3dcf1ebbf6451c07b6020dcf0000482225672eb081f1d11", - "human_verified/hgx_h100_chassis/h100_chassis_power_model.py": "6850ec92346af1864f724a41d9ea512e0d55f45d683a3d575477aa08ca89a6c8", - "human_verified/hgx_h100_chassis/plot_h100_chassis_inference_gpu_sweep.py": "4af0da4e056f2300d271dd041ffb2ff9a76d39aaf80ae227a38ce777b949d852", - "human_verified/hgx_h200_chassis/h200_chassis_power_model.py": "56b40c9f13e50d81a02a594f0762f5c498483f261f9aec7b7294b148c65d7eb9", - "human_verified/hgx_h200_chassis/plot_h200_chassis_inference_gpu_sweep.py": "326b3baff711b5a322349cb272f83eaa44c8825cdbb9cc87cce52d7433e4b695", - "human_verified/hopper_fans/hopper_fan_power_model.py": "8b656f8498b6709b8ac7392999333eef06a529e12af141cc4565a2272c589f24", - "human_verified/hopper_fans/plot_hopper_fan_power.py": "006c13f8143ad0ebc853ecba9713660f65daa9ee520ce0b55763ce06a4a436fa", - "human_verified/hopper_nvswitch/hopper_nvswitch_power_model.py": "3d3c536bc2af1e75f2cc3c246e0d5337c04809cb90d470227a0d7336b413790c", - "human_verified/hopper_nvswitch/plot_hopper_nvswitch_power.py": "c13b291df6426a2117d82f1425009ddf0f55340888a0843dd5f5888503c62e06", - "human_verified/hopper_psu/hopper_psu_power_model.py": "a7630454e128e87bc0529a02cebb11138efbe1b07d4e4cb08a9721cc49f8d58d", - "human_verified/hopper_psu/plot_hopper_psu_power.py": "4a73fffa633e0a599a8bf746e0868f719669ff59d69a102bf563fcc4596e40a4", - "human_verified/hopper_ubb/hopper_ubb_power_model.py": "3f8c6c9560c32dbe1e98d0af82dbafa3fc697efc3d32642afe303455bfd5ab74", - "human_verified/mi300x_chassis/mi300x_chassis_power_model.py": "69c4b11e860e9a174664ae040691aab9e349410040ac8524dee6a7f2102afab6", - "human_verified/mi300x_chassis/plot_mi300x_chassis_inference_gpu_sweep.py": "1793a7356b95821cd4f7390a4cae55c58ffcc37f1b8f43f873499d84937463e4", - "human_verified/mi325x_chassis/mi325x_chassis_power_model.py": "59b6ce4ff626f1c5f42e8d7b0d33318e0e493a3e358dc1f4f92011be94b16c9c", - "human_verified/mi325x_chassis/plot_mi325x_chassis_inference_gpu_sweep.py": "c0f350bc4978a18759dd108756290fcd803f30209ccfc2682988bbf7ff0feb4d", - "human_verified/mi355x_chassis/mi355x_chassis_power_model.py": "c178f71efe53b424f5a1804fd188a79f99575e821a357b9128added48f154c8d", - "human_verified/mi355x_chassis/plot_mi355x_chassis_inference_gpu_sweep.py": "5071db35ce4bb4d7754419877f1e0d9df0c48be4381316b77d80fa1e58fe46c6" + "rackAssumptions": { + "u_nvlink": 0.5, + "u_ib": 0.0, + "u_pcie": 0.05, + "pue": 1.2 }, "profiles": { "h100": { - "modelPath": "human_verified/hgx_h100_chassis/h100_chassis_power_model.py", - "functionName": "h100_chassis_power", - "configFactory": "make_h100_config", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -91,152 +26,6 @@ "u_ib": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "H100 SXM5", - "gpu_component_key": "h100_sxm_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 640.0, - "gpu_idle_example_w_per_gpu": 100.0, - "gpu_decode_example_w_per_gpu": 300.0, - "gpu_prefill_example_w_per_gpu": 520.0, - "gpu_aggregate_example_w_per_gpu": 480.0, - "gpu_peak_w_per_gpu": 700.0, - "ubb": { - "gpu_label": "H100 SXM5", - "gpu_component_key": "h100_sxm_8x_measured", - "n_gpu": 8, - "nvswitch": { - "n_asic": 4, - "nvlink_ports_per_asic": 64, - "phy_lanes_per_asic": 128, - "phy_lane_gbps": 100.0, - "serdes_class_gbps": 112.0, - "serdes_pj_per_bit": 3.0, - "serdes_floor_frac": 0.95, - "digital_max_w": 120.0, - "digital_floor_frac": 0.4 - }, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "baseboard_controller_w": 10.0, - "fpga_cpld_sequencing_w": 8.0, - "hsc_power_monitor_w": 5.0, - "clock_refclk_reset_w": 5.0, - "sensors_i2c_fru_led_w": 4.0, - "aux_rails_misc_w": 13.0, - "normal_low_w": 30.0, - "normal_high_w": 70.0, - "conservative_cap_w": 90.0 - } - }, - "cpu": { - "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)", - "n_cpu": 2, - "cores_per_cpu": 56, - "threads_per_cpu": 112, - "package_tdp_w": 350.0, - "base_ghz": 2.0, - "max_turbo_ghz": 3.8, - "l3_cache_mb": 105.0, - "memory_channels_per_cpu": 8, - "pcie_lanes_per_cpu": 80, - "package_static_w": 24.0, - "uncore_io_baseline_w": 36.0, - "memory_controller_baseline_w": 22.0, - "uncore_dynamic_max_w": 8.0, - "memory_controller_dynamic_max_w": 6.0, - "core_curve_r": 1.72, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 8, - "n_dimm": 32, - "capacity_gb_per_dimm": 64.0, - "data_rate_mtps": 4800.0, - "dimm_type": "RDIMM", - "background_w": 2.147, - "refresh_w": 0.5509999999999999, - "termination_w": 1.102, - "io_dynamic_max_w": 2.3400000000000003, - "core_dynamic_max_w": 2.86, - "activity_exponent": 1.0 - }, - "connectx_compute": { - "n_nic": 8, - "net_serdes_w": 9.5, - "pcie_serdes_w": 5.5, - "board_w": 2.0, - "digital_max_w": 10.0, - "digital_floor_frac": 0.65, - "include_optic": true, - "optic_w": 8.0 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 8, - "front_u2_idle_w": 5.0, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 10.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 8.0, - "cpld_tpm_superio_w": 5.0, - "clock_sensor_fru_w": 4.0, - "storage_backplane_idle_w": 6.0, - "front_panel_usb_led_w": 2.0, - "aux_margin_w": 2.0, - "normal_low_w": 35.0, - "normal_high_w": 70.0 - }, - "fans": { - "electrical_nameplate_w": 1100.0, - "airflow_cfm_at_normal_max_pwm": 1105.0, - "min_pwm_frac": 0.22, - "normal_max_pwm_frac": 0.8, - "full_cooling_load_w": 9500.0, - "fan_curve_exponent": 1.1 - }, - "psu": { - "n_installed_psu": 6, - "n_load_sharing_psu": 6, - "n_redundant_capacity_psu": 4, - "psu_capacity_w": 3300.0, - "redundancy": "4+2", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "storage_mgmt_network_static_w": 60.0, - "optional_dpu_idle_w": 0.0, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "hopper_nvswitch_4x": 485.8, "hopper_pcie_retimers_8x": 92.4, @@ -270,9 +59,7 @@ } }, "h200": { - "modelPath": "human_verified/hgx_h200_chassis/h200_chassis_power_model.py", - "functionName": "h200_chassis_power", - "configFactory": "make_h200_config", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -284,152 +71,6 @@ "u_ib": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "H200 SXM5", - "gpu_component_key": "h200_sxm_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 1128.0, - "gpu_idle_example_w_per_gpu": 115.0, - "gpu_decode_example_w_per_gpu": 330.0, - "gpu_prefill_example_w_per_gpu": 540.0, - "gpu_aggregate_example_w_per_gpu": 510.0, - "gpu_peak_w_per_gpu": 700.0, - "ubb": { - "gpu_label": "H200 SXM5", - "gpu_component_key": "h200_sxm_8x_measured", - "n_gpu": 8, - "nvswitch": { - "n_asic": 4, - "nvlink_ports_per_asic": 64, - "phy_lanes_per_asic": 128, - "phy_lane_gbps": 100.0, - "serdes_class_gbps": 112.0, - "serdes_pj_per_bit": 3.0, - "serdes_floor_frac": 0.95, - "digital_max_w": 120.0, - "digital_floor_frac": 0.4 - }, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "baseboard_controller_w": 10.0, - "fpga_cpld_sequencing_w": 8.0, - "hsc_power_monitor_w": 5.0, - "clock_refclk_reset_w": 5.0, - "sensors_i2c_fru_led_w": 4.0, - "aux_rails_misc_w": 13.0, - "normal_low_w": 30.0, - "normal_high_w": 70.0, - "conservative_cap_w": 90.0 - } - }, - "cpu": { - "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)", - "n_cpu": 2, - "cores_per_cpu": 56, - "threads_per_cpu": 112, - "package_tdp_w": 350.0, - "base_ghz": 2.0, - "max_turbo_ghz": 3.8, - "l3_cache_mb": 105.0, - "memory_channels_per_cpu": 8, - "pcie_lanes_per_cpu": 80, - "package_static_w": 24.0, - "uncore_io_baseline_w": 36.0, - "memory_controller_baseline_w": 22.0, - "uncore_dynamic_max_w": 8.0, - "memory_controller_dynamic_max_w": 6.0, - "core_curve_r": 1.72, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 8, - "n_dimm": 32, - "capacity_gb_per_dimm": 64.0, - "data_rate_mtps": 4800.0, - "dimm_type": "RDIMM", - "background_w": 2.147, - "refresh_w": 0.5509999999999999, - "termination_w": 1.102, - "io_dynamic_max_w": 2.3400000000000003, - "core_dynamic_max_w": 2.86, - "activity_exponent": 1.0 - }, - "connectx_compute": { - "n_nic": 8, - "net_serdes_w": 9.5, - "pcie_serdes_w": 5.5, - "board_w": 2.0, - "digital_max_w": 10.0, - "digital_floor_frac": 0.65, - "include_optic": true, - "optic_w": 8.0 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 8, - "front_u2_idle_w": 5.0, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 10.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 8.0, - "cpld_tpm_superio_w": 5.0, - "clock_sensor_fru_w": 4.0, - "storage_backplane_idle_w": 6.0, - "front_panel_usb_led_w": 2.0, - "aux_margin_w": 2.0, - "normal_low_w": 35.0, - "normal_high_w": 70.0 - }, - "fans": { - "electrical_nameplate_w": 1100.0, - "airflow_cfm_at_normal_max_pwm": 1105.0, - "min_pwm_frac": 0.22, - "normal_max_pwm_frac": 0.8, - "full_cooling_load_w": 9500.0, - "fan_curve_exponent": 1.1 - }, - "psu": { - "n_installed_psu": 6, - "n_load_sharing_psu": 6, - "n_redundant_capacity_psu": 4, - "psu_capacity_w": 3300.0, - "redundancy": "4+2", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "storage_mgmt_network_static_w": 60.0, - "optional_dpu_idle_w": 0.0, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "hopper_nvswitch_4x": 485.8, "hopper_pcie_retimers_8x": 92.4, @@ -463,9 +104,7 @@ } }, "b200": { - "modelPath": "human_verified/hgx_b200_chassis/b200_chassis_power_model.py", - "functionName": "b200_chassis_power", - "configFactory": "B200ChassisMasterConfig", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -478,147 +117,6 @@ "u_dpu": 0.0, "fan_pwm": null }, - "defaultConfig": { - "ubb": { - "n_gpu": 8, - "nvswitch": { - "n_asic": 2, - "serdes_lanes": 144, - "serdes_lane_gbps": 200.0, - "serdes_pj_per_bit": 2.5, - "serdes_floor_frac": 0.95, - "digital_max_w": 190.0, - "digital_floor_frac": 0.4 - }, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "include_management_bridge_controller": true, - "management_bridge_controller_w": 24.0, - "hmc_bmc_w": 6.0, - "fpga_cpld_w": 8.0, - "erot_security_w": 3.0, - "hsc_power_monitor_w": 5.0, - "clock_refclk_reset_w": 4.0, - "sensors_i2c_fru_led_w": 3.0, - "aux_rails_misc_w": 9.0, - "normal_low_w": 45.0, - "normal_high_w": 85.0, - "conservative_cap_w": 100.0 - } - }, - "cpu": { - "name": "2x Intel Xeon Platinum 8570 (Emerald Rapids)", - "n_cpu": 2, - "cores_per_cpu": 56, - "threads_per_cpu": 112, - "package_tdp_w": 350.0, - "base_ghz": 2.1, - "max_turbo_ghz": 4.0, - "l3_cache_mb": 300.0, - "memory_channels_per_cpu": 8, - "pcie_lanes_per_cpu": 80, - "package_static_w": 22.0, - "uncore_io_baseline_w": 36.0, - "memory_controller_baseline_w": 22.0, - "uncore_dynamic_max_w": 8.0, - "memory_controller_dynamic_max_w": 6.0, - "core_curve_r": 1.7, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "DGX-B200-like Xeon 8570, 32x64GB DDR5-5600 RDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 8, - "n_dimm": 32, - "capacity_gb_per_dimm": 64.0, - "data_rate_mtps": 5600.0, - "dimm_type": "RDIMM", - "background_w": 2.147, - "refresh_w": 0.5509999999999999, - "termination_w": 1.102, - "io_dynamic_max_w": 2.3400000000000003, - "core_dynamic_max_w": 2.86, - "activity_exponent": 1.0 - }, - "connectx7": { - "n_nic": 8, - "net_serdes_w": 9.5, - "pcie_serdes_w": 5.5, - "board_w": 2.0, - "digital_max_w": 10.0, - "digital_floor_frac": 0.65, - "include_optic": true, - "optic_w": 8.0 - }, - "dpu": { - "n_dpu": 1, - "idle_w_per_dpu": 65.0, - "public_max_power_cap_w": 150.0, - "max_modeled_u_dpu": 0.02 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 10, - "front_u2_idle_w": 5.0, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 10.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 8.0, - "cpld_tpm_superio_w": 5.0, - "clock_sensor_fru_w": 4.0, - "storage_backplane_idle_w": 6.0, - "front_panel_usb_led_w": 2.0, - "aux_margin_w": 2.0, - "normal_low_w": 35.0, - "normal_high_w": 70.0 - }, - "fans": { - "n_80mm": 15, - "rated_80mm_w": 120.0, - "n_60mm": 4, - "rated_60mm_w": 25.0, - "min_pwm_frac": 0.25, - "normal_max_pwm_frac": 0.78, - "full_cooling_load_w": 12000.0, - "fan_curve_exponent": 1.15 - }, - "psu": { - "n_installed_psu": 6, - "n_active_psu": 3, - "psu_capacity_w": 5250.0, - "redundancy": "3+3", - "efficiency_curve": { - "0.05": 0.8864, - "0.1": 0.9238, - "0.2": 0.9448, - "0.5": 0.964, - "1.0": 0.9556 - } - }, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "nvswitch5_2x": 406.4, "ubb_pcie_retimers_8x": 92.4, @@ -652,9 +150,7 @@ } }, "b300": { - "modelPath": "human_verified/hgx_b300_chassis/b300_chassis_power_model.py", - "functionName": "b300_chassis_power", - "configFactory": "B300ChassisConfig", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -667,133 +163,6 @@ "u_dpu": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "B300 Blackwell Ultra SXM", - "gpu_component_key": "b300_sxm_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 2304.0, - "gpu_idle_example_w_per_gpu": 200.0, - "gpu_decode_example_w_per_gpu": 520.0, - "gpu_prefill_example_w_per_gpu": 800.0, - "gpu_aggregate_example_w_per_gpu": 760.0, - "gpu_peak_w_per_gpu": 1100.0, - "ubb": { - "n_gpu": 8, - "gpu_label": "B300 Blackwell Ultra SXM", - "gpu_component_key": "b300_sxm_8x_measured", - "nvswitch": { - "n_asic": 2, - "serdes_lanes": 144, - "serdes_lane_gbps": 200.0, - "serdes_pj_per_bit": 2.5, - "serdes_floor_frac": 0.95, - "digital_max_w": 190.0, - "digital_floor_frac": 0.4 - }, - "connectx8": { - "n_nic": 8, - "include_optic": true, - "network_serdes_static_w": 18.0, - "pcie_switch_static_w": 28.0, - "board_mgmt_static_w": 5.0, - "digital_static_w": 12.0, - "network_dynamic_max_w": 8.0, - "pcie_switch_dynamic_max_w": 5.0, - "digital_dynamic_max_w": 10.0, - "optic_idle_w": 15.0, - "optic_dynamic_max_w": 2.0, - "max_nic_slot_power_ref_w": 75.0, - "normal_low_w_per_nic_with_optic": 70.0, - "normal_high_w_per_nic_with_optic": 100.0 - }, - "residual_static_w": 85.0, - "residual_normal_low_w": 60.0, - "residual_normal_high_w": 120.0 - }, - "cpu": { - "name": "2x Intel Xeon 6776P (Granite Rapids, DGX B300)", - "n_cpu": 2, - "cores_per_cpu": 64, - "threads_per_cpu": 128, - "package_tdp_w": 350.0, - "base_ghz": 2.3, - "max_turbo_ghz": 3.9, - "l3_cache_mb": 336.0, - "memory_channels_per_cpu": 8, - "pcie_lanes_per_cpu": 88, - "package_static_w": 24.0, - "uncore_io_baseline_w": 40.0, - "memory_controller_baseline_w": 24.0, - "uncore_dynamic_max_w": 9.0, - "memory_controller_dynamic_max_w": 7.0, - "core_curve_r": 1.68, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "DGX-B300-like Xeon 6776P, 32x64GB DDR5-6400 RDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 8, - "n_dimm": 32, - "capacity_gb_per_dimm": 64.0, - "data_rate_mtps": 6400.0, - "dimm_type": "RDIMM", - "background_w": 2.147, - "refresh_w": 0.5509999999999999, - "termination_w": 1.102, - "io_dynamic_max_w": 2.3400000000000003, - "core_dynamic_max_w": 2.86, - "activity_exponent": 1.0 - }, - "dpu": { - "n_dpu": 2, - "idle_w_per_dpu": 65.0, - "public_max_power_cap_w": 150.0, - "max_modeled_u_dpu": 0.02 - }, - "nvme": { - "n_front_u2": 8, - "front_u2_idle_w": 4.5, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 12.0, - "onboard_10gbe_w": 4.0, - "motherboard_pch_aux_w": 10.0, - "cpld_tpm_superio_w": 6.0, - "clock_sensor_fru_w": 5.0, - "storage_backplane_idle_w": 10.0, - "front_panel_usb_led_w": 3.0, - "aux_margin_w": 5.0, - "normal_low_w": 40.0, - "normal_high_w": 80.0 - }, - "fans": { - "electrical_nameplate_w": 2000.0, - "min_pwm_frac": 0.22, - "normal_max_pwm_frac": 0.85, - "full_cooling_load_w": 12000.0, - "fan_curve_exponent": 1.08 - }, - "psu": { - "n_installed_psu": 12, - "n_load_sharing_psu": 12, - "n_redundant_capacity_psu": 6, - "psu_capacity_w": 3300.0, - "system_max_w": 15000.0, - "redundancy": "N+N / 6+6", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "nvswitch5_2x": 406.4, "connectx8_8x_integrated_pcie_with_optics": 630.0, @@ -825,9 +194,7 @@ } }, "mi300x": { - "modelPath": "human_verified/mi300x_chassis/mi300x_chassis_power_model.py", - "functionName": "mi300x_chassis_power", - "configFactory": "MI300XChassisConfig", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -838,146 +205,6 @@ "u_eth": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "AMD Instinct MI300X OAM", - "gpu_component_key": "mi300x_oam_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 1536.0, - "gpu_idle_example_w_per_gpu": 150.0, - "gpu_decode_example_w_per_gpu": 380.0, - "gpu_prefill_example_w_per_gpu": 620.0, - "gpu_aggregate_example_w_per_gpu": 580.0, - "gpu_peak_w_per_gpu": 750.0, - "ubb": { - "gpu_label": "AMD Instinct OAM", - "gpu_component_key": "amd_oam_8x_measured", - "n_gpu": 8, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "management_controller_w": 10.0, - "fpga_cpld_sequencing_w": 8.0, - "hsc_power_monitor_w": 6.0, - "clock_refclk_reset_w": 5.0, - "sensors_i2c_fru_led_w": 4.0, - "aux_rails_misc_w": 12.0, - "normal_low_w": 30.0, - "normal_high_w": 70.0, - "conservative_cap_w": 90.0 - } - }, - "cpu": { - "name": "2x AMD EPYC 9654 (Genoa, MI300X host baseline)", - "n_cpu": 2, - "cores_per_cpu": 96, - "threads_per_cpu": 192, - "package_tdp_w": 360.0, - "base_ghz": 2.4, - "max_turbo_ghz": 3.7, - "l3_cache_mb": 384.0, - "memory_channels_per_cpu": 12, - "pcie_lanes_per_cpu": 128, - "package_static_w": 26.0, - "uncore_io_baseline_w": 42.0, - "memory_controller_baseline_w": 30.0, - "uncore_dynamic_max_w": 12.0, - "memory_controller_dynamic_max_w": 12.0, - "core_curve_r": 1.65, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "MI300X host, 24x96GB DDR5-4800 RDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 12, - "n_dimm": 24, - "capacity_gb_per_dimm": 96.0, - "data_rate_mtps": 4800.0, - "dimm_type": "RDIMM", - "background_w": 2.7119999999999997, - "refresh_w": 0.696, - "termination_w": 1.3920000000000001, - "io_dynamic_max_w": 2.79, - "core_dynamic_max_w": 3.41, - "activity_exponent": 1.0 - }, - "thor2": { - "n_nic": 8, - "ports_per_nic": 2, - "aggregate_gbps_per_nic": 400.0, - "pcie_generation": 5, - "pcie_lanes": 16, - "card_idle_w": 12.5, - "card_traffic_dynamic_w": 0.4, - "include_optics": true, - "optical_modules_per_nic": 2, - "optics_total_w_per_nic": 10.9, - "deployment_margin_w": 6.5, - "deployment_traffic_dynamic_w": 2.8 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 12, - "front_u2_idle_w": 4.5, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 12.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 12.0, - "cpld_tpm_superio_w": 6.0, - "clock_sensor_fru_w": 5.0, - "storage_backplane_idle_w": 12.0, - "front_panel_usb_led_w": 3.0, - "aux_margin_w": 8.0, - "normal_low_w": 45.0, - "normal_high_w": 90.0 - }, - "fans": { - "platform_label": "MI300X 8U air-cooled chassis", - "n_fan": 10, - "electrical_nameplate_w": 1800.0, - "min_pwm_frac": 0.22, - "normal_max_pwm_frac": 0.82, - "full_cooling_load_w": 8500.0, - "fan_curve_exponent": 1.08 - }, - "psu": { - "platform_label": "MI300X 8U chassis PSU bank", - "n_installed_psu": 6, - "n_load_sharing_psu": 6, - "n_redundant_capacity_psu": 3, - "psu_capacity_w": 3000.0, - "system_max_w": 9000.0, - "redundancy": "N+N / 3+3", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "pcie5_x16_retimers_8x": 92.4, "amd_oam_ubb_residual_static": 45.0, @@ -1009,9 +236,7 @@ } }, "mi325x": { - "modelPath": "human_verified/mi325x_chassis/mi325x_chassis_power_model.py", - "functionName": "mi325x_chassis_power", - "configFactory": "MI325XChassisConfig", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -1022,146 +247,6 @@ "u_eth": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "AMD Instinct MI325X OAM", - "gpu_component_key": "mi325x_oam_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 2048.0, - "gpu_idle_example_w_per_gpu": 170.0, - "gpu_decode_example_w_per_gpu": 450.0, - "gpu_prefill_example_w_per_gpu": 800.0, - "gpu_aggregate_example_w_per_gpu": 760.0, - "gpu_peak_w_per_gpu": 1000.0, - "ubb": { - "gpu_label": "AMD Instinct OAM", - "gpu_component_key": "amd_oam_8x_measured", - "n_gpu": 8, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "management_controller_w": 10.0, - "fpga_cpld_sequencing_w": 8.0, - "hsc_power_monitor_w": 6.0, - "clock_refclk_reset_w": 5.0, - "sensors_i2c_fru_led_w": 4.0, - "aux_rails_misc_w": 12.0, - "normal_low_w": 30.0, - "normal_high_w": 70.0, - "conservative_cap_w": 90.0 - } - }, - "cpu": { - "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)", - "n_cpu": 2, - "cores_per_cpu": 64, - "threads_per_cpu": 128, - "package_tdp_w": 400.0, - "base_ghz": 3.3, - "max_turbo_ghz": 4.3, - "l3_cache_mb": 384.0, - "memory_channels_per_cpu": 12, - "pcie_lanes_per_cpu": 160, - "package_static_w": 28.0, - "uncore_io_baseline_w": 48.0, - "memory_controller_baseline_w": 34.0, - "uncore_dynamic_max_w": 14.0, - "memory_controller_dynamic_max_w": 14.0, - "core_curve_r": 1.62, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "MI325X host, 24x256GB DDR5-6400 RDIMM/MRDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 12, - "n_dimm": 24, - "capacity_gb_per_dimm": 256.0, - "data_rate_mtps": 6400.0, - "dimm_type": "RDIMM/MRDIMM", - "background_w": 3.9549999999999996, - "refresh_w": 1.015, - "termination_w": 2.0300000000000002, - "io_dynamic_max_w": 4.95, - "core_dynamic_max_w": 6.05, - "activity_exponent": 1.0 - }, - "thor2": { - "n_nic": 8, - "ports_per_nic": 2, - "aggregate_gbps_per_nic": 400.0, - "pcie_generation": 5, - "pcie_lanes": 16, - "card_idle_w": 12.5, - "card_traffic_dynamic_w": 0.4, - "include_optics": true, - "optical_modules_per_nic": 2, - "optics_total_w_per_nic": 10.9, - "deployment_margin_w": 6.5, - "deployment_traffic_dynamic_w": 2.8 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 8, - "front_u2_idle_w": 4.5, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 12.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 14.0, - "cpld_tpm_superio_w": 7.0, - "clock_sensor_fru_w": 6.0, - "storage_backplane_idle_w": 14.0, - "front_panel_usb_led_w": 3.0, - "aux_margin_w": 10.0, - "normal_low_w": 50.0, - "normal_high_w": 100.0 - }, - "fans": { - "platform_label": "MI325X 8U air-cooled chassis", - "n_fan": 14, - "electrical_nameplate_w": 2600.0, - "min_pwm_frac": 0.24, - "normal_max_pwm_frac": 0.88, - "full_cooling_load_w": 12000.0, - "fan_curve_exponent": 1.06 - }, - "psu": { - "platform_label": "MI325X 8U chassis PSU bank", - "n_installed_psu": 6, - "n_load_sharing_psu": 6, - "n_redundant_capacity_psu": 3, - "psu_capacity_w": 5250.0, - "system_max_w": 15750.0, - "redundancy": "N+N / 3+3", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "pcie5_x16_retimers_8x": 92.4, "amd_oam_ubb_residual_static": 45.0, @@ -1193,9 +278,7 @@ } }, "mi355x": { - "modelPath": "human_verified/mi355x_chassis/mi355x_chassis_power_model.py", - "functionName": "mi355x_chassis_power", - "configFactory": "MI355XChassisConfig", + "modelPath": "packages/app/src/lib/system-power-model.ts", "gpuCount": 8, "assumptions": { "u_pcie": 0.05, @@ -1206,145 +289,6 @@ "u_eth": 0.0, "fan_pwm": null }, - "defaultConfig": { - "gpu_label": "AMD Instinct MI355X OAM", - "gpu_component_key": "mi355x_oam_8x_measured", - "n_gpu": 8, - "gpu_memory_total_gb": 2304.0, - "gpu_idle_example_w_per_gpu": 250.0, - "gpu_decode_example_w_per_gpu": 700.0, - "gpu_prefill_example_w_per_gpu": 1150.0, - "gpu_aggregate_example_w_per_gpu": 1080.0, - "gpu_peak_w_per_gpu": 1400.0, - "ubb": { - "gpu_label": "AMD Instinct OAM", - "gpu_component_key": "amd_oam_8x_measured", - "n_gpu": 8, - "retimer": { - "n_retimer": 8, - "lanes_per_retimer": 16, - "pcie_gtps": 32.0, - "analog_serdes_equalization_w": 9.5, - "digital_floor_w": 2.0, - "digital_variable_w": 1.0 - }, - "residual": { - "management_controller_w": 10.0, - "fpga_cpld_sequencing_w": 8.0, - "hsc_power_monitor_w": 6.0, - "clock_refclk_reset_w": 5.0, - "sensors_i2c_fru_led_w": 4.0, - "aux_rails_misc_w": 12.0, - "normal_low_w": 30.0, - "normal_high_w": 70.0, - "conservative_cap_w": 90.0 - } - }, - "cpu": { - "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)", - "n_cpu": 2, - "cores_per_cpu": 64, - "threads_per_cpu": 128, - "package_tdp_w": 400.0, - "base_ghz": 3.3, - "max_turbo_ghz": 4.3, - "l3_cache_mb": 384.0, - "memory_channels_per_cpu": 12, - "pcie_lanes_per_cpu": 160, - "package_static_w": 28.0, - "uncore_io_baseline_w": 48.0, - "memory_controller_baseline_w": 34.0, - "uncore_dynamic_max_w": 14.0, - "memory_controller_dynamic_max_w": 14.0, - "core_curve_r": 1.62, - "vrm_efficiency": 0.92 - }, - "dram": { - "label": "MI355X host, 24x256GB DDR5-6400 RDIMM/MRDIMM", - "status": "DRAFT - pending human verification", - "n_sockets": 2, - "channels_per_socket": 12, - "n_dimm": 24, - "capacity_gb_per_dimm": 256.0, - "data_rate_mtps": 6400.0, - "dimm_type": "RDIMM/MRDIMM", - "background_w": 3.9549999999999996, - "refresh_w": 1.015, - "termination_w": 2.0300000000000002, - "io_dynamic_max_w": 4.95, - "core_dynamic_max_w": 6.05, - "activity_exponent": 1.0 - }, - "pollara400": { - "n_nic": 8, - "aggregate_gbps_per_nic": 400.0, - "pcie_generation": 5, - "pcie_lanes": 16, - "net_serdes_w": 10.5, - "pcie_serdes_w": 6.0, - "board_mgmt_w": 2.5, - "packet_engine_max_w": 12.5, - "packet_engine_floor_frac": 0.64, - "include_optic": true, - "optic_w": 8.0 - }, - "pcie_switch": { - "n_switch": 4, - "lanes_per_switch": 144, - "ports_per_switch": 72, - "pcie_gtps": 32.0, - "serdes_phy_floor_w": 30.5, - "control_leakage_clock_w": 6.0, - "fabric_datapath_floor_w": 4.0, - "fabric_datapath_variable_w": 8.5, - "stress_cap_per_switch_w": 70.0 - }, - "nvme": { - "n_front_u2": 8, - "front_u2_idle_w": 4.5, - "n_boot_m2": 2, - "boot_m2_idle_w": 2.0, - "max_modeled_u_nvme": 0.02 - }, - "residual": { - "bmc_ipmi_w": 12.0, - "onboard_10gbe_w": 8.0, - "motherboard_pch_aux_w": 16.0, - "cpld_tpm_superio_w": 8.0, - "clock_sensor_fru_w": 6.0, - "storage_backplane_idle_w": 16.0, - "front_panel_usb_led_w": 4.0, - "aux_margin_w": 12.0, - "normal_low_w": 60.0, - "normal_high_w": 120.0 - }, - "fans": { - "platform_label": "MI355X 10U air-cooled chassis", - "n_fan": 19, - "electrical_nameplate_w": 3600.0, - "min_pwm_frac": 0.25, - "normal_max_pwm_frac": 0.92, - "full_cooling_load_w": 15500.0, - "fan_curve_exponent": 1.04 - }, - "psu": { - "platform_label": "MI355X 10U chassis PSU bank", - "n_installed_psu": 6, - "n_load_sharing_psu": 6, - "n_redundant_capacity_psu": 4, - "psu_capacity_w": 6600.0, - "system_max_w": 26400.0, - "redundancy": "4+2", - "efficiency_curve": { - "0.05": 0.885, - "0.1": 0.92, - "0.2": 0.94, - "0.5": 0.96, - "1.0": 0.955 - } - }, - "pue": 1.2 - }, "fixedComponentsDcWatts": { "pcie5_x16_retimers_8x": 92.4, "amd_oam_ubb_residual_static": 45.0, @@ -1375,5 +319,161 @@ ] } } + }, + "rackProfiles": { + "gb200": { + "modelPath": "packages/app/src/lib/system-power-model.ts", + "topology": "nvl72-rack", + "gpuCount": 72, + "computeTrayCount": 18, + "gpusPerComputeTray": 4, + "graceSocketsPerComputeTray": 2, + "nvswitchTrayCount": 9, + "assumptions": { + "u_nvlink": 0.5, + "u_ib": 0.0, + "u_pcie": 0.05, + "pue": 1.2 + }, + "computeTrayStaticDcWatts": { + "connectx7_nics_with_optics": 126.0, + "bluefield3_dpu_idle": 130.0, + "nvme_idle": 22.0, + "compute_tray_fans_unverified": 130.0, + "compute_tray_board_residual_unverified": 40.0 + }, + "computeTrayStaticDetails": { + "nic_generation": "connectx7", + "per_nic_w": 31.5, + "n_nic": 4, + "n_dpu": 2, + "dpu_model": "idle_only", + "nvme_model": "idle_only" + }, + "nvswitchTraySiliconWatts": 406.4, + "nvswitchTrayResidualWatts": 50.0, + "managementSwitchCount": 2, + "managementSwitchWatts": 100.0, + "trayInputConversionEfficiency": 0.9725, + "regulatorLossFracOfTdp": 0.15, + "regulatorAllowanceIncludesGrace": false, + "powerShelf": { + "installedCapacityWatts": 264000.0, + "redundantCapacityWatts": 132000.0, + "efficiencyCurve": [ + [0.1, 0.9], + [0.2, 0.94], + [0.3, 0.965], + [1.0, 0.965] + ] + }, + "unverifiedParameters": { + "compute_tray_fans_w": { + "low": 40.0, + "high": 220.0, + "default": 130.0, + "unit": "W per compute tray", + "why": "8 dual-rotor 40x56 fans per tray are documented (Lenovo LP2357, SemiAnalysis) but no tray fan power is published. High end = 8 x 27.24 W rated Delta GFC0412DS-SM06B0G; low end = deep PWM on a liquid-cooled tray where only NICs/DPU/NVMe/PDB are air cooled." + }, + "compute_tray_board_residual_w": { + "low": 20.0, + "high": 60.0, + "default": 40.0, + "unit": "W per compute tray", + "why": "BMC module, CPLD/ERoT, clocks, sensors, leak detection and PDB housekeeping. No public rail. Anchored to the repo chassis residual note (35-70 W for a full HGX chassis) scaled to a 1RU tray." + }, + "tray_input_conversion_efficiency": { + "low": 0.96, + "high": 0.985, + "default": 0.9725, + "unit": "fraction", + "why": "Trays take 50 V DC from the busbar and feed the boards at 12 V (SemiAnalysis: 4x RapidLock 12 V connectors per Bianca board). The 50 V to 12 V stage sits outside the module sensor. No efficiency is published for it." + }, + "nvswitch_tray_residual_w": { + "low": 20.0, + "high": 80.0, + "default": 50.0, + "unit": "W per NVLink switch tray", + "why": "Switch tray BMC, two OOB 1GbE ports, console, CPLDs, clocks and NVSwitch VR loss outside the ASIC figure. Lenovo front/rear views show coolant, busbar and cartridge connectors and no fans. No public rail." + } + } + }, + "gb300": { + "modelPath": "packages/app/src/lib/system-power-model.ts", + "topology": "nvl72-rack", + "gpuCount": 72, + "computeTrayCount": 18, + "gpusPerComputeTray": 4, + "graceSocketsPerComputeTray": 2, + "nvswitchTrayCount": 9, + "assumptions": { + "u_nvlink": 0.5, + "u_ib": 0.0, + "u_pcie": 0.05, + "pue": 1.2 + }, + "computeTrayStaticDcWatts": { + "connectx8_nics_integrated_pcie_with_optics": 315.0, + "bluefield3_dpu_idle": 130.0, + "nvme_idle": 22.0, + "compute_tray_fans_unverified": 130.0, + "compute_tray_board_residual_unverified": 40.0 + }, + "computeTrayStaticDetails": { + "nic_generation": "connectx8", + "per_nic_w": 78.8, + "n_nic": 4, + "n_dpu": 2, + "dpu_model": "idle_only", + "nvme_model": "idle_only" + }, + "nvswitchTraySiliconWatts": 406.4, + "nvswitchTrayResidualWatts": 50.0, + "managementSwitchCount": 2, + "managementSwitchWatts": 100.0, + "trayInputConversionEfficiency": 0.9725, + "regulatorLossFracOfTdp": 0.15, + "regulatorAllowanceIncludesGrace": false, + "powerShelf": { + "installedCapacityWatts": 264000.0, + "redundantCapacityWatts": 132000.0, + "efficiencyCurve": [ + [0.1, 0.9], + [0.2, 0.94], + [0.3, 0.965], + [1.0, 0.965] + ] + }, + "unverifiedParameters": { + "compute_tray_fans_w": { + "low": 40.0, + "high": 220.0, + "default": 130.0, + "unit": "W per compute tray", + "why": "8 dual-rotor 40x56 fans per tray are documented (Lenovo LP2357, SemiAnalysis) but no tray fan power is published. High end = 8 x 27.24 W rated Delta GFC0412DS-SM06B0G; low end = deep PWM on a liquid-cooled tray where only NICs/DPU/NVMe/PDB are air cooled." + }, + "compute_tray_board_residual_w": { + "low": 20.0, + "high": 60.0, + "default": 40.0, + "unit": "W per compute tray", + "why": "BMC module, CPLD/ERoT, clocks, sensors, leak detection and PDB housekeeping. No public rail. Anchored to the repo chassis residual note (35-70 W for a full HGX chassis) scaled to a 1RU tray." + }, + "tray_input_conversion_efficiency": { + "low": 0.96, + "high": 0.985, + "default": 0.9725, + "unit": "fraction", + "why": "Trays take 50 V DC from the busbar and feed the boards at 12 V (SemiAnalysis: 4x RapidLock 12 V connectors per Bianca board). The 50 V to 12 V stage sits outside the module sensor. No efficiency is published for it." + }, + "nvswitch_tray_residual_w": { + "low": 20.0, + "high": 80.0, + "default": 50.0, + "unit": "W per NVLink switch tray", + "why": "Switch tray BMC, two OOB 1GbE ports, console, CPLDs, clocks and NVSwitch VR loss outside the ASIC figure. Lenovo front/rear views show coolant, busbar and cartridge connectors and no fans. No public rail." + } + } + } } } diff --git a/packages/app/src/lib/system-power-model.provenance.json b/packages/app/src/lib/system-power-model.provenance.json new file mode 100644 index 000000000..2ffb30b26 --- /dev/null +++ b/packages/app/src/lib/system-power-model.provenance.json @@ -0,0 +1,13 @@ +{ + "modelRevision": "app-sha256:99b252a0422c01934973760dbf5779c10aa87cc76e89807d1643258ef586e58e", + "modelRevisionStatus": "App-owned TypeScript equations, parameters and admission/PUE policy", + "source": "https://github.com/SemiAnalysisAI/InferenceX-app", + "status": "DRAFT / pending human verification", + "assumptionsSource": "packages/app/src/lib/system-power-model.profiles.json#/assumptions", + "rackAssumptionsSource": "packages/app/src/lib/system-power-model.profiles.json#/rackAssumptions", + "sourceSha256": { + "packages/app/src/lib/modeled-system-power.ts": "96c6016a5afabb860e01420038944422e3418befe347dc08d3602107878147fe", + "packages/app/src/lib/system-power-model.profiles.json": "9aaf0b8dc6365a86f099504e2fe7bbd612c3aef5606752059cd6062173aa4e20", + "packages/app/src/lib/system-power-model.ts": "9f0ff5dd1d1e85a6a09e1e20e9d9f682ece5f4f8eae7e5fe813e799329e966c4" + } +} diff --git a/packages/app/src/lib/system-power-model.reference.json b/packages/app/src/lib/system-power-model.reference.json index bf47f2777..d081154cc 100644 --- a/packages/app/src/lib/system-power-model.reference.json +++ b/packages/app/src/lib/system-power-model.reference.json @@ -1,7 +1,1278 @@ { - "modelRevision": "ca4403aa527069857351ad8047dbb726844b3382", + "modelRevision": "6fcc086b77576d4cecb9d0c79637d6daf980308c", + "modelRevisionStatus": "Historical local Python snapshot; unpublished at capture; retained only as migration-baseline provenance.", "source": "https://github.com/SemiAnalysisAI/inferencex_power_model", "status": "DRAFT / pending human verification", + "fixtureRole": "Frozen historical Python migration baseline; independent expected values, not the current app model version.", + "generatorSha256": "c33e50aac2b1ce6360df1895bd9f9c5496a3f783313c0243563c6d3c37f8a5bc", + "sourceSha256": { + "human_verified/amd_oam_fans/amd_oam_fan_power_model.py": "7ea6f1c65b685311277e6f2c53992d588ec79b598180f2fca300e904ad6dc19a", + "human_verified/amd_oam_fans/plot_amd_oam_fan_power.py": "e38118fa4d94412af268eaa1667d2c4086d1dcced4da07c951e761d54ea55b46", + "human_verified/amd_oam_psu/amd_oam_psu_power_model.py": "badcf849a82da35bc9bd621d42cd955ec7f8e3879dbc1bc29c7c03a3e0217af1", + "human_verified/amd_oam_psu/plot_amd_oam_psu_power.py": "39a1c3eac5a1e6337425948bfd252e422021b67de2ec5c963be1f8db3a8292fc", + "human_verified/amd_oam_ubb/amd_oam_ubb_power_model.py": "2f2f685e595d980c3fc353c624ca7183ee4b7f34c52d649066ef266664bccab4", + "human_verified/b200_fans/b200_fan_power_model.py": "163c6dd95f460cf6bde1b60cec061cce9be11e69cac1c6fe000b01a7ace81975", + "human_verified/b200_fans/plot_b200_fan_power.py": "8172b706c4a882ebb1fed220b3bcd602ea9db49019c4b1ac5a862d83be07bedd", + "human_verified/b200_psu/b200_psu_power_model.py": "58e3f81c7fbf722848184ddd73ed94edb065cfa96d6382d74bf42e8f7e15bbe8", + "human_verified/b200_psu/plot_b200_psu_power.py": "c96127308dab7105243ca21dcef07dce4fad8f674d9273006e3f4b1c29656b10", + "human_verified/b200_ubb/b200_ubb_power_model.py": "e3d3acf16a5ef9ca318ba49dc8b58d40823f8b76144d2999dffc9973db96648a", + "human_verified/b200_ubb/b200_ubb_residual_power_model.py": "85ca624b8d85d504f4e8a245f05aac07ab804b3fb262c2f00eb50ed35e22d273", + "human_verified/b200_ubb/plot_b200_ubb_residual_power.py": "a0d7a7e5f44948c59262bf62f4f088cb64fb17c29926f7b3327a18fb695903ed", + "human_verified/b300_fans/b300_fan_power_model.py": "9c5f85d0b7c9d4ee9fb8d233bb3d150a69643902970a68b475bb266918a495c7", + "human_verified/b300_fans/plot_b300_fan_power.py": "761a8c1c80f8efa0fe49bb749897da20764c1a207efe4d77c22a43b001b319c4", + "human_verified/b300_psu/b300_psu_power_model.py": "930f30f3a720b964623b35a7ee2d510a03f39ec6751dda5b480c7a8cfb0bd99e", + "human_verified/b300_psu/plot_b300_psu_power.py": "241880ab01f8d4012a6d2b358f85648bebb229b4966b1af081ecd331997cefc8", + "human_verified/b300_ubb/b300_ubb_power_model.py": "8cd694f1306ca09de53dcc9590dcc47a1b9bd5c10af9837bb875dc417fe74837", + "human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7", + "human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db", + "human_verified/chassis_plot_utils.py": "0f1c809a527788c8c25490420e29e737e6b7f83cb914dc022c3a4e3701b80643", + "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "b4640f94c8f9e6be50eb18deff9728bb58559ba193b54577859cafabec4ccc33", + "human_verified/gb200_nvl72_rack/plot_gb200_nvl72_rack_inference_gpu_sweep.py": "d0f6903df0348ccb2790e973b0dc10408c6f6022695ce039aad0b1e1fcd39519", + "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "631bd7fd44b84260f16175a814d093795a4cdd0a91a0d19dd572a5ea6c088314", + "human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f", + "human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c", + "human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb", + "human_verified/generic/connectx8/plot_connectx8_power.py": "97921625373a479f03ad3930c8542e86da6c4ee4521654f77bcb965d93759433", + "human_verified/generic/cpu/cpu_power_model.py": "ab4315df415d70474ae4fb5c700fdf5b6de2a8a89f33823fbc8664efa01bf72f", + "human_verified/generic/cpu/plot_cpu_power.py": "bc5866234e7af605c8b43664cca1c1e96116d4fefb88c791cbe5e35f6234e821", + "human_verified/generic/dpu/dpu_power_model.py": "677cce41b990213c96f73e016c8bf24da9841ef995a7522e1038c832d502e398", + "human_verified/generic/dpu/plot_dpu_power.py": "d52357acdabb6a69b2ac87037cc339208e70b1bf7bdd31483b98dbedad06f05b", + "human_verified/generic/dram/dram_power_model.py": "b8ace82de182715a003c747710fc62275f2392d883aadb57c05f19beb1367b67", + "human_verified/generic/dram/plot_dram_power.py": "8a3d59fbf31492f9de04d8f3f80ddc33719018b6f47e8b0e671d3f63e29acead", + "human_verified/generic/nvme/nvme_power_model.py": "53fa1f59edf4427a62c946758d4f3df55b1591f8d43f824e07fd61e6d4196743", + "human_verified/generic/nvme/plot_nvme_power.py": "276d013eb26ff13c68e60cbf48bddc88fd6f5ff6c0a59a9854aec11da6ebf8b7", + "human_verified/generic/pcie_switches/pcie5_144lane_switch_power_model.py": "23d23322126c4251cb2ad5bee4448b0f9ad71b531607e8353cd96a90410cc28d", + "human_verified/generic/pcie_switches/plot_pcie_switch_power.py": "3cf3139b18495efc320c1f3d2754d8f732b16f1775959f66ff222ef97b64de90", + "human_verified/generic/pollara400/plot_pollara400_power.py": "22196c3037330b07903109e0f9a6917fd2e9e8e56ee942df552a84efd06a470e", + "human_verified/generic/pollara400/pollara400_power_model.py": "8dd5e674d9bfa5e19cb68c5684eb717df61063762afa555cd1a9c43f1d24c5d8", + "human_verified/generic/retimers/pcie5_x16_retimer_power_model.py": "393381bb40ef12cb81b681eff742129453136c6ba6c4b6950682e6cdc63a5335", + "human_verified/generic/retimers/plot_pcie5_x16_retimer_power.py": "c8f3a9aede8876cebab9f0674926b59f591f71ec1e58e86bdebe3ec63f61a208", + "human_verified/generic/thor2/plot_thor2_power.py": "7f03cd9af3dfd4e787e0e9d429a5758e24f4b62990af6d7d7e22a1a5224a78c7", + "human_verified/generic/thor2/thor2_power_model.py": "67b8e91a01dccd3abd7bdd37dbef86d1194d398c9326b8774eede1488962c54c", + "human_verified/hgx_b200_chassis/b200_chassis_power_model.py": "89d94969ce1acfeee67784ad269c431415995f9b18996d4b28d213d34393864a", + "human_verified/hgx_b200_chassis/plot_chassis_inference_gpu_sweep.py": "ed9378b471bf502adda4ac8f2467d803e48a5347ae329e118bfeebc9400ecc79", + "human_verified/hgx_b200_chassis_residual/b200_chassis_residual_power_model.py": "d23fe72c6039545e51d6271ddef28bee5d69f9ee7e4f8032a5def60021f973eb", + "human_verified/hgx_b200_chassis_residual/plot_chassis_residual_power.py": "a46d81768e0af6982bdb9d146b7e42fdbb86f1afcdb581d808657b5508a65414", + "human_verified/hgx_b300_chassis/b300_chassis_power_model.py": "68af8917ead2472cb0f6473784a5a75a24233cb766df8fa2684421a07d660fbc", + "human_verified/hgx_b300_chassis/plot_b300_chassis_inference_gpu_sweep.py": "bfe74771dd61b6dbb3dcf1ebbf6451c07b6020dcf0000482225672eb081f1d11", + "human_verified/hgx_h100_chassis/h100_chassis_power_model.py": "6850ec92346af1864f724a41d9ea512e0d55f45d683a3d575477aa08ca89a6c8", + "human_verified/hgx_h100_chassis/plot_h100_chassis_inference_gpu_sweep.py": "4af0da4e056f2300d271dd041ffb2ff9a76d39aaf80ae227a38ce777b949d852", + "human_verified/hgx_h200_chassis/h200_chassis_power_model.py": "56b40c9f13e50d81a02a594f0762f5c498483f261f9aec7b7294b148c65d7eb9", + "human_verified/hgx_h200_chassis/plot_h200_chassis_inference_gpu_sweep.py": "326b3baff711b5a322349cb272f83eaa44c8825cdbb9cc87cce52d7433e4b695", + "human_verified/hopper_fans/hopper_fan_power_model.py": "8b656f8498b6709b8ac7392999333eef06a529e12af141cc4565a2272c589f24", + "human_verified/hopper_fans/plot_hopper_fan_power.py": "006c13f8143ad0ebc853ecba9713660f65daa9ee520ce0b55763ce06a4a436fa", + "human_verified/hopper_nvswitch/hopper_nvswitch_power_model.py": "3d3c536bc2af1e75f2cc3c246e0d5337c04809cb90d470227a0d7336b413790c", + "human_verified/hopper_nvswitch/plot_hopper_nvswitch_power.py": "c13b291df6426a2117d82f1425009ddf0f55340888a0843dd5f5888503c62e06", + "human_verified/hopper_psu/hopper_psu_power_model.py": "a7630454e128e87bc0529a02cebb11138efbe1b07d4e4cb08a9721cc49f8d58d", + "human_verified/hopper_psu/plot_hopper_psu_power.py": "4a73fffa633e0a599a8bf746e0868f719669ff59d69a102bf563fcc4596e40a4", + "human_verified/hopper_ubb/hopper_ubb_power_model.py": "3f8c6c9560c32dbe1e98d0af82dbafa3fc697efc3d32642afe303455bfd5ab74", + "human_verified/mi300x_chassis/mi300x_chassis_power_model.py": "69c4b11e860e9a174664ae040691aab9e349410040ac8524dee6a7f2102afab6", + "human_verified/mi300x_chassis/plot_mi300x_chassis_inference_gpu_sweep.py": "1793a7356b95821cd4f7390a4cae55c58ffcc37f1b8f43f873499d84937463e4", + "human_verified/mi325x_chassis/mi325x_chassis_power_model.py": "59b6ce4ff626f1c5f42e8d7b0d33318e0e493a3e358dc1f4f92011be94b16c9c", + "human_verified/mi325x_chassis/plot_mi325x_chassis_inference_gpu_sweep.py": "c0f350bc4978a18759dd108756290fcd803f30209ccfc2682988bbf7ff0feb4d", + "human_verified/mi355x_chassis/mi355x_chassis_power_model.py": "c178f71efe53b424f5a1804fd188a79f99575e821a357b9128added48f154c8d", + "human_verified/mi355x_chassis/plot_mi355x_chassis_inference_gpu_sweep.py": "5071db35ce4bb4d7754419877f1e0d9df0c48be4381316b77d80fa1e58fe46c6" + }, + "modelPaths": { + "h100": "human_verified/hgx_h100_chassis/h100_chassis_power_model.py", + "h200": "human_verified/hgx_h200_chassis/h200_chassis_power_model.py", + "b200": "human_verified/hgx_b200_chassis/b200_chassis_power_model.py", + "b300": "human_verified/hgx_b300_chassis/b300_chassis_power_model.py", + "mi300x": "human_verified/mi300x_chassis/mi300x_chassis_power_model.py", + "mi325x": "human_verified/mi325x_chassis/mi325x_chassis_power_model.py", + "mi355x": "human_verified/mi355x_chassis/mi355x_chassis_power_model.py", + "gb200": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py", + "gb300": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py" + }, + "pythonConfigurations": { + "h100": { + "functionName": "h100_chassis_power", + "configFactory": "make_h100_config", + "defaultConfig": { + "gpu_label": "H100 SXM5", + "gpu_component_key": "h100_sxm_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 640.0, + "gpu_idle_example_w_per_gpu": 100.0, + "gpu_decode_example_w_per_gpu": 300.0, + "gpu_prefill_example_w_per_gpu": 520.0, + "gpu_aggregate_example_w_per_gpu": 480.0, + "gpu_peak_w_per_gpu": 700.0, + "ubb": { + "gpu_label": "H100 SXM5", + "gpu_component_key": "h100_sxm_8x_measured", + "n_gpu": 8, + "nvswitch": { + "n_asic": 4, + "nvlink_ports_per_asic": 64, + "phy_lanes_per_asic": 128, + "phy_lane_gbps": 100.0, + "serdes_class_gbps": 112.0, + "serdes_pj_per_bit": 3.0, + "serdes_floor_frac": 0.95, + "digital_max_w": 120.0, + "digital_floor_frac": 0.4 + }, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "baseboard_controller_w": 10.0, + "fpga_cpld_sequencing_w": 8.0, + "hsc_power_monitor_w": 5.0, + "clock_refclk_reset_w": 5.0, + "sensors_i2c_fru_led_w": 4.0, + "aux_rails_misc_w": 13.0, + "normal_low_w": 30.0, + "normal_high_w": 70.0, + "conservative_cap_w": 90.0 + } + }, + "cpu": { + "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)", + "n_cpu": 2, + "cores_per_cpu": 56, + "threads_per_cpu": 112, + "package_tdp_w": 350.0, + "base_ghz": 2.0, + "max_turbo_ghz": 3.8, + "l3_cache_mb": 105.0, + "memory_channels_per_cpu": 8, + "pcie_lanes_per_cpu": 80, + "package_static_w": 24.0, + "uncore_io_baseline_w": 36.0, + "memory_controller_baseline_w": 22.0, + "uncore_dynamic_max_w": 8.0, + "memory_controller_dynamic_max_w": 6.0, + "core_curve_r": 1.72, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 8, + "n_dimm": 32, + "capacity_gb_per_dimm": 64.0, + "data_rate_mtps": 4800.0, + "dimm_type": "RDIMM", + "background_w": 2.147, + "refresh_w": 0.5509999999999999, + "termination_w": 1.102, + "io_dynamic_max_w": 2.3400000000000003, + "core_dynamic_max_w": 2.86, + "activity_exponent": 1.0 + }, + "connectx_compute": { + "n_nic": 8, + "net_serdes_w": 9.5, + "pcie_serdes_w": 5.5, + "board_w": 2.0, + "digital_max_w": 10.0, + "digital_floor_frac": 0.65, + "include_optic": true, + "optic_w": 8.0 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 8, + "front_u2_idle_w": 5.0, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 10.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 8.0, + "cpld_tpm_superio_w": 5.0, + "clock_sensor_fru_w": 4.0, + "storage_backplane_idle_w": 6.0, + "front_panel_usb_led_w": 2.0, + "aux_margin_w": 2.0, + "normal_low_w": 35.0, + "normal_high_w": 70.0 + }, + "fans": { + "electrical_nameplate_w": 1100.0, + "airflow_cfm_at_normal_max_pwm": 1105.0, + "min_pwm_frac": 0.22, + "normal_max_pwm_frac": 0.8, + "full_cooling_load_w": 9500.0, + "fan_curve_exponent": 1.1 + }, + "psu": { + "n_installed_psu": 6, + "n_load_sharing_psu": 6, + "n_redundant_capacity_psu": 4, + "psu_capacity_w": 3300.0, + "redundancy": "4+2", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "storage_mgmt_network_static_w": 60.0, + "optional_dpu_idle_w": 0.0, + "pue": 1.2 + } + }, + "h200": { + "functionName": "h200_chassis_power", + "configFactory": "make_h200_config", + "defaultConfig": { + "gpu_label": "H200 SXM5", + "gpu_component_key": "h200_sxm_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 1128.0, + "gpu_idle_example_w_per_gpu": 115.0, + "gpu_decode_example_w_per_gpu": 330.0, + "gpu_prefill_example_w_per_gpu": 540.0, + "gpu_aggregate_example_w_per_gpu": 510.0, + "gpu_peak_w_per_gpu": 700.0, + "ubb": { + "gpu_label": "H200 SXM5", + "gpu_component_key": "h200_sxm_8x_measured", + "n_gpu": 8, + "nvswitch": { + "n_asic": 4, + "nvlink_ports_per_asic": 64, + "phy_lanes_per_asic": 128, + "phy_lane_gbps": 100.0, + "serdes_class_gbps": 112.0, + "serdes_pj_per_bit": 3.0, + "serdes_floor_frac": 0.95, + "digital_max_w": 120.0, + "digital_floor_frac": 0.4 + }, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "baseboard_controller_w": 10.0, + "fpga_cpld_sequencing_w": 8.0, + "hsc_power_monitor_w": 5.0, + "clock_refclk_reset_w": 5.0, + "sensors_i2c_fru_led_w": 4.0, + "aux_rails_misc_w": 13.0, + "normal_low_w": 30.0, + "normal_high_w": 70.0, + "conservative_cap_w": 90.0 + } + }, + "cpu": { + "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)", + "n_cpu": 2, + "cores_per_cpu": 56, + "threads_per_cpu": 112, + "package_tdp_w": 350.0, + "base_ghz": 2.0, + "max_turbo_ghz": 3.8, + "l3_cache_mb": 105.0, + "memory_channels_per_cpu": 8, + "pcie_lanes_per_cpu": 80, + "package_static_w": 24.0, + "uncore_io_baseline_w": 36.0, + "memory_controller_baseline_w": 22.0, + "uncore_dynamic_max_w": 8.0, + "memory_controller_dynamic_max_w": 6.0, + "core_curve_r": 1.72, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 8, + "n_dimm": 32, + "capacity_gb_per_dimm": 64.0, + "data_rate_mtps": 4800.0, + "dimm_type": "RDIMM", + "background_w": 2.147, + "refresh_w": 0.5509999999999999, + "termination_w": 1.102, + "io_dynamic_max_w": 2.3400000000000003, + "core_dynamic_max_w": 2.86, + "activity_exponent": 1.0 + }, + "connectx_compute": { + "n_nic": 8, + "net_serdes_w": 9.5, + "pcie_serdes_w": 5.5, + "board_w": 2.0, + "digital_max_w": 10.0, + "digital_floor_frac": 0.65, + "include_optic": true, + "optic_w": 8.0 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 8, + "front_u2_idle_w": 5.0, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 10.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 8.0, + "cpld_tpm_superio_w": 5.0, + "clock_sensor_fru_w": 4.0, + "storage_backplane_idle_w": 6.0, + "front_panel_usb_led_w": 2.0, + "aux_margin_w": 2.0, + "normal_low_w": 35.0, + "normal_high_w": 70.0 + }, + "fans": { + "electrical_nameplate_w": 1100.0, + "airflow_cfm_at_normal_max_pwm": 1105.0, + "min_pwm_frac": 0.22, + "normal_max_pwm_frac": 0.8, + "full_cooling_load_w": 9500.0, + "fan_curve_exponent": 1.1 + }, + "psu": { + "n_installed_psu": 6, + "n_load_sharing_psu": 6, + "n_redundant_capacity_psu": 4, + "psu_capacity_w": 3300.0, + "redundancy": "4+2", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "storage_mgmt_network_static_w": 60.0, + "optional_dpu_idle_w": 0.0, + "pue": 1.2 + } + }, + "b200": { + "functionName": "b200_chassis_power", + "configFactory": "B200ChassisMasterConfig", + "defaultConfig": { + "ubb": { + "n_gpu": 8, + "nvswitch": { + "n_asic": 2, + "serdes_lanes": 144, + "serdes_lane_gbps": 200.0, + "serdes_pj_per_bit": 2.5, + "serdes_floor_frac": 0.95, + "digital_max_w": 190.0, + "digital_floor_frac": 0.4 + }, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "include_management_bridge_controller": true, + "management_bridge_controller_w": 24.0, + "hmc_bmc_w": 6.0, + "fpga_cpld_w": 8.0, + "erot_security_w": 3.0, + "hsc_power_monitor_w": 5.0, + "clock_refclk_reset_w": 4.0, + "sensors_i2c_fru_led_w": 3.0, + "aux_rails_misc_w": 9.0, + "normal_low_w": 45.0, + "normal_high_w": 85.0, + "conservative_cap_w": 100.0 + } + }, + "cpu": { + "name": "2x Intel Xeon Platinum 8570 (Emerald Rapids)", + "n_cpu": 2, + "cores_per_cpu": 56, + "threads_per_cpu": 112, + "package_tdp_w": 350.0, + "base_ghz": 2.1, + "max_turbo_ghz": 4.0, + "l3_cache_mb": 300.0, + "memory_channels_per_cpu": 8, + "pcie_lanes_per_cpu": 80, + "package_static_w": 22.0, + "uncore_io_baseline_w": 36.0, + "memory_controller_baseline_w": 22.0, + "uncore_dynamic_max_w": 8.0, + "memory_controller_dynamic_max_w": 6.0, + "core_curve_r": 1.7, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "DGX-B200-like Xeon 8570, 32x64GB DDR5-5600 RDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 8, + "n_dimm": 32, + "capacity_gb_per_dimm": 64.0, + "data_rate_mtps": 5600.0, + "dimm_type": "RDIMM", + "background_w": 2.147, + "refresh_w": 0.5509999999999999, + "termination_w": 1.102, + "io_dynamic_max_w": 2.3400000000000003, + "core_dynamic_max_w": 2.86, + "activity_exponent": 1.0 + }, + "connectx7": { + "n_nic": 8, + "net_serdes_w": 9.5, + "pcie_serdes_w": 5.5, + "board_w": 2.0, + "digital_max_w": 10.0, + "digital_floor_frac": 0.65, + "include_optic": true, + "optic_w": 8.0 + }, + "dpu": { + "n_dpu": 1, + "idle_w_per_dpu": 65.0, + "public_max_power_cap_w": 150.0, + "max_modeled_u_dpu": 0.02 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 10, + "front_u2_idle_w": 5.0, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 10.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 8.0, + "cpld_tpm_superio_w": 5.0, + "clock_sensor_fru_w": 4.0, + "storage_backplane_idle_w": 6.0, + "front_panel_usb_led_w": 2.0, + "aux_margin_w": 2.0, + "normal_low_w": 35.0, + "normal_high_w": 70.0 + }, + "fans": { + "n_80mm": 15, + "rated_80mm_w": 120.0, + "n_60mm": 4, + "rated_60mm_w": 25.0, + "min_pwm_frac": 0.25, + "normal_max_pwm_frac": 0.78, + "full_cooling_load_w": 12000.0, + "fan_curve_exponent": 1.15 + }, + "psu": { + "n_installed_psu": 6, + "n_active_psu": 3, + "psu_capacity_w": 5250.0, + "redundancy": "3+3", + "efficiency_curve": { + "0.05": 0.8864, + "0.1": 0.9238, + "0.2": 0.9448, + "0.5": 0.964, + "1.0": 0.9556 + } + }, + "pue": 1.2 + } + }, + "b300": { + "functionName": "b300_chassis_power", + "configFactory": "B300ChassisConfig", + "defaultConfig": { + "gpu_label": "B300 Blackwell Ultra SXM", + "gpu_component_key": "b300_sxm_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 2304.0, + "gpu_idle_example_w_per_gpu": 200.0, + "gpu_decode_example_w_per_gpu": 520.0, + "gpu_prefill_example_w_per_gpu": 800.0, + "gpu_aggregate_example_w_per_gpu": 760.0, + "gpu_peak_w_per_gpu": 1100.0, + "ubb": { + "n_gpu": 8, + "gpu_label": "B300 Blackwell Ultra SXM", + "gpu_component_key": "b300_sxm_8x_measured", + "nvswitch": { + "n_asic": 2, + "serdes_lanes": 144, + "serdes_lane_gbps": 200.0, + "serdes_pj_per_bit": 2.5, + "serdes_floor_frac": 0.95, + "digital_max_w": 190.0, + "digital_floor_frac": 0.4 + }, + "connectx8": { + "n_nic": 8, + "include_optic": true, + "network_serdes_static_w": 18.0, + "pcie_switch_static_w": 28.0, + "board_mgmt_static_w": 5.0, + "digital_static_w": 12.0, + "network_dynamic_max_w": 8.0, + "pcie_switch_dynamic_max_w": 5.0, + "digital_dynamic_max_w": 10.0, + "optic_idle_w": 15.0, + "optic_dynamic_max_w": 2.0, + "max_nic_slot_power_ref_w": 75.0, + "normal_low_w_per_nic_with_optic": 70.0, + "normal_high_w_per_nic_with_optic": 100.0 + }, + "residual_static_w": 85.0, + "residual_normal_low_w": 60.0, + "residual_normal_high_w": 120.0 + }, + "cpu": { + "name": "2x Intel Xeon 6776P (Granite Rapids, DGX B300)", + "n_cpu": 2, + "cores_per_cpu": 64, + "threads_per_cpu": 128, + "package_tdp_w": 350.0, + "base_ghz": 2.3, + "max_turbo_ghz": 3.9, + "l3_cache_mb": 336.0, + "memory_channels_per_cpu": 8, + "pcie_lanes_per_cpu": 88, + "package_static_w": 24.0, + "uncore_io_baseline_w": 40.0, + "memory_controller_baseline_w": 24.0, + "uncore_dynamic_max_w": 9.0, + "memory_controller_dynamic_max_w": 7.0, + "core_curve_r": 1.68, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "DGX-B300-like Xeon 6776P, 32x64GB DDR5-6400 RDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 8, + "n_dimm": 32, + "capacity_gb_per_dimm": 64.0, + "data_rate_mtps": 6400.0, + "dimm_type": "RDIMM", + "background_w": 2.147, + "refresh_w": 0.5509999999999999, + "termination_w": 1.102, + "io_dynamic_max_w": 2.3400000000000003, + "core_dynamic_max_w": 2.86, + "activity_exponent": 1.0 + }, + "dpu": { + "n_dpu": 2, + "idle_w_per_dpu": 65.0, + "public_max_power_cap_w": 150.0, + "max_modeled_u_dpu": 0.02 + }, + "nvme": { + "n_front_u2": 8, + "front_u2_idle_w": 4.5, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 12.0, + "onboard_10gbe_w": 4.0, + "motherboard_pch_aux_w": 10.0, + "cpld_tpm_superio_w": 6.0, + "clock_sensor_fru_w": 5.0, + "storage_backplane_idle_w": 10.0, + "front_panel_usb_led_w": 3.0, + "aux_margin_w": 5.0, + "normal_low_w": 40.0, + "normal_high_w": 80.0 + }, + "fans": { + "electrical_nameplate_w": 2000.0, + "min_pwm_frac": 0.22, + "normal_max_pwm_frac": 0.85, + "full_cooling_load_w": 12000.0, + "fan_curve_exponent": 1.08 + }, + "psu": { + "n_installed_psu": 12, + "n_load_sharing_psu": 12, + "n_redundant_capacity_psu": 6, + "psu_capacity_w": 3300.0, + "system_max_w": 15000.0, + "redundancy": "N+N / 6+6", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "pue": 1.2 + } + }, + "mi300x": { + "functionName": "mi300x_chassis_power", + "configFactory": "MI300XChassisConfig", + "defaultConfig": { + "gpu_label": "AMD Instinct MI300X OAM", + "gpu_component_key": "mi300x_oam_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 1536.0, + "gpu_idle_example_w_per_gpu": 150.0, + "gpu_decode_example_w_per_gpu": 380.0, + "gpu_prefill_example_w_per_gpu": 620.0, + "gpu_aggregate_example_w_per_gpu": 580.0, + "gpu_peak_w_per_gpu": 750.0, + "ubb": { + "gpu_label": "AMD Instinct OAM", + "gpu_component_key": "amd_oam_8x_measured", + "n_gpu": 8, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "management_controller_w": 10.0, + "fpga_cpld_sequencing_w": 8.0, + "hsc_power_monitor_w": 6.0, + "clock_refclk_reset_w": 5.0, + "sensors_i2c_fru_led_w": 4.0, + "aux_rails_misc_w": 12.0, + "normal_low_w": 30.0, + "normal_high_w": 70.0, + "conservative_cap_w": 90.0 + } + }, + "cpu": { + "name": "2x AMD EPYC 9654 (Genoa, MI300X host baseline)", + "n_cpu": 2, + "cores_per_cpu": 96, + "threads_per_cpu": 192, + "package_tdp_w": 360.0, + "base_ghz": 2.4, + "max_turbo_ghz": 3.7, + "l3_cache_mb": 384.0, + "memory_channels_per_cpu": 12, + "pcie_lanes_per_cpu": 128, + "package_static_w": 26.0, + "uncore_io_baseline_w": 42.0, + "memory_controller_baseline_w": 30.0, + "uncore_dynamic_max_w": 12.0, + "memory_controller_dynamic_max_w": 12.0, + "core_curve_r": 1.65, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "MI300X host, 24x96GB DDR5-4800 RDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 12, + "n_dimm": 24, + "capacity_gb_per_dimm": 96.0, + "data_rate_mtps": 4800.0, + "dimm_type": "RDIMM", + "background_w": 2.7119999999999997, + "refresh_w": 0.696, + "termination_w": 1.3920000000000001, + "io_dynamic_max_w": 2.79, + "core_dynamic_max_w": 3.41, + "activity_exponent": 1.0 + }, + "thor2": { + "n_nic": 8, + "ports_per_nic": 2, + "aggregate_gbps_per_nic": 400.0, + "pcie_generation": 5, + "pcie_lanes": 16, + "card_idle_w": 12.5, + "card_traffic_dynamic_w": 0.4, + "include_optics": true, + "optical_modules_per_nic": 2, + "optics_total_w_per_nic": 10.9, + "deployment_margin_w": 6.5, + "deployment_traffic_dynamic_w": 2.8 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 12, + "front_u2_idle_w": 4.5, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 12.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 12.0, + "cpld_tpm_superio_w": 6.0, + "clock_sensor_fru_w": 5.0, + "storage_backplane_idle_w": 12.0, + "front_panel_usb_led_w": 3.0, + "aux_margin_w": 8.0, + "normal_low_w": 45.0, + "normal_high_w": 90.0 + }, + "fans": { + "platform_label": "MI300X 8U air-cooled chassis", + "n_fan": 10, + "electrical_nameplate_w": 1800.0, + "min_pwm_frac": 0.22, + "normal_max_pwm_frac": 0.82, + "full_cooling_load_w": 8500.0, + "fan_curve_exponent": 1.08 + }, + "psu": { + "platform_label": "MI300X 8U chassis PSU bank", + "n_installed_psu": 6, + "n_load_sharing_psu": 6, + "n_redundant_capacity_psu": 3, + "psu_capacity_w": 3000.0, + "system_max_w": 9000.0, + "redundancy": "N+N / 3+3", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "pue": 1.2 + } + }, + "mi325x": { + "functionName": "mi325x_chassis_power", + "configFactory": "MI325XChassisConfig", + "defaultConfig": { + "gpu_label": "AMD Instinct MI325X OAM", + "gpu_component_key": "mi325x_oam_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 2048.0, + "gpu_idle_example_w_per_gpu": 170.0, + "gpu_decode_example_w_per_gpu": 450.0, + "gpu_prefill_example_w_per_gpu": 800.0, + "gpu_aggregate_example_w_per_gpu": 760.0, + "gpu_peak_w_per_gpu": 1000.0, + "ubb": { + "gpu_label": "AMD Instinct OAM", + "gpu_component_key": "amd_oam_8x_measured", + "n_gpu": 8, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "management_controller_w": 10.0, + "fpga_cpld_sequencing_w": 8.0, + "hsc_power_monitor_w": 6.0, + "clock_refclk_reset_w": 5.0, + "sensors_i2c_fru_led_w": 4.0, + "aux_rails_misc_w": 12.0, + "normal_low_w": 30.0, + "normal_high_w": 70.0, + "conservative_cap_w": 90.0 + } + }, + "cpu": { + "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)", + "n_cpu": 2, + "cores_per_cpu": 64, + "threads_per_cpu": 128, + "package_tdp_w": 400.0, + "base_ghz": 3.3, + "max_turbo_ghz": 4.3, + "l3_cache_mb": 384.0, + "memory_channels_per_cpu": 12, + "pcie_lanes_per_cpu": 160, + "package_static_w": 28.0, + "uncore_io_baseline_w": 48.0, + "memory_controller_baseline_w": 34.0, + "uncore_dynamic_max_w": 14.0, + "memory_controller_dynamic_max_w": 14.0, + "core_curve_r": 1.62, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "MI325X host, 24x256GB DDR5-6400 RDIMM/MRDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 12, + "n_dimm": 24, + "capacity_gb_per_dimm": 256.0, + "data_rate_mtps": 6400.0, + "dimm_type": "RDIMM/MRDIMM", + "background_w": 3.9549999999999996, + "refresh_w": 1.015, + "termination_w": 2.0300000000000002, + "io_dynamic_max_w": 4.95, + "core_dynamic_max_w": 6.05, + "activity_exponent": 1.0 + }, + "thor2": { + "n_nic": 8, + "ports_per_nic": 2, + "aggregate_gbps_per_nic": 400.0, + "pcie_generation": 5, + "pcie_lanes": 16, + "card_idle_w": 12.5, + "card_traffic_dynamic_w": 0.4, + "include_optics": true, + "optical_modules_per_nic": 2, + "optics_total_w_per_nic": 10.9, + "deployment_margin_w": 6.5, + "deployment_traffic_dynamic_w": 2.8 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 8, + "front_u2_idle_w": 4.5, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 12.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 14.0, + "cpld_tpm_superio_w": 7.0, + "clock_sensor_fru_w": 6.0, + "storage_backplane_idle_w": 14.0, + "front_panel_usb_led_w": 3.0, + "aux_margin_w": 10.0, + "normal_low_w": 50.0, + "normal_high_w": 100.0 + }, + "fans": { + "platform_label": "MI325X 8U air-cooled chassis", + "n_fan": 14, + "electrical_nameplate_w": 2600.0, + "min_pwm_frac": 0.24, + "normal_max_pwm_frac": 0.88, + "full_cooling_load_w": 12000.0, + "fan_curve_exponent": 1.06 + }, + "psu": { + "platform_label": "MI325X 8U chassis PSU bank", + "n_installed_psu": 6, + "n_load_sharing_psu": 6, + "n_redundant_capacity_psu": 3, + "psu_capacity_w": 5250.0, + "system_max_w": 15750.0, + "redundancy": "N+N / 3+3", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "pue": 1.2 + } + }, + "mi355x": { + "functionName": "mi355x_chassis_power", + "configFactory": "MI355XChassisConfig", + "defaultConfig": { + "gpu_label": "AMD Instinct MI355X OAM", + "gpu_component_key": "mi355x_oam_8x_measured", + "n_gpu": 8, + "gpu_memory_total_gb": 2304.0, + "gpu_idle_example_w_per_gpu": 250.0, + "gpu_decode_example_w_per_gpu": 700.0, + "gpu_prefill_example_w_per_gpu": 1150.0, + "gpu_aggregate_example_w_per_gpu": 1080.0, + "gpu_peak_w_per_gpu": 1400.0, + "ubb": { + "gpu_label": "AMD Instinct OAM", + "gpu_component_key": "amd_oam_8x_measured", + "n_gpu": 8, + "retimer": { + "n_retimer": 8, + "lanes_per_retimer": 16, + "pcie_gtps": 32.0, + "analog_serdes_equalization_w": 9.5, + "digital_floor_w": 2.0, + "digital_variable_w": 1.0 + }, + "residual": { + "management_controller_w": 10.0, + "fpga_cpld_sequencing_w": 8.0, + "hsc_power_monitor_w": 6.0, + "clock_refclk_reset_w": 5.0, + "sensors_i2c_fru_led_w": 4.0, + "aux_rails_misc_w": 12.0, + "normal_low_w": 30.0, + "normal_high_w": 70.0, + "conservative_cap_w": 90.0 + } + }, + "cpu": { + "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)", + "n_cpu": 2, + "cores_per_cpu": 64, + "threads_per_cpu": 128, + "package_tdp_w": 400.0, + "base_ghz": 3.3, + "max_turbo_ghz": 4.3, + "l3_cache_mb": 384.0, + "memory_channels_per_cpu": 12, + "pcie_lanes_per_cpu": 160, + "package_static_w": 28.0, + "uncore_io_baseline_w": 48.0, + "memory_controller_baseline_w": 34.0, + "uncore_dynamic_max_w": 14.0, + "memory_controller_dynamic_max_w": 14.0, + "core_curve_r": 1.62, + "vrm_efficiency": 0.92 + }, + "dram": { + "label": "MI355X host, 24x256GB DDR5-6400 RDIMM/MRDIMM", + "status": "DRAFT - pending human verification", + "n_sockets": 2, + "channels_per_socket": 12, + "n_dimm": 24, + "capacity_gb_per_dimm": 256.0, + "data_rate_mtps": 6400.0, + "dimm_type": "RDIMM/MRDIMM", + "background_w": 3.9549999999999996, + "refresh_w": 1.015, + "termination_w": 2.0300000000000002, + "io_dynamic_max_w": 4.95, + "core_dynamic_max_w": 6.05, + "activity_exponent": 1.0 + }, + "pollara400": { + "n_nic": 8, + "aggregate_gbps_per_nic": 400.0, + "pcie_generation": 5, + "pcie_lanes": 16, + "net_serdes_w": 10.5, + "pcie_serdes_w": 6.0, + "board_mgmt_w": 2.5, + "packet_engine_max_w": 12.5, + "packet_engine_floor_frac": 0.64, + "include_optic": true, + "optic_w": 8.0 + }, + "pcie_switch": { + "n_switch": 4, + "lanes_per_switch": 144, + "ports_per_switch": 72, + "pcie_gtps": 32.0, + "serdes_phy_floor_w": 30.5, + "control_leakage_clock_w": 6.0, + "fabric_datapath_floor_w": 4.0, + "fabric_datapath_variable_w": 8.5, + "stress_cap_per_switch_w": 70.0 + }, + "nvme": { + "n_front_u2": 8, + "front_u2_idle_w": 4.5, + "n_boot_m2": 2, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "residual": { + "bmc_ipmi_w": 12.0, + "onboard_10gbe_w": 8.0, + "motherboard_pch_aux_w": 16.0, + "cpld_tpm_superio_w": 8.0, + "clock_sensor_fru_w": 6.0, + "storage_backplane_idle_w": 16.0, + "front_panel_usb_led_w": 4.0, + "aux_margin_w": 12.0, + "normal_low_w": 60.0, + "normal_high_w": 120.0 + }, + "fans": { + "platform_label": "MI355X 10U air-cooled chassis", + "n_fan": 19, + "electrical_nameplate_w": 3600.0, + "min_pwm_frac": 0.25, + "normal_max_pwm_frac": 0.92, + "full_cooling_load_w": 15500.0, + "fan_curve_exponent": 1.04 + }, + "psu": { + "platform_label": "MI355X 10U chassis PSU bank", + "n_installed_psu": 6, + "n_load_sharing_psu": 6, + "n_redundant_capacity_psu": 4, + "psu_capacity_w": 6600.0, + "system_max_w": 26400.0, + "redundancy": "4+2", + "efficiency_curve": { + "0.05": 0.885, + "0.1": 0.92, + "0.2": 0.94, + "0.5": 0.96, + "1.0": 0.955 + } + }, + "pue": 1.2 + } + }, + "gb200": { + "functionName": "gb200_nvl72_rack_power", + "configFactory": "gb200_nvl72_rack_config", + "defaultConfig": { + "variant": "gb200", + "gpu_label": "GB200 Blackwell (NVL72)", + "n_compute_trays": 18, + "n_nvswitch_trays": 9, + "compute_tray": { + "n_gpu": 4, + "n_grace": 2, + "nic_generation": "connectx7", + "connectx7": { + "n_nic": 4, + "net_serdes_w": 9.5, + "pcie_serdes_w": 5.5, + "board_w": 2.0, + "digital_max_w": 10.0, + "digital_floor_frac": 0.65, + "include_optic": true, + "optic_w": 8.0 + }, + "connectx8": { + "n_nic": 4, + "include_optic": true, + "network_serdes_static_w": 18.0, + "pcie_switch_static_w": 28.0, + "board_mgmt_static_w": 5.0, + "digital_static_w": 12.0, + "network_dynamic_max_w": 8.0, + "pcie_switch_dynamic_max_w": 5.0, + "digital_dynamic_max_w": 10.0, + "optic_idle_w": 15.0, + "optic_dynamic_max_w": 2.0, + "max_nic_slot_power_ref_w": 75.0, + "normal_low_w_per_nic_with_optic": 70.0, + "normal_high_w_per_nic_with_optic": 100.0 + }, + "dpu": { + "n_dpu": 2, + "idle_w_per_dpu": 65.0, + "public_max_power_cap_w": 150.0, + "max_modeled_u_dpu": 0.02 + }, + "nvme": { + "n_front_u2": 4, + "front_u2_idle_w": 5.0, + "n_boot_m2": 1, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "fans_w": 130.0, + "board_residual_w": 40.0 + }, + "nvswitch": { + "n_asic": 2, + "serdes_lanes": 144, + "serdes_lane_gbps": 200.0, + "serdes_pj_per_bit": 2.5, + "serdes_floor_frac": 0.95, + "digital_max_w": 190.0, + "digital_floor_frac": 0.4 + }, + "nvswitch_tray_residual_w": 50.0, + "tray_input_conversion_efficiency": 0.9725, + "n_management_switches": 2, + "management_switch_w": 100.0, + "power_shelf": { + "n_shelves": 8, + "n_psu_per_shelf": 6, + "psu_capacity_w": 5500.0, + "redundancy": "N+N", + "busbar_nominal_v": 50.0, + "efficiency_curve": { + "0.1": 0.9, + "0.2": 0.94, + "0.3": 0.965, + "1.0": 0.965 + }, + "peak_efficiency_ref": 0.975 + }, + "regulator_loss_frac_of_tdp": 0.15, + "regulator_allowance_includes_grace": false, + "pue": 1.2, + "cooling": "direct_liquid", + "gpu_tdp_example_w_per_gpu": 1200.0, + "grace_tdp_example_w_per_socket": 300.0, + "module_tdp_example_w_per_superchip": 2700.0 + } + }, + "gb300": { + "functionName": "gb200_nvl72_rack_power", + "configFactory": "gb300_nvl72_rack_config", + "defaultConfig": { + "variant": "gb300", + "gpu_label": "GB300 Blackwell Ultra (NVL72)", + "n_compute_trays": 18, + "n_nvswitch_trays": 9, + "compute_tray": { + "n_gpu": 4, + "n_grace": 2, + "nic_generation": "connectx8", + "connectx7": { + "n_nic": 4, + "net_serdes_w": 9.5, + "pcie_serdes_w": 5.5, + "board_w": 2.0, + "digital_max_w": 10.0, + "digital_floor_frac": 0.65, + "include_optic": true, + "optic_w": 8.0 + }, + "connectx8": { + "n_nic": 4, + "include_optic": true, + "network_serdes_static_w": 18.0, + "pcie_switch_static_w": 28.0, + "board_mgmt_static_w": 5.0, + "digital_static_w": 12.0, + "network_dynamic_max_w": 8.0, + "pcie_switch_dynamic_max_w": 5.0, + "digital_dynamic_max_w": 10.0, + "optic_idle_w": 15.0, + "optic_dynamic_max_w": 2.0, + "max_nic_slot_power_ref_w": 75.0, + "normal_low_w_per_nic_with_optic": 70.0, + "normal_high_w_per_nic_with_optic": 100.0 + }, + "dpu": { + "n_dpu": 2, + "idle_w_per_dpu": 65.0, + "public_max_power_cap_w": 150.0, + "max_modeled_u_dpu": 0.02 + }, + "nvme": { + "n_front_u2": 4, + "front_u2_idle_w": 5.0, + "n_boot_m2": 1, + "boot_m2_idle_w": 2.0, + "max_modeled_u_nvme": 0.02 + }, + "fans_w": 130.0, + "board_residual_w": 40.0 + }, + "nvswitch": { + "n_asic": 2, + "serdes_lanes": 144, + "serdes_lane_gbps": 200.0, + "serdes_pj_per_bit": 2.5, + "serdes_floor_frac": 0.95, + "digital_max_w": 190.0, + "digital_floor_frac": 0.4 + }, + "nvswitch_tray_residual_w": 50.0, + "tray_input_conversion_efficiency": 0.9725, + "n_management_switches": 2, + "management_switch_w": 100.0, + "power_shelf": { + "n_shelves": 8, + "n_psu_per_shelf": 6, + "psu_capacity_w": 5500.0, + "redundancy": "N+N", + "busbar_nominal_v": 50.0, + "efficiency_curve": { + "0.1": 0.9, + "0.2": 0.94, + "0.3": 0.965, + "1.0": 0.965 + }, + "peak_efficiency_ref": 0.975 + }, + "regulator_loss_frac_of_tdp": 0.15, + "regulator_allowance_includes_grace": false, + "pue": 1.2, + "cooling": "direct_liquid", + "gpu_tdp_example_w_per_gpu": 1400.0, + "grace_tdp_example_w_per_socket": 300.0, + "module_tdp_example_w_per_superchip": 0.0 + } + } + }, "cases": [ { "hardware": "h100", @@ -3909,5 +5180,4015 @@ "expected": null, "referenceError": "dc_load_w=30559.8 exceeds modeled MI355X 10U chassis PSU bank capacity 26400.0 W" } + ], + "rackCases": [ + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.2, + "rackDcWatts": 12716.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1413.0, + "rackAcWatts": 14129.7, + "facilityWatts": 14129.7, + "perGpuAcWatts": 196.2, + "perGpuFacilityWatts": 196.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.2, + "rackDcWatts": 12716.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1413.0, + "rackAcWatts": 14129.7, + "facilityWatts": 15542.7, + "perGpuAcWatts": 196.2, + "perGpuFacilityWatts": 215.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.2, + "rackDcWatts": 12716.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1413.0, + "rackAcWatts": 14129.7, + "facilityWatts": 16955.6, + "perGpuAcWatts": 196.2, + "perGpuFacilityWatts": 235.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.8, + "rackDcWatts": 12738.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1415.4, + "rackAcWatts": 14154.4, + "facilityWatts": 14154.4, + "perGpuAcWatts": 196.6, + "perGpuFacilityWatts": 196.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.8, + "rackDcWatts": 12738.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1415.4, + "rackAcWatts": 14154.4, + "facilityWatts": 15569.8, + "perGpuAcWatts": 196.6, + "perGpuFacilityWatts": 216.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.8, + "rackDcWatts": 12738.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1415.4, + "rackAcWatts": 14154.4, + "facilityWatts": 16985.3, + "perGpuAcWatts": 196.6, + "perGpuFacilityWatts": 235.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.125, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 13304.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 29329.2, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 407.4 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.125, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 13304.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 32262.1, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 448.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.125, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 13304.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 35195.0, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 488.8 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.325, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 13307.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 29333.3, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 407.4 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.325, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 13307.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 32266.6, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 448.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.325, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 13307.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 35200.0, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 488.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.525, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 13311.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 29337.2, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 407.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.525, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 13311.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 32270.9, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 448.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 739.525, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 13311.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 35204.6, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 489.0 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 853.3, + "rackDcWatts": 31229.4, + "powerShelfEfficiency": 0.9073, + "powerShelfLossWatts": 3190.1, + "rackAcWatts": 34419.5, + "facilityWatts": 34419.5, + "perGpuAcWatts": 478.0, + "perGpuFacilityWatts": 478.0 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 853.3, + "rackDcWatts": 31229.4, + "powerShelfEfficiency": 0.9073, + "powerShelfLossWatts": 3190.1, + "rackAcWatts": 34419.5, + "facilityWatts": 37861.5, + "perGpuAcWatts": 478.0, + "perGpuFacilityWatts": 525.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 853.3, + "rackDcWatts": 31229.4, + "powerShelfEfficiency": 0.9073, + "powerShelfLossWatts": 3190.1, + "rackAcWatts": 34419.5, + "facilityWatts": 41303.4, + "perGpuAcWatts": 478.0, + "perGpuFacilityWatts": 573.7 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1362.2, + "rackDcWatts": 49733.8, + "powerShelfEfficiency": 0.9354, + "powerShelfLossWatts": 3437.3, + "rackAcWatts": 53171.1, + "facilityWatts": 53171.1, + "perGpuAcWatts": 738.5, + "perGpuFacilityWatts": 738.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1362.2, + "rackDcWatts": 49733.8, + "powerShelfEfficiency": 0.9354, + "powerShelfLossWatts": 3437.3, + "rackAcWatts": 53171.1, + "facilityWatts": 58488.2, + "perGpuAcWatts": 738.5, + "perGpuFacilityWatts": 812.3 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1362.2, + "rackDcWatts": 49733.8, + "powerShelfEfficiency": 0.9354, + "powerShelfLossWatts": 3437.3, + "rackAcWatts": 53171.1, + "facilityWatts": 63805.3, + "perGpuAcWatts": 738.5, + "perGpuFacilityWatts": 886.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.458, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 38978.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 56166.6, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 780.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.458, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 38978.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 61783.3, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 858.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.458, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 38978.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 67399.9, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 936.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.658, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 38981.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 56170.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 780.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.658, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 38981.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 61787.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 858.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.658, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 38981.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 67404.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 936.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.858, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 38985.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 56173.9, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 780.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.858, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 38985.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 61791.3, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 858.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 2165.858, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 38985.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 67408.7, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 936.2 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1871.6, + "rackDcWatts": 68256.7, + "powerShelfEfficiency": 0.9546, + "powerShelfLossWatts": 3243.5, + "rackAcWatts": 71500.1, + "facilityWatts": 71500.1, + "perGpuAcWatts": 993.1, + "perGpuFacilityWatts": 993.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1871.6, + "rackDcWatts": 68256.7, + "powerShelfEfficiency": 0.9546, + "powerShelfLossWatts": 3243.5, + "rackAcWatts": 71500.1, + "facilityWatts": 78650.1, + "perGpuAcWatts": 993.1, + "perGpuFacilityWatts": 1092.4 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1871.6, + "rackDcWatts": 68256.7, + "powerShelfEfficiency": 0.9546, + "powerShelfLossWatts": 3243.5, + "rackAcWatts": 71500.1, + "facilityWatts": 85800.1, + "perGpuAcWatts": 993.1, + "perGpuFacilityWatts": 1191.7 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.792, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 64652.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 82069.0, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1139.8 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.792, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 64652.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 90275.9, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1253.8 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.792, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 64652.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 98482.8, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1367.8 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.992, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 64655.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 82072.5, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1139.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.992, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 64655.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 90279.8, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1253.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3591.992, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 64655.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 98487.0, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1367.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3592.192, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 64659.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 82076.3, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1139.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3592.192, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 64659.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 90283.9, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1253.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 3592.192, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 64659.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 98491.6, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1367.9 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2380.2, + "rackDcWatts": 86751.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3146.4, + "rackAcWatts": 89898.2, + "facilityWatts": 89898.2, + "perGpuAcWatts": 1248.6, + "perGpuFacilityWatts": 1248.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2380.2, + "rackDcWatts": 86751.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3146.4, + "rackAcWatts": 89898.2, + "facilityWatts": 98888.0, + "perGpuAcWatts": 1248.6, + "perGpuFacilityWatts": 1373.4 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2380.2, + "rackDcWatts": 86751.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3146.4, + "rackAcWatts": 89898.2, + "facilityWatts": 107877.8, + "perGpuAcWatts": 1248.6, + "perGpuFacilityWatts": 1498.3 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3092.8, + "rackDcWatts": 112664.4, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4086.3, + "rackAcWatts": 116750.6, + "facilityWatts": 116750.6, + "perGpuAcWatts": 1621.5, + "perGpuFacilityWatts": 1621.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3092.8, + "rackDcWatts": 112664.4, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4086.3, + "rackAcWatts": 116750.6, + "facilityWatts": 128425.7, + "perGpuAcWatts": 1621.5, + "perGpuFacilityWatts": 1783.7 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3092.8, + "rackDcWatts": 112664.4, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4086.3, + "rackAcWatts": 116750.6, + "facilityWatts": 140100.7, + "perGpuAcWatts": 1621.5, + "perGpuFacilityWatts": 1945.8 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3398.2, + "rackDcWatts": 123769.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4489.1, + "rackAcWatts": 128258.8, + "facilityWatts": 128258.8, + "perGpuAcWatts": 1781.4, + "perGpuFacilityWatts": 1781.4 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3398.2, + "rackDcWatts": 123769.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4489.1, + "rackAcWatts": 128258.8, + "facilityWatts": 141084.7, + "perGpuAcWatts": 1781.4, + "perGpuFacilityWatts": 1959.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3398.2, + "rackDcWatts": 123769.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4489.1, + "rackAcWatts": 128258.8, + "facilityWatts": 153910.6, + "perGpuAcWatts": 1781.4, + "perGpuFacilityWatts": 2137.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4009.0, + "rackDcWatts": 145980.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5294.6, + "rackAcWatts": 151275.2, + "facilityWatts": 151275.2, + "perGpuAcWatts": 2101.0, + "perGpuFacilityWatts": 2101.0 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4009.0, + "rackDcWatts": 145980.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5294.6, + "rackAcWatts": 151275.2, + "facilityWatts": 166402.7, + "perGpuAcWatts": 2101.0, + "perGpuFacilityWatts": 2311.1 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4009.0, + "rackDcWatts": 145980.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5294.6, + "rackAcWatts": 151275.2, + "facilityWatts": 181530.2, + "perGpuAcWatts": 2101.0, + "perGpuFacilityWatts": 2521.3 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.125, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 244370.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 273571.2, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 3799.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.125, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 244370.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 300928.3, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 4179.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.125, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 244370.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 328285.4, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 4559.5 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.325, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 244373.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 273575.1, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 3799.7 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.325, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 244373.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 300932.6, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 4179.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.325, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 244373.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 328290.1, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 4559.6 + } + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.525, + "pue": 1.0, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.525, + "pue": 1.1, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 13576.525, + "pue": 1.2, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.0, + "expected": null, + "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.1, + "expected": null, + "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.2, + "expected": null, + "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 2.0, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.2, + "rackDcWatts": 12717.8, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1413.1, + "rackAcWatts": 14130.9, + "facilityWatts": 14130.9, + "perGpuAcWatts": 196.3, + "perGpuFacilityWatts": 196.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 2.0, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 344.2, + "rackDcWatts": 12717.8, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1413.1, + "rackAcWatts": 14130.9, + "facilityWatts": 15544.0, + "perGpuAcWatts": 196.3, + "perGpuFacilityWatts": 215.9 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 5401.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 496.9, + "rackDcWatts": 18269.6, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2030.0, + "rackAcWatts": 20299.5, + "facilityWatts": 20299.5, + "perGpuAcWatts": 281.9, + "perGpuFacilityWatts": 281.9 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 5401.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 496.9, + "rackDcWatts": 18269.6, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2030.0, + "rackAcWatts": 20299.5, + "facilityWatts": 22329.5, + "perGpuAcWatts": 281.9, + "perGpuFacilityWatts": 310.1 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 10810.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 649.9, + "rackDcWatts": 23831.5, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2647.9, + "rackAcWatts": 26479.5, + "facilityWatts": 26479.5, + "perGpuAcWatts": 367.8, + "perGpuFacilityWatts": 367.8 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 10810.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 649.9, + "rackDcWatts": 23831.5, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2647.9, + "rackAcWatts": 26479.5, + "facilityWatts": 29127.5, + "perGpuAcWatts": 367.8, + "perGpuFacilityWatts": 404.5 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 27.4, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 345.0, + "rackDcWatts": 12743.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1416.0, + "rackAcWatts": 14159.9, + "facilityWatts": 14159.9, + "perGpuAcWatts": 196.7, + "perGpuFacilityWatts": 196.7 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 27.4, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 345.0, + "rackDcWatts": 12743.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1416.0, + "rackAcWatts": 14159.9, + "facilityWatts": 15575.9, + "perGpuAcWatts": 196.7, + "perGpuFacilityWatts": 216.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 5426.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 497.6, + "rackDcWatts": 18295.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2032.9, + "rackAcWatts": 20328.6, + "facilityWatts": 20328.6, + "perGpuAcWatts": 282.3, + "perGpuFacilityWatts": 282.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 5426.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 497.6, + "rackDcWatts": 18295.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2032.9, + "rackAcWatts": 20328.6, + "facilityWatts": 22361.5, + "perGpuAcWatts": 282.3, + "perGpuFacilityWatts": 310.6 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 10835.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 650.6, + "rackDcWatts": 23857.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2650.9, + "rackAcWatts": 26508.5, + "facilityWatts": 26508.5, + "perGpuAcWatts": 368.2, + "perGpuFacilityWatts": 368.2 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 10835.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 650.6, + "rackDcWatts": 23857.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2650.9, + "rackAcWatts": 26508.5, + "facilityWatts": 29159.4, + "perGpuAcWatts": 368.2, + "perGpuFacilityWatts": 405.0 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 42353.8, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1541.9, + "rackDcWatts": 56267.3, + "powerShelfEfficiency": 0.9433, + "powerShelfLossWatts": 3383.2, + "rackAcWatts": 59650.5, + "facilityWatts": 59650.5, + "perGpuAcWatts": 828.5, + "perGpuFacilityWatts": 828.5 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 42353.8, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1541.9, + "rackDcWatts": 56267.3, + "powerShelfEfficiency": 0.9433, + "powerShelfLossWatts": 3383.2, + "rackAcWatts": 59650.5, + "facilityWatts": 65615.6, + "perGpuAcWatts": 828.5, + "perGpuFacilityWatts": 911.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 47752.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1694.5, + "rackDcWatts": 61819.1, + "powerShelfEfficiency": 0.9485, + "powerShelfLossWatts": 3353.7, + "rackAcWatts": 65172.8, + "facilityWatts": 65172.8, + "perGpuAcWatts": 905.2, + "perGpuFacilityWatts": 905.2 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 47752.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1694.5, + "rackDcWatts": 61819.1, + "powerShelfEfficiency": 0.9485, + "powerShelfLossWatts": 3353.7, + "rackAcWatts": 65172.8, + "facilityWatts": 71690.1, + "perGpuAcWatts": 905.2, + "perGpuFacilityWatts": 995.7 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 53161.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1847.5, + "rackDcWatts": 67381.0, + "powerShelfEfficiency": 0.9538, + "powerShelfLossWatts": 3263.2, + "rackAcWatts": 70644.2, + "facilityWatts": 70644.2, + "perGpuAcWatts": 981.2, + "perGpuFacilityWatts": 981.2 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 53161.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1847.5, + "rackDcWatts": 67381.0, + "powerShelfEfficiency": 0.9538, + "powerShelfLossWatts": 3263.2, + "rackAcWatts": 70644.2, + "facilityWatts": 77708.6, + "perGpuAcWatts": 981.2, + "perGpuFacilityWatts": 1079.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 63535.6, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2140.8, + "rackDcWatts": 78048.0, + "powerShelfEfficiency": 0.9639, + "powerShelfLossWatts": 2922.3, + "rackAcWatts": 80970.3, + "facilityWatts": 80970.3, + "perGpuAcWatts": 1124.6, + "perGpuFacilityWatts": 1124.6 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 63535.6, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2140.8, + "rackDcWatts": 78048.0, + "powerShelfEfficiency": 0.9639, + "powerShelfLossWatts": 2922.3, + "rackAcWatts": 80970.3, + "facilityWatts": 89067.3, + "perGpuAcWatts": 1124.6, + "perGpuFacilityWatts": 1237.0 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 68934.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2293.5, + "rackDcWatts": 83599.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3032.1, + "rackAcWatts": 86631.9, + "facilityWatts": 86631.9, + "perGpuAcWatts": 1203.2, + "perGpuFacilityWatts": 1203.2 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 68934.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2293.5, + "rackDcWatts": 83599.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3032.1, + "rackAcWatts": 86631.9, + "facilityWatts": 95295.1, + "perGpuAcWatts": 1203.2, + "perGpuFacilityWatts": 1323.5 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 74343.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2446.4, + "rackDcWatts": 89161.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3233.8, + "rackAcWatts": 92395.6, + "facilityWatts": 92395.6, + "perGpuAcWatts": 1283.3, + "perGpuFacilityWatts": 1283.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 74343.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2446.4, + "rackDcWatts": 89161.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3233.8, + "rackAcWatts": 92395.6, + "facilityWatts": 101635.2, + "perGpuAcWatts": 1283.3, + "perGpuFacilityWatts": 1411.6 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 101648.0, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3218.5, + "rackDcWatts": 117238.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4252.2, + "rackAcWatts": 121490.3, + "facilityWatts": 121490.3, + "perGpuAcWatts": 1687.4, + "perGpuFacilityWatts": 1687.4 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 101648.0, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3218.5, + "rackDcWatts": 117238.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4252.2, + "rackAcWatts": 121490.3, + "facilityWatts": 133639.3, + "perGpuAcWatts": 1687.4, + "perGpuFacilityWatts": 1856.1 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 107047.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3371.2, + "rackDcWatts": 122789.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4453.5, + "rackAcWatts": 127243.4, + "facilityWatts": 127243.4, + "perGpuAcWatts": 1767.3, + "perGpuFacilityWatts": 1767.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 107047.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3371.2, + "rackDcWatts": 122789.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4453.5, + "rackAcWatts": 127243.4, + "facilityWatts": 139967.7, + "perGpuAcWatts": 1767.3, + "perGpuFacilityWatts": 1944.0 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 112456.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3524.2, + "rackDcWatts": 128351.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4655.2, + "rackAcWatts": 133007.1, + "facilityWatts": 133007.1, + "perGpuAcWatts": 1847.3, + "perGpuFacilityWatts": 1847.3 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 112456.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3524.2, + "rackDcWatts": 128351.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4655.2, + "rackAcWatts": 133007.1, + "facilityWatts": 146307.8, + "perGpuAcWatts": 1847.3, + "perGpuFacilityWatts": 2032.1 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 118589.1, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3697.6, + "rackDcWatts": 134658.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4884.0, + "rackAcWatts": 139542.3, + "facilityWatts": 139542.3, + "perGpuAcWatts": 1938.1, + "perGpuFacilityWatts": 1938.1 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 118589.1, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3697.6, + "rackDcWatts": 134658.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4884.0, + "rackAcWatts": 139542.3, + "facilityWatts": 153496.5, + "perGpuAcWatts": 1938.1, + "perGpuFacilityWatts": 2131.9 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 123988.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3850.3, + "rackDcWatts": 140210.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5085.3, + "rackAcWatts": 145295.5, + "facilityWatts": 145295.5, + "perGpuAcWatts": 2018.0, + "perGpuFacilityWatts": 2018.0 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 123988.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3850.3, + "rackDcWatts": 140210.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5085.3, + "rackAcWatts": 145295.5, + "facilityWatts": 159825.1, + "perGpuAcWatts": 2018.0, + "perGpuFacilityWatts": 2219.8 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 129397.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4003.2, + "rackDcWatts": 145772.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5287.1, + "rackAcWatts": 151059.1, + "facilityWatts": 151059.1, + "perGpuAcWatts": 2098.0, + "perGpuFacilityWatts": 2098.0 + } + }, + { + "hardware": "gb200", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 129397.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 8064.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4003.2, + "rackDcWatts": 145772.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5287.1, + "rackAcWatts": 151059.1, + "facilityWatts": 166165.0, + "perGpuAcWatts": 2098.0, + "perGpuFacilityWatts": 2307.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 440.4, + "rackDcWatts": 16214.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1801.7, + "rackAcWatts": 18016.6, + "facilityWatts": 18016.6, + "perGpuAcWatts": 250.2, + "perGpuFacilityWatts": 250.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 440.4, + "rackDcWatts": 16214.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1801.7, + "rackAcWatts": 18016.6, + "facilityWatts": 19818.3, + "perGpuAcWatts": 250.2, + "perGpuFacilityWatts": 275.3 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 0.05, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 0.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 440.4, + "rackDcWatts": 16214.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1801.7, + "rackAcWatts": 18016.6, + "facilityWatts": 21619.9, + "perGpuAcWatts": 250.2, + "perGpuFacilityWatts": 300.3 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 441.0, + "rackDcWatts": 16237.1, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1804.1, + "rackAcWatts": 18041.2, + "facilityWatts": 18041.2, + "perGpuAcWatts": 250.6, + "perGpuFacilityWatts": 250.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 441.0, + "rackDcWatts": 16237.1, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1804.1, + "rackAcWatts": 18041.2, + "facilityWatts": 19845.3, + "perGpuAcWatts": 250.6, + "perGpuFacilityWatts": 275.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1.25, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 22.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 441.0, + "rackDcWatts": 16237.1, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1804.1, + "rackAcWatts": 18041.2, + "facilityWatts": 21649.4, + "perGpuAcWatts": 250.6, + "perGpuFacilityWatts": 300.7 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.125, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 9902.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 29329.2, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 407.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.125, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 9902.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 32262.1, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 448.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.125, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 9902.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.4, + "rackDcWatts": 26396.2, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2932.9, + "rackAcWatts": 29329.2, + "facilityWatts": 35195.0, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 488.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.325, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 9905.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 29333.3, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 407.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.325, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 9905.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 32266.6, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 448.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.325, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 9905.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.5, + "rackDcWatts": 26399.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.3, + "rackAcWatts": 29333.3, + "facilityWatts": 35200.0, + "perGpuAcWatts": 407.4, + "perGpuFacilityWatts": 488.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.525, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 9909.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 29337.2, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 407.5 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.525, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 9909.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 32270.9, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 448.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 550.525, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 9909.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 720.6, + "rackDcWatts": 26403.7, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2933.6, + "rackAcWatts": 29337.2, + "facilityWatts": 35204.6, + "perGpuAcWatts": 407.5, + "perGpuFacilityWatts": 489.0 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 949.5, + "rackDcWatts": 34727.6, + "powerShelfEfficiency": 0.9126, + "powerShelfLossWatts": 3325.1, + "rackAcWatts": 38052.8, + "facilityWatts": 38052.8, + "perGpuAcWatts": 528.5, + "perGpuFacilityWatts": 528.5 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 949.5, + "rackDcWatts": 34727.6, + "powerShelfEfficiency": 0.9126, + "powerShelfLossWatts": 3325.1, + "rackAcWatts": 38052.8, + "facilityWatts": 41858.1, + "perGpuAcWatts": 528.5, + "perGpuFacilityWatts": 581.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1000.25, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 18004.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 949.5, + "rackDcWatts": 34727.6, + "powerShelfEfficiency": 0.9126, + "powerShelfLossWatts": 3325.1, + "rackAcWatts": 38052.8, + "facilityWatts": 45663.4, + "perGpuAcWatts": 528.5, + "perGpuFacilityWatts": 634.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.458, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 35576.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 56166.6, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 780.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.458, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 35576.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 61783.3, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 858.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.458, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 35576.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.4, + "rackDcWatts": 52796.2, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.3, + "rackAcWatts": 56166.6, + "facilityWatts": 67399.9, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 936.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.658, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 35579.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 56170.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 780.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.658, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 35579.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 61787.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 858.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.658, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 35579.8, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.5, + "rackDcWatts": 52799.9, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56170.2, + "facilityWatts": 67404.2, + "perGpuAcWatts": 780.1, + "perGpuFacilityWatts": 936.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.858, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 35583.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 56173.9, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 780.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.858, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 35583.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 61791.3, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 858.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 1976.858, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 35583.4, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1446.6, + "rackDcWatts": 52803.6, + "powerShelfEfficiency": 0.94, + "powerShelfLossWatts": 3370.2, + "rackAcWatts": 56173.9, + "facilityWatts": 67408.7, + "perGpuAcWatts": 780.2, + "perGpuFacilityWatts": 936.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1458.4, + "rackDcWatts": 53232.0, + "powerShelfEfficiency": 0.9404, + "powerShelfLossWatts": 3373.2, + "rackAcWatts": 56605.1, + "facilityWatts": 56605.1, + "perGpuAcWatts": 786.2, + "perGpuFacilityWatts": 786.2 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1458.4, + "rackDcWatts": 53232.0, + "powerShelfEfficiency": 0.9404, + "powerShelfLossWatts": 3373.2, + "rackAcWatts": 56605.1, + "facilityWatts": 62265.6, + "perGpuAcWatts": 786.2, + "perGpuFacilityWatts": 864.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 2000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 36000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1458.4, + "rackDcWatts": 53232.0, + "powerShelfEfficiency": 0.9404, + "powerShelfLossWatts": 3373.2, + "rackAcWatts": 56605.1, + "facilityWatts": 67926.1, + "perGpuAcWatts": 786.2, + "perGpuFacilityWatts": 943.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1967.8, + "rackDcWatts": 71754.9, + "powerShelfEfficiency": 0.9579, + "powerShelfLossWatts": 3149.8, + "rackAcWatts": 74904.6, + "facilityWatts": 74904.6, + "perGpuAcWatts": 1040.3, + "perGpuFacilityWatts": 1040.3 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1967.8, + "rackDcWatts": 71754.9, + "powerShelfEfficiency": 0.9579, + "powerShelfLossWatts": 3149.8, + "rackAcWatts": 74904.6, + "facilityWatts": 82395.1, + "perGpuAcWatts": 1040.3, + "perGpuFacilityWatts": 1144.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3000.75, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 54013.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1967.8, + "rackDcWatts": 71754.9, + "powerShelfEfficiency": 0.9579, + "powerShelfLossWatts": 3149.8, + "rackAcWatts": 74904.6, + "facilityWatts": 89885.5, + "perGpuAcWatts": 1040.3, + "perGpuFacilityWatts": 1248.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.792, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 61250.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 82069.0, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1139.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.792, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 61250.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 90275.9, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1253.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.792, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 61250.3, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.4, + "rackDcWatts": 79196.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82069.0, + "facilityWatts": 98482.8, + "perGpuAcWatts": 1139.8, + "perGpuFacilityWatts": 1367.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.992, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 61253.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 82072.5, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1139.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.992, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 61253.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 90279.8, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1253.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3402.992, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 61253.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.5, + "rackDcWatts": 79200.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.5, + "rackAcWatts": 82072.5, + "facilityWatts": 98487.0, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1367.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3403.192, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 61257.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 82076.3, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1139.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3403.192, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 61257.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 90283.9, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1253.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 3403.192, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 61257.5, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2172.6, + "rackDcWatts": 79203.7, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2872.7, + "rackAcWatts": 82076.3, + "facilityWatts": 98491.6, + "perGpuAcWatts": 1139.9, + "perGpuFacilityWatts": 1367.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2476.4, + "rackDcWatts": 90250.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3273.3, + "rackAcWatts": 93523.3, + "facilityWatts": 93523.3, + "perGpuAcWatts": 1298.9, + "perGpuFacilityWatts": 1298.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2476.4, + "rackDcWatts": 90250.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3273.3, + "rackAcWatts": 93523.3, + "facilityWatts": 102875.6, + "perGpuAcWatts": 1298.9, + "perGpuFacilityWatts": 1428.8 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 4000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 72000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2476.4, + "rackDcWatts": 90250.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3273.3, + "rackAcWatts": 93523.3, + "facilityWatts": 112228.0, + "perGpuAcWatts": 1298.9, + "perGpuFacilityWatts": 1558.7 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3189.0, + "rackDcWatts": 116162.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4213.2, + "rackAcWatts": 120375.7, + "facilityWatts": 120375.7, + "perGpuAcWatts": 1671.9, + "perGpuFacilityWatts": 1671.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3189.0, + "rackDcWatts": 116162.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4213.2, + "rackAcWatts": 120375.7, + "facilityWatts": 132413.3, + "perGpuAcWatts": 1671.9, + "perGpuFacilityWatts": 1839.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 5400.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 97200.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3189.0, + "rackDcWatts": 116162.6, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4213.2, + "rackAcWatts": 120375.7, + "facilityWatts": 144450.8, + "perGpuAcWatts": 1671.9, + "perGpuFacilityWatts": 2006.3 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3494.4, + "rackDcWatts": 127268.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4615.9, + "rackAcWatts": 131883.9, + "facilityWatts": 131883.9, + "perGpuAcWatts": 1831.7, + "perGpuFacilityWatts": 1831.7 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3494.4, + "rackDcWatts": 127268.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4615.9, + "rackAcWatts": 131883.9, + "facilityWatts": 145072.3, + "perGpuAcWatts": 1831.7, + "perGpuFacilityWatts": 2014.9 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 6000.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 108000.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3494.4, + "rackDcWatts": 127268.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4615.9, + "rackAcWatts": 131883.9, + "facilityWatts": 158260.7, + "perGpuAcWatts": 1831.7, + "perGpuFacilityWatts": 2198.1 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4105.2, + "rackDcWatts": 149478.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5421.5, + "rackAcWatts": 154900.3, + "facilityWatts": 154900.3, + "perGpuAcWatts": 2151.4, + "perGpuFacilityWatts": 2151.4 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4105.2, + "rackDcWatts": 149478.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5421.5, + "rackAcWatts": 154900.3, + "facilityWatts": 170390.3, + "perGpuAcWatts": 2151.4, + "perGpuFacilityWatts": 2366.5 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 7200.0, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 129600.0, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4105.2, + "rackDcWatts": 149478.8, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5421.5, + "rackAcWatts": 154900.3, + "facilityWatts": 185880.4, + "perGpuAcWatts": 2151.4, + "perGpuFacilityWatts": 2581.7 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.125, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 240968.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 273571.2, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 3799.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.125, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 240968.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 300928.3, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 4179.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.125, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 240968.2, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.4, + "rackDcWatts": 263996.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.0, + "rackAcWatts": 273571.2, + "facilityWatts": 328285.4, + "perGpuAcWatts": 3799.6, + "perGpuFacilityWatts": 4559.5 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.325, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 240971.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 273575.1, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 3799.7 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.325, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 240971.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 300932.6, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 4179.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.325, + "pue": 1.2, + "expected": { + "computeModulesDcWatts": 240971.9, + "regulatorAllowanceWatts": 0.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 7254.5, + "rackDcWatts": 263999.9, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 9575.1, + "rackAcWatts": 273575.1, + "facilityWatts": 328290.1, + "perGpuAcWatts": 3799.7, + "perGpuFacilityWatts": 4559.6 + } + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.525, + "pue": 1.0, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.525, + "pue": 1.1, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 13387.525, + "pue": 1.2, + "expected": null, + "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.0, + "expected": null, + "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.1, + "expected": null, + "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "module", + "moduleWattsPerTray": 14666.666666666666, + "pue": 1.2, + "expected": null, + "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W" + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 2.0, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 440.4, + "rackDcWatts": 16216.0, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1801.8, + "rackAcWatts": 18017.8, + "facilityWatts": 18017.8, + "perGpuAcWatts": 250.2, + "perGpuFacilityWatts": 250.2 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 2.0, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 440.4, + "rackDcWatts": 16216.0, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1801.8, + "rackAcWatts": 18017.8, + "facilityWatts": 19819.6, + "perGpuAcWatts": 250.2, + "perGpuFacilityWatts": 275.3 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 5401.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 593.1, + "rackDcWatts": 21767.8, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2418.6, + "rackAcWatts": 24186.4, + "facilityWatts": 24186.4, + "perGpuAcWatts": 335.9, + "perGpuFacilityWatts": 335.9 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 5401.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 593.1, + "rackDcWatts": 21767.8, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2418.6, + "rackAcWatts": 24186.4, + "facilityWatts": 26605.0, + "perGpuAcWatts": 335.9, + "perGpuFacilityWatts": 369.5 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 10810.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 746.1, + "rackDcWatts": 27329.7, + "powerShelfEfficiency": 0.9014, + "powerShelfLossWatts": 2989.2, + "rackAcWatts": 30318.9, + "facilityWatts": 30318.9, + "perGpuAcWatts": 421.1, + "perGpuFacilityWatts": 421.1 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 0.05, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 10810.1, + "regulatorAllowanceWatts": 0.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 746.1, + "rackDcWatts": 27329.7, + "powerShelfEfficiency": 0.9014, + "powerShelfLossWatts": 2989.2, + "rackAcWatts": 30318.9, + "facilityWatts": 33350.8, + "perGpuAcWatts": 421.1, + "perGpuFacilityWatts": 463.2 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 27.4, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 441.2, + "rackDcWatts": 16242.1, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1804.7, + "rackAcWatts": 18046.8, + "facilityWatts": 18046.8, + "perGpuAcWatts": 250.6, + "perGpuFacilityWatts": 250.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 27.4, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 441.2, + "rackDcWatts": 16242.1, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 1804.7, + "rackAcWatts": 18046.8, + "facilityWatts": 19851.5, + "perGpuAcWatts": 250.6, + "perGpuFacilityWatts": 275.7 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 5426.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 593.8, + "rackDcWatts": 21793.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2421.5, + "rackAcWatts": 24215.4, + "facilityWatts": 24215.4, + "perGpuAcWatts": 336.3, + "perGpuFacilityWatts": 336.3 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 5426.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 593.8, + "rackDcWatts": 21793.9, + "powerShelfEfficiency": 0.9, + "powerShelfLossWatts": 2421.5, + "rackAcWatts": 24215.4, + "facilityWatts": 26636.9, + "perGpuAcWatts": 336.3, + "perGpuFacilityWatts": 370.0 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 10835.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 746.8, + "rackDcWatts": 27355.9, + "powerShelfEfficiency": 0.9014, + "powerShelfLossWatts": 2990.7, + "rackAcWatts": 30346.6, + "facilityWatts": 30346.6, + "perGpuAcWatts": 421.5, + "perGpuFacilityWatts": 421.5 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 1.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 10835.5, + "regulatorAllowanceWatts": 4.0, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 746.8, + "rackDcWatts": 27355.9, + "powerShelfEfficiency": 0.9014, + "powerShelfLossWatts": 2990.7, + "rackAcWatts": 30346.6, + "facilityWatts": 33381.3, + "perGpuAcWatts": 421.5, + "perGpuFacilityWatts": 463.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 42353.8, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1638.1, + "rackDcWatts": 59765.5, + "powerShelfEfficiency": 0.9466, + "powerShelfLossWatts": 3371.8, + "rackAcWatts": 63137.3, + "facilityWatts": 63137.3, + "perGpuAcWatts": 876.9, + "perGpuFacilityWatts": 876.9 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 42353.8, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1638.1, + "rackDcWatts": 59765.5, + "powerShelfEfficiency": 0.9466, + "powerShelfLossWatts": 3371.8, + "rackAcWatts": 63137.3, + "facilityWatts": 69451.0, + "perGpuAcWatts": 876.9, + "perGpuFacilityWatts": 964.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 47752.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1790.7, + "rackDcWatts": 65317.3, + "powerShelfEfficiency": 0.9519, + "powerShelfLossWatts": 3303.9, + "rackAcWatts": 68621.1, + "facilityWatts": 68621.1, + "perGpuAcWatts": 953.1, + "perGpuFacilityWatts": 953.1 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 47752.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1790.7, + "rackDcWatts": 65317.3, + "powerShelfEfficiency": 0.9519, + "powerShelfLossWatts": 3303.9, + "rackAcWatts": 68621.1, + "facilityWatts": 75483.2, + "perGpuAcWatts": 953.1, + "perGpuFacilityWatts": 1048.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 53161.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1943.7, + "rackDcWatts": 70879.2, + "powerShelfEfficiency": 0.9571, + "powerShelfLossWatts": 3175.4, + "rackAcWatts": 74054.6, + "facilityWatts": 74054.6, + "perGpuAcWatts": 1028.5, + "perGpuFacilityWatts": 1028.5 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 2000.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 53161.9, + "regulatorAllowanceWatts": 6352.9, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 1943.7, + "rackDcWatts": 70879.2, + "powerShelfEfficiency": 0.9571, + "powerShelfLossWatts": 3175.4, + "rackAcWatts": 74054.6, + "facilityWatts": 81460.1, + "perGpuAcWatts": 1028.5, + "perGpuFacilityWatts": 1131.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 63535.6, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2237.0, + "rackDcWatts": 81546.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2957.6, + "rackAcWatts": 84503.9, + "facilityWatts": 84503.9, + "perGpuAcWatts": 1173.7, + "perGpuFacilityWatts": 1173.7 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 63535.6, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2237.0, + "rackDcWatts": 81546.2, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 2957.6, + "rackAcWatts": 84503.9, + "facilityWatts": 92954.3, + "perGpuAcWatts": 1173.7, + "perGpuFacilityWatts": 1291.0 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 68934.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2389.7, + "rackDcWatts": 87098.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3159.0, + "rackAcWatts": 90257.0, + "facilityWatts": 90257.0, + "perGpuAcWatts": 1253.6, + "perGpuFacilityWatts": 1253.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 68934.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2389.7, + "rackDcWatts": 87098.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3159.0, + "rackAcWatts": 90257.0, + "facilityWatts": 99282.7, + "perGpuAcWatts": 1253.6, + "perGpuFacilityWatts": 1378.9 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 74343.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2542.6, + "rackDcWatts": 92660.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3360.7, + "rackAcWatts": 96020.7, + "facilityWatts": 96020.7, + "perGpuAcWatts": 1333.6, + "perGpuFacilityWatts": 1333.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 3000.25, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 74343.7, + "regulatorAllowanceWatts": 9530.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 2542.6, + "rackDcWatts": 92660.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 3360.7, + "rackAcWatts": 96020.7, + "facilityWatts": 105622.8, + "perGpuAcWatts": 1333.6, + "perGpuFacilityWatts": 1467.0 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 101648.0, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3314.7, + "rackDcWatts": 120736.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4379.0, + "rackAcWatts": 125115.3, + "facilityWatts": 125115.3, + "perGpuAcWatts": 1737.7, + "perGpuFacilityWatts": 1737.7 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 101648.0, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3314.7, + "rackDcWatts": 120736.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4379.0, + "rackAcWatts": 125115.3, + "facilityWatts": 137626.8, + "perGpuAcWatts": 1737.7, + "perGpuFacilityWatts": 1911.5 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 107047.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3467.4, + "rackDcWatts": 126288.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4580.4, + "rackAcWatts": 130868.5, + "facilityWatts": 130868.5, + "perGpuAcWatts": 1817.6, + "perGpuFacilityWatts": 1817.6 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 107047.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3467.4, + "rackDcWatts": 126288.1, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4580.4, + "rackAcWatts": 130868.5, + "facilityWatts": 143955.4, + "perGpuAcWatts": 1817.6, + "perGpuFacilityWatts": 1999.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 112456.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3620.4, + "rackDcWatts": 131850.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4782.1, + "rackAcWatts": 136632.2, + "facilityWatts": 136632.2, + "perGpuAcWatts": 1897.7, + "perGpuFacilityWatts": 1897.7 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 4800.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 112456.1, + "regulatorAllowanceWatts": 15247.1, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3620.4, + "rackDcWatts": 131850.0, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 4782.1, + "rackAcWatts": 136632.2, + "facilityWatts": 150295.4, + "perGpuAcWatts": 1897.7, + "perGpuFacilityWatts": 2087.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 118589.1, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3793.8, + "rackDcWatts": 138156.5, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5010.9, + "rackAcWatts": 143167.4, + "facilityWatts": 143167.4, + "perGpuAcWatts": 1988.4, + "perGpuFacilityWatts": 1988.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 0.05, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 118589.1, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3793.8, + "rackDcWatts": 138156.5, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5010.9, + "rackAcWatts": 143167.4, + "facilityWatts": 157484.1, + "perGpuAcWatts": 1988.4, + "perGpuFacilityWatts": 2187.3 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 123988.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3946.5, + "rackDcWatts": 143708.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5212.2, + "rackAcWatts": 148920.5, + "facilityWatts": 148920.5, + "perGpuAcWatts": 2068.3, + "perGpuFacilityWatts": 2068.3 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 300.0, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 123988.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 3946.5, + "rackDcWatts": 143708.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5212.2, + "rackAcWatts": 148920.5, + "facilityWatts": 163812.6, + "perGpuAcWatts": 2068.3, + "perGpuFacilityWatts": 2275.2 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.0, + "expected": { + "computeModulesDcWatts": 129397.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4099.4, + "rackDcWatts": 149270.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5413.9, + "rackAcWatts": 154684.2, + "facilityWatts": 154684.2, + "perGpuAcWatts": 2148.4, + "perGpuFacilityWatts": 2148.4 + } + }, + { + "hardware": "gb300", + "basis": "gpu-plus-grace", + "gpuBoardWattsPerTray": 5600.0, + "graceSocketWattsPerTray": 600.5, + "pue": 1.1, + "expected": { + "computeModulesDcWatts": 129397.2, + "regulatorAllowanceWatts": 17788.2, + "trayStaticDcWatts": 11466.0, + "nvswitchTraysDcWatts": 4107.6, + "trayConversionLossWatts": 4099.4, + "rackDcWatts": 149270.3, + "powerShelfEfficiency": 0.965, + "powerShelfLossWatts": 5413.9, + "rackAcWatts": 154684.2, + "facilityWatts": 170152.6, + "perGpuAcWatts": 2148.4, + "perGpuFacilityWatts": 2363.2 + } + } ] } diff --git a/packages/app/src/lib/system-power-model.test.ts b/packages/app/src/lib/system-power-model.test.ts index ef04e7d5c..62c9521fa 100644 --- a/packages/app/src/lib/system-power-model.test.ts +++ b/packages/app/src/lib/system-power-model.test.ts @@ -1,15 +1,34 @@ +import { GPU_KEYS } from '@semianalysisai/inferencex-constants'; +import { createHash } from 'node:crypto'; +import { mkdtemp, mkdir, readFile, rm, writeFile } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import { dirname, resolve } from 'node:path'; import { describe, expect, it } from 'vitest'; +import { buildSystemPowerProvenance } from '../../scripts/update-system-power-provenance'; import { estimateChassisPower, + estimateRackPower, + type RackMeasuredInput, SUPPORTED_SYSTEM_POWER_HARDWARE, + SUPPORTED_SYSTEM_POWER_RACK_HARDWARE, SYSTEM_POWER_MODEL_REVISION, + SYSTEM_POWER_MODEL_METADATA, + systemPowerSourceSha256, } from './system-power-model'; import reference from './system-power-model.reference.json'; +const rackInput = (row: (typeof reference.rackCases)[number]): RackMeasuredInput => + row.moduleWattsPerTray === undefined + ? { + basis: 'gpu-plus-grace', + gpuBoardWattsPerTray: row.gpuBoardWattsPerTray, + graceSocketWattsPerTray: row.graceSocketWattsPerTray, + } + : { basis: 'module', moduleWattsPerTray: row.moduleWattsPerTray }; + describe('fixed 8k1k chassis model', () => { - it('matches the pinned Python implementation across platforms, fan/PSU boundaries, and PUE', () => { - expect(reference.modelRevision).toBe(SYSTEM_POWER_MODEL_REVISION); + it('matches the historical Python baseline across platforms, fan/PSU boundaries, and PUE', () => { expect(new Set(reference.cases.map((row) => row.hardware))).toEqual( new Set(SUPPORTED_SYSTEM_POWER_HARDWARE), ); @@ -36,15 +55,141 @@ describe('fixed 8k1k chassis model', () => { } }); - it('keeps measured input, modeled chassis AC, and post-AC facility power separate', () => { - const chassis = estimateChassisPower('H100', 4000, 1)!; - const facility = estimateChassisPower('h100', 4000, 1.2)!; - expect(chassis.measuredGpuWatts).toBe(4000); - expect(chassis.chassisAcWatts).toBe(6228.2); - expect(chassis.facilityWatts).toBe(chassis.chassisAcWatts); - expect(facility.chassisAcWatts).toBe(chassis.chassisAcWatts); - expect(facility.facilityWatts).toBe(7473.8); - // Fixed chassis overhead and nonlinear fan/PSU behavior forbid proportional GPU scaling. - expect(estimateChassisPower('h100', 8000)!.chassisAcWatts).not.toBe(chassis.chassisAcWatts * 2); + it('names every chassis and rack profile by its hardware registry key', () => { + for (const hardware of [ + ...SUPPORTED_SYSTEM_POWER_HARDWARE, + ...SUPPORTED_SYSTEM_POWER_RACK_HARDWARE, + ]) { + expect(GPU_KEYS.has(hardware), hardware).toBe(true); + } + expect(SUPPORTED_SYSTEM_POWER_RACK_HARDWARE).toEqual(['gb200', 'gb300']); + }); +}); + +describe('NVL72 rack model with measured compute-module input', () => { + it('matches the historical Python baseline across variants, bases, shelf knots, and PUE', () => { + expect(new Set(reference.rackCases.map((row) => row.hardware))).toEqual( + new Set(SUPPORTED_SYSTEM_POWER_RACK_HARDWARE), + ); + expect(new Set(reference.rackCases.map((row) => row.basis))).toEqual( + new Set(['module', 'gpu-plus-grace']), + ); + for (const row of reference.rackCases) { + const actual = estimateRackPower(row.hardware, rackInput(row), row.pue); + const context = `${row.hardware}: ${JSON.stringify(rackInput(row))}, PUE=${row.pue}`; + if (row.expected === null) { + expect(actual, context).toBeNull(); + } else { + expect(actual, context).toMatchObject(row.expected); + } + } + }); + + it('rejects unavailable measurements, PUE, chassis hardware, and shelf overflow', () => { + for (const watts of [0, -1, NaN, Infinity, -Infinity, Number.MAX_VALUE]) { + expect(estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: watts })).toBeNull(); + expect( + estimateRackPower('gb200', { + basis: 'gpu-plus-grace', + gpuBoardWattsPerTray: watts, + graceSocketWattsPerTray: 300, + }), + ).toBeNull(); + // A zero Grace socket reading means the socket was not measured, never that it drew nothing. + expect( + estimateRackPower('gb200', { + basis: 'gpu-plus-grace', + gpuBoardWattsPerTray: 3000, + graceSocketWattsPerTray: watts, + }), + ).toBeNull(); + } + for (const pue of [0, 0.99, NaN, Infinity, -Infinity, Number.MAX_VALUE]) { + expect( + estimateRackPower('gb300', { basis: 'module', moduleWattsPerTray: 4000 }, pue), + ).toBeNull(); + } + for (const hardware of [ + 'b200', + 'b300', + 'h100', + 'gb200-nvl', + 'gb200_dynamo-trt', + '', + 'toString', + ]) { + expect(estimateRackPower(hardware, { basis: 'module', moduleWattsPerTray: 4000 })).toBeNull(); + } + expect(estimateChassisPower('gb200', 4000)).toBeNull(); + }); +}); + +describe('app-owned model provenance', () => { + it('identifies the actual app equations, parameters and admission/PUE policy', async () => { + const root = resolve(import.meta.dirname, '../../../..'); + const metadata = SYSTEM_POWER_MODEL_METADATA; + expect(metadata).toMatchObject(await buildSystemPowerProvenance(root)); + expect(metadata.source).toBe('https://github.com/SemiAnalysisAI/InferenceX-app'); + expect(metadata.status).toBe('DRAFT / pending human verification'); + expect(SYSTEM_POWER_MODEL_REVISION).toMatch(/^app-sha256:[0-9a-f]{64}$/u); + expect(Object.keys(metadata.sourceSha256)).toEqual([ + 'packages/app/src/lib/modeled-system-power.ts', + 'packages/app/src/lib/system-power-model.profiles.json', + 'packages/app/src/lib/system-power-model.ts', + ]); + for (const [path, expectedHash] of Object.entries(metadata.sourceSha256)) { + const bytes = await readFile(resolve(root, path)); + expect(expectedHash, path).toBe(createHash('sha256').update(bytes).digest('hex')); + } + for (const profile of Object.values({ ...metadata.profiles, ...metadata.rackProfiles })) { + expect(profile.modelPath).toBe('packages/app/src/lib/system-power-model.ts'); + expect(profile).not.toHaveProperty('defaultConfig'); + expect(profile).not.toHaveProperty('functionName'); + expect(profile).not.toHaveProperty('configFactory'); + expect(systemPowerSourceSha256(profile.modelPath)).toBe( + metadata.sourceSha256['packages/app/src/lib/system-power-model.ts'], + ); + } + expect(systemPowerSourceSha256('toString')).toBeNull(); + expect(systemPowerSourceSha256(reference.modelPaths.gb200)).toBeNull(); + }); + + it('changes the version when profile parameters or admission/PUE policy change', async () => { + const root = resolve(import.meta.dirname, '../../../..'); + const copy = await mkdtemp(resolve(tmpdir(), 'system-power-provenance-')); + try { + for (const path of Object.keys(SYSTEM_POWER_MODEL_METADATA.sourceSha256)) { + await mkdir(dirname(resolve(copy, path)), { recursive: true }); + await writeFile(resolve(copy, path), await readFile(resolve(root, path))); + } + const baseline = await buildSystemPowerProvenance(copy); + expect(baseline.modelRevision).toBe(SYSTEM_POWER_MODEL_REVISION); + const profilePath = resolve(copy, 'packages/app/src/lib/system-power-model.profiles.json'); + const profiles = JSON.parse(await readFile(profilePath, 'utf8')); + profiles.rackProfiles.gb200.computeTrayStaticDcWatts.compute_tray_fans_unverified += 1; + await writeFile(profilePath, JSON.stringify(profiles)); + const parametersChanged = await buildSystemPowerProvenance(copy); + expect(parametersChanged.modelRevision).not.toBe(baseline.modelRevision); + const adapterPath = resolve(copy, 'packages/app/src/lib/modeled-system-power.ts'); + const adapter = await readFile(adapterPath, 'utf8'); + await writeFile(adapterPath, adapter.replace('DLC_SYSTEM_PUE = 1.1', 'DLC_SYSTEM_PUE = 1.2')); + const policyChanged = await buildSystemPowerProvenance(copy); + expect(policyChanged.modelRevision).not.toBe(parametersChanged.modelRevision); + } finally { + await rm(copy, { recursive: true, force: true }); + } + }); + + it('keeps all independent historical expected values and their original lineage', () => { + expect(reference.modelRevision).toBe('6fcc086b77576d4cecb9d0c79637d6daf980308c'); + expect(reference.source).toBe('https://github.com/SemiAnalysisAI/inferencex_power_model'); + expect(reference.fixtureRole).toContain('Frozen historical Python migration baseline'); + expect(reference.pythonConfigurations.gb200.defaultConfig.pue).toBe(1.2); + expect(reference.cases.length + reference.rackCases.length).toBe(496); + expect( + createHash('sha256') + .update(JSON.stringify([reference.cases, reference.rackCases])) + .digest('hex'), + ).toBe('98cb77c182c510b0e73fefaea1242f20b732e8623cc2cfcc8b6f367d9412d268'); }); }); diff --git a/packages/app/src/lib/system-power-model.ts b/packages/app/src/lib/system-power-model.ts index 68592f4cd..6f1accfe4 100644 --- a/packages/app/src/lib/system-power-model.ts +++ b/packages/app/src/lib/system-power-model.ts @@ -1,12 +1,66 @@ import profileData from './system-power-model.profiles.json'; +import provenance from './system-power-model.provenance.json'; -export const SYSTEM_POWER_MODEL_REVISION = profileData.modelRevision; +export const SYSTEM_POWER_MODEL_REVISION = provenance.modelRevision; +export const SYSTEM_POWER_MODEL_SOURCE = provenance.source; +export const SYSTEM_POWER_MODEL_METADATA = { ...provenance, ...profileData }; export const SYSTEM_POWER_ASSUMPTIONS = profileData.assumptions; export const SYSTEM_POWER_PROFILES = profileData.profiles; export type SystemPowerHardware = keyof typeof SYSTEM_POWER_PROFILES; export const SUPPORTED_SYSTEM_POWER_HARDWARE = Object.keys( SYSTEM_POWER_PROFILES, ) as SystemPowerHardware[]; +export const SYSTEM_POWER_RACK_ASSUMPTIONS = profileData.rackAssumptions; +export const SYSTEM_POWER_RACK_PROFILES = profileData.rackProfiles; +export type SystemPowerRackHardware = keyof typeof SYSTEM_POWER_RACK_PROFILES; +export const SUPPORTED_SYSTEM_POWER_RACK_HARDWARE = Object.keys( + SYSTEM_POWER_RACK_PROFILES, +) as SystemPowerRackHardware[]; + +export type RackMeasuredBasis = 'module' | 'gpu-plus-grace'; + +/** The full revision also covers profile parameters and the admission/PUE adapter. */ +export function systemPowerSourceSha256(modelPath: string): string | null { + const hashes: Readonly> = provenance.sourceSha256; + return Object.hasOwn(hashes, modelPath) ? hashes[modelPath] : null; +} + +/** + * Measured compute-module input for every tray of one NVL72 rack. `module` is the + * sum of the two Module Power sensors per tray (Grace + 2 Blackwell + HBM + LPDDR5X + + * regulator loss). `gpu-plus-grace` is four GPU-board readings plus two Grace socket + * readings; the model then adds the sourced regulator-loss allowance on the GPU share. + */ +export type RackMeasuredInput = + | { basis: 'module'; moduleWattsPerTray: number } + | { basis: 'gpu-plus-grace'; gpuBoardWattsPerTray: number; graceSocketWattsPerTray: number }; + +export interface RackPowerEstimate { + hardware: SystemPowerRackHardware; + basis: RackMeasuredBasis; + /** Measured watts handed to the model for every compute tray, before any allowance. */ + measuredWattsPerTray: number; + /** Sourced regulator-loss allowance on the GPU-board share; zero on the module basis. */ + regulatorAllowanceWattsPerTray: number; + computeModulesDcWatts: number; + regulatorAllowanceWatts: number; + trayStaticDcWatts: number; + nvswitchTraysDcWatts: number; + trayConversionLossWatts: number; + rackDcWatts: number; + powerShelfEfficiency: number; + powerShelfLossWatts: number; + rackAcWatts: number; + facilityWatts: number; + perGpuAcWatts: number; + perGpuFacilityWatts: number; + pue: number; + computeTrayCount: number; + /** GPUs in the modeled rack (72); the rack figures are amortised over all of them. */ + gpuCount: number; + modelRevision: string; + modelPath: string; +} export interface ChassisPowerEstimate { hardware: SystemPowerHardware; @@ -36,11 +90,24 @@ function pythonRound(value: number, digits = 1): number { return Number(value.toFixed(digits)); } +/** Clamping avoids extrapolating beyond the profile's efficiency knots. */ +function interpolateEfficiency(loadFraction: number, curve: number[][]): number { + const [first, last] = [curve[0], curve.at(-1)!]; + if (loadFraction <= first[0]) return first[1]; + if (loadFraction >= last[0]) return last[1]; + for (let i = 1; i < curve.length; i++) { + const [x0, y0] = curve[i - 1]; + const [x1, y1] = curve[i]; + if (x0 <= loadFraction && loadFraction <= x1) { + return y0 + ((y1 - y0) * (loadFraction - x0)) / (x1 - x0); + } + } + return last[1]; +} + /** - * Fixed README inference sweep at the pinned revision, for one complete 8-GPU chassis. - * Profiles are generated by executing the original component models. Only their - * load-dependent fan curve and PSU interpolation execute here; other components - * stay fixed at the recorded assumptions. This function does no allocation scaling. + * One complete 8-GPU chassis. Only fan and PSU loads vary at runtime; the remaining + * component parameters stay fixed at the profile's recorded inference assumptions. */ export function estimateChassisPower( hardware: string, @@ -77,16 +144,7 @@ export function estimateChassisPower( const psu = profile.psu; if (dc > psu.maxDcWatts) return null; - const fraction = dc / psu.loadSharingCapacityWatts; - const curve = psu.efficiencyCurve; - let efficiency = curve[0][1]; - for (let i = 1; i < curve.length; i++) { - const [x0, y0] = curve[i - 1]; - const [x1, y1] = curve[i]; - if (fraction <= x0) break; - efficiency = fraction < x1 ? y0 + ((y1 - y0) * (fraction - x0)) / (x1 - x0) : y1; - if (fraction <= x1) break; - } + const efficiency = interpolateEfficiency(dc / psu.loadSharingCapacityWatts, psu.efficiencyCurve); const ac = pythonRound(dc / efficiency); // The Python chassis wrappers apply PUE to the already-rounded PSU AC output. const facility = pythonRound(ac + ac * (pue - 1)); @@ -106,3 +164,94 @@ export function estimateChassisPower( modelPath: profile.modelPath, }; } + +/** + * One NVL72 rack whose 18 compute trays all carry the given measured compute-module + * input. Only the power-shelf efficiency curve is load dependent; switch trays, tray + * static electronics, management switches, and the tray input-conversion stage stay + * fixed at the recorded assumptions. The Grace CPU and LPDDR5X are never modelled: + * they are inside the measured reading. Rounding follows the source: rack AC is + * rounded before PUE, and every reported total is rounded once at the end. + */ +export function estimateRackPower( + hardware: string, + input: RackMeasuredInput, + pue = SYSTEM_POWER_RACK_ASSUMPTIONS.pue, +): RackPowerEstimate | null { + const key = hardware.toLowerCase(); + const gpuShare = input.basis === 'module' ? 0 : input.gpuBoardWattsPerTray; + const graceShare = input.basis === 'module' ? 0 : input.graceSocketWattsPerTray; + const measured = input.basis === 'module' ? [input.moduleWattsPerTray] : [gpuShare, graceShare]; + if ( + !Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, key) || + measured.some((watts) => !Number.isFinite(watts) || watts <= 0) || + !Number.isFinite(pue) || + pue < 1 + ) { + return null; + } + const canonicalHardware = key as SystemPowerRackHardware; + const profile = SYSTEM_POWER_RACK_PROFILES[canonicalHardware]; + const trays = profile.computeTrayCount; + const measuredWattsPerTray = + input.basis === 'module' + ? input.moduleWattsPerTray + : input.gpuBoardWattsPerTray + input.graceSocketWattsPerTray; + + // Grace tuning guide: regulator loss is 15% of the TDP limit, so loss / delivered = + // f / (1 - f) on the GPU-board share. The Grace socket reading already includes its own. + const frac = profile.regulatorLossFracOfTdp; + const allowanceBase = gpuShare + (profile.regulatorAllowanceIncludesGrace ? graceShare : 0); + const allowancePerTray = + input.basis === 'gpu-plus-grace' ? (allowanceBase * frac) / (1 - frac) : 0; + const computeModulesDc = trays * measuredWattsPerTray + trays * allowancePerTray; + + // Per-tray static blocks keep the source's summation order. + const trayStaticPerTray = Object.values(profile.computeTrayStaticDcWatts).reduce( + (sum, watts) => sum + watts, + 0, + ); + const trayStaticDc = trays * trayStaticPerTray; + const nvswitchTraysDc = + profile.nvswitchTrayCount * + (profile.nvswitchTraySiliconWatts + profile.nvswitchTrayResidualWatts); + const trayLoads = computeModulesDc + trayStaticDc + nvswitchTraysDc; + const trayConversionLoss = trayLoads * (1 / profile.trayInputConversionEfficiency - 1); + const managementDc = profile.managementSwitchCount * profile.managementSwitchWatts; + const rackDc = trayLoads + trayConversionLoss + managementDc; + + const shelf = profile.powerShelf; + if (rackDc > shelf.installedCapacityWatts) return null; + const efficiency = interpolateEfficiency( + rackDc / shelf.installedCapacityWatts, + shelf.efficiencyCurve, + ); + const ac = rackDc / efficiency; + const rackAc = pythonRound(ac); + // PUE applies once, to the already-rounded shelf AC output. + const facility = pythonRound(rackAc + rackAc * (pue - 1)); + if (!Number.isFinite(facility)) return null; + return { + hardware: canonicalHardware, + basis: input.basis, + measuredWattsPerTray, + regulatorAllowanceWattsPerTray: allowancePerTray, + computeModulesDcWatts: pythonRound(computeModulesDc), + regulatorAllowanceWatts: pythonRound(trays * allowancePerTray), + trayStaticDcWatts: pythonRound(trayStaticDc), + nvswitchTraysDcWatts: pythonRound(nvswitchTraysDc), + trayConversionLossWatts: pythonRound(trayConversionLoss), + rackDcWatts: pythonRound(rackDc), + powerShelfEfficiency: pythonRound(efficiency, 4), + powerShelfLossWatts: pythonRound(ac - rackDc), + rackAcWatts: rackAc, + facilityWatts: facility, + perGpuAcWatts: pythonRound(rackAc / profile.gpuCount), + perGpuFacilityWatts: pythonRound(facility / profile.gpuCount), + pue, + computeTrayCount: trays, + gpuCount: profile.gpuCount, + modelRevision: SYSTEM_POWER_MODEL_REVISION, + modelPath: profile.modelPath, + }; +} diff --git a/packages/app/src/lib/views-api/calculator-extensions.ts b/packages/app/src/lib/views-api/calculator-extensions.ts index ba33414f7..1373907e9 100644 --- a/packages/app/src/lib/views-api/calculator-extensions.ts +++ b/packages/app/src/lib/views-api/calculator-extensions.ts @@ -4,13 +4,15 @@ import { parseFirstTokenCaps, selectFirstTokenWinners, } from '@/components/calculator/first-token-limits'; -import { interpolateForGPU } from '@/components/calculator/interpolation'; import { DEFAULT_UTILIZATION_PCT, listPricingToTokenRevenuePricing, profitModelDefaults, } from '@/components/calculator/profit-estimator'; -import { estimateProfitByPower } from '@/components/calculator/profit-power'; +import { + estimateProfitByPower, + interpolateProfitForGPU, +} from '@/components/calculator/profit-power'; import type { TokenRevenuePricing } from '@/components/inference/types'; import { fetchOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing'; import { cachedJson } from '@/lib/api-cache'; @@ -183,11 +185,12 @@ export function calculatorExtension(view: CalculatorExtension, request: NextRequ } as const; function estimate(group: typeof groups.official) { const results = Object.entries(group.grouped).flatMap(([key, points]) => { - const result = interpolateForGPU( + const result = interpolateProfitForGPU( points, target, 'interactivity_to_throughput', costProvider === 'custom' ? 'costh' : costProvider, + basis === 'gw-year' ? powerBasis : 'provisioned', ); return result && result.value > 0 ? [{ ...result, ...group.groupMeta[key], resultKey: key }] diff --git a/packages/app/src/lib/views-api/docs/extensions.ts b/packages/app/src/lib/views-api/docs/extensions.ts index b4ef43114..aa0bd9d98 100644 --- a/packages/app/src/lib/views-api/docs/extensions.ts +++ b/packages/app/src/lib/views-api/docs/extensions.ts @@ -122,8 +122,8 @@ const PARAMETER_NOTES: Record = { '模型许可或收入分成百分比,范围 0 至 100,默认值随模型变化。', ], powerBasis: [ - 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. All in Measured uses measured GPU power plus modeled unmeasured components and PUE; it is not measured wall power. Eligible measured source rows are required; powerLabel identifies paired estimates and full-chassis extrapolation. Missing coverage is not zero.', - 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。整体实测功耗采用 GPU 实测值,加上未实测组件的功耗估算和 PUE,并非墙上电表读数。该估算需要符合条件的实测数据行;powerLabel 标明对比方式和整机外推。缺失数据不按零处理。', + 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. For GW-year estimates, modeled filters source points to valid system-power inputs before building the curve at the same target, without extrapolation or historical substitution. compare uses identical valid-curve throughput for paired budgets; when no valid curve reaches the target, it retains the original provisioned estimate. provisioned-only keeps the original performance curve. All in Measured adds modeled components and PUE to measured inputs; it is not measured wall power. NVL72 requires complete validated GPU and Grace/module telemetry; CPU rail alone is insufficient, and sensor bases must be compatible. powerSource records topology, measured basis, sensor, PUE, model revision and source hash. powerLabel marks paired estimates and full-chassis extrapolation; skipped reasons preserve missing coverage.', + 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。按 GW 年估算时,modeled 先筛选满足系统功耗要求的数据点,再在同一目标值上构建曲线,不外推,也不借用历史数据。compare 的两种功耗方案使用同一条有效曲线的吞吐量;没有有效曲线覆盖目标时,保留原曲线的预配估算。provisioned 单独使用时仍沿用原性能曲线。整体实测功耗在实测输入上叠加组件估算和 PUE,并非墙上电表读数。NVL72 需要完整且通过验证的 GPU 与 Grace/模块遥测,仅 CPU rail 读数不足,传感器口径也必须兼容。powerSource 记录拓扑、实测口径、传感器、PUE、模型版本和源码哈希;powerLabel 标明配对估算和整机外推,skipped 保留覆盖缺失原因。', ], power: [ 'Comma-separated certified and/or legacy power tiers. Omit for all tiers.', diff --git a/packages/app/src/lib/views-api/docs/inference.ts b/packages/app/src/lib/views-api/docs/inference.ts index 5530d2b22..bdbed2ac9 100644 --- a/packages/app/src/lib/views-api/docs/inference.ts +++ b/packages/app/src/lib/views-api/docs/inference.ts @@ -293,8 +293,8 @@ const parameters: readonly ApiParameter[] = [ required: false, type: 'enum', description: text( - 'Response encoding. csv returns one flat row per plotted point; serviceCompare, roleShare and powerFit analytical results require JSON and return 400 with CSV.', - '响应编码。csv 为每个图表点返回一行平面数据;serviceCompare、roleShare 和 powerFit 分析结果仅支持 JSON,与 CSV 同用时返回 400。', + 'Response encoding. csv returns plotted points, or tableRows for All in Measured with unavailable y cells blank. serviceCompare, roleShare and powerFit analytical results require JSON and return 400 with CSV.', + '响应编码。csv 返回图表数据点;整体实测指标则导出 tableRows,不可用的 y 留空。serviceCompare、roleShare 和 powerFit 分析结果仅支持 JSON,与 CSV 同用时返回 400。', ), schema: { type: 'string', enum: ['json', 'csv'], default: 'json' }, example: 'csv', @@ -345,6 +345,68 @@ const seriesSchema = objectSchema( ], ); +const tableRowSchema = objectSchema( + { + ...pointSchema.properties, + id: { type: ['integer', 'null'] }, + hwKey: stringSchema, + gpu: stringSchema, + framework: stringSchema, + specMethod: stringSchema, + label: stringSchema, + vendor: stringSchema, + deployment: stringSchema, + kvOffload: booleanSchema, + x: { type: ['number', 'null'], description: 'Selected service axis; null when unavailable.' }, + y: { type: ['number', 'null'], description: 'Selected all-in metric; null when unavailable.' }, + measuredGpuWatts: { ...numberSchema, description: 'Validated measured mean W/GPU.' }, + status: { type: 'string', enum: ['available', 'unavailable'] }, + unavailableReason: { + oneOf: [ + { + type: 'string', + enum: [ + 'workload', + 'hardware', + 'telemetry', + 'cpu-telemetry', + 'gpu-count', + 'topology', + 'role-power', + 'model-domain', + 'energy', + 'model-unavailable', + ], + }, + { type: 'null' }, + ], + }, + }, + [ + 'id', + 'precision', + 'hwKey', + 'gpu', + 'framework', + 'specMethod', + 'label', + 'deployment', + 'kvOffload', + 'x', + 'y', + 'concurrency', + 'topologyKey', + 'tp', + 'date', + 'frontier', + 'bestPerSku', + 'metrics', + 'measuredGpuWatts', + 'status', + 'unavailableReason', + ], +); + const sourceSchema = objectSchema({ key: stringSchema, label: stringSchema }, ['key', 'label']); const identitySchema = objectSchema({ id: { type: ['integer', 'null'] }, @@ -459,6 +521,11 @@ const responseSchema = objectSchema( ]), ), series: arraySchema(seriesSchema), + tableRows: { + ...arraySchema(tableRowSchema), + description: + 'All in Measured only: GPU-valid observations in the selected scope, including unavailable estimates. Honors hardware/best selection, but not optimal or chart axis clipping. Also returned per comparison and overlay. count remains the numeric series point count.', + }, count: integerSchema, serviceSources: arraySchema(sourceSchema), equalServiceComparison: { ...comparisonSchema, type: ['object', 'null'] }, @@ -546,8 +613,16 @@ const responseSchema = objectSchema( }), ), pricing: { type: ['object', 'null'], additionalProperties: true }, - comparisons: arraySchema({ type: 'object', additionalProperties: true }), - overlays: arraySchema({ type: 'object', additionalProperties: true }), + comparisons: arraySchema({ + type: 'object', + properties: { tableRows: arraySchema(tableRowSchema) }, + additionalProperties: true, + }), + overlays: arraySchema({ + type: 'object', + properties: { tableRows: arraySchema(tableRowSchema) }, + additionalProperties: true, + }), }, ['view', 'apiVersion', 'params', 'metric', 'xAxis', 'frontier', 'series', 'count'], ); @@ -685,8 +760,8 @@ export const operations: ApiOperation[] = [ path: '/api/v1/views/inference', summary: text('Get the main inference chart view', '获取主推理图表视图'), description: text( - 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection.', - '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。', + 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection. All in Measured watts and energy support 8K/1K and AgentX with validated telemetry and supported topology; AgentX reuses the model without independent workload calibration. The standalone Modeled Chassis AC metric remains limited to 8K/1K. All in Measured adds tableRows with every GPU-valid observation in the selected scope and best-series selection, before optimal or axis clipping; unavailable estimates have y=null and a reason. series and count remain numeric chart points.', + '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。整体实测功耗及能耗支持遥测已验证、拓扑受支持的 8K/1K 和 AgentX 数据;AgentX 复用同一模型,尚未针对该工作负载单独校准。单独列出的每 GPU 分摊的机箱交流功耗估算指标仍仅支持 8K/1K。整体实测指标另返回 tableRows,保留当前筛选范围和最优曲线选择内所有 GPU 遥测有效的观测点,不按 optimal 或坐标轴显示范围裁剪;估算不可用时 y 为 null,并给出原因。series 和 count 仍仅包含可绘制的数值点。', ), audience: 'public', stability: 'beta', diff --git a/packages/app/src/lib/views-api/series.test.ts b/packages/app/src/lib/views-api/series.test.ts index 78a4c5c7a..e25c68d6b 100644 --- a/packages/app/src/lib/views-api/series.test.ts +++ b/packages/app/src/lib/views-api/series.test.ts @@ -116,6 +116,127 @@ function powerSweepRows(): BenchmarkRow[] { } describe('buildInferenceSeries', () => { + it('retains GPU-valid all-in table rows without adding unavailable estimates to chart series', () => { + const measured = metrics({ + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + pp: 1, + pcp_size: 1, + power_valid: 1, + power_metric_schema_version: 2, + joules_per_output_token: 2, + }); + const supported = makeRow({ metrics: measured }); + const noCpu = makeRow({ hardware: 'gb200', metrics: measured }); + const invalid = makeRow({ hardware: 'b200', metrics: { ...measured, power_valid: 0 } }); + const otherPrecision = makeRow({ hardware: 'gb300', precision: 'fp4', metrics: measured }); + const result = buildInferenceSeries([supported, noCpu, invalid, otherPrecision], { + ...BASE_OPTIONS, + metricConfigKey: 'y_utilityModeledWatts', + optimal: true, + }); + + expect(result.count).toBe(1); + expect(result.series.flatMap((series) => series.points.map((point) => point.id))).toEqual([ + supported.id, + ]); + expect(result.tableRows).toHaveLength(2); + expect(result.tableRows?.find((point) => point.id === supported.id)).toMatchObject({ + y: result.series[0].points[0].y, + measuredGpuWatts: 500, + status: 'available', + unavailableReason: null, + }); + expect(result.tableRows?.find((point) => point.id === noCpu.id)).toMatchObject({ + hwKey: 'gb200_trt', + y: null, + measuredGpuWatts: 500, + status: 'unavailable', + unavailableReason: 'cpu-telemetry', + frontier: false, + }); + }); + + it('keeps all-in energy missing, honors selected hardware, and leaves ordinary views unchanged', () => { + const row = makeRow({ + metrics: metrics({ + avg_power_w: 500, + avg_total_gpu_power_w: 4000, + pp: 1, + pcp_size: 1, + power_valid: 1, + power_metric_schema_version: 2, + }), + }); + const result = buildInferenceSeries([row], { + ...BASE_OPTIONS, + metricConfigKey: 'y_utilityModeledJPerOutputToken', + }); + expect(result.series).toEqual([]); + expect(result.tableRows).toMatchObject([ + { id: row.id, y: null, status: 'unavailable', unavailableReason: 'energy' }, + ]); + expect( + buildInferenceSeries([row], { + ...BASE_OPTIONS, + metricConfigKey: 'y_utilityModeledWatts', + gpus: ['gb200'], + }).tableRows, + ).toEqual([]); + expect(buildInferenceSeries([row], BASE_OPTIONS)).not.toHaveProperty('tableRows'); + }); + + it('keeps non-frontier transient rows in the table without assigning them frontier flags', () => { + const rows = [ + [1, 10, 400], + [2, 20, 500], + ].map(([conc, x, watts]) => { + const row = makeRow({ + conc, + metrics: metrics({ + median_intvty: x, + avg_power_w: watts, + avg_total_gpu_power_w: watts * 8, + power_valid: 1, + power_metric_schema_version: 2, + pp: 1, + pcp_size: 1, + }), + }); + Reflect.deleteProperty(row, 'id'); + return row; + }); + const result = buildInferenceSeries(rows, { + ...BASE_OPTIONS, + metricConfigKey: 'y_utilityModeledWatts', + optimal: true, + }); + expect(result.count).toBe(1); + expect(result.tableRows).toHaveLength(2); + expect(result.tableRows?.map((point) => [point.concurrency, point.frontier])).toEqual([ + [1, false], + [2, true], + ]); + }); + + it('returns null for missing or invalid normalized axes without dropping table rows', () => { + const row = makeRow({ + benchmark_type: 'agentic_traces', + metrics: metrics({ avg_power_w: 500, power_valid: 1, power_metric_schema_version: 2 }), + }); + const result = buildInferenceSeries([row], { + ...BASE_OPTIONS, + sequence: Sequence.AgenticTraces, + metricConfigKey: 'y_utilityModeledWatts', + xmode: 'e2e-normalized-interactivity', + derivedMetrics: { + [row.id]: { id: row.id, p75_e2e_norm_intvty: null, p90_e2e_norm_intvty: 0 }, + }, + }); + expect(result.series).toEqual([]); + expect(result.tableRows).toMatchObject([{ id: row.id, x: null, y: null }]); + }); + it('assembles one series per hardware config with x-sorted points', () => { const result = buildInferenceSeries(fixtureRows(), BASE_OPTIONS); diff --git a/packages/app/src/lib/views-api/series.ts b/packages/app/src/lib/views-api/series.ts index 54cc033e4..40fb3fa18 100644 --- a/packages/app/src/lib/views-api/series.ts +++ b/packages/app/src/lib/views-api/series.ts @@ -9,6 +9,7 @@ import { type RooflineDirection, } from '@/components/inference/hooks/chart-data-core'; import chartDefinitions, { + isAllInMeasuredConfigKey, METRIC_REGISTRY, tokenMetricTypeForConfigKey, type MetricConfigKey, @@ -25,6 +26,10 @@ import type { } from '@/components/inference/types'; import { partitionChartDataByLimits } from '@/components/inference/utils'; import { bestSeriesPerSku } from '@/components/inference/utils/best-series-per-sku'; +import { + allInMeasuredTableData, + allInMeasuredUnavailableReason, +} from '@/components/inference/utils/inference-table-data'; import { isMeasuredPowerCurveMetric, upperPowerEnvelope, @@ -127,6 +132,22 @@ export interface InferenceSeriesEntry { readonly points: readonly InferenceSeriesPoint[]; } +export interface InferenceTableRow extends Omit { + readonly hwKey: string; + readonly gpu: string; + readonly framework: string; + readonly specMethod: string; + readonly label: string; + readonly vendor?: string; + readonly deployment: string; + readonly kvOffload: boolean; + readonly x: number | null; + readonly y: number | null; + readonly measuredGpuWatts: number; + readonly status: 'available' | 'unavailable'; + readonly unavailableReason: ReturnType; +} + export interface InferenceSeriesMetricMeta { readonly key: MetricKey; readonly configKey: MetricConfigKey; @@ -139,6 +160,8 @@ export interface InferenceSeriesMetricMeta { export interface InferenceSeriesResult { readonly series: readonly InferenceSeriesEntry[]; + /** All in Measured table population, including unavailable system estimates. */ + readonly tableRows?: readonly InferenceTableRow[]; readonly hardware: readonly { key: string; label: string; vendor?: string }[]; readonly frontier: { direction: ParetoDirection | null; points: number }; readonly metric: InferenceSeriesMetricMeta; @@ -408,9 +431,58 @@ export function buildInferenceSeries( } const count = series.reduce((total, entry) => total + entry.points.length, 0); + // Remapping retains nested metric objects, including for transient rows without IDs. + const frontierMetrics = new Set([...frontierPoints].map((point) => point[metricKey])); + const tableRows = isAllInMeasuredConfigKey(metricConfigKey) + ? allInMeasuredTableData(scoped, metricKey, resolved.xAxisField) + .filter((point) => !best || bestHwKeys.size === 0 || bestHwKeys.has(point.hwKey)) + .map((point): InferenceTableRow => { + const derived = point.id === undefined ? undefined : options.derivedMetrics?.[point.id]; + const x = + xmode === 'e2e-normalized-interactivity' + ? percentile === 'p75' + ? derived?.p75_e2e_norm_intvty + : derived?.p90_e2e_norm_intvty + : point.x; + const unavailableReason = allInMeasuredUnavailableReason(point, metricKey); + return { + id: point.id ?? null, + precision: point.precision, + hwKey: point.hwKey, + gpu: point.hwKey.split('_')[0], + framework: point.framework ?? '', + specMethod: point.spec_decoding ?? 'none', + label: hardwareLegendLabel(point.hwKey, point.model), + vendor: GPU_VENDORS[point.hwKey.split('_')[0]], + deployment: pointDeploymentMode(point), + kvOffload: isKvOffloadEnabled(point), + x: + typeof x === 'number' && + Number.isFinite(x) && + (xmode !== 'e2e-normalized-interactivity' || x > 0) + ? x + : null, + y: Number.isFinite(point.y) ? point.y : null, + concurrency: point.conc ?? 0, + topologyKey: pointTopologyKey(point), + tp: point.tp ?? 0, + date: point.date ?? '', + ...(runIdFromUrl(point.run_url) === undefined + ? {} + : { runId: runIdFromUrl(point.run_url) }), + frontier: unavailableReason === null && frontierMetrics.has(point[metricKey]), + bestPerSku: bestHwKeys.has(point.hwKey), + metrics: pointMetrics(point, metricKey), + measuredGpuWatts: point.measuredAvgPower!.y, + status: unavailableReason === null ? 'available' : 'unavailable', + unavailableReason, + }; + }) + : undefined; return { series, + ...(tableRows === undefined ? {} : { tableRows }), hardware: series.map((entry) => ({ key: entry.hwKey, label: entry.label, diff --git a/packages/app/timings.json b/packages/app/timings.json index df735201a..976d0b235 100644 --- a/packages/app/timings.json +++ b/packages/app/timings.json @@ -248,6 +248,10 @@ "spec": "cypress/e2e/profit-estimator.cy.ts", "duration": 39450 }, + { + "spec": "cypress/e2e/profit-power-curves.cy.ts", + "duration": 3568 + }, { "spec": "cypress/e2e/reliability-chart.cy.ts", "duration": 5511 diff --git a/packages/constants/src/metric-keys.test.ts b/packages/constants/src/metric-keys.test.ts index 570da7c5e..923504ff0 100644 --- a/packages/constants/src/metric-keys.test.ts +++ b/packages/constants/src/metric-keys.test.ts @@ -1,12 +1,34 @@ import { describe, expect, it } from 'vitest'; import { + CPU_SIDE_POWER_METRIC_KEY_LIST, + CPU_SIDE_POWER_METRIC_KEYS, MEASURED_POWER_METRIC_KEY_LIST, MEASURED_POWER_METRIC_KEYS, METRIC_KEYS, POWER_METRIC_KEYS, } from './metric-keys'; +describe('CPU_SIDE_POWER_METRIC_KEYS', () => { + it('names exactly the NVL72 Grace-side and compute-module keys, all of them measured keys', () => { + expect(new Set(CPU_SIDE_POWER_METRIC_KEY_LIST)).toEqual( + new Set([ + 'avg_cpu_socket_power_w', + 'avg_total_cpu_power_w', + 'total_cpu_energy_j', + 'avg_total_module_power_w', + 'total_module_energy_j', + ]), + ); + expect(CPU_SIDE_POWER_METRIC_KEYS.size).toBe(5); + for (const key of CPU_SIDE_POWER_METRIC_KEYS) { + expect(MEASURED_POWER_METRIC_KEYS.has(key)).toBe(true); + } + // The verdict itself is a discriminator, not a measurement. + expect(CPU_SIDE_POWER_METRIC_KEYS.has('cpu_power_valid')).toBe(false); + }); +}); + describe('MEASURED_POWER_METRIC_KEYS', () => { it('is a subset of METRIC_KEYS', () => { for (const key of MEASURED_POWER_METRIC_KEYS) { @@ -36,6 +58,7 @@ describe('MEASURED_POWER_METRIC_KEYS', () => { 'peak_temp_c', 'avg_util_pct', 'avg_mem_used_mb', + // NVL72 Grace-side and compute-module measurements (same window as GPU energy). 'avg_cpu_socket_power_w', 'avg_total_cpu_power_w', 'total_cpu_energy_j', diff --git a/packages/db/src/etl/benchmark-mapper.test.ts b/packages/db/src/etl/benchmark-mapper.test.ts index 5982348d4..3614806ff 100644 --- a/packages/db/src/etl/benchmark-mapper.test.ts +++ b/packages/db/src/etl/benchmark-mapper.test.ts @@ -1,5 +1,8 @@ import { describe, it, expect, vi } from 'vitest'; -import { MEASURED_POWER_METRIC_KEYS } from '@semianalysisai/inferencex-constants'; +import { + CPU_SIDE_POWER_METRIC_KEYS, + MEASURED_POWER_METRIC_KEYS, +} from '@semianalysisai/inferencex-constants'; import { extractPowerAudit, extractPowerInvalidReasons, @@ -88,6 +91,13 @@ function dirtyPowerPayload(): Record { peak_temp_c: 79.2, avg_util_pct: 88.5, avg_mem_used_mb: 71234.5, + // NVL72 CPU-side measurements share the GPU window but follow their own + // cpu_power_valid verdict; tests that want them kept must supply it. + avg_cpu_socket_power_w: 250.5, + avg_total_cpu_power_w: 1002, + total_cpu_energy_j: 601200, + avg_total_module_power_w: 17203, + total_module_energy_j: 10321800, workers: [ { role: 'prefill', worker_idx: 0, hosts: ['pn0'], num_gpus: 4, avg_power_w: 612.3 }, { role: 'decode', worker_idx: 0, hosts: ['dn0'], num_gpus: 8, avg_power_w: 701.5 }, @@ -95,6 +105,11 @@ function dirtyPowerPayload(): Record { }; } +const CPU_SIDE_KEYS = [...CPU_SIDE_POWER_METRIC_KEYS]; +const GPU_SIDE_KEYS = [...MEASURED_POWER_METRIC_KEYS].filter( + (key) => !CPU_SIDE_POWER_METRIC_KEYS.has(key), +); + describe('mapBenchmarkRow', () => { describe('v1 schema', () => { it('maps a valid v1 row to BenchmarkParams', () => { @@ -339,13 +354,71 @@ describe('mapBenchmarkRow', () => { expect(result!.metrics.median_ttft).toBe(50.2); }); - it('keeps every measured key and the workers payload on a valid verdict', () => { + // Producer contract: cpu_power_valid is independent of power_valid. Each leg + // withholds only its own keys; worker telemetry belongs to the GPU leg. + it.each([ + { power_valid: 1, cpu_power_valid: 1, gpuKept: true, cpuKept: true }, + { power_valid: 1, cpu_power_valid: 0, gpuKept: true, cpuKept: false }, + { power_valid: 0, cpu_power_valid: 1, gpuKept: false, cpuKept: true }, + { power_valid: 0, cpu_power_valid: 0, gpuKept: false, cpuKept: false }, + ])( + 'withholds each leg on its own verdict: power_valid=$power_valid cpu_power_valid=$cpu_power_valid', + ({ power_valid, cpu_power_valid, gpuKept, cpuKept }) => { + const tracker = createSkipTracker(); + const dirty = dirtyPowerPayload(); + const result = mapBenchmarkRow( + makeV2Row({ power_valid, power_metric_schema_version: 2, cpu_power_valid, ...dirty }), + tracker, + ); + + expect(result!.metrics.power_valid).toBe(power_valid); + expect(result!.metrics.cpu_power_valid).toBe(cpu_power_valid); + for (const key of GPU_SIDE_KEYS) { + if (gpuKept) expect(result!.metrics[key]).toBe(dirty[key]); + else expect(result!.metrics).not.toHaveProperty(key); + } + for (const key of CPU_SIDE_KEYS) { + if (cpuKept) expect(result!.metrics[key]).toBe(dirty[key]); + else expect(result!.metrics).not.toHaveProperty(key); + } + if (gpuKept) expect(result!.workers).toHaveLength(2); + else expect(result!.workers).toBeUndefined(); + }, + ); + + it('withholds CPU-side keys that arrive without a cpu_power_valid verdict or with a malformed one', () => { + // No legacy rows predate cpu_power_valid, so absence is out of contract and fails closed. const tracker = createSkipTracker(); const dirty = dirtyPowerPayload(); - const result = mapBenchmarkRow( + const absent = mapBenchmarkRow( makeV2Row({ power_valid: 1, power_metric_schema_version: 2, ...dirty }), tracker, ); + expect(absent!.metrics).not.toHaveProperty('cpu_power_valid'); + for (const key of CPU_SIDE_KEYS) expect(absent!.metrics).not.toHaveProperty(key); + for (const key of GPU_SIDE_KEYS) expect(absent!.metrics[key]).toBe(dirty[key]); + + const malformed = mapBenchmarkRow( + makeV2Row({ + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 'garbage', + ...dirty, + }), + tracker, + ); + expect(malformed!.metrics.cpu_power_valid).toBe(0); + for (const key of CPU_SIDE_KEYS) expect(malformed!.metrics).not.toHaveProperty(key); + for (const key of GPU_SIDE_KEYS) expect(malformed!.metrics[key]).toBe(dirty[key]); + }); + + it('keeps every measured key and the workers payload on valid verdicts', () => { + const tracker = createSkipTracker(); + const dirty = dirtyPowerPayload(); + const result = mapBenchmarkRow( + makeV2Row({ power_valid: 1, power_metric_schema_version: 2, cpu_power_valid: 1, ...dirty }), + tracker, + ); expect(result!.metrics.power_valid).toBe(1); for (const key of MEASURED_POWER_METRIC_KEYS) { @@ -354,13 +427,13 @@ describe('mapBenchmarkRow', () => { expect(result!.workers).toHaveLength(2); }); - it('leaves legacy rows without a verdict untouched (historical measurements kept)', () => { + it('leaves legacy rows without a GPU verdict untouched (historical measurements kept)', () => { const tracker = createSkipTracker(); const dirty = dirtyPowerPayload(); const result = mapBenchmarkRow(makeV2Row(dirty), tracker); expect(result!.metrics).not.toHaveProperty('power_valid'); - for (const key of MEASURED_POWER_METRIC_KEYS) { + for (const key of GPU_SIDE_KEYS) { expect(result!.metrics[key]).toBe(dirty[key]); } expect(result!.workers).toHaveLength(2); @@ -828,6 +901,41 @@ describe('mapBenchmarkRow', () => { expect(result!.metrics).not.toHaveProperty('workers'); }); + it('captures the NVL72 CPU-side keys without an unknown-key warning', async () => { + // The warning fires once per process per key, so a fresh module instance is + // the only way to observe whether these keys are known. + vi.resetModules(); + const fresh = await import('./benchmark-mapper'); + const warn = vi.spyOn(console, 'warn').mockImplementation(() => {}); + try { + const tracker = createSkipTracker(); + const result = fresh.mapBenchmarkRow( + makeV2Row({ + power_valid: 1, + power_metric_schema_version: 2, + cpu_power_valid: 1, + avg_cpu_socket_power_w: 250.5, + avg_total_cpu_power_w: 1002, + total_cpu_energy_j: 601200, + avg_total_module_power_w: 17203, + total_module_energy_j: 10321800, + }), + tracker, + ); + expect(result!.metrics).toMatchObject({ + cpu_power_valid: 1, + avg_cpu_socket_power_w: 250.5, + avg_total_cpu_power_w: 1002, + total_cpu_energy_j: 601200, + avg_total_module_power_w: 17203, + total_module_energy_j: 10321800, + }); + expect(warn).not.toHaveBeenCalled(); + } finally { + warn.mockRestore(); + } + }); + it('captures new cluster-wide temp / util / mem scalars into metrics', () => { // These are flat scalars on the agg row (sibling of avg_power_w), so // the auto-capture path must store them under their raw keys without @@ -883,14 +991,57 @@ describe('scrubWithheldPowerMetrics (direct — supplemental ingest path)', () = expect(metrics.tput_per_gpu).toBe(567.8); }); - it('leaves power_valid=1 and legacy no-verdict records untouched', () => { - for (const metrics of [supplementalMetrics({ power_valid: 1 }), supplementalMetrics()]) { + it('leaves power_valid=1 and legacy no-GPU-verdict records untouched', () => { + for (const metrics of [ + supplementalMetrics({ power_valid: 1, cpu_power_valid: 1 }), + supplementalMetrics({ cpu_power_valid: 1 }), + ]) { const before = { ...metrics }; expect(scrubWithheldPowerMetrics(metrics)).toBe(false); expect(metrics).toEqual(before); } }); + it('withholds only the CPU-side keys on cpu_power_valid=0 and leaves GPU power published', () => { + const metrics = supplementalMetrics({ power_valid: 1, cpu_power_valid: 0 }); + const before = { ...metrics }; + expect(scrubWithheldPowerMetrics(metrics)).toBe(false); + expect(metrics.cpu_power_valid).toBe(0); + for (const key of CPU_SIDE_KEYS) expect(metrics).not.toHaveProperty(key); + for (const key of GPU_SIDE_KEYS) expect(metrics[key]).toBe(before[key]); + expect(metrics.avg_power_w).toBe(685.5); + }); + + it('keeps the CPU-side keys on power_valid=0 when cpu_power_valid=1', () => { + const metrics = supplementalMetrics({ power_valid: 0, cpu_power_valid: 1 }); + expect(scrubWithheldPowerMetrics(metrics)).toBe(true); + for (const key of GPU_SIDE_KEYS) expect(metrics).not.toHaveProperty(key); + for (const key of CPU_SIDE_KEYS) expect(metrics).toHaveProperty(key); + expect(metrics.avg_total_module_power_w).toBe(17203); + }); + + it('normalizes cpu_power_valid as a verdict, independent of power_valid', () => { + const metrics = supplementalMetrics({ power_valid: 1, cpu_power_valid: '1' }); + normalizePowerContractMetrics(metrics, metrics); + expect(metrics.cpu_power_valid).toBe(1); + expect(scrubWithheldPowerMetrics(metrics)).toBe(false); + expect(metrics.avg_total_cpu_power_w).toBe(1002); + + const malformed = supplementalMetrics({ power_valid: 1, cpu_power_valid: 2 }); + normalizePowerContractMetrics(malformed, malformed); + expect(malformed.cpu_power_valid).toBe(0); + expect(malformed.avg_total_cpu_power_w).toBe(1002); + expect(scrubWithheldPowerMetrics(malformed)).toBe(false); + expect(malformed).not.toHaveProperty('avg_total_cpu_power_w'); + expect(malformed.avg_power_w).toBe(685.5); + + const absent = supplementalMetrics({ power_valid: 1 }); + normalizePowerContractMetrics(absent, absent); + expect(absent).not.toHaveProperty('cpu_power_valid'); + expect(scrubWithheldPowerMetrics(absent)).toBe(false); + for (const key of CPU_SIDE_KEYS) expect(absent).not.toHaveProperty(key); + }); + it('fails closed on a malformed verdict when composed with normalization', () => { const metrics = supplementalMetrics({ power_valid: 2 }); normalizePowerContractMetrics(metrics, metrics); @@ -1158,7 +1309,43 @@ describe('extractPowerAudit', () => { }); }); - it('drops unknown keys (fixed 8-key shape bounds the stored object)', () => { + it('keeps the bounded CPU-side audit block and drops malformed CPU fields', () => { + const cpu = { + sensor_kind: 'module', + source: 'acpi', + expected_sockets: 4, + observed_sockets: 4, + sample_row_count: 2400, + reason_codes: [], + }; + expect(extractPowerAudit({ ...fullAudit, cpu })).toEqual({ ...fullAudit, cpu }); + expect( + extractPowerAudit({ + sample_count: 1, + cpu: { + sensor_kind: 'thermocouple', + source: 'x'.repeat(33), + expected_sockets: -1, + observed_sockets: Number.NaN, + sample_row_count: '12', + reason_codes: ['cpu_socket_count_mismatch', 'cpu_socket_count_mismatch', '', 7], + }, + }), + ).toEqual({ + sample_count: 1, + producer_sha: null, + exporter_image_sha256: null, + cpu: { sample_row_count: 12, reason_codes: ['cpu_socket_count_mismatch'] }, + }); + expect(extractPowerAudit({ sample_count: 1, cpu: {} })).toEqual({ + sample_count: 1, + producer_sha: null, + exporter_image_sha256: null, + }); + expect(extractPowerAudit({ sample_count: 1, cpu: 'acpi' })).not.toHaveProperty('cpu'); + }); + + it('drops unknown keys (fixed 9-key shape bounds the stored object)', () => { expect(extractPowerAudit({ sample_count: 3, integration_method: 'trapezoid' })).toEqual({ sample_count: 3, producer_sha: null, diff --git a/packages/db/src/etl/power-publication.ts b/packages/db/src/etl/power-publication.ts index 16c40a305..9245a4aca 100644 --- a/packages/db/src/etl/power-publication.ts +++ b/packages/db/src/etl/power-publication.ts @@ -33,6 +33,7 @@ const IDENTITY_FIELDS = [ 'image', 'run_url', ] as const; +// The CPU-side keys are withheld on cpu_power_valid, so the manifest carries that verdict too. const POWER_FIELDS = [ ...MEASURED_POWER_METRIC_KEYS, 'power_valid', diff --git a/packages/skills/skills/inferencex-api/integrity.json b/packages/skills/skills/inferencex-api/integrity.json index 576650ce5..01e661cce 100644 --- a/packages/skills/skills/inferencex-api/integrity.json +++ b/packages/skills/skills/inferencex-api/integrity.json @@ -8,7 +8,7 @@ "references/cli-contract.md": "fa44faa38d889b4fbdee5ba42752b47758ee3706fab150513bd4e6db87a86cb0", "references/cli.md": "96b228f34cb3600f4548f4dc84df8506531747286763c56ca834c77bee05e1eb", "references/collectivex.md": "eb794f9c28d4a27bec4db80c42c4685ff3b1d204fd9b12a6788511fb258a789f", - "references/dashboard-views.md": "42906d29813e31fe863dc46ba3fdb0ab1d647f98249d82858d231c8d094c16e6", + "references/dashboard-views.md": "47fb1611effb6d3c12467e693e1659ff1a91891c7979d36898fb7194fa7caf19", "references/offline-exports.md": "95aa565dcd4c9e592159baa560a9a58217bc61e395de1c6daf0576795c7f4ec9", "references/pareto.md": "1b4d2d163f982e3f2d97addce310789eae9c1e5badd82dd7501708c9e0385027", "references/powerx.md": "ae01164107aec35247f729e425379b350d2e7ecdd8d4668fc68dbd879f3d965d", diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md index 3df908d57..7ea5896d6 100644 --- a/packages/skills/skills/inferencex-api/references/dashboard-views.md +++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md @@ -96,6 +96,35 @@ wall power. Metric IDs and API selector values are unchanged. Profit `powerBasis still accepts `provisioned`, `modeled`, or `compare`; `powerLabel` is display text. Expanding assumptions or unavailable-estimate details does not change returned data. +All in Measured watts and energy accept validated 8K/1K and AgentX rows through the +shared chart/API transform, including historical and unofficial rows. AgentX reuses +the chassis or rack model without independent workload calibration. Telemetry and +topology gates still apply; NVL72 needs complete Grace or module power. The standalone +Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope. + +For All in Measured, `tableRows` retains every GPU-valid observation in the selected +scope and best-series selection, including axis-clipped and non-frontier points. +Missing system estimates use `y: null`, `status: "unavailable"`, and +`unavailableReason`; `measuredGpuWatts` remains available. CSV exports these rows +with blank missing values. Numeric `series` and `count` are unchanged. Each date +comparison and unofficial overlay has its own `tableRows`; latest does not pool history. + +NVL72 estimates require valid GPU power plus validated Grace-socket or compute-module +power with complete socket coverage. CPU-rail-only readings do not establish the +Grace/LPDDR boundary. A module reading already includes GPU power; do not add GPU +watts again. Read `powerSource` for topology, measured basis, sensor, PUE, model +content revision, app TypeScript source path and source hash. Equations and parameters +are maintained in InferenceX-app; the revision is a content digest, not a private-repository +Git commit. Model-only updates recalculate retained valid measurements after deployment; +they do not require telemetry backfill. For GW-year estimates, `modeled` first selects +points with valid system-power inputs, then builds the curve at the requested target. +It does not extrapolate or substitute historical snapshots. `compare` uses the same +valid-curve throughput for both budgets; if no valid curve covers the target, it keeps +the original provisioned estimate. `provisioned` alone retains the original performance +curve. Official, comparison and unofficial scopes are evaluated independently. +`skipped.reason` distinguishes missing CPU power and incompatible sensor bases; +modeled estimates never substitute provisioned watts. + Prefer equal-service comparisons for article-facing hardware analysis. Use `xstat=mean` only for fixed-sequence service axes when that statistic is intended: streaming speed then means **1 / mean TPOT**, not arithmetic mean request speed. @@ -151,8 +180,7 @@ n, x-range, registry `tdpWatts` and point identities. Fewer than three distinct rates return `fit: null`, `reason: "too-few-points"`. Call `P₀` an extrapolated intercept, not idle power, and do not read the line outside its x-range. -These analytical results require JSON; enabling any with CSV returns 400. Existing CSV remains -a plotted-point export. +These analytical results require JSON; enabling any with CSV returns 400. CSV exports plotted points except for All in Measured, which exports `tableRows`. 同等服务对比与同并发诊断仅通过 API 提供,查询参数为 `serviceCompare`、`serviceBaseline`、 `serviceComparator` 和 `serviceTarget`;仪表板没有对应的服务对比控件、来源选择、目标值输入、