From 374b4e14382f7b0f5419cc72c3b03ba750dc9363 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 14:27:50 -0700
Subject: [PATCH 01/22] feat: add NVL72 rack profiles and tray-based measured
system power estimate
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Pin inferencex_power_model 963ead8b (feat/gb200-nvl72-rack-model) and emit
gb200/gb300 rack profiles plus 204 Python parity cases; the x86 chassis
profiles and their 292 cases regenerate byte-identically. Port the rack
evaluation to TypeScript with the source rounding order (rack AC rounded
before PUE). modelSystemPower admits NVL72 rows as compute trays (one host,
four GPUs, two Grace sockets) whose module or GPU-board + Grace-socket watts
are measured: cpu_power_valid=1 and the Grace-side keys are required, the
Grace side is never modelled, DLC PUE 1.1 applies once at rack AC, a partial
tray extrapolates only the GPU-board share, and the result carries
topologyBasis 'nvl72-trays', measuredBasis, sensorKind and the new
cpu-telemetry reason. x86 rows are unchanged.
中文:固定 inferencex_power_model 963ead8b(feat/gb200-nvl72-rack-model),
生成 gb200/gb300 机架 profile 和 204 条 Python 对照用例,x86 机箱 profile 及
其 292 条用例逐字节不变。TypeScript 按原模型的舍入顺序移植机架计算(机架交流
功率先舍入再乘 PUE)。modelSystemPower 将 NVL72 行按计算 tray(单主机、4 张
GPU、2 个 Grace socket)接纳,输入为实测模块功耗或 GPU 板卡 + Grace socket
功耗:要求 cpu_power_valid=1 及 Grace 侧指标,Grace 侧从不建模,DLC PUE 1.1
在机架交流侧只应用一次,部分分配的 tray 仅外推 GPU 板卡份额,结果新增
topologyBasis 'nvl72-trays'、measuredBasis、sensorKind 和 cpu-telemetry
原因。x86 行为不变。
---
docs/powerx-system-power.md | 45 +-
.../generate-system-power-reference.py | 135 +-
.../inference/utils/tooltipUtils.ts | 3 +
.../lib/modeled-system-power-export.test.ts | 3 +-
.../app/src/lib/modeled-system-power.test.ts | 244 +-
packages/app/src/lib/modeled-system-power.ts | 275 +-
.../src/lib/system-power-model.profiles.json | 349 +-
.../src/lib/system-power-model.reference.json | 4013 ++++++++++++++++-
.../app/src/lib/system-power-model.test.ts | 119 +
packages/app/src/lib/system-power-model.ts | 162 +-
10 files changed, 5249 insertions(+), 99 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index efee20866..1a3399e89 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -29,8 +29,9 @@ and `u_nvme=0.0`. The pinned Python model defaults to PUE `1.2`; PowerX uses
power × PUE (`1.3` air, `1.1` DLC).
The factor applies after chassis AC; measured GPU power and chassis AC do not change.
Cooling describes the modeled chassis, not verified benchmark-site cooling.
-The current profiles do not model DLC; `--pue` remains an explicit facility-factor
-override and does not convert an air-cooled chassis model into a DLC model.
+The chassis profiles do not model DLC; `--pue` remains an explicit facility-factor
+override and does not convert an air-cooled chassis model into a DLC model. The
+NVL72 rack profiles below are direct-liquid-cooled and default to `1.1`.
Platform-specific network assumptions,
fan control, component counts, and chassis defaults are preserved in the
generated profile; every JSON export includes that profile and every CSV row
@@ -47,9 +48,29 @@ not measured CPU/DRAM utilization.
| `mi325x` | `human_verified/mi325x_chassis/mi325x_chassis_power_model.py` |
| `mi355x` | `human_verified/mi355x_chassis/mi355x_chassis_power_model.py` |
-All listed profiles describe a complete eight-GPU chassis. GB200 and GB300 have
-no matching model and are unsupported. Their rack topology is not substituted
-with B200 or B300.
+All listed profiles describe a complete eight-GPU chassis. Their topology is not
+substituted onto GB200 or GB300, which use the NVL72 rack profiles instead:
+
+| Hardware identity | Source rack implementation |
+| ----------------- | ----------------------------------------------------------------- |
+| `gb200` | `human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py` |
+| `gb300` | same module, `gb300_nvl72_rack_config` |
+
+The rack profiles (`rackProfiles`) take **measured** compute-module watts per tray
+as their input: the module sensor total (`avg_total_module_power_w`) when the
+producer publishes it, otherwise GPU-board watts plus the Grace-socket total
+(`avg_total_cpu_power_w`) with the source's regulator-loss allowance on the GPU
+share. The Grace CPU and LPDDR5X are never modelled; rows without
+`cpu_power_valid=1` and the Grace-side keys stay unavailable (`cpu-telemetry`).
+Each measured worker host is one compute tray (four GPUs, two Grace sockets), and
+a tray's estimate is its 1/18 share of a rack of identical trays, so NVSwitch
+trays, power shelves, and management switches are amortised over 72 GPUs. The
+result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`.
+A partially allocated tray extrapolates only the GPU-board share (a module reading
+already covers the whole tray) and is labeled `extrapolated`. The pinned revision
+`963ead8b` is the power-model repo's `feat/gb200-nvl72-rack-model` branch, pending
+push upstream. The Profit Estimator gate below still accepts only eight-GPU
+single-node chassis.
A partially allocated chassis (one to seven measured GPUs on one host) is
modeled at measured per-GPU power × 8. That is the same `n_gpu × W/GPU` input
@@ -200,15 +221,23 @@ PowerX 的系统功耗结果以实测 GPU 功率为输入,使用固定版本
固定版本的 Python 模型默认 PUE 为 1.2;PowerX 对当前风冷机箱模型
采用 1.3。市电侧功率 = IT 负载功率 × PUE,风冷取 1.3,直接液冷(DLC)
取 1.1。PUE 仅作用于机箱交流功率,不改变 GPU 实测功率或机箱交流功率。这里的
-冷却方式指建模机箱,并非已核实的测试站点配置。当前模型不支持 DLC;`--pue` 仅
-覆盖设施功率系数,不会把风冷机箱模型转换为液冷模型。
+冷却方式指建模机箱,并非已核实的测试站点配置。机箱模型不支持 DLC;`--pue` 仅
+覆盖设施功率系数,不会把风冷机箱模型转换为液冷模型。NVL72 机架 profile 为直接
+液冷,默认 PUE 取 1.1。
仅使用部分 GPU 的机箱(单台主机上实测 1–7 张 GPU)按实测每卡功率 × 8 建模,
与模型源码 sweep 脚本喂给各机箱模型的 `n_gpu × W/GPU` 输入一致,并假设未实测的
GPU 运行相同负载。结果标记为 `chassisBasis: 'extrapolated'`:每卡数值按建模机箱
的 GPU 总数分摊,`deploymentAcWatts` 只保留实测 GPU 在各机箱中的份额。这不是把
部分分配的机箱按比例分摊:固定组件、风扇曲线和 PSU 效率都在满机箱负载点求值。
-GB200、GB300 没有匹配模型,也不能套用 B200、B300 模型。缺失、无效和不支持的
+GB200、GB300 使用单独的 NVL72 机架 profile,不套用 B200、B300 机箱模型:每台实测
+worker 主机视为一个计算 tray(4 张 GPU、2 个 Grace socket)。输入为实测模块功耗
+(`avg_total_module_power_w`);缺失时改用 GPU 板卡功耗加 Grace socket 功耗
+(`avg_total_cpu_power_w`),并按来源模型计入 GPU 份额的稳压损耗余量。Grace CPU 与
+LPDDR5X 从不建模,缺少 `cpu_power_valid=1` 和 Grace 侧指标的行保持不可用
+(`cpu-telemetry`)。单个 tray 的估算取由相同 tray 组成的整机架的 1/18,因此 NVSwitch
+tray、电源架和管理交换机按 72 张 GPU 分摊;部分分配的 tray 只外推 GPU 板卡份额
+(模块读数本身已覆盖整个 tray),并标记为 `extrapolated`。缺失、无效和不支持的
情况保持不可用。纯 CPU frontend worker 不计入 GPU 机箱数;独立的纯 CPU
frontend/router 主机不在估算范围内,GPU 机箱内的 CPU 功率仍按 20% 利用率计算。
diff --git a/packages/app/scripts/generate-system-power-reference.py b/packages/app/scripts/generate-system-power-reference.py
index 1c303cfe2..e149c4c66 100644
--- a/packages/app/scripts/generate-system-power-reference.py
+++ b/packages/app/scripts/generate-system-power-reference.py
@@ -4,6 +4,11 @@
Usage: python3 packages/app/scripts/generate-system-power-reference.py /path/to/inferencex_power_model
Only Python's standard library is required. No telemetry, dependencies, or GPUs are fetched.
Apply the repository formatter to generated JSON before committing.
+
+Chassis profiles (`profiles`) describe one eight-GPU HGX/OAM system whose GPU watts
+are the only measured input. Rack profiles (`rackProfiles`) describe one NVL72 rack
+whose compute-module watts per tray are measured (module sensor, or GPU board plus
+Grace socket); the Grace CPU and LPDDR5X are never modelled.
"""
import argparse
@@ -16,7 +21,9 @@
import sys
sys.dont_write_bytecode = True
-REVISION = "ca4403aa527069857351ad8047dbb726844b3382"
+# feat/gb200-nvl72-rack-model on top of ca4403aa (PR #10 merge); pending push to the upstream repo.
+REVISION = "963ead8b20a722595c34f7f4a0041259501cf019"
+REVISION_STATUS = "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream"
SOURCE = "https://github.com/SemiAnalysisAI/inferencex_power_model"
MODELS = {
"h100": ("hgx_h100_chassis/h100_chassis_power_model.py", "h100_chassis_power", "make_h100_config"),
@@ -28,6 +35,31 @@
"mi355x": ("mi355x_chassis/mi355x_chassis_power_model.py", "mi355x_chassis_power", "MI355XChassisConfig"),
}
ASSUMPTIONS = {"u_pcie": 0.05, "u_cpu": 0.20, "u_ram": 0.20, "u_nvme": 0.0, "pue": 1.20}
+RACK_MODEL_PATH = "gb200_nvl72_rack/gb200_nvl72_rack_power_model.py"
+RACK_MODELS = {
+ "gb200": (RACK_MODEL_PATH, "gb200_nvl72_rack_power", "gb200_nvl72_rack_config"),
+ "gb300": (RACK_MODEL_PATH, "gb200_nvl72_rack_power", "gb300_nvl72_rack_config"),
+}
+# Same fixed network utilization as the chassis sweep; the rack model has no CPU/DRAM inputs.
+RACK_ASSUMPTIONS = {"u_nvlink": 0.50, "u_ib": 0.0, "u_pcie": 0.05, "pue": 1.20}
+RACK_BASES = {"module": "module", "gpu_plus_grace": "gpu-plus-grace"}
+
+
+def rack_expected(result):
+ return {
+ "computeModulesDcWatts": result["compute_modules_dc_w"],
+ "regulatorAllowanceWatts": result["regulator_allowance_w"],
+ "trayStaticDcWatts": result["tray_static_dc_w"],
+ "nvswitchTraysDcWatts": result["nvswitch_trays_dc_w"],
+ "trayConversionLossWatts": result["tray_conversion_loss_w"],
+ "rackDcWatts": result["rack_dc_w"],
+ "powerShelfEfficiency": result["power_shelf_efficiency"],
+ "powerShelfLossWatts": result["power_shelf_loss_w"],
+ "rackAcWatts": result["rack_ac_w"],
+ "facilityWatts": result["facility_w"],
+ "perGpuAcWatts": result["per_gpu_ac_w"],
+ "perGpuFacilityWatts": result["per_gpu_facility_w"],
+ }
def git(repo, *args):
@@ -130,25 +162,120 @@ def evaluate(gpu, pue=1.2):
case["referenceError"] = str(error)
cases.append(case)
+ rack_profiles, rack_cases = {}, []
+ for hardware, (model_path, function_name, config_factory) in RACK_MODELS.items():
+ path = repo / "human_verified" / model_path
+ sys.path.insert(0, str(path.parent))
+ module = importlib.import_module(path.stem)
+ function, cfg = getattr(module, function_name), getattr(module, config_factory)()
+ utilization = {key: RACK_ASSUMPTIONS[key] for key in ("u_nvlink", "u_ib", "u_pcie")}
+ # Everything except the measured compute modules and the shelf curve is fixed at
+ # these utilizations. Keep the source's per-tray block order: the app re-sums it.
+ tray_blocks, tray_details = module._compute_tray_static(
+ cfg.compute_tray, u_ib=utilization["u_ib"], u_pcie=utilization["u_pcie"])
+ nvswitch = module.nvswitch5_power(utilization["u_nvlink"], cfg.nvswitch)
+ shelf = cfg.power_shelf
+ rack_profiles[hardware] = {
+ "modelPath": "human_verified/" + model_path,
+ "functionName": function_name,
+ "configFactory": config_factory,
+ "topology": "nvl72-rack",
+ "gpuCount": cfg.n_gpu,
+ "computeTrayCount": cfg.n_compute_trays,
+ "gpusPerComputeTray": cfg.compute_tray.n_gpu,
+ "graceSocketsPerComputeTray": cfg.compute_tray.n_grace,
+ "nvswitchTrayCount": cfg.n_nvswitch_trays,
+ "assumptions": dict(RACK_ASSUMPTIONS),
+ "defaultConfig": asdict(cfg),
+ "computeTrayStaticDcWatts": tray_blocks,
+ "computeTrayStaticDetails": tray_details,
+ "nvswitchTraySiliconWatts": nvswitch["pair_w"],
+ "nvswitchTrayResidualWatts": cfg.nvswitch_tray_residual_w,
+ "managementSwitchCount": cfg.n_management_switches,
+ "managementSwitchWatts": cfg.management_switch_w,
+ "trayInputConversionEfficiency": cfg.tray_input_conversion_efficiency,
+ "regulatorLossFracOfTdp": cfg.regulator_loss_frac_of_tdp,
+ "regulatorAllowanceIncludesGrace": cfg.regulator_allowance_includes_grace,
+ "powerShelf": {
+ "installedCapacityWatts": shelf.installed_capacity_w,
+ "redundantCapacityWatts": shelf.redundant_capacity_w,
+ "efficiencyCurve": sorted(shelf.efficiency_curve.items()),
+ },
+ "unverifiedParameters": module.UNVERIFIED_PARAMETERS,
+ }
+
+ def evaluate_rack(basis, pue=1.2, **watts):
+ return function(basis=basis, cfg=cfg, **{**utilization, "pue": pue}, **watts)
+
+ installed, n_trays = shelf.installed_capacity_w, cfg.n_compute_trays
+ # 5400 W/tray is the GB200 module TDP anchor (2 x 2700 W) from the research note.
+ module_samples = {0.05, 1.25, 1000.25, 2000.0, 3000.75, 4000.0, 5400.0, 6000.0, 7200.0}
+ baseline_dc = evaluate_rack("module", module_w_per_tray=0.0)["rack_dc_w"]
+ # Both sides of every reachable shelf-efficiency knot, on rack DC load.
+ for fraction in sorted(shelf.efficiency_curve):
+ target = fraction * installed
+ if not baseline_dc < target <= installed:
+ continue
+ lo, hi = 0.0, installed / n_trays
+ for _ in range(50):
+ mid = (lo + hi) / 2
+ try:
+ below = evaluate_rack("module", module_w_per_tray=mid)["rack_dc_w"] < target
+ except ValueError:
+ below = False
+ if below:
+ lo = mid
+ else:
+ hi = mid
+ module_samples.update(round(hi + offset, 3) for offset in (-0.2, 0.0, 0.2))
+ module_samples.add(installed / n_trays) # Shelf overflow must be unavailable, never clamped.
+ for module_w in sorted(module_samples):
+ for pue in (1.0, 1.1, 1.2):
+ case = {"hardware": hardware, "basis": RACK_BASES["module"],
+ "moduleWattsPerTray": module_w, "pue": pue}
+ try:
+ case["expected"] = rack_expected(evaluate_rack("module", pue, module_w_per_tray=module_w))
+ except ValueError as error:
+ case["expected"] = None
+ case["referenceError"] = str(error)
+ rack_cases.append(case)
+ for gpu_w in (0.05, 1.25, 2000.0, 3000.25, 4800.0, 5600.0):
+ for grace_w in (0.05, 300.0, 600.5):
+ for pue in (1.0, 1.1):
+ result = evaluate_rack("gpu_plus_grace", pue, gpu_board_w_per_tray=gpu_w,
+ grace_socket_w_per_tray=grace_w)
+ rack_cases.append({"hardware": hardware, "basis": RACK_BASES["gpu_plus_grace"],
+ "gpuBoardWattsPerTray": gpu_w, "graceSocketWattsPerTray": grace_w,
+ "pue": pue, "expected": rack_expected(result)})
+
source_hashes = {}
# Include the complete pinned Python implementation and plot entry points.
for relative in git(repo, "ls-files", "*.py").splitlines():
source_hashes[relative] = hashlib.sha256((repo / relative).read_bytes()).hexdigest()
- provenance = {"modelRevision": REVISION, "source": SOURCE, "status": "DRAFT / pending human verification"}
+ provenance = {
+ "modelRevision": REVISION,
+ "modelRevisionStatus": REVISION_STATUS,
+ "source": SOURCE,
+ "status": "DRAFT / pending human verification",
+ }
outputs = {
"system-power-model.profiles.json": {
**provenance,
"assumptionsSource": f"{SOURCE}/blob/{REVISION}/README.md#chassis-models",
"assumptions": ASSUMPTIONS,
+ "rackAssumptionsSource": f"{SOURCE}/blob/{REVISION}/human_verified/{RACK_MODEL_PATH}",
+ "rackAssumptions": RACK_ASSUMPTIONS,
"sourceSha256": dict(sorted(source_hashes.items())),
"profiles": profiles,
+ "rackProfiles": rack_profiles,
},
- "system-power-model.reference.json": {**provenance, "cases": cases},
+ "system-power-model.reference.json": {**provenance, "cases": cases, "rackCases": rack_cases},
}
args.output_dir.mkdir(parents=True, exist_ok=True)
for filename, payload in outputs.items():
(args.output_dir / filename).write_text(json.dumps(payload, indent=2, allow_nan=False) + "\n")
- print(f"Generated {len(profiles)} profiles and {len(cases)} Python reference cases at {REVISION}")
+ print(f"Generated {len(profiles)} chassis profiles, {len(rack_profiles)} rack profiles, "
+ f"{len(cases)} chassis and {len(rack_cases)} rack Python reference cases at {REVISION}")
if __name__ == "__main__":
diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts
index 314941a89..d3fecb9fb 100644
--- a/packages/app/src/components/inference/utils/tooltipUtils.ts
+++ b/packages/app/src/components/inference/utils/tooltipUtils.ts
@@ -240,6 +240,8 @@ const SYSTEM_POWER_STRINGS = {
topology: 'The available topology does not establish chassis placement.',
'role-power': 'Valid measured power and topology are required for every GPU worker role.',
'model-domain': 'The measured input is outside the source model’s supported range.',
+ 'cpu-telemetry':
+ 'Validated measured Grace-side (CPU) power is required for NVL72 compute trays.',
} satisfies Record,
},
zh: {
@@ -269,6 +271,7 @@ const SYSTEM_POWER_STRINGS = {
topology: '现有拓扑信息无法确认 GPU 所在的机箱。',
'role-power': '每个 GPU worker 角色都需要有效的实测功耗和拓扑信息。',
'model-domain': '实测输入超出功耗模型的支持范围。',
+ 'cpu-telemetry': 'NVL72 计算 tray 需要通过验证的 Grace 侧(CPU)实测功耗。',
} satisfies Record,
},
} as const;
diff --git a/packages/app/src/lib/modeled-system-power-export.test.ts b/packages/app/src/lib/modeled-system-power-export.test.ts
index 59f6c57fa..dd221dce9 100644
--- a/packages/app/src/lib/modeled-system-power-export.test.ts
+++ b/packages/app/src/lib/modeled-system-power-export.test.ts
@@ -223,10 +223,11 @@ describe('offline modeled PowerX comparisons', () => {
assumptions: { u_cpu: 0.2 },
model_path: 'human_verified/hgx_h200_chassis/h200_chassis_power_model.py',
});
+ // NVL72 rows need the schema-v2 contract; the unversioned exception is x86 single-node only.
source.rows[0].benchmark.hardware = 'gb200';
expect(buildComparison(source).rows[0].modeled).toMatchObject({
status: 'unsupported',
- reason: 'hardware',
+ reason: 'telemetry',
});
expect(csv([{ a: null, b: 0, c: 'a,"b"\nc' }])).toBe('"a","b","c"\r\n,"0","a,""b""\nc"\r\n');
expect(() => buildComparison({ ...source, rows: [source.rows[0], source.rows[0]] })).toThrow(
diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts
index c968537ae..b89d7e590 100644
--- a/packages/app/src/lib/modeled-system-power.test.ts
+++ b/packages/app/src/lib/modeled-system-power.test.ts
@@ -3,7 +3,7 @@ import { describe, expect, it } from 'vitest';
import type { BenchmarkRow } from '@/lib/api';
import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform';
import { modelSystemPower } from '@/lib/modeled-system-power';
-import { estimateChassisPower } from '@/lib/system-power-model';
+import { estimateChassisPower, estimateRackPower } from '@/lib/system-power-model';
// Qwen3.5 B200 c1, run 34175132645: actual rounded telemetry, eight GPUs.
function row(overrides: Partial = {}): BenchmarkRow {
@@ -49,6 +49,36 @@ function row(overrides: Partial = {}): BenchmarkRow {
};
}
+// One GB200 NVL72 compute tray: four GPUs on one host, two Grace sockets. The
+// CPU-side keys follow the ticket-01 contract (sums over every socket, same window).
+// Watts are controlled inputs, not published constants.
+const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 };
+function nvl72Row(
+ metrics: Record = {},
+ overrides: Partial = {},
+): BenchmarkRow {
+ return row({
+ hardware: 'gb200',
+ framework: 'dynamo-trt',
+ prefill_tp: 4,
+ decode_tp: 4,
+ num_prefill_gpu: 4,
+ num_decode_gpu: 4,
+ metrics: {
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ avg_power_w: 900.25,
+ avg_total_gpu_power_w: 3601,
+ ...GRACE,
+ pp: 1,
+ pcp_size: 1,
+ ...(metrics as Record),
+ },
+ ...overrides,
+ });
+}
+
describe('modeled system power admission and accounting', () => {
it('defaults air-cooled chassis to PUE 1.3 and preserves explicit facility overrides', () => {
// Pinned Python b200_chassis_power, fixed README utilization inputs.
@@ -88,16 +118,29 @@ describe('modeled system power admission and accounting', () => {
expect(source.metrics.joules_per_output_token).toBe(12.937902);
});
- it.each(['gb200', 'gb300', 'rtx6000pro', 'tpuv7', 'b200-nvl'])(
- 'does not substitute for %s',
+ it.each(['rtx6000pro', 'tpuv7', 'b200-nvl'])('does not substitute for %s', (hardware) => {
+ expect(modelSystemPower(row({ hardware }))).toMatchObject({
+ status: 'unsupported',
+ reason: 'hardware',
+ });
+ });
+
+ it.each(['gb200', 'gb300'])(
+ 'never models the Grace side of %s from GPU-only telemetry',
(hardware) => {
expect(modelSystemPower(row({ hardware }))).toMatchObject({
status: 'unsupported',
- reason: 'hardware',
+ reason: 'cpu-telemetry',
});
},
);
+ it('ignores CPU-side keys on x86 chassis rows', () => {
+ const source = row();
+ Object.assign(source.metrics, GRACE, { cpu_power_valid: 1, avg_total_module_power_w: 4300 });
+ expect(modelSystemPower(source)).toEqual(modelSystemPower(row()));
+ });
+
it.each([{ benchmark_type: 'agentic_traces' }, { isl: 1024 }, { osl: 8192 }, { isl: null }])(
'keeps non-8k1k workloads unavailable: %j',
(overrides) => {
@@ -471,3 +514,196 @@ describe('modeled system power admission and accounting', () => {
}
});
});
+
+describe('NVL72 trays with measured compute-module power', () => {
+ const MODULE = { avg_total_module_power_w: 4300.75 };
+
+ it('models a full GB200 tray on the module basis with the DLC PUE applied once', () => {
+ const result = modelSystemPower(nvl72Row(MODULE));
+ const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!;
+ expect(result).toMatchObject({
+ status: 'supported',
+ hardware: 'gb200',
+ modelPath: rack.modelPath,
+ gpuCount: 4,
+ chassisCount: 1,
+ modeledGpuCount: 4,
+ measuredGpuWattsPerGpu: 900.25,
+ chassisAcWatts: rack.rackAcWatts / 18,
+ chassisAcWattsPerGpu: rack.rackAcWatts / 18 / 4,
+ facilityWatts: rack.facilityWatts / 18,
+ deploymentAcWatts: rack.rackAcWatts / 18,
+ deploymentFacilityWatts: rack.facilityWatts / 18,
+ pue: 1.1,
+ telemetryBasis: 'validated-v2',
+ topologyBasis: 'nvl72-trays',
+ chassisBasis: 'full',
+ measuredBasis: 'module',
+ sensorKind: 'module',
+ });
+ const noPue = modelSystemPower(nvl72Row(MODULE), 1);
+ expect(noPue).toMatchObject({ pue: 1, chassisAcWatts: rack.rackAcWatts / 18 });
+ expect(noPue.status === 'supported' && noPue.facilityWatts).toBe(rack.rackAcWatts / 18);
+ expect(modelSystemPower(nvl72Row(MODULE), 1.3)).toMatchObject({ pue: 1.3 });
+ // The measured compute module is the input; the rack residual is added on top.
+ expect(result.status === 'supported' && result.deploymentAcWatts).toBeGreaterThan(4300.75);
+ });
+
+ it('falls back to GPU board plus Grace socket for GB300 trays without module keys', () => {
+ const source = nvl72Row(
+ {
+ avg_power_w: 900,
+ avg_total_gpu_power_w: 7200,
+ prefill_avg_power_w: 950,
+ decode_avg_power_w: 850,
+ avg_cpu_socket_power_w: 260,
+ avg_total_cpu_power_w: 1040,
+ },
+ {
+ hardware: 'gb300',
+ disagg: true,
+ is_multinode: true,
+ workers: [
+ { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 950 },
+ { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 850 },
+ ],
+ },
+ );
+ // Grace-side watts are a deployment total; each tray receives the two-socket mean.
+ const prefill = estimateRackPower(
+ 'gb300',
+ { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 3800, graceSocketWattsPerTray: 520 },
+ 1.1,
+ )!;
+ const decode = estimateRackPower(
+ 'gb300',
+ { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 3400, graceSocketWattsPerTray: 520 },
+ 1.1,
+ )!;
+ expect(modelSystemPower(source)).toMatchObject({
+ status: 'supported',
+ hardware: 'gb300',
+ gpuCount: 8,
+ chassisCount: 2,
+ modeledGpuCount: 8,
+ chassisAcWatts: prefill.rackAcWatts / 18 + decode.rackAcWatts / 18,
+ facilityWatts: prefill.facilityWatts / 18 + decode.facilityWatts / 18,
+ deploymentAcWatts: prefill.rackAcWatts / 18 + decode.rackAcWatts / 18,
+ pue: 1.1,
+ topologyBasis: 'nvl72-trays',
+ chassisBasis: 'full',
+ measuredBasis: 'gpu-plus-grace',
+ sensorKind: 'grace-socket',
+ });
+ // Two hosts carry four Grace sockets; any other socket count is not a tray topology.
+ source.metrics.avg_total_cpu_power_w = 260 * 3;
+ expect(modelSystemPower(source)).toMatchObject({ reason: 'cpu-telemetry' });
+ source.metrics.avg_total_cpu_power_w = 1040;
+ source.workers![1].num_gpus = 5;
+ expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' });
+ source.workers![1].num_gpus = 4;
+ source.workers![1].hosts = ['tray-b', 'tray-c'];
+ expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' });
+ });
+
+ it('extrapolates a partially measured tray on the GPU-board share and keeps the measured share', () => {
+ const partial = nvl72Row(
+ { avg_total_gpu_power_w: 2700.75 },
+ { prefill_tp: 3, decode_tp: 3, num_prefill_gpu: 3, num_decode_gpu: 3 },
+ );
+ // Both Grace sockets are measured regardless of allocation; only the GPU board
+ // share is the tray's per-GPU mean × 4, mirroring the partial-chassis rule.
+ const rack = estimateRackPower(
+ 'gb200',
+ { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 900.25 * 4, graceSocketWattsPerTray: 501 },
+ 1.1,
+ )!;
+ expect(modelSystemPower(partial)).toMatchObject({
+ status: 'supported',
+ gpuCount: 3,
+ chassisCount: 1,
+ modeledGpuCount: 4,
+ measuredGpuWattsPerGpu: 900.25,
+ chassisAcWatts: rack.rackAcWatts / 18,
+ chassisAcWattsPerGpu: rack.rackAcWatts / 18 / 4,
+ deploymentAcWatts: ((rack.rackAcWatts / 18) * 3) / 4,
+ deploymentFacilityWatts: ((rack.facilityWatts / 18) * 3) / 4,
+ topologyBasis: 'nvl72-trays',
+ chassisBasis: 'extrapolated',
+ measuredBasis: 'gpu-plus-grace',
+ });
+ // Module sensors cover the whole tray, idle GPUs included, so that reading is
+ // never scaled; the label still records the modeled-versus-measured count.
+ const partialModule = nvl72Row(
+ { avg_total_gpu_power_w: 2700.75, ...MODULE },
+ { prefill_tp: 3, decode_tp: 3, num_prefill_gpu: 3, num_decode_gpu: 3 },
+ );
+ const moduleRack = estimateRackPower(
+ 'gb200',
+ { basis: 'module', moduleWattsPerTray: 4300.75 },
+ 1.1,
+ )!;
+ expect(modelSystemPower(partialModule)).toMatchObject({
+ gpuCount: 3,
+ modeledGpuCount: 4,
+ chassisAcWatts: moduleRack.rackAcWatts / 18,
+ deploymentAcWatts: ((moduleRack.rackAcWatts / 18) * 3) / 4,
+ chassisBasis: 'extrapolated',
+ measuredBasis: 'module',
+ });
+ // Four measured GPUs cannot establish a TP8 width; eight cannot sit on one tray.
+ partial.prefill_tp = 8;
+ partial.decode_tp = 8;
+ expect(modelSystemPower(partial)).toMatchObject({ reason: 'gpu-count' });
+ const twoTrays = nvl72Row(
+ { avg_total_gpu_power_w: 7202, avg_total_cpu_power_w: 1002 },
+ { prefill_tp: 8, decode_tp: 8, num_prefill_gpu: 8, num_decode_gpu: 8 },
+ );
+ expect(modelSystemPower(twoTrays)).toMatchObject({ reason: 'topology' });
+ });
+
+ it.each([
+ { cpu_power_valid: undefined },
+ { cpu_power_valid: 0 },
+ { cpu_power_valid: '1' },
+ { avg_total_cpu_power_w: undefined },
+ { avg_total_cpu_power_w: 0 },
+ { avg_total_cpu_power_w: -1 },
+ { avg_cpu_socket_power_w: undefined },
+ { avg_cpu_socket_power_w: 0 },
+ { avg_total_module_power_w: 0 },
+ { avg_total_module_power_w: -1 },
+ { avg_total_module_power_w: NaN },
+ { avg_total_module_power_w: '4300' },
+ ])('keeps NVL72 rows without valid CPU-side telemetry unavailable: %j', (overrides) => {
+ const source = nvl72Row(MODULE);
+ Object.assign(source.metrics, overrides);
+ expect(modelSystemPower(source)).toMatchObject({
+ status: 'unsupported',
+ reason: 'cpu-telemetry',
+ });
+ });
+
+ it('requires schema-v2 GPU telemetry and stays within the shelf and facility domain', () => {
+ const legacy = nvl72Row({ ...MODULE, power_metric_schema_version: undefined });
+ expect(modelSystemPower(legacy)).toMatchObject({ status: 'unsupported', reason: 'telemetry' });
+ expect(
+ modelSystemPower(nvl72Row({ avg_total_module_power_w: Number.MAX_VALUE })),
+ ).toMatchObject({
+ reason: 'model-domain',
+ });
+ expect(modelSystemPower(nvl72Row(MODULE), 1e304)).toMatchObject({ reason: 'model-domain' });
+ expect(modelSystemPower(nvl72Row(MODULE), 0.9)).toMatchObject({ reason: 'model-domain' });
+ });
+
+ it('plots the amortised rack AC per GPU through the shared transform', () => {
+ const source = nvl72Row(MODULE);
+ const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!;
+ const { chartData } = transformBenchmarkRows([source]);
+ expect(chartData.length).toBeGreaterThan(0);
+ for (const points of chartData) {
+ expect(points[0].modeledChassisPowerPerGpu?.y).toBe(rack.rackAcWatts / 18 / 4);
+ expect(points[0].measuredAvgPower?.y).toBe(900.25);
+ }
+ });
+});
diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts
index 550fe23fc..1e96423f6 100644
--- a/packages/app/src/lib/modeled-system-power.ts
+++ b/packages/app/src/lib/modeled-system-power.ts
@@ -1,12 +1,19 @@
import type { BenchmarkRow } from '@/lib/api';
import {
estimateChassisPower,
+ estimateRackPower,
+ type RackMeasuredBasis,
+ type RackMeasuredInput,
SUPPORTED_SYSTEM_POWER_HARDWARE,
SYSTEM_POWER_MODEL_REVISION,
+ SYSTEM_POWER_RACK_PROFILES,
+ type SystemPowerRackHardware,
} from '@/lib/system-power-model';
// Application policy for the air-cooled chassis profiles; the pinned Python default stays 1.2.
export const AIR_COOLED_SYSTEM_PUE = 1.3;
+// Application policy for the direct-liquid-cooled NVL72 rack profiles (docs/powerx-system-power.md).
+export const DLC_SYSTEM_PUE = 1.1;
/** Every supported chassis model describes one complete eight-GPU HGX/OAM system. */
const CHASSIS_GPU_COUNT = 8;
@@ -15,50 +22,72 @@ export type SystemPowerUnsupportedReason =
| 'workload'
| 'hardware'
| 'telemetry'
+ /** NVL72 only: the Grace side is measured or the row stays unavailable. */
+ | 'cpu-telemetry'
| 'gpu-count'
| 'topology'
| 'role-power'
| 'model-domain';
+/** Producer sensor behind the CPU-side keys; the module sensor whenever its keys are present. */
+export type SystemPowerSensorKind = 'module' | 'grace-socket';
+
+interface SupportedSystemPowerEstimate {
+ status: 'supported';
+ hardware: string;
+ modelRevision: string;
+ modelPath: string;
+ /** Physical GPUs covered by the validated telemetry. */
+ gpuCount: number;
+ /** Modeled units: eight-GPU chassis, or NVL72 compute trays. */
+ chassisCount: number;
+ /**
+ * GPUs the units were evaluated for: chassisCount × 8 for chassis, × 4 for trays.
+ * Exceeds gpuCount when extrapolated.
+ */
+ modeledGpuCount: number;
+ measuredGpuWattsPerGpu: number;
+ /**
+ * Modeled AC for every full unit, summed. A tray's AC is its 1/18 share of a rack
+ * whose trays all match it, so the switch trays, shelves, and management switches
+ * are amortised over all 72 GPUs.
+ */
+ chassisAcWatts: number;
+ /** chassisAcWatts ÷ modeledGpuCount: the plotted metric. */
+ chassisAcWattsPerGpu: number;
+ facilityWatts: number;
+ /** Share of the modeled units attributable to the measured GPUs; equals the totals for full units. */
+ deploymentAcWatts: number;
+ deploymentFacilityWatts: number;
+ pue: number;
+ telemetryBasis: 'validated-v2' | 'validated-unversioned-single-node';
+ /**
+ * 'full': every unit had all of its GPUs measured. 'extrapolated': at least one
+ * unit was partially allocated. For chassis and for the tray GPU-board share, the
+ * model input is the measured per-GPU power × the unit's GPU count, assuming the
+ * unmeasured GPUs run the same workload (the source README sweep's own n_gpu × W/GPU
+ * input, not a proportional share of a unit evaluated at partial load). A module
+ * sensor already covers the whole tray, idle GPUs included, so that reading is
+ * never scaled; the label then records the modeled-versus-measured GPU count.
+ */
+ chassisBasis: 'full' | 'extrapolated';
+}
+
export type SystemPowerEstimate =
| { status: 'unsupported'; reason: SystemPowerUnsupportedReason; modelRevision: string }
- | {
- status: 'supported';
- hardware: string;
- modelRevision: string;
- modelPath: string;
- /** Physical GPUs covered by the validated telemetry. */
- gpuCount: number;
- chassisCount: number;
- /** GPUs the chassis models were evaluated for: chassisCount × 8. Exceeds gpuCount when extrapolated. */
- modeledGpuCount: number;
- measuredGpuWattsPerGpu: number;
- /** Modeled AC for every full chassis, summed. */
- chassisAcWatts: number;
- /** chassisAcWatts ÷ modeledGpuCount: the plotted metric. */
- chassisAcWattsPerGpu: number;
- facilityWatts: number;
- /** Share of the modeled chassis attributable to the measured GPUs; equals the totals for full chassis. */
- deploymentAcWatts: number;
- deploymentFacilityWatts: number;
- pue: number;
- telemetryBasis: 'validated-v2' | 'validated-unversioned-single-node';
- topologyBasis: 'single-node' | 'worker-hosts';
- /**
- * 'full': every chassis had all eight GPUs measured. 'extrapolated': at least
- * one chassis was partially allocated; its model input is the measured per-GPU
- * power × 8, assuming the unmeasured GPUs run the same workload. This is the
- * source README sweep's own n_gpu × W/GPU input, not a proportional share of a
- * chassis evaluated at partial load.
- */
- chassisBasis: 'full' | 'extrapolated';
- };
+ | (SupportedSystemPowerEstimate & { topologyBasis: 'single-node' | 'worker-hosts' })
+ | (SupportedSystemPowerEstimate & {
+ topologyBasis: 'nvl72-trays';
+ /** Which measured reading fed every tray: the module sensor, or GPU board + Grace socket. */
+ measuredBasis: RackMeasuredBasis;
+ sensorKind: SystemPowerSensorKind;
+ });
-interface MeasuredChassis {
- /** GPUs on this chassis covered by telemetry (1–8). */
+interface MeasuredUnit {
+ /** GPUs on this chassis or tray covered by telemetry. */
measuredGpus: number;
- /** Full-chassis GPU watts handed to the source model. */
- modelInputWatts: number;
+ /** Full-unit GPU-board watts handed to the source model. */
+ gpuBoardWatts: number;
}
interface MeasuredWorker {
@@ -67,12 +96,26 @@ interface MeasuredWorker {
watts: number;
}
+interface CpuSideTelemetry {
+ basis: RackMeasuredBasis;
+ socketCount: number;
+ /** Deployment totals over every Grace socket; each tray receives the mean. */
+ graceTotalWatts: number;
+ moduleTotalWatts: number | undefined;
+}
+
+interface UnitPower {
+ acWatts: number;
+ facilityWatts: number;
+ modelPath: string;
+ modelRevision: string;
+}
+
const sumGpus = (items: MeasuredWorker[]) => items.reduce((sum, c) => sum + c.gpus, 0);
const sumWatts = (items: MeasuredWorker[]) => items.reduce((sum, c) => sum + c.watts, 0);
const positive = (n: unknown): n is number => typeof n === 'number' && Number.isFinite(n) && n > 0;
const count = (n: unknown): n is number => positive(n) && Number.isSafeInteger(n);
-const chassisShare = (n: unknown): n is number => count(n) && n <= CHASSIS_GPU_COUNT;
// The schema-v2 producer rounds each watts field to 0.001 W. This bound
// accounts for both the aggregate rounding and every multiplied mean.
@@ -84,14 +127,59 @@ function unavailable(reason: SystemPowerUnsupportedReason): SystemPowerEstimate
return { status: 'unsupported', reason, modelRevision: SYSTEM_POWER_MODEL_REVISION };
}
+function modelChassis(hardware: string, unit: MeasuredUnit, pue: number): UnitPower | null {
+ const chassis = estimateChassisPower(hardware, unit.gpuBoardWatts, pue);
+ return (
+ chassis && {
+ acWatts: chassis.chassisAcWatts,
+ facilityWatts: chassis.facilityWatts,
+ modelPath: chassis.modelPath,
+ modelRevision: chassis.modelRevision,
+ }
+ );
+}
+
+/**
+ * One tray's share of a rack whose trays all match it. The producer publishes the
+ * Grace-side and module readings as deployment totals, so every tray receives the mean.
+ * The Grace CPU and LPDDR5X are never modelled: they are inside the measured reading.
+ */
+function modelTray(
+ hardware: string,
+ profile: (typeof SYSTEM_POWER_RACK_PROFILES)[SystemPowerRackHardware],
+ unit: MeasuredUnit,
+ cpu: CpuSideTelemetry,
+ trayCount: number,
+ pue: number,
+): UnitPower | null {
+ const input: RackMeasuredInput =
+ cpu.moduleTotalWatts === undefined
+ ? {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: unit.gpuBoardWatts,
+ graceSocketWattsPerTray: cpu.graceTotalWatts / trayCount,
+ }
+ : { basis: 'module', moduleWattsPerTray: cpu.moduleTotalWatts / trayCount };
+ const rack = estimateRackPower(hardware, input, pue);
+ return (
+ rack && {
+ acWatts: rack.rackAcWatts / profile.computeTrayCount,
+ facilityWatts: rack.facilityWatts / profile.computeTrayCount,
+ modelPath: rack.modelPath,
+ modelRevision: rack.modelRevision,
+ }
+ );
+}
+
/**
- * Model the mean GPU telemetry on known eight-GPU chassis.
- * This is f(mean GPU power), not a time-integrated wall-power measurement.
- * Do not use display counts here: legacy ingest can encode TP * EP twice.
+ * Model the mean GPU telemetry on known eight-GPU chassis, or on NVL72 compute trays
+ * whose Grace side is measured. This is f(mean power), not a time-integrated wall-power
+ * measurement. Do not use display counts here: legacy ingest can encode TP * EP twice.
*/
export function modelSystemPower(
row: BenchmarkRow,
- pue: number = AIR_COOLED_SYSTEM_PUE,
+ /** Facility PUE; defaults to 1.3 for air-cooled chassis and 1.1 for DLC NVL72 racks. */
+ pue?: number,
/** Opt in so AgentX estimates do not widen the ordinary 8K/1K chart policy. */
allowAgenticPreview = false,
): SystemPowerEstimate {
@@ -103,9 +191,15 @@ export function modelSystemPower(
}
if (typeof row.hardware !== 'string') return unavailable('hardware');
const hardware = row.hardware.toLowerCase();
- if (!(SUPPORTED_SYSTEM_POWER_HARDWARE as readonly string[]).includes(hardware)) {
+ const rack = Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, hardware)
+ ? SYSTEM_POWER_RACK_PROFILES[hardware as SystemPowerRackHardware]
+ : null;
+ if (!rack && !(SUPPORTED_SYSTEM_POWER_HARDWARE as readonly string[]).includes(hardware)) {
return unavailable('hardware');
}
+ const facilityPue = pue ?? (rack ? DLC_SYSTEM_PUE : AIR_COOLED_SYSTEM_PUE);
+ const unitGpuCount = rack ? rack.gpusPerComputeTray : CHASSIS_GPU_COUNT;
+ const unitShare = (n: unknown): n is number => count(n) && n <= unitGpuCount;
if (typeof row.disagg !== 'boolean' || typeof row.is_multinode !== 'boolean') {
return unavailable('topology');
}
@@ -113,8 +207,10 @@ export function modelSystemPower(
// The original validated single-node producer already defines these two
// watts fields identically (InferenceX bf4461db, aggregate_power.py). It
// predates the schema version marker; retain that distinction. Unversioned
- // disaggregated/multinode telemetry is not admitted through this exception.
+ // disaggregated/multinode telemetry is not admitted through this exception,
+ // and no NVL72 row predates it.
const unversionedSingleNode =
+ !rack &&
row.disagg === false &&
row.is_multinode === false &&
m?.power_metric_schema_version === undefined;
@@ -137,11 +233,41 @@ export function modelSystemPower(
return unavailable('gpu-count');
}
- const chassis: MeasuredChassis[] = [];
+ // NVL72: the CPU-side keys are sums over every Grace socket in the deployment,
+ // integrated over the GPU window. The socket mean recovers the socket count the
+ // same way the GPU mean recovers the GPU count. A module key that is present but
+ // invalid never falls back to the Grace socket silently.
+ let cpu: CpuSideTelemetry | null = null;
+ if (rack) {
+ const moduleTotal = m.avg_total_module_power_w;
+ if (
+ m.cpu_power_valid !== 1 ||
+ !positive(m.avg_total_cpu_power_w) ||
+ !positive(m.avg_cpu_socket_power_w) ||
+ (moduleTotal !== undefined && !positive(moduleTotal))
+ ) {
+ return unavailable('cpu-telemetry');
+ }
+ const socketCount = Math.round(m.avg_total_cpu_power_w / m.avg_cpu_socket_power_w);
+ if (
+ !count(socketCount) ||
+ !matchingWatts(m.avg_total_cpu_power_w, m.avg_cpu_socket_power_w * socketCount, socketCount)
+ ) {
+ return unavailable('cpu-telemetry');
+ }
+ cpu = {
+ basis: moduleTotal === undefined ? 'gpu-plus-grace' : 'module',
+ socketCount,
+ graceTotalWatts: m.avg_total_cpu_power_w,
+ moduleTotalWatts: moduleTotal,
+ };
+ }
+
+ const units: MeasuredUnit[] = [];
let topologyBasis: 'single-node' | 'worker-hosts';
if (row.disagg === false && row.is_multinode === false) {
- // One host cannot hold more than one chassis.
- if (gpuCount > CHASSIS_GPU_COUNT) return unavailable('topology');
+ // One host cannot hold more than one chassis or tray.
+ if (gpuCount > unitGpuCount) return unavailable('topology');
// The producer's physical width is TP * PP * PCP; EP partitions that
// width. Check the populated aggregate side, not summed role aliases.
const tp = row.decode_tp > 0 ? row.decode_tp : row.prefill_tp;
@@ -164,19 +290,17 @@ export function modelSystemPower(
return unavailable('gpu-count');
}
topologyBasis = 'single-node';
- chassis.push({
+ units.push({
measuredGpus: gpuCount,
- // The producer's exact total avoids re-rounding a full chassis through
+ // The producer's exact total avoids re-rounding a full unit through
// the per-GPU mean.
- modelInputWatts:
- gpuCount === CHASSIS_GPU_COUNT
- ? m.avg_total_gpu_power_w
- : m.avg_power_w * CHASSIS_GPU_COUNT,
+ gpuBoardWatts:
+ gpuCount === unitGpuCount ? m.avg_total_gpu_power_w : m.avg_power_w * unitGpuCount,
});
} else {
// A role average across several hosts is insufficient for nonlinear
- // fan/PSU evaluation. Require one chassis per measured worker and a
- // distinct host for every worker.
+ // fan/PSU or shelf evaluation. Require one chassis or tray per measured
+ // worker and a distinct host for every worker.
if (!Array.isArray(row.workers) || row.workers.length === 0) {
return unavailable('topology');
}
@@ -186,7 +310,7 @@ export function modelSystemPower(
// CPU-only frontends are outside the modeled GPU-chassis boundary.
if (worker.role === 'frontend' && worker.num_gpus === 0) continue;
if (
- !chassisShare(worker.num_gpus) ||
+ !unitShare(worker.num_gpus) ||
!Array.isArray(worker.hosts) ||
worker.hosts.length !== 1 ||
typeof worker.hosts[0] !== 'string' ||
@@ -206,9 +330,9 @@ export function modelSystemPower(
hosts.add(worker.hosts[0]);
const role = row.disagg ? worker.role : 'aggregate';
measured.push({ role, gpus: worker.num_gpus, watts: worker.avg_power_w * worker.num_gpus });
- chassis.push({
+ units.push({
measuredGpus: worker.num_gpus,
- modelInputWatts: worker.avg_power_w * CHASSIS_GPU_COUNT,
+ gpuBoardWatts: worker.avg_power_w * unitGpuCount,
});
}
if (sumGpus(measured) !== gpuCount) return unavailable('gpu-count');
@@ -232,31 +356,38 @@ export function modelSystemPower(
}
topologyBasis = 'worker-hosts';
}
+ // Every compute tray carries two Grace sockets; a different count is not a tray topology.
+ if (rack && cpu && cpu.socketCount !== units.length * rack.graceSocketsPerComputeTray) {
+ return unavailable('cpu-telemetry');
+ }
- const results = chassis.map((c) => ({
- ...c,
- model: estimateChassisPower(hardware, c.modelInputWatts, pue),
+ const results = units.map((unit) => ({
+ ...unit,
+ model:
+ rack && cpu
+ ? modelTray(hardware, rack, unit, cpu, units.length, facilityPue)
+ : modelChassis(hardware, unit, facilityPue),
}));
if (results.some((r) => r.model === null)) return unavailable('model-domain');
const first = results[0].model!;
- const modeledGpuCount = chassis.length * CHASSIS_GPU_COUNT;
- const extrapolated = chassis.some((c) => c.measuredGpus !== CHASSIS_GPU_COUNT);
- const chassisAcWatts = results.reduce((sum, r) => sum + r.model!.chassisAcWatts, 0);
+ const modeledGpuCount = units.length * unitGpuCount;
+ const extrapolated = units.some((c) => c.measuredGpus !== unitGpuCount);
+ const chassisAcWatts = results.reduce((sum, r) => sum + r.model!.acWatts, 0);
const facilityWatts = results.reduce((sum, r) => sum + r.model!.facilityWatts, 0);
const share = (watts: (r: (typeof results)[number]) => number) =>
- results.reduce((sum, r) => sum + (watts(r) * r.measuredGpus) / CHASSIS_GPU_COUNT, 0);
- const deploymentAcWatts = extrapolated ? share((r) => r.model!.chassisAcWatts) : chassisAcWatts;
+ results.reduce((sum, r) => sum + (watts(r) * r.measuredGpus) / unitGpuCount, 0);
+ const deploymentAcWatts = extrapolated ? share((r) => r.model!.acWatts) : chassisAcWatts;
const deploymentFacilityWatts = extrapolated
? share((r) => r.model!.facilityWatts)
: facilityWatts;
if (!positive(chassisAcWatts) || !positive(facilityWatts)) return unavailable('model-domain');
- return {
+ const supported: SupportedSystemPowerEstimate = {
status: 'supported',
hardware,
modelRevision: first.modelRevision,
modelPath: first.modelPath,
gpuCount,
- chassisCount: chassis.length,
+ chassisCount: units.length,
modeledGpuCount,
measuredGpuWattsPerGpu: m.avg_power_w,
chassisAcWatts,
@@ -264,9 +395,17 @@ export function modelSystemPower(
facilityWatts,
deploymentAcWatts,
deploymentFacilityWatts,
- pue,
+ pue: facilityPue,
telemetryBasis: unversionedSingleNode ? 'validated-unversioned-single-node' : 'validated-v2',
- topologyBasis,
chassisBasis: extrapolated ? 'extrapolated' : 'full',
};
+ if (rack && cpu) {
+ return {
+ ...supported,
+ topologyBasis: 'nvl72-trays',
+ measuredBasis: cpu.basis,
+ sensorKind: cpu.basis === 'module' ? 'module' : 'grace-socket',
+ };
+ }
+ return { ...supported, topologyBasis };
}
diff --git a/packages/app/src/lib/system-power-model.profiles.json b/packages/app/src/lib/system-power-model.profiles.json
index 136f769fb..302cc781f 100644
--- a/packages/app/src/lib/system-power-model.profiles.json
+++ b/packages/app/src/lib/system-power-model.profiles.json
@@ -1,8 +1,9 @@
{
- "modelRevision": "ca4403aa527069857351ad8047dbb726844b3382",
+ "modelRevision": "963ead8b20a722595c34f7f4a0041259501cf019",
+ "modelRevisionStatus": "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream",
"source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
"status": "DRAFT / pending human verification",
- "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/ca4403aa527069857351ad8047dbb726844b3382/README.md#chassis-models",
+ "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/963ead8b20a722595c34f7f4a0041259501cf019/README.md#chassis-models",
"assumptions": {
"u_pcie": 0.05,
"u_cpu": 0.2,
@@ -10,6 +11,13 @@
"u_nvme": 0.0,
"pue": 1.2
},
+ "rackAssumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/963ead8b20a722595c34f7f4a0041259501cf019/human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
+ "rackAssumptions": {
+ "u_nvlink": 0.5,
+ "u_ib": 0.0,
+ "u_pcie": 0.05,
+ "pue": 1.2
+ },
"sourceSha256": {
"human_verified/amd_oam_fans/amd_oam_fan_power_model.py": "7ea6f1c65b685311277e6f2c53992d588ec79b598180f2fca300e904ad6dc19a",
"human_verified/amd_oam_fans/plot_amd_oam_fan_power.py": "e38118fa4d94412af268eaa1667d2c4086d1dcced4da07c951e761d54ea55b46",
@@ -30,7 +38,10 @@
"human_verified/b300_ubb/b300_ubb_power_model.py": "8cd694f1306ca09de53dcc9590dcc47a1b9bd5c10af9837bb875dc417fe74837",
"human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7",
"human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db",
- "human_verified/chassis_plot_utils.py": "66f5d76f5ae5342433963595e58f665bdd0955eabdb31041a34c2f8180839af5",
+ "human_verified/chassis_plot_utils.py": "0f1c809a527788c8c25490420e29e737e6b7f83cb914dc022c3a4e3701b80643",
+ "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "f907916213a9176a4d82a05b905ed3c7b595035184150a352dc8049952ffd340",
+ "human_verified/gb200_nvl72_rack/plot_gb200_nvl72_rack_inference_gpu_sweep.py": "d0f6903df0348ccb2790e973b0dc10408c6f6022695ce039aad0b1e1fcd39519",
+ "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "4b33288998d4355d1f4812b9f2e014ab4cf2a588fb9fc94ac56fea9ae8922950",
"human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f",
"human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c",
"human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb",
@@ -1375,5 +1386,337 @@
]
}
}
+ },
+ "rackProfiles": {
+ "gb200": {
+ "modelPath": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
+ "functionName": "gb200_nvl72_rack_power",
+ "configFactory": "gb200_nvl72_rack_config",
+ "topology": "nvl72-rack",
+ "gpuCount": 72,
+ "computeTrayCount": 18,
+ "gpusPerComputeTray": 4,
+ "graceSocketsPerComputeTray": 2,
+ "nvswitchTrayCount": 9,
+ "assumptions": {
+ "u_nvlink": 0.5,
+ "u_ib": 0.0,
+ "u_pcie": 0.05,
+ "pue": 1.2
+ },
+ "defaultConfig": {
+ "variant": "gb200",
+ "gpu_label": "GB200 Blackwell (NVL72)",
+ "n_compute_trays": 18,
+ "n_nvswitch_trays": 9,
+ "compute_tray": {
+ "n_gpu": 4,
+ "n_grace": 2,
+ "nic_generation": "connectx7",
+ "connectx7": {
+ "n_nic": 4,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "connectx8": {
+ "n_nic": 4,
+ "include_optic": true,
+ "network_serdes_static_w": 18.0,
+ "pcie_switch_static_w": 28.0,
+ "board_mgmt_static_w": 5.0,
+ "digital_static_w": 12.0,
+ "network_dynamic_max_w": 8.0,
+ "pcie_switch_dynamic_max_w": 5.0,
+ "digital_dynamic_max_w": 10.0,
+ "optic_idle_w": 15.0,
+ "optic_dynamic_max_w": 2.0,
+ "max_nic_slot_power_ref_w": 75.0,
+ "normal_low_w_per_nic_with_optic": 70.0,
+ "normal_high_w_per_nic_with_optic": 100.0
+ },
+ "dpu": {
+ "n_dpu": 2,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "nvme": {
+ "n_front_u2": 4,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 1,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "fans_w": 130.0,
+ "board_residual_w": 40.0
+ },
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "nvswitch_tray_residual_w": 50.0,
+ "tray_input_conversion_efficiency": 0.9725,
+ "n_management_switches": 2,
+ "management_switch_w": 100.0,
+ "power_shelf": {
+ "n_shelves": 8,
+ "n_psu_per_shelf": 6,
+ "psu_capacity_w": 5500.0,
+ "redundancy": "N+N",
+ "busbar_nominal_v": 50.0,
+ "efficiency_curve": {
+ "0.1": 0.9,
+ "0.2": 0.94,
+ "0.3": 0.965,
+ "1.0": 0.965
+ },
+ "peak_efficiency_ref": 0.975
+ },
+ "regulator_loss_frac_of_tdp": 0.15,
+ "regulator_allowance_includes_grace": false,
+ "pue": 1.2,
+ "cooling": "direct_liquid",
+ "gpu_tdp_example_w_per_gpu": 1200.0,
+ "grace_tdp_example_w_per_socket": 300.0,
+ "module_tdp_example_w_per_superchip": 2700.0
+ },
+ "computeTrayStaticDcWatts": {
+ "connectx7_nics_with_optics": 126.0,
+ "bluefield3_dpu_idle": 130.0,
+ "nvme_idle": 22.0,
+ "compute_tray_fans_unverified": 130.0,
+ "compute_tray_board_residual_unverified": 40.0
+ },
+ "computeTrayStaticDetails": {
+ "nic_generation": "connectx7",
+ "per_nic_w": 31.5,
+ "n_nic": 4,
+ "n_dpu": 2,
+ "dpu_model": "idle_only",
+ "nvme_model": "idle_only"
+ },
+ "nvswitchTraySiliconWatts": 406.4,
+ "nvswitchTrayResidualWatts": 50.0,
+ "managementSwitchCount": 2,
+ "managementSwitchWatts": 100.0,
+ "trayInputConversionEfficiency": 0.9725,
+ "regulatorLossFracOfTdp": 0.15,
+ "regulatorAllowanceIncludesGrace": false,
+ "powerShelf": {
+ "installedCapacityWatts": 264000.0,
+ "redundantCapacityWatts": 132000.0,
+ "efficiencyCurve": [
+ [0.1, 0.9],
+ [0.2, 0.94],
+ [0.3, 0.965],
+ [1.0, 0.965]
+ ]
+ },
+ "unverifiedParameters": {
+ "compute_tray_fans_w": {
+ "low": 40.0,
+ "high": 220.0,
+ "default": 130.0,
+ "unit": "W per compute tray",
+ "why": "8 dual-rotor 40x56 fans per tray are documented (Lenovo LP2357, SemiAnalysis) but no tray fan power is published. High end = 8 x 27.24 W rated Delta GFC0412DS-SM06B0G; low end = deep PWM on a liquid-cooled tray where only NICs/DPU/NVMe/PDB are air cooled."
+ },
+ "compute_tray_board_residual_w": {
+ "low": 20.0,
+ "high": 60.0,
+ "default": 40.0,
+ "unit": "W per compute tray",
+ "why": "BMC module, CPLD/ERoT, clocks, sensors, leak detection and PDB housekeeping. No public rail. Anchored to the repo chassis residual note (35-70 W for a full HGX chassis) scaled to a 1RU tray."
+ },
+ "tray_input_conversion_efficiency": {
+ "low": 0.96,
+ "high": 0.985,
+ "default": 0.9725,
+ "unit": "fraction",
+ "why": "Trays take 50 V DC from the busbar and feed the boards at 12 V (SemiAnalysis: 4x RapidLock 12 V connectors per Bianca board). The 50 V to 12 V stage sits outside the module sensor. No efficiency is published for it."
+ },
+ "nvswitch_tray_residual_w": {
+ "low": 20.0,
+ "high": 80.0,
+ "default": 50.0,
+ "unit": "W per NVLink switch tray",
+ "why": "Switch tray BMC, two OOB 1GbE ports, console, CPLDs, clocks and NVSwitch VR loss outside the ASIC figure. Lenovo front/rear views show coolant, busbar and cartridge connectors and no fans. No public rail."
+ }
+ }
+ },
+ "gb300": {
+ "modelPath": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
+ "functionName": "gb200_nvl72_rack_power",
+ "configFactory": "gb300_nvl72_rack_config",
+ "topology": "nvl72-rack",
+ "gpuCount": 72,
+ "computeTrayCount": 18,
+ "gpusPerComputeTray": 4,
+ "graceSocketsPerComputeTray": 2,
+ "nvswitchTrayCount": 9,
+ "assumptions": {
+ "u_nvlink": 0.5,
+ "u_ib": 0.0,
+ "u_pcie": 0.05,
+ "pue": 1.2
+ },
+ "defaultConfig": {
+ "variant": "gb300",
+ "gpu_label": "GB300 Blackwell Ultra (NVL72)",
+ "n_compute_trays": 18,
+ "n_nvswitch_trays": 9,
+ "compute_tray": {
+ "n_gpu": 4,
+ "n_grace": 2,
+ "nic_generation": "connectx8",
+ "connectx7": {
+ "n_nic": 4,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "connectx8": {
+ "n_nic": 4,
+ "include_optic": true,
+ "network_serdes_static_w": 18.0,
+ "pcie_switch_static_w": 28.0,
+ "board_mgmt_static_w": 5.0,
+ "digital_static_w": 12.0,
+ "network_dynamic_max_w": 8.0,
+ "pcie_switch_dynamic_max_w": 5.0,
+ "digital_dynamic_max_w": 10.0,
+ "optic_idle_w": 15.0,
+ "optic_dynamic_max_w": 2.0,
+ "max_nic_slot_power_ref_w": 75.0,
+ "normal_low_w_per_nic_with_optic": 70.0,
+ "normal_high_w_per_nic_with_optic": 100.0
+ },
+ "dpu": {
+ "n_dpu": 2,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "nvme": {
+ "n_front_u2": 4,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 1,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "fans_w": 130.0,
+ "board_residual_w": 40.0
+ },
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "nvswitch_tray_residual_w": 50.0,
+ "tray_input_conversion_efficiency": 0.9725,
+ "n_management_switches": 2,
+ "management_switch_w": 100.0,
+ "power_shelf": {
+ "n_shelves": 8,
+ "n_psu_per_shelf": 6,
+ "psu_capacity_w": 5500.0,
+ "redundancy": "N+N",
+ "busbar_nominal_v": 50.0,
+ "efficiency_curve": {
+ "0.1": 0.9,
+ "0.2": 0.94,
+ "0.3": 0.965,
+ "1.0": 0.965
+ },
+ "peak_efficiency_ref": 0.975
+ },
+ "regulator_loss_frac_of_tdp": 0.15,
+ "regulator_allowance_includes_grace": false,
+ "pue": 1.2,
+ "cooling": "direct_liquid",
+ "gpu_tdp_example_w_per_gpu": 1400.0,
+ "grace_tdp_example_w_per_socket": 300.0,
+ "module_tdp_example_w_per_superchip": 0.0
+ },
+ "computeTrayStaticDcWatts": {
+ "connectx8_nics_integrated_pcie_with_optics": 315.0,
+ "bluefield3_dpu_idle": 130.0,
+ "nvme_idle": 22.0,
+ "compute_tray_fans_unverified": 130.0,
+ "compute_tray_board_residual_unverified": 40.0
+ },
+ "computeTrayStaticDetails": {
+ "nic_generation": "connectx8",
+ "per_nic_w": 78.8,
+ "n_nic": 4,
+ "n_dpu": 2,
+ "dpu_model": "idle_only",
+ "nvme_model": "idle_only"
+ },
+ "nvswitchTraySiliconWatts": 406.4,
+ "nvswitchTrayResidualWatts": 50.0,
+ "managementSwitchCount": 2,
+ "managementSwitchWatts": 100.0,
+ "trayInputConversionEfficiency": 0.9725,
+ "regulatorLossFracOfTdp": 0.15,
+ "regulatorAllowanceIncludesGrace": false,
+ "powerShelf": {
+ "installedCapacityWatts": 264000.0,
+ "redundantCapacityWatts": 132000.0,
+ "efficiencyCurve": [
+ [0.1, 0.9],
+ [0.2, 0.94],
+ [0.3, 0.965],
+ [1.0, 0.965]
+ ]
+ },
+ "unverifiedParameters": {
+ "compute_tray_fans_w": {
+ "low": 40.0,
+ "high": 220.0,
+ "default": 130.0,
+ "unit": "W per compute tray",
+ "why": "8 dual-rotor 40x56 fans per tray are documented (Lenovo LP2357, SemiAnalysis) but no tray fan power is published. High end = 8 x 27.24 W rated Delta GFC0412DS-SM06B0G; low end = deep PWM on a liquid-cooled tray where only NICs/DPU/NVMe/PDB are air cooled."
+ },
+ "compute_tray_board_residual_w": {
+ "low": 20.0,
+ "high": 60.0,
+ "default": 40.0,
+ "unit": "W per compute tray",
+ "why": "BMC module, CPLD/ERoT, clocks, sensors, leak detection and PDB housekeeping. No public rail. Anchored to the repo chassis residual note (35-70 W for a full HGX chassis) scaled to a 1RU tray."
+ },
+ "tray_input_conversion_efficiency": {
+ "low": 0.96,
+ "high": 0.985,
+ "default": 0.9725,
+ "unit": "fraction",
+ "why": "Trays take 50 V DC from the busbar and feed the boards at 12 V (SemiAnalysis: 4x RapidLock 12 V connectors per Bianca board). The 50 V to 12 V stage sits outside the module sensor. No efficiency is published for it."
+ },
+ "nvswitch_tray_residual_w": {
+ "low": 20.0,
+ "high": 80.0,
+ "default": 50.0,
+ "unit": "W per NVLink switch tray",
+ "why": "Switch tray BMC, two OOB 1GbE ports, console, CPLDs, clocks and NVSwitch VR loss outside the ASIC figure. Lenovo front/rear views show coolant, busbar and cartridge connectors and no fans. No public rail."
+ }
+ }
+ }
}
}
diff --git a/packages/app/src/lib/system-power-model.reference.json b/packages/app/src/lib/system-power-model.reference.json
index bf47f2777..5e0f49c77 100644
--- a/packages/app/src/lib/system-power-model.reference.json
+++ b/packages/app/src/lib/system-power-model.reference.json
@@ -1,5 +1,6 @@
{
- "modelRevision": "ca4403aa527069857351ad8047dbb726844b3382",
+ "modelRevision": "963ead8b20a722595c34f7f4a0041259501cf019",
+ "modelRevisionStatus": "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream",
"source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
"status": "DRAFT / pending human verification",
"cases": [
@@ -3909,5 +3910,4015 @@
"expected": null,
"referenceError": "dc_load_w=30559.8 exceeds modeled MI355X 10U chassis PSU bank capacity 26400.0 W"
}
+ ],
+ "rackCases": [
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.2,
+ "rackDcWatts": 12716.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1413.0,
+ "rackAcWatts": 14129.7,
+ "facilityWatts": 14129.7,
+ "perGpuAcWatts": 196.2,
+ "perGpuFacilityWatts": 196.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.2,
+ "rackDcWatts": 12716.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1413.0,
+ "rackAcWatts": 14129.7,
+ "facilityWatts": 15542.7,
+ "perGpuAcWatts": 196.2,
+ "perGpuFacilityWatts": 215.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.2,
+ "rackDcWatts": 12716.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1413.0,
+ "rackAcWatts": 14129.7,
+ "facilityWatts": 16955.6,
+ "perGpuAcWatts": 196.2,
+ "perGpuFacilityWatts": 235.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.8,
+ "rackDcWatts": 12738.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1415.4,
+ "rackAcWatts": 14154.4,
+ "facilityWatts": 14154.4,
+ "perGpuAcWatts": 196.6,
+ "perGpuFacilityWatts": 196.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.8,
+ "rackDcWatts": 12738.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1415.4,
+ "rackAcWatts": 14154.4,
+ "facilityWatts": 15569.8,
+ "perGpuAcWatts": 196.6,
+ "perGpuFacilityWatts": 216.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.8,
+ "rackDcWatts": 12738.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1415.4,
+ "rackAcWatts": 14154.4,
+ "facilityWatts": 16985.3,
+ "perGpuAcWatts": 196.6,
+ "perGpuFacilityWatts": 235.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.125,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 13304.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 29329.2,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 407.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.125,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 13304.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 32262.1,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 448.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.125,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 13304.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 35195.0,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 488.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.325,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 13307.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 29333.3,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 407.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.325,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 13307.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 32266.6,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 448.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.325,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 13307.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 35200.0,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 488.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.525,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 13311.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 29337.2,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 407.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.525,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 13311.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 32270.9,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 448.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 739.525,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 13311.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 35204.6,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 489.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 853.3,
+ "rackDcWatts": 31229.4,
+ "powerShelfEfficiency": 0.9073,
+ "powerShelfLossWatts": 3190.1,
+ "rackAcWatts": 34419.5,
+ "facilityWatts": 34419.5,
+ "perGpuAcWatts": 478.0,
+ "perGpuFacilityWatts": 478.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 853.3,
+ "rackDcWatts": 31229.4,
+ "powerShelfEfficiency": 0.9073,
+ "powerShelfLossWatts": 3190.1,
+ "rackAcWatts": 34419.5,
+ "facilityWatts": 37861.5,
+ "perGpuAcWatts": 478.0,
+ "perGpuFacilityWatts": 525.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 853.3,
+ "rackDcWatts": 31229.4,
+ "powerShelfEfficiency": 0.9073,
+ "powerShelfLossWatts": 3190.1,
+ "rackAcWatts": 34419.5,
+ "facilityWatts": 41303.4,
+ "perGpuAcWatts": 478.0,
+ "perGpuFacilityWatts": 573.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1362.2,
+ "rackDcWatts": 49733.8,
+ "powerShelfEfficiency": 0.9354,
+ "powerShelfLossWatts": 3437.3,
+ "rackAcWatts": 53171.1,
+ "facilityWatts": 53171.1,
+ "perGpuAcWatts": 738.5,
+ "perGpuFacilityWatts": 738.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1362.2,
+ "rackDcWatts": 49733.8,
+ "powerShelfEfficiency": 0.9354,
+ "powerShelfLossWatts": 3437.3,
+ "rackAcWatts": 53171.1,
+ "facilityWatts": 58488.2,
+ "perGpuAcWatts": 738.5,
+ "perGpuFacilityWatts": 812.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1362.2,
+ "rackDcWatts": 49733.8,
+ "powerShelfEfficiency": 0.9354,
+ "powerShelfLossWatts": 3437.3,
+ "rackAcWatts": 53171.1,
+ "facilityWatts": 63805.3,
+ "perGpuAcWatts": 738.5,
+ "perGpuFacilityWatts": 886.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.458,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 38978.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 56166.6,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 780.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.458,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 38978.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 61783.3,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 858.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.458,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 38978.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 67399.9,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 936.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.658,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 38981.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 56170.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 780.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.658,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 38981.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 61787.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 858.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.658,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 38981.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 67404.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 936.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.858,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 38985.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 56173.9,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 780.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.858,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 38985.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 61791.3,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 858.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 2165.858,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 38985.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 67408.7,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 936.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1871.6,
+ "rackDcWatts": 68256.7,
+ "powerShelfEfficiency": 0.9546,
+ "powerShelfLossWatts": 3243.5,
+ "rackAcWatts": 71500.1,
+ "facilityWatts": 71500.1,
+ "perGpuAcWatts": 993.1,
+ "perGpuFacilityWatts": 993.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1871.6,
+ "rackDcWatts": 68256.7,
+ "powerShelfEfficiency": 0.9546,
+ "powerShelfLossWatts": 3243.5,
+ "rackAcWatts": 71500.1,
+ "facilityWatts": 78650.1,
+ "perGpuAcWatts": 993.1,
+ "perGpuFacilityWatts": 1092.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1871.6,
+ "rackDcWatts": 68256.7,
+ "powerShelfEfficiency": 0.9546,
+ "powerShelfLossWatts": 3243.5,
+ "rackAcWatts": 71500.1,
+ "facilityWatts": 85800.1,
+ "perGpuAcWatts": 993.1,
+ "perGpuFacilityWatts": 1191.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.792,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 64652.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 82069.0,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1139.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.792,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 64652.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 90275.9,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1253.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.792,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 64652.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 98482.8,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1367.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.992,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 64655.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 82072.5,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1139.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.992,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 64655.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 90279.8,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1253.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3591.992,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 64655.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 98487.0,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1367.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3592.192,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 64659.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 82076.3,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1139.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3592.192,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 64659.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 90283.9,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1253.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 3592.192,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 64659.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 98491.6,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1367.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2380.2,
+ "rackDcWatts": 86751.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3146.4,
+ "rackAcWatts": 89898.2,
+ "facilityWatts": 89898.2,
+ "perGpuAcWatts": 1248.6,
+ "perGpuFacilityWatts": 1248.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2380.2,
+ "rackDcWatts": 86751.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3146.4,
+ "rackAcWatts": 89898.2,
+ "facilityWatts": 98888.0,
+ "perGpuAcWatts": 1248.6,
+ "perGpuFacilityWatts": 1373.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2380.2,
+ "rackDcWatts": 86751.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3146.4,
+ "rackAcWatts": 89898.2,
+ "facilityWatts": 107877.8,
+ "perGpuAcWatts": 1248.6,
+ "perGpuFacilityWatts": 1498.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3092.8,
+ "rackDcWatts": 112664.4,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4086.3,
+ "rackAcWatts": 116750.6,
+ "facilityWatts": 116750.6,
+ "perGpuAcWatts": 1621.5,
+ "perGpuFacilityWatts": 1621.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3092.8,
+ "rackDcWatts": 112664.4,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4086.3,
+ "rackAcWatts": 116750.6,
+ "facilityWatts": 128425.7,
+ "perGpuAcWatts": 1621.5,
+ "perGpuFacilityWatts": 1783.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3092.8,
+ "rackDcWatts": 112664.4,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4086.3,
+ "rackAcWatts": 116750.6,
+ "facilityWatts": 140100.7,
+ "perGpuAcWatts": 1621.5,
+ "perGpuFacilityWatts": 1945.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3398.2,
+ "rackDcWatts": 123769.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4489.1,
+ "rackAcWatts": 128258.8,
+ "facilityWatts": 128258.8,
+ "perGpuAcWatts": 1781.4,
+ "perGpuFacilityWatts": 1781.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3398.2,
+ "rackDcWatts": 123769.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4489.1,
+ "rackAcWatts": 128258.8,
+ "facilityWatts": 141084.7,
+ "perGpuAcWatts": 1781.4,
+ "perGpuFacilityWatts": 1959.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3398.2,
+ "rackDcWatts": 123769.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4489.1,
+ "rackAcWatts": 128258.8,
+ "facilityWatts": 153910.6,
+ "perGpuAcWatts": 1781.4,
+ "perGpuFacilityWatts": 2137.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4009.0,
+ "rackDcWatts": 145980.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5294.6,
+ "rackAcWatts": 151275.2,
+ "facilityWatts": 151275.2,
+ "perGpuAcWatts": 2101.0,
+ "perGpuFacilityWatts": 2101.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4009.0,
+ "rackDcWatts": 145980.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5294.6,
+ "rackAcWatts": 151275.2,
+ "facilityWatts": 166402.7,
+ "perGpuAcWatts": 2101.0,
+ "perGpuFacilityWatts": 2311.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4009.0,
+ "rackDcWatts": 145980.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5294.6,
+ "rackAcWatts": 151275.2,
+ "facilityWatts": 181530.2,
+ "perGpuAcWatts": 2101.0,
+ "perGpuFacilityWatts": 2521.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.125,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 244370.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 273571.2,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 3799.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.125,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 244370.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 300928.3,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 4179.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.125,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 244370.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 328285.4,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 4559.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.325,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 244373.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 273575.1,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 3799.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.325,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 244373.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 300932.6,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 4179.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.325,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 244373.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 328290.1,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 4559.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.525,
+ "pue": 1.0,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.525,
+ "pue": 1.1,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 13576.525,
+ "pue": 1.2,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.0,
+ "expected": null,
+ "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.1,
+ "expected": null,
+ "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.2,
+ "expected": null,
+ "referenceError": "dc_load_w=284181.1 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 2.0,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.2,
+ "rackDcWatts": 12717.8,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1413.1,
+ "rackAcWatts": 14130.9,
+ "facilityWatts": 14130.9,
+ "perGpuAcWatts": 196.3,
+ "perGpuFacilityWatts": 196.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 2.0,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 344.2,
+ "rackDcWatts": 12717.8,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1413.1,
+ "rackAcWatts": 14130.9,
+ "facilityWatts": 15544.0,
+ "perGpuAcWatts": 196.3,
+ "perGpuFacilityWatts": 215.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 5401.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 496.9,
+ "rackDcWatts": 18269.6,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2030.0,
+ "rackAcWatts": 20299.5,
+ "facilityWatts": 20299.5,
+ "perGpuAcWatts": 281.9,
+ "perGpuFacilityWatts": 281.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 5401.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 496.9,
+ "rackDcWatts": 18269.6,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2030.0,
+ "rackAcWatts": 20299.5,
+ "facilityWatts": 22329.5,
+ "perGpuAcWatts": 281.9,
+ "perGpuFacilityWatts": 310.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 10810.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 649.9,
+ "rackDcWatts": 23831.5,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2647.9,
+ "rackAcWatts": 26479.5,
+ "facilityWatts": 26479.5,
+ "perGpuAcWatts": 367.8,
+ "perGpuFacilityWatts": 367.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 10810.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 649.9,
+ "rackDcWatts": 23831.5,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2647.9,
+ "rackAcWatts": 26479.5,
+ "facilityWatts": 29127.5,
+ "perGpuAcWatts": 367.8,
+ "perGpuFacilityWatts": 404.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 27.4,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 345.0,
+ "rackDcWatts": 12743.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1416.0,
+ "rackAcWatts": 14159.9,
+ "facilityWatts": 14159.9,
+ "perGpuAcWatts": 196.7,
+ "perGpuFacilityWatts": 196.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 27.4,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 345.0,
+ "rackDcWatts": 12743.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1416.0,
+ "rackAcWatts": 14159.9,
+ "facilityWatts": 15575.9,
+ "perGpuAcWatts": 196.7,
+ "perGpuFacilityWatts": 216.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 5426.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 497.6,
+ "rackDcWatts": 18295.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2032.9,
+ "rackAcWatts": 20328.6,
+ "facilityWatts": 20328.6,
+ "perGpuAcWatts": 282.3,
+ "perGpuFacilityWatts": 282.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 5426.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 497.6,
+ "rackDcWatts": 18295.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2032.9,
+ "rackAcWatts": 20328.6,
+ "facilityWatts": 22361.5,
+ "perGpuAcWatts": 282.3,
+ "perGpuFacilityWatts": 310.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 10835.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 650.6,
+ "rackDcWatts": 23857.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2650.9,
+ "rackAcWatts": 26508.5,
+ "facilityWatts": 26508.5,
+ "perGpuAcWatts": 368.2,
+ "perGpuFacilityWatts": 368.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 10835.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 650.6,
+ "rackDcWatts": 23857.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2650.9,
+ "rackAcWatts": 26508.5,
+ "facilityWatts": 29159.4,
+ "perGpuAcWatts": 368.2,
+ "perGpuFacilityWatts": 405.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 42353.8,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1541.9,
+ "rackDcWatts": 56267.3,
+ "powerShelfEfficiency": 0.9433,
+ "powerShelfLossWatts": 3383.2,
+ "rackAcWatts": 59650.5,
+ "facilityWatts": 59650.5,
+ "perGpuAcWatts": 828.5,
+ "perGpuFacilityWatts": 828.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 42353.8,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1541.9,
+ "rackDcWatts": 56267.3,
+ "powerShelfEfficiency": 0.9433,
+ "powerShelfLossWatts": 3383.2,
+ "rackAcWatts": 59650.5,
+ "facilityWatts": 65615.6,
+ "perGpuAcWatts": 828.5,
+ "perGpuFacilityWatts": 911.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 47752.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1694.5,
+ "rackDcWatts": 61819.1,
+ "powerShelfEfficiency": 0.9485,
+ "powerShelfLossWatts": 3353.7,
+ "rackAcWatts": 65172.8,
+ "facilityWatts": 65172.8,
+ "perGpuAcWatts": 905.2,
+ "perGpuFacilityWatts": 905.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 47752.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1694.5,
+ "rackDcWatts": 61819.1,
+ "powerShelfEfficiency": 0.9485,
+ "powerShelfLossWatts": 3353.7,
+ "rackAcWatts": 65172.8,
+ "facilityWatts": 71690.1,
+ "perGpuAcWatts": 905.2,
+ "perGpuFacilityWatts": 995.7
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 53161.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1847.5,
+ "rackDcWatts": 67381.0,
+ "powerShelfEfficiency": 0.9538,
+ "powerShelfLossWatts": 3263.2,
+ "rackAcWatts": 70644.2,
+ "facilityWatts": 70644.2,
+ "perGpuAcWatts": 981.2,
+ "perGpuFacilityWatts": 981.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 53161.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1847.5,
+ "rackDcWatts": 67381.0,
+ "powerShelfEfficiency": 0.9538,
+ "powerShelfLossWatts": 3263.2,
+ "rackAcWatts": 70644.2,
+ "facilityWatts": 77708.6,
+ "perGpuAcWatts": 981.2,
+ "perGpuFacilityWatts": 1079.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 63535.6,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2140.8,
+ "rackDcWatts": 78048.0,
+ "powerShelfEfficiency": 0.9639,
+ "powerShelfLossWatts": 2922.3,
+ "rackAcWatts": 80970.3,
+ "facilityWatts": 80970.3,
+ "perGpuAcWatts": 1124.6,
+ "perGpuFacilityWatts": 1124.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 63535.6,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2140.8,
+ "rackDcWatts": 78048.0,
+ "powerShelfEfficiency": 0.9639,
+ "powerShelfLossWatts": 2922.3,
+ "rackAcWatts": 80970.3,
+ "facilityWatts": 89067.3,
+ "perGpuAcWatts": 1124.6,
+ "perGpuFacilityWatts": 1237.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 68934.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2293.5,
+ "rackDcWatts": 83599.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3032.1,
+ "rackAcWatts": 86631.9,
+ "facilityWatts": 86631.9,
+ "perGpuAcWatts": 1203.2,
+ "perGpuFacilityWatts": 1203.2
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 68934.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2293.5,
+ "rackDcWatts": 83599.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3032.1,
+ "rackAcWatts": 86631.9,
+ "facilityWatts": 95295.1,
+ "perGpuAcWatts": 1203.2,
+ "perGpuFacilityWatts": 1323.5
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 74343.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2446.4,
+ "rackDcWatts": 89161.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3233.8,
+ "rackAcWatts": 92395.6,
+ "facilityWatts": 92395.6,
+ "perGpuAcWatts": 1283.3,
+ "perGpuFacilityWatts": 1283.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 74343.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2446.4,
+ "rackDcWatts": 89161.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3233.8,
+ "rackAcWatts": 92395.6,
+ "facilityWatts": 101635.2,
+ "perGpuAcWatts": 1283.3,
+ "perGpuFacilityWatts": 1411.6
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 101648.0,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3218.5,
+ "rackDcWatts": 117238.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4252.2,
+ "rackAcWatts": 121490.3,
+ "facilityWatts": 121490.3,
+ "perGpuAcWatts": 1687.4,
+ "perGpuFacilityWatts": 1687.4
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 101648.0,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3218.5,
+ "rackDcWatts": 117238.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4252.2,
+ "rackAcWatts": 121490.3,
+ "facilityWatts": 133639.3,
+ "perGpuAcWatts": 1687.4,
+ "perGpuFacilityWatts": 1856.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 107047.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3371.2,
+ "rackDcWatts": 122789.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4453.5,
+ "rackAcWatts": 127243.4,
+ "facilityWatts": 127243.4,
+ "perGpuAcWatts": 1767.3,
+ "perGpuFacilityWatts": 1767.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 107047.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3371.2,
+ "rackDcWatts": 122789.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4453.5,
+ "rackAcWatts": 127243.4,
+ "facilityWatts": 139967.7,
+ "perGpuAcWatts": 1767.3,
+ "perGpuFacilityWatts": 1944.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 112456.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3524.2,
+ "rackDcWatts": 128351.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4655.2,
+ "rackAcWatts": 133007.1,
+ "facilityWatts": 133007.1,
+ "perGpuAcWatts": 1847.3,
+ "perGpuFacilityWatts": 1847.3
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 112456.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3524.2,
+ "rackDcWatts": 128351.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4655.2,
+ "rackAcWatts": 133007.1,
+ "facilityWatts": 146307.8,
+ "perGpuAcWatts": 1847.3,
+ "perGpuFacilityWatts": 2032.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 118589.1,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3697.6,
+ "rackDcWatts": 134658.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4884.0,
+ "rackAcWatts": 139542.3,
+ "facilityWatts": 139542.3,
+ "perGpuAcWatts": 1938.1,
+ "perGpuFacilityWatts": 1938.1
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 118589.1,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3697.6,
+ "rackDcWatts": 134658.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4884.0,
+ "rackAcWatts": 139542.3,
+ "facilityWatts": 153496.5,
+ "perGpuAcWatts": 1938.1,
+ "perGpuFacilityWatts": 2131.9
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 123988.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3850.3,
+ "rackDcWatts": 140210.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5085.3,
+ "rackAcWatts": 145295.5,
+ "facilityWatts": 145295.5,
+ "perGpuAcWatts": 2018.0,
+ "perGpuFacilityWatts": 2018.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 123988.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3850.3,
+ "rackDcWatts": 140210.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5085.3,
+ "rackAcWatts": 145295.5,
+ "facilityWatts": 159825.1,
+ "perGpuAcWatts": 2018.0,
+ "perGpuFacilityWatts": 2219.8
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 129397.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4003.2,
+ "rackDcWatts": 145772.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5287.1,
+ "rackAcWatts": 151059.1,
+ "facilityWatts": 151059.1,
+ "perGpuAcWatts": 2098.0,
+ "perGpuFacilityWatts": 2098.0
+ }
+ },
+ {
+ "hardware": "gb200",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 129397.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 8064.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4003.2,
+ "rackDcWatts": 145772.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5287.1,
+ "rackAcWatts": 151059.1,
+ "facilityWatts": 166165.0,
+ "perGpuAcWatts": 2098.0,
+ "perGpuFacilityWatts": 2307.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 440.4,
+ "rackDcWatts": 16214.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1801.7,
+ "rackAcWatts": 18016.6,
+ "facilityWatts": 18016.6,
+ "perGpuAcWatts": 250.2,
+ "perGpuFacilityWatts": 250.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 440.4,
+ "rackDcWatts": 16214.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1801.7,
+ "rackAcWatts": 18016.6,
+ "facilityWatts": 19818.3,
+ "perGpuAcWatts": 250.2,
+ "perGpuFacilityWatts": 275.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 0.05,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 0.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 440.4,
+ "rackDcWatts": 16214.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1801.7,
+ "rackAcWatts": 18016.6,
+ "facilityWatts": 21619.9,
+ "perGpuAcWatts": 250.2,
+ "perGpuFacilityWatts": 300.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 441.0,
+ "rackDcWatts": 16237.1,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1804.1,
+ "rackAcWatts": 18041.2,
+ "facilityWatts": 18041.2,
+ "perGpuAcWatts": 250.6,
+ "perGpuFacilityWatts": 250.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 441.0,
+ "rackDcWatts": 16237.1,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1804.1,
+ "rackAcWatts": 18041.2,
+ "facilityWatts": 19845.3,
+ "perGpuAcWatts": 250.6,
+ "perGpuFacilityWatts": 275.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1.25,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 22.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 441.0,
+ "rackDcWatts": 16237.1,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1804.1,
+ "rackAcWatts": 18041.2,
+ "facilityWatts": 21649.4,
+ "perGpuAcWatts": 250.6,
+ "perGpuFacilityWatts": 300.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.125,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 9902.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 29329.2,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 407.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.125,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 9902.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 32262.1,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 448.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.125,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 9902.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.4,
+ "rackDcWatts": 26396.2,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2932.9,
+ "rackAcWatts": 29329.2,
+ "facilityWatts": 35195.0,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 488.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.325,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 9905.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 29333.3,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 407.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.325,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 9905.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 32266.6,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 448.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.325,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 9905.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.5,
+ "rackDcWatts": 26399.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.3,
+ "rackAcWatts": 29333.3,
+ "facilityWatts": 35200.0,
+ "perGpuAcWatts": 407.4,
+ "perGpuFacilityWatts": 488.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.525,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 9909.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 29337.2,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 407.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.525,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 9909.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 32270.9,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 448.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 550.525,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 9909.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 720.6,
+ "rackDcWatts": 26403.7,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2933.6,
+ "rackAcWatts": 29337.2,
+ "facilityWatts": 35204.6,
+ "perGpuAcWatts": 407.5,
+ "perGpuFacilityWatts": 489.0
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 949.5,
+ "rackDcWatts": 34727.6,
+ "powerShelfEfficiency": 0.9126,
+ "powerShelfLossWatts": 3325.1,
+ "rackAcWatts": 38052.8,
+ "facilityWatts": 38052.8,
+ "perGpuAcWatts": 528.5,
+ "perGpuFacilityWatts": 528.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 949.5,
+ "rackDcWatts": 34727.6,
+ "powerShelfEfficiency": 0.9126,
+ "powerShelfLossWatts": 3325.1,
+ "rackAcWatts": 38052.8,
+ "facilityWatts": 41858.1,
+ "perGpuAcWatts": 528.5,
+ "perGpuFacilityWatts": 581.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1000.25,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 18004.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 949.5,
+ "rackDcWatts": 34727.6,
+ "powerShelfEfficiency": 0.9126,
+ "powerShelfLossWatts": 3325.1,
+ "rackAcWatts": 38052.8,
+ "facilityWatts": 45663.4,
+ "perGpuAcWatts": 528.5,
+ "perGpuFacilityWatts": 634.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.458,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 35576.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 56166.6,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 780.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.458,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 35576.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 61783.3,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 858.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.458,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 35576.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.4,
+ "rackDcWatts": 52796.2,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.3,
+ "rackAcWatts": 56166.6,
+ "facilityWatts": 67399.9,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 936.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.658,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 35579.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 56170.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 780.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.658,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 35579.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 61787.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 858.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.658,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 35579.8,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.5,
+ "rackDcWatts": 52799.9,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56170.2,
+ "facilityWatts": 67404.2,
+ "perGpuAcWatts": 780.1,
+ "perGpuFacilityWatts": 936.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.858,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 35583.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 56173.9,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 780.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.858,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 35583.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 61791.3,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 858.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 1976.858,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 35583.4,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1446.6,
+ "rackDcWatts": 52803.6,
+ "powerShelfEfficiency": 0.94,
+ "powerShelfLossWatts": 3370.2,
+ "rackAcWatts": 56173.9,
+ "facilityWatts": 67408.7,
+ "perGpuAcWatts": 780.2,
+ "perGpuFacilityWatts": 936.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1458.4,
+ "rackDcWatts": 53232.0,
+ "powerShelfEfficiency": 0.9404,
+ "powerShelfLossWatts": 3373.2,
+ "rackAcWatts": 56605.1,
+ "facilityWatts": 56605.1,
+ "perGpuAcWatts": 786.2,
+ "perGpuFacilityWatts": 786.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1458.4,
+ "rackDcWatts": 53232.0,
+ "powerShelfEfficiency": 0.9404,
+ "powerShelfLossWatts": 3373.2,
+ "rackAcWatts": 56605.1,
+ "facilityWatts": 62265.6,
+ "perGpuAcWatts": 786.2,
+ "perGpuFacilityWatts": 864.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 2000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 36000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1458.4,
+ "rackDcWatts": 53232.0,
+ "powerShelfEfficiency": 0.9404,
+ "powerShelfLossWatts": 3373.2,
+ "rackAcWatts": 56605.1,
+ "facilityWatts": 67926.1,
+ "perGpuAcWatts": 786.2,
+ "perGpuFacilityWatts": 943.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1967.8,
+ "rackDcWatts": 71754.9,
+ "powerShelfEfficiency": 0.9579,
+ "powerShelfLossWatts": 3149.8,
+ "rackAcWatts": 74904.6,
+ "facilityWatts": 74904.6,
+ "perGpuAcWatts": 1040.3,
+ "perGpuFacilityWatts": 1040.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1967.8,
+ "rackDcWatts": 71754.9,
+ "powerShelfEfficiency": 0.9579,
+ "powerShelfLossWatts": 3149.8,
+ "rackAcWatts": 74904.6,
+ "facilityWatts": 82395.1,
+ "perGpuAcWatts": 1040.3,
+ "perGpuFacilityWatts": 1144.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3000.75,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 54013.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1967.8,
+ "rackDcWatts": 71754.9,
+ "powerShelfEfficiency": 0.9579,
+ "powerShelfLossWatts": 3149.8,
+ "rackAcWatts": 74904.6,
+ "facilityWatts": 89885.5,
+ "perGpuAcWatts": 1040.3,
+ "perGpuFacilityWatts": 1248.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.792,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 61250.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 82069.0,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1139.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.792,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 61250.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 90275.9,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1253.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.792,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 61250.3,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.4,
+ "rackDcWatts": 79196.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82069.0,
+ "facilityWatts": 98482.8,
+ "perGpuAcWatts": 1139.8,
+ "perGpuFacilityWatts": 1367.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.992,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 61253.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 82072.5,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1139.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.992,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 61253.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 90279.8,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1253.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3402.992,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 61253.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.5,
+ "rackDcWatts": 79200.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.5,
+ "rackAcWatts": 82072.5,
+ "facilityWatts": 98487.0,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1367.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3403.192,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 61257.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 82076.3,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1139.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3403.192,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 61257.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 90283.9,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1253.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 3403.192,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 61257.5,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2172.6,
+ "rackDcWatts": 79203.7,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2872.7,
+ "rackAcWatts": 82076.3,
+ "facilityWatts": 98491.6,
+ "perGpuAcWatts": 1139.9,
+ "perGpuFacilityWatts": 1367.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2476.4,
+ "rackDcWatts": 90250.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3273.3,
+ "rackAcWatts": 93523.3,
+ "facilityWatts": 93523.3,
+ "perGpuAcWatts": 1298.9,
+ "perGpuFacilityWatts": 1298.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2476.4,
+ "rackDcWatts": 90250.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3273.3,
+ "rackAcWatts": 93523.3,
+ "facilityWatts": 102875.6,
+ "perGpuAcWatts": 1298.9,
+ "perGpuFacilityWatts": 1428.8
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 4000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 72000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2476.4,
+ "rackDcWatts": 90250.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3273.3,
+ "rackAcWatts": 93523.3,
+ "facilityWatts": 112228.0,
+ "perGpuAcWatts": 1298.9,
+ "perGpuFacilityWatts": 1558.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3189.0,
+ "rackDcWatts": 116162.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4213.2,
+ "rackAcWatts": 120375.7,
+ "facilityWatts": 120375.7,
+ "perGpuAcWatts": 1671.9,
+ "perGpuFacilityWatts": 1671.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3189.0,
+ "rackDcWatts": 116162.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4213.2,
+ "rackAcWatts": 120375.7,
+ "facilityWatts": 132413.3,
+ "perGpuAcWatts": 1671.9,
+ "perGpuFacilityWatts": 1839.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 5400.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 97200.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3189.0,
+ "rackDcWatts": 116162.6,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4213.2,
+ "rackAcWatts": 120375.7,
+ "facilityWatts": 144450.8,
+ "perGpuAcWatts": 1671.9,
+ "perGpuFacilityWatts": 2006.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3494.4,
+ "rackDcWatts": 127268.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4615.9,
+ "rackAcWatts": 131883.9,
+ "facilityWatts": 131883.9,
+ "perGpuAcWatts": 1831.7,
+ "perGpuFacilityWatts": 1831.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3494.4,
+ "rackDcWatts": 127268.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4615.9,
+ "rackAcWatts": 131883.9,
+ "facilityWatts": 145072.3,
+ "perGpuAcWatts": 1831.7,
+ "perGpuFacilityWatts": 2014.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 6000.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 108000.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3494.4,
+ "rackDcWatts": 127268.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4615.9,
+ "rackAcWatts": 131883.9,
+ "facilityWatts": 158260.7,
+ "perGpuAcWatts": 1831.7,
+ "perGpuFacilityWatts": 2198.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4105.2,
+ "rackDcWatts": 149478.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5421.5,
+ "rackAcWatts": 154900.3,
+ "facilityWatts": 154900.3,
+ "perGpuAcWatts": 2151.4,
+ "perGpuFacilityWatts": 2151.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4105.2,
+ "rackDcWatts": 149478.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5421.5,
+ "rackAcWatts": 154900.3,
+ "facilityWatts": 170390.3,
+ "perGpuAcWatts": 2151.4,
+ "perGpuFacilityWatts": 2366.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 7200.0,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 129600.0,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4105.2,
+ "rackDcWatts": 149478.8,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5421.5,
+ "rackAcWatts": 154900.3,
+ "facilityWatts": 185880.4,
+ "perGpuAcWatts": 2151.4,
+ "perGpuFacilityWatts": 2581.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.125,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 240968.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 273571.2,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 3799.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.125,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 240968.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 300928.3,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 4179.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.125,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 240968.2,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.4,
+ "rackDcWatts": 263996.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.0,
+ "rackAcWatts": 273571.2,
+ "facilityWatts": 328285.4,
+ "perGpuAcWatts": 3799.6,
+ "perGpuFacilityWatts": 4559.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.325,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 240971.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 273575.1,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 3799.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.325,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 240971.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 300932.6,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 4179.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.325,
+ "pue": 1.2,
+ "expected": {
+ "computeModulesDcWatts": 240971.9,
+ "regulatorAllowanceWatts": 0.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 7254.5,
+ "rackDcWatts": 263999.9,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 9575.1,
+ "rackAcWatts": 273575.1,
+ "facilityWatts": 328290.1,
+ "perGpuAcWatts": 3799.7,
+ "perGpuFacilityWatts": 4559.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.525,
+ "pue": 1.0,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.525,
+ "pue": 1.1,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 13387.525,
+ "pue": 1.2,
+ "expected": null,
+ "referenceError": "dc_load_w=264003.7 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.0,
+ "expected": null,
+ "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.1,
+ "expected": null,
+ "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "module",
+ "moduleWattsPerTray": 14666.666666666666,
+ "pue": 1.2,
+ "expected": null,
+ "referenceError": "dc_load_w=287679.3 exceeds installed power-shelf capacity 264000.0 W"
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 2.0,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 440.4,
+ "rackDcWatts": 16216.0,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1801.8,
+ "rackAcWatts": 18017.8,
+ "facilityWatts": 18017.8,
+ "perGpuAcWatts": 250.2,
+ "perGpuFacilityWatts": 250.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 2.0,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 440.4,
+ "rackDcWatts": 16216.0,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1801.8,
+ "rackAcWatts": 18017.8,
+ "facilityWatts": 19819.6,
+ "perGpuAcWatts": 250.2,
+ "perGpuFacilityWatts": 275.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 5401.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 593.1,
+ "rackDcWatts": 21767.8,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2418.6,
+ "rackAcWatts": 24186.4,
+ "facilityWatts": 24186.4,
+ "perGpuAcWatts": 335.9,
+ "perGpuFacilityWatts": 335.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 5401.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 593.1,
+ "rackDcWatts": 21767.8,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2418.6,
+ "rackAcWatts": 24186.4,
+ "facilityWatts": 26605.0,
+ "perGpuAcWatts": 335.9,
+ "perGpuFacilityWatts": 369.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 10810.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 746.1,
+ "rackDcWatts": 27329.7,
+ "powerShelfEfficiency": 0.9014,
+ "powerShelfLossWatts": 2989.2,
+ "rackAcWatts": 30318.9,
+ "facilityWatts": 30318.9,
+ "perGpuAcWatts": 421.1,
+ "perGpuFacilityWatts": 421.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 0.05,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 10810.1,
+ "regulatorAllowanceWatts": 0.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 746.1,
+ "rackDcWatts": 27329.7,
+ "powerShelfEfficiency": 0.9014,
+ "powerShelfLossWatts": 2989.2,
+ "rackAcWatts": 30318.9,
+ "facilityWatts": 33350.8,
+ "perGpuAcWatts": 421.1,
+ "perGpuFacilityWatts": 463.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 27.4,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 441.2,
+ "rackDcWatts": 16242.1,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1804.7,
+ "rackAcWatts": 18046.8,
+ "facilityWatts": 18046.8,
+ "perGpuAcWatts": 250.6,
+ "perGpuFacilityWatts": 250.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 27.4,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 441.2,
+ "rackDcWatts": 16242.1,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 1804.7,
+ "rackAcWatts": 18046.8,
+ "facilityWatts": 19851.5,
+ "perGpuAcWatts": 250.6,
+ "perGpuFacilityWatts": 275.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 5426.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 593.8,
+ "rackDcWatts": 21793.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2421.5,
+ "rackAcWatts": 24215.4,
+ "facilityWatts": 24215.4,
+ "perGpuAcWatts": 336.3,
+ "perGpuFacilityWatts": 336.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 5426.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 593.8,
+ "rackDcWatts": 21793.9,
+ "powerShelfEfficiency": 0.9,
+ "powerShelfLossWatts": 2421.5,
+ "rackAcWatts": 24215.4,
+ "facilityWatts": 26636.9,
+ "perGpuAcWatts": 336.3,
+ "perGpuFacilityWatts": 370.0
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 10835.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 746.8,
+ "rackDcWatts": 27355.9,
+ "powerShelfEfficiency": 0.9014,
+ "powerShelfLossWatts": 2990.7,
+ "rackAcWatts": 30346.6,
+ "facilityWatts": 30346.6,
+ "perGpuAcWatts": 421.5,
+ "perGpuFacilityWatts": 421.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 1.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 10835.5,
+ "regulatorAllowanceWatts": 4.0,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 746.8,
+ "rackDcWatts": 27355.9,
+ "powerShelfEfficiency": 0.9014,
+ "powerShelfLossWatts": 2990.7,
+ "rackAcWatts": 30346.6,
+ "facilityWatts": 33381.3,
+ "perGpuAcWatts": 421.5,
+ "perGpuFacilityWatts": 463.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 42353.8,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1638.1,
+ "rackDcWatts": 59765.5,
+ "powerShelfEfficiency": 0.9466,
+ "powerShelfLossWatts": 3371.8,
+ "rackAcWatts": 63137.3,
+ "facilityWatts": 63137.3,
+ "perGpuAcWatts": 876.9,
+ "perGpuFacilityWatts": 876.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 42353.8,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1638.1,
+ "rackDcWatts": 59765.5,
+ "powerShelfEfficiency": 0.9466,
+ "powerShelfLossWatts": 3371.8,
+ "rackAcWatts": 63137.3,
+ "facilityWatts": 69451.0,
+ "perGpuAcWatts": 876.9,
+ "perGpuFacilityWatts": 964.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 47752.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1790.7,
+ "rackDcWatts": 65317.3,
+ "powerShelfEfficiency": 0.9519,
+ "powerShelfLossWatts": 3303.9,
+ "rackAcWatts": 68621.1,
+ "facilityWatts": 68621.1,
+ "perGpuAcWatts": 953.1,
+ "perGpuFacilityWatts": 953.1
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 47752.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1790.7,
+ "rackDcWatts": 65317.3,
+ "powerShelfEfficiency": 0.9519,
+ "powerShelfLossWatts": 3303.9,
+ "rackAcWatts": 68621.1,
+ "facilityWatts": 75483.2,
+ "perGpuAcWatts": 953.1,
+ "perGpuFacilityWatts": 1048.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 53161.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1943.7,
+ "rackDcWatts": 70879.2,
+ "powerShelfEfficiency": 0.9571,
+ "powerShelfLossWatts": 3175.4,
+ "rackAcWatts": 74054.6,
+ "facilityWatts": 74054.6,
+ "perGpuAcWatts": 1028.5,
+ "perGpuFacilityWatts": 1028.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 2000.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 53161.9,
+ "regulatorAllowanceWatts": 6352.9,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 1943.7,
+ "rackDcWatts": 70879.2,
+ "powerShelfEfficiency": 0.9571,
+ "powerShelfLossWatts": 3175.4,
+ "rackAcWatts": 74054.6,
+ "facilityWatts": 81460.1,
+ "perGpuAcWatts": 1028.5,
+ "perGpuFacilityWatts": 1131.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 63535.6,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2237.0,
+ "rackDcWatts": 81546.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2957.6,
+ "rackAcWatts": 84503.9,
+ "facilityWatts": 84503.9,
+ "perGpuAcWatts": 1173.7,
+ "perGpuFacilityWatts": 1173.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 63535.6,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2237.0,
+ "rackDcWatts": 81546.2,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 2957.6,
+ "rackAcWatts": 84503.9,
+ "facilityWatts": 92954.3,
+ "perGpuAcWatts": 1173.7,
+ "perGpuFacilityWatts": 1291.0
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 68934.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2389.7,
+ "rackDcWatts": 87098.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3159.0,
+ "rackAcWatts": 90257.0,
+ "facilityWatts": 90257.0,
+ "perGpuAcWatts": 1253.6,
+ "perGpuFacilityWatts": 1253.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 68934.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2389.7,
+ "rackDcWatts": 87098.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3159.0,
+ "rackAcWatts": 90257.0,
+ "facilityWatts": 99282.7,
+ "perGpuAcWatts": 1253.6,
+ "perGpuFacilityWatts": 1378.9
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 74343.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2542.6,
+ "rackDcWatts": 92660.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3360.7,
+ "rackAcWatts": 96020.7,
+ "facilityWatts": 96020.7,
+ "perGpuAcWatts": 1333.6,
+ "perGpuFacilityWatts": 1333.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 3000.25,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 74343.7,
+ "regulatorAllowanceWatts": 9530.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 2542.6,
+ "rackDcWatts": 92660.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 3360.7,
+ "rackAcWatts": 96020.7,
+ "facilityWatts": 105622.8,
+ "perGpuAcWatts": 1333.6,
+ "perGpuFacilityWatts": 1467.0
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 101648.0,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3314.7,
+ "rackDcWatts": 120736.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4379.0,
+ "rackAcWatts": 125115.3,
+ "facilityWatts": 125115.3,
+ "perGpuAcWatts": 1737.7,
+ "perGpuFacilityWatts": 1737.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 101648.0,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3314.7,
+ "rackDcWatts": 120736.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4379.0,
+ "rackAcWatts": 125115.3,
+ "facilityWatts": 137626.8,
+ "perGpuAcWatts": 1737.7,
+ "perGpuFacilityWatts": 1911.5
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 107047.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3467.4,
+ "rackDcWatts": 126288.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4580.4,
+ "rackAcWatts": 130868.5,
+ "facilityWatts": 130868.5,
+ "perGpuAcWatts": 1817.6,
+ "perGpuFacilityWatts": 1817.6
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 107047.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3467.4,
+ "rackDcWatts": 126288.1,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4580.4,
+ "rackAcWatts": 130868.5,
+ "facilityWatts": 143955.4,
+ "perGpuAcWatts": 1817.6,
+ "perGpuFacilityWatts": 1999.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 112456.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3620.4,
+ "rackDcWatts": 131850.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4782.1,
+ "rackAcWatts": 136632.2,
+ "facilityWatts": 136632.2,
+ "perGpuAcWatts": 1897.7,
+ "perGpuFacilityWatts": 1897.7
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 4800.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 112456.1,
+ "regulatorAllowanceWatts": 15247.1,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3620.4,
+ "rackDcWatts": 131850.0,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 4782.1,
+ "rackAcWatts": 136632.2,
+ "facilityWatts": 150295.4,
+ "perGpuAcWatts": 1897.7,
+ "perGpuFacilityWatts": 2087.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 118589.1,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3793.8,
+ "rackDcWatts": 138156.5,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5010.9,
+ "rackAcWatts": 143167.4,
+ "facilityWatts": 143167.4,
+ "perGpuAcWatts": 1988.4,
+ "perGpuFacilityWatts": 1988.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 0.05,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 118589.1,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3793.8,
+ "rackDcWatts": 138156.5,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5010.9,
+ "rackAcWatts": 143167.4,
+ "facilityWatts": 157484.1,
+ "perGpuAcWatts": 1988.4,
+ "perGpuFacilityWatts": 2187.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 123988.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3946.5,
+ "rackDcWatts": 143708.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5212.2,
+ "rackAcWatts": 148920.5,
+ "facilityWatts": 148920.5,
+ "perGpuAcWatts": 2068.3,
+ "perGpuFacilityWatts": 2068.3
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 300.0,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 123988.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 3946.5,
+ "rackDcWatts": 143708.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5212.2,
+ "rackAcWatts": 148920.5,
+ "facilityWatts": 163812.6,
+ "perGpuAcWatts": 2068.3,
+ "perGpuFacilityWatts": 2275.2
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.0,
+ "expected": {
+ "computeModulesDcWatts": 129397.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4099.4,
+ "rackDcWatts": 149270.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5413.9,
+ "rackAcWatts": 154684.2,
+ "facilityWatts": 154684.2,
+ "perGpuAcWatts": 2148.4,
+ "perGpuFacilityWatts": 2148.4
+ }
+ },
+ {
+ "hardware": "gb300",
+ "basis": "gpu-plus-grace",
+ "gpuBoardWattsPerTray": 5600.0,
+ "graceSocketWattsPerTray": 600.5,
+ "pue": 1.1,
+ "expected": {
+ "computeModulesDcWatts": 129397.2,
+ "regulatorAllowanceWatts": 17788.2,
+ "trayStaticDcWatts": 11466.0,
+ "nvswitchTraysDcWatts": 4107.6,
+ "trayConversionLossWatts": 4099.4,
+ "rackDcWatts": 149270.3,
+ "powerShelfEfficiency": 0.965,
+ "powerShelfLossWatts": 5413.9,
+ "rackAcWatts": 154684.2,
+ "facilityWatts": 170152.6,
+ "perGpuAcWatts": 2148.4,
+ "perGpuFacilityWatts": 2363.2
+ }
+ }
]
}
diff --git a/packages/app/src/lib/system-power-model.test.ts b/packages/app/src/lib/system-power-model.test.ts
index ef04e7d5c..d9d9ddbdc 100644
--- a/packages/app/src/lib/system-power-model.test.ts
+++ b/packages/app/src/lib/system-power-model.test.ts
@@ -1,12 +1,25 @@
+import { GPU_KEYS } from '@semianalysisai/inferencex-constants';
import { describe, expect, it } from 'vitest';
import {
estimateChassisPower,
+ estimateRackPower,
+ type RackMeasuredInput,
SUPPORTED_SYSTEM_POWER_HARDWARE,
+ SUPPORTED_SYSTEM_POWER_RACK_HARDWARE,
SYSTEM_POWER_MODEL_REVISION,
} from './system-power-model';
import reference from './system-power-model.reference.json';
+const rackInput = (row: (typeof reference.rackCases)[number]): RackMeasuredInput =>
+ row.moduleWattsPerTray === undefined
+ ? {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: row.gpuBoardWattsPerTray,
+ graceSocketWattsPerTray: row.graceSocketWattsPerTray,
+ }
+ : { basis: 'module', moduleWattsPerTray: row.moduleWattsPerTray };
+
describe('fixed 8k1k chassis model', () => {
it('matches the pinned Python implementation across platforms, fan/PSU boundaries, and PUE', () => {
expect(reference.modelRevision).toBe(SYSTEM_POWER_MODEL_REVISION);
@@ -47,4 +60,110 @@ describe('fixed 8k1k chassis model', () => {
// Fixed chassis overhead and nonlinear fan/PSU behavior forbid proportional GPU scaling.
expect(estimateChassisPower('h100', 8000)!.chassisAcWatts).not.toBe(chassis.chassisAcWatts * 2);
});
+
+ it('names every chassis and rack profile by its hardware registry key', () => {
+ for (const hardware of [
+ ...SUPPORTED_SYSTEM_POWER_HARDWARE,
+ ...SUPPORTED_SYSTEM_POWER_RACK_HARDWARE,
+ ]) {
+ expect(GPU_KEYS.has(hardware), hardware).toBe(true);
+ }
+ expect(SUPPORTED_SYSTEM_POWER_RACK_HARDWARE).toEqual(['gb200', 'gb300']);
+ });
+});
+
+describe('NVL72 rack model with measured compute-module input', () => {
+ it('matches the pinned Python rack model across variants, bases, shelf knots, and PUE', () => {
+ expect(new Set(reference.rackCases.map((row) => row.hardware))).toEqual(
+ new Set(SUPPORTED_SYSTEM_POWER_RACK_HARDWARE),
+ );
+ expect(new Set(reference.rackCases.map((row) => row.basis))).toEqual(
+ new Set(['module', 'gpu-plus-grace']),
+ );
+ for (const row of reference.rackCases) {
+ const actual = estimateRackPower(row.hardware, rackInput(row), row.pue);
+ const context = `${row.hardware}: ${JSON.stringify(rackInput(row))}, PUE=${row.pue}`;
+ if (row.expected === null) {
+ expect(actual, context).toBeNull();
+ } else {
+ expect(actual, context).toMatchObject(row.expected);
+ }
+ }
+ });
+
+ it('rejects unavailable measurements, PUE, chassis hardware, and shelf overflow', () => {
+ for (const watts of [0, -1, NaN, Infinity, -Infinity, Number.MAX_VALUE]) {
+ expect(estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: watts })).toBeNull();
+ expect(
+ estimateRackPower('gb200', {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: watts,
+ graceSocketWattsPerTray: 300,
+ }),
+ ).toBeNull();
+ // A zero Grace socket reading means the socket was not measured, never that it drew nothing.
+ expect(
+ estimateRackPower('gb200', {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: 3000,
+ graceSocketWattsPerTray: watts,
+ }),
+ ).toBeNull();
+ }
+ for (const pue of [0, 0.99, NaN, Infinity, -Infinity, Number.MAX_VALUE]) {
+ expect(
+ estimateRackPower('gb300', { basis: 'module', moduleWattsPerTray: 4000 }, pue),
+ ).toBeNull();
+ }
+ for (const hardware of [
+ 'b200',
+ 'b300',
+ 'h100',
+ 'gb200-nvl',
+ 'gb200_dynamo-trt',
+ '',
+ 'toString',
+ ]) {
+ expect(estimateRackPower(hardware, { basis: 'module', moduleWattsPerTray: 4000 })).toBeNull();
+ }
+ expect(estimateChassisPower('gb200', 4000)).toBeNull();
+ });
+
+ it('applies PUE once to the rounded rack AC and amortises the rack over all 72 GPUs', () => {
+ const input: RackMeasuredInput = { basis: 'module', moduleWattsPerTray: 5400 };
+ const rack = estimateRackPower('GB200', input, 1)!;
+ const facility = estimateRackPower('gb200', input, 1.1)!;
+ expect(rack.hardware).toBe('gb200');
+ expect(rack.facilityWatts).toBe(rack.rackAcWatts);
+ expect(facility.rackAcWatts).toBe(rack.rackAcWatts);
+ expect(facility.facilityWatts).toBeCloseTo(rack.rackAcWatts * 1.1, 0);
+ expect(facility.perGpuAcWatts).toBeCloseTo(rack.rackAcWatts / 72, 0);
+ expect(facility.perGpuFacilityWatts).toBeCloseTo(facility.facilityWatts / 72, 0);
+ expect(facility.gpuCount).toBe(72);
+ expect(facility.computeTrayCount).toBe(18);
+ // Static trays, switch trays, and the shelf curve forbid proportional scaling.
+ expect(
+ estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 2700 }, 1)!.rackAcWatts * 2,
+ ).not.toBe(rack.rackAcWatts);
+ });
+
+ it('adds the sourced regulator allowance only on the GPU-board share of the split basis', () => {
+ const split = estimateRackPower('gb300', {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: 4800,
+ graceSocketWattsPerTray: 600,
+ })!;
+ const module = estimateRackPower('gb300', {
+ basis: 'module',
+ moduleWattsPerTray: split.measuredWattsPerTray + split.regulatorAllowanceWattsPerTray,
+ })!;
+ expect(split.basis).toBe('gpu-plus-grace');
+ expect(split.measuredWattsPerTray).toBe(5400);
+ expect(split.regulatorAllowanceWattsPerTray).toBeGreaterThan(0);
+ expect(module.basis).toBe('module');
+ expect(module.regulatorAllowanceWattsPerTray).toBe(0);
+ // The electrical equivalent module reading reproduces the split-basis rack.
+ expect(module.rackAcWatts).toBeCloseTo(split.rackAcWatts, 0);
+ expect(module.facilityWatts).toBeCloseTo(split.facilityWatts, 0);
+ });
});
diff --git a/packages/app/src/lib/system-power-model.ts b/packages/app/src/lib/system-power-model.ts
index 68592f4cd..8cf6e1fd2 100644
--- a/packages/app/src/lib/system-power-model.ts
+++ b/packages/app/src/lib/system-power-model.ts
@@ -7,6 +7,51 @@ export type SystemPowerHardware = keyof typeof SYSTEM_POWER_PROFILES;
export const SUPPORTED_SYSTEM_POWER_HARDWARE = Object.keys(
SYSTEM_POWER_PROFILES,
) as SystemPowerHardware[];
+export const SYSTEM_POWER_RACK_ASSUMPTIONS = profileData.rackAssumptions;
+export const SYSTEM_POWER_RACK_PROFILES = profileData.rackProfiles;
+export type SystemPowerRackHardware = keyof typeof SYSTEM_POWER_RACK_PROFILES;
+export const SUPPORTED_SYSTEM_POWER_RACK_HARDWARE = Object.keys(
+ SYSTEM_POWER_RACK_PROFILES,
+) as SystemPowerRackHardware[];
+
+export type RackMeasuredBasis = 'module' | 'gpu-plus-grace';
+
+/**
+ * Measured compute-module input for every tray of one NVL72 rack. `module` is the
+ * sum of the two Module Power sensors per tray (Grace + 2 Blackwell + HBM + LPDDR5X +
+ * regulator loss). `gpu-plus-grace` is four GPU-board readings plus two Grace socket
+ * readings; the model then adds the sourced regulator-loss allowance on the GPU share.
+ */
+export type RackMeasuredInput =
+ | { basis: 'module'; moduleWattsPerTray: number }
+ | { basis: 'gpu-plus-grace'; gpuBoardWattsPerTray: number; graceSocketWattsPerTray: number };
+
+export interface RackPowerEstimate {
+ hardware: SystemPowerRackHardware;
+ basis: RackMeasuredBasis;
+ /** Measured watts handed to the model for every compute tray, before any allowance. */
+ measuredWattsPerTray: number;
+ /** Sourced regulator-loss allowance on the GPU-board share; zero on the module basis. */
+ regulatorAllowanceWattsPerTray: number;
+ computeModulesDcWatts: number;
+ regulatorAllowanceWatts: number;
+ trayStaticDcWatts: number;
+ nvswitchTraysDcWatts: number;
+ trayConversionLossWatts: number;
+ rackDcWatts: number;
+ powerShelfEfficiency: number;
+ powerShelfLossWatts: number;
+ rackAcWatts: number;
+ facilityWatts: number;
+ perGpuAcWatts: number;
+ perGpuFacilityWatts: number;
+ pue: number;
+ computeTrayCount: number;
+ /** GPUs in the modeled rack (72); the rack figures are amortised over all of them. */
+ gpuCount: number;
+ modelRevision: string;
+ modelPath: string;
+}
export interface ChassisPowerEstimate {
hardware: SystemPowerHardware;
@@ -36,6 +81,21 @@ function pythonRound(value: number, digits = 1): number {
return Number(value.toFixed(digits));
}
+/** Port of the source's `_interp_efficiency`: clamp outside the knots, linear between them. */
+function interpolateEfficiency(loadFraction: number, curve: number[][]): number {
+ const [first, last] = [curve[0], curve.at(-1)!];
+ if (loadFraction <= first[0]) return first[1];
+ if (loadFraction >= last[0]) return last[1];
+ for (let i = 1; i < curve.length; i++) {
+ const [x0, y0] = curve[i - 1];
+ const [x1, y1] = curve[i];
+ if (x0 <= loadFraction && loadFraction <= x1) {
+ return y0 + ((y1 - y0) * (loadFraction - x0)) / (x1 - x0);
+ }
+ }
+ return last[1];
+}
+
/**
* Fixed README inference sweep at the pinned revision, for one complete 8-GPU chassis.
* Profiles are generated by executing the original component models. Only their
@@ -77,16 +137,7 @@ export function estimateChassisPower(
const psu = profile.psu;
if (dc > psu.maxDcWatts) return null;
- const fraction = dc / psu.loadSharingCapacityWatts;
- const curve = psu.efficiencyCurve;
- let efficiency = curve[0][1];
- for (let i = 1; i < curve.length; i++) {
- const [x0, y0] = curve[i - 1];
- const [x1, y1] = curve[i];
- if (fraction <= x0) break;
- efficiency = fraction < x1 ? y0 + ((y1 - y0) * (fraction - x0)) / (x1 - x0) : y1;
- if (fraction <= x1) break;
- }
+ const efficiency = interpolateEfficiency(dc / psu.loadSharingCapacityWatts, psu.efficiencyCurve);
const ac = pythonRound(dc / efficiency);
// The Python chassis wrappers apply PUE to the already-rounded PSU AC output.
const facility = pythonRound(ac + ac * (pue - 1));
@@ -106,3 +157,94 @@ export function estimateChassisPower(
modelPath: profile.modelPath,
};
}
+
+/**
+ * One NVL72 rack whose 18 compute trays all carry the given measured compute-module
+ * input. Only the power-shelf efficiency curve is load dependent; switch trays, tray
+ * static electronics, management switches, and the tray input-conversion stage stay
+ * fixed at the recorded assumptions. The Grace CPU and LPDDR5X are never modelled:
+ * they are inside the measured reading. Rounding follows the source: rack AC is
+ * rounded before PUE, and every reported total is rounded once at the end.
+ */
+export function estimateRackPower(
+ hardware: string,
+ input: RackMeasuredInput,
+ pue = SYSTEM_POWER_RACK_ASSUMPTIONS.pue,
+): RackPowerEstimate | null {
+ const key = hardware.toLowerCase();
+ const gpuShare = input.basis === 'module' ? 0 : input.gpuBoardWattsPerTray;
+ const graceShare = input.basis === 'module' ? 0 : input.graceSocketWattsPerTray;
+ const measured = input.basis === 'module' ? [input.moduleWattsPerTray] : [gpuShare, graceShare];
+ if (
+ !Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, key) ||
+ measured.some((watts) => !Number.isFinite(watts) || watts <= 0) ||
+ !Number.isFinite(pue) ||
+ pue < 1
+ ) {
+ return null;
+ }
+ const canonicalHardware = key as SystemPowerRackHardware;
+ const profile = SYSTEM_POWER_RACK_PROFILES[canonicalHardware];
+ const trays = profile.computeTrayCount;
+ const measuredWattsPerTray =
+ input.basis === 'module'
+ ? input.moduleWattsPerTray
+ : input.gpuBoardWattsPerTray + input.graceSocketWattsPerTray;
+
+ // Grace tuning guide: regulator loss is 15% of the TDP limit, so loss / delivered =
+ // f / (1 - f) on the GPU-board share. The Grace socket reading already includes its own.
+ const frac = profile.regulatorLossFracOfTdp;
+ const allowanceBase = gpuShare + (profile.regulatorAllowanceIncludesGrace ? graceShare : 0);
+ const allowancePerTray =
+ input.basis === 'gpu-plus-grace' ? (allowanceBase * frac) / (1 - frac) : 0;
+ const computeModulesDc = trays * measuredWattsPerTray + trays * allowancePerTray;
+
+ // Per-tray static blocks keep the source's summation order.
+ const trayStaticPerTray = Object.values(profile.computeTrayStaticDcWatts).reduce(
+ (sum, watts) => sum + watts,
+ 0,
+ );
+ const trayStaticDc = trays * trayStaticPerTray;
+ const nvswitchTraysDc =
+ profile.nvswitchTrayCount *
+ (profile.nvswitchTraySiliconWatts + profile.nvswitchTrayResidualWatts);
+ const trayLoads = computeModulesDc + trayStaticDc + nvswitchTraysDc;
+ const trayConversionLoss = trayLoads * (1 / profile.trayInputConversionEfficiency - 1);
+ const managementDc = profile.managementSwitchCount * profile.managementSwitchWatts;
+ const rackDc = trayLoads + trayConversionLoss + managementDc;
+
+ const shelf = profile.powerShelf;
+ if (rackDc > shelf.installedCapacityWatts) return null;
+ const efficiency = interpolateEfficiency(
+ rackDc / shelf.installedCapacityWatts,
+ shelf.efficiencyCurve,
+ );
+ const ac = rackDc / efficiency;
+ const rackAc = pythonRound(ac);
+ // PUE applies once, to the already-rounded shelf AC output.
+ const facility = pythonRound(rackAc + rackAc * (pue - 1));
+ if (!Number.isFinite(facility)) return null;
+ return {
+ hardware: canonicalHardware,
+ basis: input.basis,
+ measuredWattsPerTray,
+ regulatorAllowanceWattsPerTray: allowancePerTray,
+ computeModulesDcWatts: pythonRound(computeModulesDc),
+ regulatorAllowanceWatts: pythonRound(trays * allowancePerTray),
+ trayStaticDcWatts: pythonRound(trayStaticDc),
+ nvswitchTraysDcWatts: pythonRound(nvswitchTraysDc),
+ trayConversionLossWatts: pythonRound(trayConversionLoss),
+ rackDcWatts: pythonRound(rackDc),
+ powerShelfEfficiency: pythonRound(efficiency, 4),
+ powerShelfLossWatts: pythonRound(ac - rackDc),
+ rackAcWatts: rackAc,
+ facilityWatts: facility,
+ perGpuAcWatts: pythonRound(rackAc / profile.gpuCount),
+ perGpuFacilityWatts: pythonRound(facility / profile.gpuCount),
+ pue,
+ computeTrayCount: trays,
+ gpuCount: profile.gpuCount,
+ modelRevision: SYSTEM_POWER_MODEL_REVISION,
+ modelPath: profile.modelPath,
+ };
+}
From 0bf119f7048643b1b27ca347dcb319df784db414 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 15:11:23 -0700
Subject: [PATCH 02/22] feat: accept fully measured NVL72 trays in the profit
planning gate and name the power basis
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Planning kW/GPU now admits nvl72-trays estimates whose trays are all fully
measured next to the eight-GPU single-node chassis; partial trays stay
rejected. Two frontier knots must share the same measured basis and sensor
kind or the bracket stays unavailable. Modeled rows carry powerSource
(topology, measured basis, sensor kind, PUE, pinned profile path, revision
and source SHA-256); the bar tooltip, a caption line and three CSV columns
name it per row. English strings for x86 rows are byte-identical.
中文:规划 kW/GPU 除八卡单节点机箱外,接受全部 tray 完整实测的 NVL72 估算,
部分 tray 仍被拒绝;两个前沿数据点的实测口径或传感器类型不同时保持不可用。
估算行携带 powerSource(拓扑、实测口径、传感器类型、PUE、固定 profile 路径、
版本和源文件 SHA-256),柱形提示、标注行和三列 CSV 逐行标出;x86 行英文字符串
逐字节不变。
---
.../calculator/ProfitEstimatorChart.tsx | 22 +++
.../calculator/ProfitEstimatorDisplay.tsx | 61 ++++++-
.../components/calculator/profit-estimator.ts | 3 +
.../calculator/profit-power.test.ts | 172 +++++++++++++++++-
.../src/components/calculator/profit-power.ts | 119 +++++++++---
packages/app/src/lib/system-power-model.ts | 6 +
6 files changed, 354 insertions(+), 29 deletions(-)
diff --git a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
index afab4e98b..5b9ff0a7d 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
@@ -28,6 +28,7 @@ import {
type ProfitEstimatorAssumptions,
type ProfitEstimatorRow,
} from './profit-estimator';
+import type { ProfitPowerSource } from './profit-power';
export type ProfitSegmentKind = 'tco' | 'labCut' | 'profit' | 'loss';
@@ -265,6 +266,13 @@ const STRINGS = {
ofRevenue: 'of revenue',
dismiss: 'Click anywhere to dismiss',
runDate: 'Run date',
+ powerBasis: 'Power basis',
+ powerBasisLabels: {
+ chassis: 'measured GPU board; CPU, DRAM, networking, storage, board, fans and PSU modeled',
+ module: 'measured module (GPU + HBM + Grace + LPDDR5X; module sensor)',
+ 'grace-socket':
+ 'measured GPU board + Grace socket (Grace socket sensor), regulator loss modeled',
+ },
noData: 'No SKU can be priced for the current selection.',
},
zh: {
@@ -291,6 +299,13 @@ const STRINGS = {
ofRevenue: '(占收入)',
dismiss: '点击任意位置关闭',
runDate: '运行日期',
+ powerBasis: '功耗口径',
+ powerBasisLabels: {
+ chassis: '实测 GPU 板卡功耗;CPU、DRAM、网络、存储、主板、风扇和 PSU 由模型估算',
+ module: '实测模块功耗(GPU + HBM + Grace + LPDDR5X;模块传感器)',
+ 'grace-socket':
+ '实测 GPU 板卡 + Grace socket 功耗(Grace socket 传感器),稳压损耗由模型估算',
+ },
noData: '当前选择下没有可定价的 SKU。',
},
} as const;
@@ -299,6 +314,12 @@ export function profitEstimatorChartStrings(locale: Locale) {
return STRINGS[locale];
}
+/** Name what a measured + modeled row's power was measured on. */
+export function powerBasisLabel(source: ProfitPowerSource, locale: Locale): string {
+ const labels = STRINGS[locale].powerBasisLabels;
+ return labels[source.topology === 'chassis' ? 'chassis' : source.sensorKind];
+}
+
/** Segment label lines that fit a given pixel height: name and amount, amount only, or none. */
export function segmentLabelLines(
kind: ProfitSegmentKind,
@@ -630,6 +651,7 @@ export function generateProfitTooltipHTML(
${isPinned ? `${t.dismiss}
` : ''}
${label}
${row.date ? line(t.runDate, escapeHtml(row.dateLabel ?? row.date)) : ''}
+ ${row.powerSource ? line(t.powerBasis, escapeHtml(powerBasisLabel(row.powerSource, locale))) : ''}
${line(`${t.revenue} (${t.utilization} ${assumptions.utilizationPct}%)`, usd(row.revenue))}
${line(t.tco, usd(row.tco), skuColor, TCO_OPACITY)}
${line(t.grossMargin, usd(row.grossMargin))}
diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
index a3bef4ea9..242938208 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
@@ -92,8 +92,8 @@ import {
type ProfitEstimatorRow,
type ProfitEstimatorSkipReason,
} from './profit-estimator';
-import { profitEstimatorChartStrings, rowLabel } from './ProfitEstimatorChart';
-import { estimateProfitByPower, type ProfitPowerBasis } from './profit-power';
+import { powerBasisLabel, profitEstimatorChartStrings, rowLabel } from './ProfitEstimatorChart';
+import { estimateProfitByPower, powerSourceKey, type ProfitPowerBasis } from './profit-power';
import {
buildProfitHistoryResults,
historyFadeShare,
@@ -214,6 +214,9 @@ const STRINGS = {
powerBarLabels: { provisioned: 'Provisioned', modeled: 'Measured + modeled' },
powerPreview:
'PowerX estimate · Same target, throughput, pricing and unit costs. GPU power comes from the same serving-frontier points; power between them is estimated linearly. Server overhead is modeled, with PUE 1.3 and 10% headroom. AgentX system power is not yet qualified.',
+ powerNvl72Note: (hardware: string, basis: string, pue: number) =>
+ `${hardware}: ${basis}. Modeled: NVSwitch trays, NICs/DPUs, NVMe, power shelves, DLC PUE ${pue}.`,
+ csvPowerHeaders: ['Power basis', 'Power sensor', 'System power profile'],
pricingGroup: 'Pricing Config',
costProviderLabel: 'Cost Provider',
costProviderTooltip:
@@ -329,6 +332,9 @@ const STRINGS = {
powerBarLabels: { provisioned: '预配功耗', modeled: '实测 + 估算' },
powerPreview:
'PowerX 估算 · 两种方式采用相同的目标交互性、吞吐量、价格和单位成本。GPU 功耗取自同一组性能前沿数据点,点间功耗采用线性估算。服务器开销由模型估算,PUE 为 1.3,功耗余量为 10%。AgentX 系统功耗模型尚未完成验证。',
+ powerNvl72Note: (hardware: string, basis: string, pue: number) =>
+ `${hardware}:${basis}。建模部分:NVSwitch tray、网卡/DPU、NVMe、电源架,液冷 PUE ${pue}。`,
+ csvPowerHeaders: ['功耗口径', '功耗传感器', '系统功耗 profile'],
pricingGroup: '定价配置',
costProviderLabel: '成本供应商',
costProviderTooltip:
@@ -1356,6 +1362,28 @@ function ProfitEstimatorInner({
[fullEstimate.skipped, hardwareConfig, historyEntryLabel, t],
);
+ // One line per NVL72 hardware whose bars price a measured compute module, so the
+ // reader sees which share is measured and which is modeled; x86 chassis rows keep
+ // the generic preview line.
+ const powerBasisNotes = useMemo(() => {
+ const notes = new Map();
+ for (const row of estimate.rows) {
+ const source = row.powerSource;
+ if (source?.topology !== 'nvl72-trays') continue;
+ const key = `${row.hwKey}|${powerSourceKey(source)}`;
+ if (notes.has(key)) continue;
+ notes.set(
+ key,
+ t.powerNvl72Note(
+ rowLabel({ hwKey: row.hwKey }, hardwareConfig),
+ powerBasisLabel(source, locale),
+ source.pue,
+ ),
+ );
+ }
+ return [...notes.values()];
+ }, [estimate.rows, hardwareConfig, locale, t]);
+
// Rendered as the chart's figcaption so it is part of the PNG export.
const caption = useMemo(() => {
if (!pricing) return null;
@@ -1378,6 +1406,11 @@ function ProfitEstimatorInner({
{powerBasis !== 'provisioned' && <>. {t.powerPreview}>}
)}
+ {powerBasisNotes.length > 0 && (
+
+ {powerBasisNotes.join(' ')}
+
+ )}
{basis === 'gw-year' && powerBasis !== 'provisioned' && fullEstimate.skipped.length > 0 && (
{powerUnavailable}
@@ -1458,6 +1491,7 @@ function ProfitEstimatorInner({
}, [
pricing,
powerBasis,
+ powerBasisNotes,
powerControlsEnabled,
powerUnavailable,
fullEstimate.skipped,
@@ -1493,6 +1527,18 @@ function ProfitEstimatorInner({
const handleExportCsv = useCallback(() => {
// Whole dollars are plenty per GW-year; per chip-hour the cents are the figure.
const usd = (value: number) => (basis === 'gw-year' ? Math.round(value) : value.toFixed(4));
+ // Measured + modeled rows name their basis, sensor, and pinned profile so a
+ // spreadsheet can tell a measured module from a modeled chassis per row.
+ const includeBasis = powerControlsEnabled && powerBasis !== 'provisioned';
+ const basisColumns = (row: ProfitEstimatorRow) => {
+ const source = row.powerSource;
+ if (!source) return [t.powerBarLabels.provisioned, '', ''];
+ return [
+ powerBasisLabel(source, locale),
+ source.topology === 'nvl72-trays' ? source.sensorKind : '',
+ `${source.modelPath} @ ${source.modelRevision}${source.profileSha256 ? ` sha256:${source.profileSha256}` : ''}`,
+ ];
+ };
const rows = estimate.rows.map((row) => [
rowLabel({ ...row, date: undefined }, hardwareConfig),
row.precision?.toUpperCase() ?? '',
@@ -1506,9 +1552,17 @@ function ProfitEstimatorInner({
row.revenuePerGpuHour.toFixed(4),
// GPU-hours is 1 per chip-hour, so that basis has no column for it.
...(basis === 'gw-year' ? [Math.round(row.gpuHours)] : []),
+ ...(includeBasis ? basisColumns(row) : []),
]);
const [sku, precision, ...rest] = t.csvHeaders[basis];
- exportToCsv(exportFileName, [sku, precision, t.csvDateHeader, ...rest], rows, [
+ const headers = [
+ sku,
+ precision,
+ t.csvDateHeader,
+ ...rest,
+ ...(includeBasis ? t.csvPowerHeaders : []),
+ ];
+ exportToCsv(exportFileName, headers, rows, [
t.captionFormula[basis](assumptions.utilizationPct, assumptions.labCutPct),
...(powerControlsEnabled
? [
@@ -1522,6 +1576,7 @@ function ProfitEstimatorInner({
hardwareConfig,
exportFileName,
t,
+ locale,
assumptions,
basis,
selectedRunDate,
diff --git a/packages/app/src/components/calculator/profit-estimator.ts b/packages/app/src/components/calculator/profit-estimator.ts
index 1fc102ad0..71a66289a 100644
--- a/packages/app/src/components/calculator/profit-estimator.ts
+++ b/packages/app/src/components/calculator/profit-estimator.ts
@@ -24,6 +24,7 @@ import { tokenRevenueFromRatesPerGpuHour } from '@/components/inference/token-re
import type { TokenRevenuePricing } from '@/components/inference/types';
import { Model } from '@/lib/data-mappings';
+import type { ProfitPowerSource } from './profit-power';
import type { InterpolatedResult } from './types';
/** Calendar hours in a year (365 x 24). */
@@ -189,6 +190,8 @@ export interface ProfitEstimatorAssumptions {
export interface ProfitEstimatorRow {
/** Distinguish identical hardware bars when both power budgets are shown. */
powerLabel?: string;
+ /** What the measured + modeled power budget was measured on; unset for provisioned rows. */
+ powerSource?: ProfitPowerSource;
hwKey: string;
resultKey: string;
precision?: string;
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index ce13042d8..57d2afd54 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -2,10 +2,15 @@ import { describe, expect, it } from 'vitest';
import type { BenchmarkRow } from '@/lib/api';
import { modelSystemPower } from '@/lib/modeled-system-power';
+import { estimateRackPower } from '@/lib/system-power-model';
import { Percentile, Sequence } from '@/lib/data-mappings';
import { buildGpuGroups, interpolateForGPU } from './useThroughputData';
import { estimateProfitRows } from './profit-estimator';
-import { estimateProfitByPower, modeledPowerAtTarget } from './profit-power';
+import {
+ estimateProfitByPower,
+ modeledPowerAtTarget,
+ type ProfitPowerSource,
+} from './profit-power';
import type { GPUDataPoint, InterpolatedResult } from './types';
// Power telemetry from MI355X Kimi K3 source row 441385; the target/rates below
@@ -91,6 +96,40 @@ const assumptions = { basis: 'gw-year' as const, utilizationPct: 60, labCutPct:
const specs = () => ({ powerKwPerGpu: 2.09, costPerGpuHour: 1.5 });
const labels = { provisioned: 'Provisioned', modeled: 'Measured + modeled' };
+// One GB200 NVL72 compute tray on the AgentX workload: four GPUs on one host, two
+// Grace sockets, module sensor present. Watts are controlled inputs, not published
+// constants; the CPU-side keys follow the ticket-01 contract.
+const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 };
+const traySource: BenchmarkRow = {
+ ...source,
+ hardware: 'gb200',
+ framework: 'sglang',
+ prefill_tp: 4,
+ decode_tp: 4,
+ num_prefill_gpu: 4,
+ num_decode_gpu: 4,
+ metrics: {
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ avg_power_w: 900.25,
+ avg_total_gpu_power_w: 3601,
+ ...GRACE,
+ avg_total_module_power_w: 4300.75,
+ },
+};
+const trayPoint: GPUDataPoint = { ...point, sourceRow: traySource, hwKey: 'gb200_sglang', tp: 4 };
+const trayResult: InterpolatedResult = {
+ ...result,
+ hwKey: trayPoint.hwKey,
+ resultKey: trayPoint.hwKey,
+ nearestPoints: [trayPoint],
+};
+const withPoints = (base: InterpolatedResult, points: GPUDataPoint[]): InterpolatedResult => ({
+ ...base,
+ nearestPoints: points,
+});
+
describe('profit power basis preview', () => {
it('keeps raw power attached through official and run-keyed frontier construction', () => {
const row = {
@@ -118,7 +157,7 @@ describe('profit power basis preview', () => {
}
});
- it('rejects partial chassis and unsupported rack hardware despite valid telemetry', () => {
+ it('rejects partial chassis and NVL72 rows without CPU-side telemetry despite valid GPU telemetry', () => {
for (const hardware of ['gb200', 'gb300']) {
expect(
modeledPowerAtTarget(
@@ -142,6 +181,135 @@ describe('profit power basis preview', () => {
).toBeNull();
});
+ it('accepts fully measured NVL72 trays and records the measured basis behind the estimate', () => {
+ // 1.1 × the tray's amortised facility watts per GPU from the pinned GB200 rack profile.
+ const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!;
+ const kw = modeledPowerAtTarget(trayResult, 45)!;
+ expect(kw).toBeCloseTo((rack.facilityWatts / rack.gpuCount / 1000) * 1.1, 8);
+ expect(kw).toBeCloseTo(1.6077325, 7);
+ const output = estimateProfitByPower(
+ [trayResult],
+ specs,
+ pricing,
+ assumptions,
+ 'compare',
+ 45,
+ labels,
+ );
+ expect(output.skipped).toEqual([]);
+ const [provisioned, modeled] = output.rows;
+ expect(provisioned.powerSource).toBeUndefined();
+ expect(modeled.powerSource).toEqual({
+ topology: 'nvl72-trays',
+ measuredBasis: 'module',
+ sensorKind: 'module',
+ pue: 1.1,
+ modelPath: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ modelRevision: rack.modelRevision,
+ profileSha256: expect.stringMatching(/^[0-9a-f]{64}$/u),
+ } satisfies ProfitPowerSource);
+ expect(modeled.gpuHours / provisioned.gpuHours).toBeCloseTo(2.09 / kw, 10);
+
+ // Two fully measured GB300 trays on distinct hosts, Grace-socket basis.
+ const twoTrays: BenchmarkRow = {
+ ...traySource,
+ hardware: 'gb300',
+ disagg: true,
+ is_multinode: true,
+ prefill_tp: 4,
+ decode_tp: 4,
+ metrics: {
+ ...traySource.metrics,
+ avg_power_w: 900,
+ avg_total_gpu_power_w: 7200,
+ prefill_avg_power_w: 950,
+ decode_avg_power_w: 850,
+ avg_cpu_socket_power_w: 260,
+ avg_total_cpu_power_w: 1040,
+ },
+ workers: [
+ { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 950 },
+ { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 850 },
+ ],
+ };
+ delete (twoTrays.metrics as Record).avg_total_module_power_w;
+ const estimate = modelSystemPower(twoTrays, undefined, true);
+ expect(estimate).toMatchObject({ status: 'supported', chassisBasis: 'full', gpuCount: 8 });
+ const rows = estimateProfitByPower(
+ [withPoints(trayResult, [{ ...trayPoint, sourceRow: twoTrays }])],
+ specs,
+ pricing,
+ assumptions,
+ 'modeled',
+ 45,
+ labels,
+ ).rows;
+ expect(rows).toHaveLength(1);
+ expect(rows[0].powerSource).toMatchObject({
+ topology: 'nvl72-trays',
+ measuredBasis: 'gpu-plus-grace',
+ sensorKind: 'grace-socket',
+ pue: 1.1,
+ });
+ });
+
+ it('rejects partially measured trays and never mixes measured bases between knots', () => {
+ const partialTray: BenchmarkRow = {
+ ...traySource,
+ prefill_tp: 2,
+ decode_tp: 2,
+ metrics: {
+ ...traySource.metrics,
+ avg_total_gpu_power_w: 1800.5,
+ avg_total_module_power_w: 4300.75,
+ },
+ };
+ expect(modelSystemPower(partialTray, undefined, true)).toMatchObject({
+ status: 'supported',
+ chassisBasis: 'extrapolated',
+ });
+ expect(
+ modeledPowerAtTarget(withPoints(trayResult, [{ ...trayPoint, sourceRow: partialTray }]), 45),
+ ).toBeNull();
+
+ const graceOnly: BenchmarkRow = { ...traySource, metrics: { ...traySource.metrics } };
+ delete (graceOnly.metrics as Record).avg_total_module_power_w;
+ const mixed = withPoints(trayResult, [
+ { ...trayPoint, interactivity: 30 },
+ { ...trayPoint, interactivity: 60, sourceRow: graceOnly },
+ ]);
+ expect(modeledPowerAtTarget(mixed, 30)).not.toBeNull();
+ expect(modeledPowerAtTarget(mixed, 60)).not.toBeNull();
+ expect(modeledPowerAtTarget(mixed, 45)).toBeNull();
+ const same = withPoints(trayResult, [
+ { ...trayPoint, interactivity: 30 },
+ { ...trayPoint, interactivity: 60 },
+ ]);
+ expect(modeledPowerAtTarget(same, 45)).toBeCloseTo(modeledPowerAtTarget(trayResult, 45)!, 10);
+ });
+
+ it('labels x86 chassis estimates with the air-cooled profile and leaves provisioned rows unlabeled', () => {
+ const [provisioned, modeled] = estimateProfitByPower(
+ [result],
+ specs,
+ pricing,
+ assumptions,
+ 'compare',
+ 45,
+ labels,
+ ).rows;
+ expect(provisioned.powerSource).toBeUndefined();
+ expect(modeled.powerSource).toMatchObject({
+ topology: 'chassis',
+ pue: 1.3,
+ modelPath: 'human_verified/mi355x_chassis/mi355x_chassis_power_model.py',
+ });
+ expect(
+ estimateProfitByPower([result], specs, pricing, assumptions, 'provisioned', 45, labels)
+ .rows[0].powerSource,
+ ).toBeUndefined();
+ });
+
it('leaves the default estimator and default AgentX model gate unchanged', () => {
expect(
estimateProfitByPower([result], specs, pricing, assumptions, 'provisioned', 45, labels),
diff --git a/packages/app/src/components/calculator/profit-power.ts b/packages/app/src/components/calculator/profit-power.ts
index 086322574..9470e31cf 100644
--- a/packages/app/src/components/calculator/profit-power.ts
+++ b/packages/app/src/components/calculator/profit-power.ts
@@ -1,4 +1,5 @@
-import { modelSystemPower } from '@/lib/modeled-system-power';
+import { modelSystemPower, type SystemPowerSensorKind } from '@/lib/modeled-system-power';
+import { type RackMeasuredBasis, systemPowerSourceSha256 } from '@/lib/system-power-model';
import type { TokenRevenuePricing } from '@/components/inference/types';
import {
@@ -13,34 +14,103 @@ import type { GPUDataPoint, InterpolatedResult } from './types';
export type ProfitPowerBasis = 'provisioned' | 'modeled' | 'compare';
-function planningKwPerGpu(point: GPUDataPoint): number | null {
+interface ProfitPowerProfile {
+ /** Facility PUE the estimate applied once at the chassis or rack AC boundary. */
+ pue: number;
+ modelPath: string;
+ modelRevision: string;
+ /** SHA-256 of the pinned source file behind `modelPath`; null if the profile lacks one. */
+ profileSha256: string | null;
+}
+
+/**
+ * What the measured + modeled budget was measured on. An eight-GPU x86 chassis
+ * measures the GPU boards and models the rest; an NVL72 tray measures the compute
+ * module (or GPU board + Grace socket) and models only the rack residual.
+ */
+export type ProfitPowerSource =
+ | (ProfitPowerProfile & { topology: 'chassis' })
+ | (ProfitPowerProfile & {
+ topology: 'nvl72-trays';
+ measuredBasis: RackMeasuredBasis;
+ sensorKind: SystemPowerSensorKind;
+ });
+
+export interface ProfitPlanningPower {
+ kwPerGpu: number;
+ source: ProfitPowerSource;
+}
+
+/** Identity of a source for deduplicating notes and refusing to mix bases between knots. */
+export function powerSourceKey(source: ProfitPowerSource): string {
+ return [
+ source.topology,
+ source.pue,
+ source.modelPath,
+ source.modelRevision,
+ source.topology === 'nvl72-trays' ? `${source.measuredBasis}/${source.sensorKind}` : '',
+ ].join('|');
+}
+
+/**
+ * Planning kW/GPU: facility watts per measured GPU plus the 10% planning margin.
+ * Accepted: one complete eight-GPU chassis, or NVL72 compute trays that are all
+ * fully measured. A partial unit would price its unmeasured GPUs at a modeled share.
+ */
+function planningPower(point: GPUDataPoint): ProfitPlanningPower | null {
const row = point.sourceRow;
if (!row || row.metrics.power_metric_schema_version !== 2) return null;
const estimate = modelSystemPower(row, undefined, true);
- if (
- estimate.status !== 'supported' ||
- estimate.gpuCount !== 8 ||
- estimate.topologyBasis !== 'single-node' ||
- estimate.chassisBasis !== 'full'
- )
- return null;
- return (estimate.deploymentFacilityWatts / estimate.gpuCount / 1000) * 1.1;
+ if (estimate.status !== 'supported' || estimate.chassisBasis !== 'full') return null;
+ const profile: ProfitPowerProfile = {
+ pue: estimate.pue,
+ modelPath: estimate.modelPath,
+ modelRevision: estimate.modelRevision,
+ profileSha256: systemPowerSourceSha256(estimate.modelPath),
+ };
+ let source: ProfitPowerSource;
+ if (estimate.topologyBasis === 'nvl72-trays') {
+ source = {
+ ...profile,
+ topology: 'nvl72-trays',
+ measuredBasis: estimate.measuredBasis,
+ sensorKind: estimate.sensorKind,
+ };
+ } else if (estimate.topologyBasis === 'single-node' && estimate.gpuCount === 8) {
+ source = { ...profile, topology: 'chassis' };
+ } else return null;
+ return {
+ kwPerGpu: (estimate.deploymentFacilityWatts / estimate.gpuCount / 1000) * 1.1,
+ source,
+ };
}
/** Reusing the original frontier prevents the power choice from changing throughput. */
-export function modeledPowerAtTarget(result: InterpolatedResult, target: number): number | null {
+export function modeledPlanningPowerAtTarget(
+ result: InterpolatedResult,
+ target: number,
+): ProfitPlanningPower | null {
if (result.clamped) return null;
const exact = result.nearestPoints.find((p) => Math.abs(p.interactivity - target) < 1e-9);
- if (exact) return planningKwPerGpu(exact);
+ if (exact) return planningPower(exact);
const [left, right] = result.nearestPoints;
if (!left || !right || target < left.interactivity || target > right.interactivity) return null;
- const lower = planningKwPerGpu(left),
- upper = planningKwPerGpu(right);
- if (lower === null || upper === null || right.interactivity <= left.interactivity) return null;
- return (
- lower +
- ((upper - lower) * (target - left.interactivity)) / (right.interactivity - left.interactivity)
- );
+ const lower = planningPower(left),
+ upper = planningPower(right);
+ if (!lower || !upper || right.interactivity <= left.interactivity) return null;
+ // Two knots measured on different bases would price one bar on two sensors.
+ if (powerSourceKey(lower.source) !== powerSourceKey(upper.source)) return null;
+ return {
+ kwPerGpu:
+ lower.kwPerGpu +
+ ((upper.kwPerGpu - lower.kwPerGpu) * (target - left.interactivity)) /
+ (right.interactivity - left.interactivity),
+ source: lower.source,
+ };
+}
+
+export function modeledPowerAtTarget(result: InterpolatedResult, target: number): number | null {
+ return modeledPlanningPowerAtTarget(result, target)?.kwPerGpu ?? null;
}
export function estimateProfitByPower(
@@ -63,7 +133,7 @@ export function estimateProfitByPower(
output.skipped.push(baseline);
continue;
}
- const power = modeledPowerAtTarget(result, target);
+ const power = modeledPlanningPowerAtTarget(result, target);
if (power === null) {
output.skipped.push({
hwKey: result.hwKey,
@@ -74,16 +144,17 @@ export function estimateProfitByPower(
});
continue;
}
- const modeled = estimateSkuProfit(
+ const estimated = estimateSkuProfit(
result,
- { ...specs, powerKwPerGpu: power },
+ { ...specs, powerKwPerGpu: power.kwPerGpu },
pricing,
assumptions,
);
- if (!isProfitEstimatorRow(modeled)) {
- output.skipped.push(modeled);
+ if (!isProfitEstimatorRow(estimated)) {
+ output.skipped.push(estimated);
continue;
}
+ const modeled = { ...estimated, powerSource: power.source };
if (powerBasis === 'compare') {
output.rows.push(
{
diff --git a/packages/app/src/lib/system-power-model.ts b/packages/app/src/lib/system-power-model.ts
index 8cf6e1fd2..15d7f3f43 100644
--- a/packages/app/src/lib/system-power-model.ts
+++ b/packages/app/src/lib/system-power-model.ts
@@ -16,6 +16,12 @@ export const SUPPORTED_SYSTEM_POWER_RACK_HARDWARE = Object.keys(
export type RackMeasuredBasis = 'module' | 'gpu-plus-grace';
+/** SHA-256 of the pinned source file behind a profile's `modelPath`, for export provenance. */
+export function systemPowerSourceSha256(modelPath: string): string | null {
+ const hashes: Readonly> = profileData.sourceSha256;
+ return Object.hasOwn(hashes, modelPath) ? hashes[modelPath] : null;
+}
+
/**
* Measured compute-module input for every tray of one NVL72 rack. `module` is the
* sum of the two Module Power sensors per tray (Grace + 2 Blackwell + HBM + LPDDR5X +
From e4d95e1bf9dcb246020dd8053fa452f70204be25 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 15:11:33 -0700
Subject: [PATCH 03/22] feat: register NVL72 CPU-side power keys and
power_audit.cpu across constants, ingest and the API reference
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
avg_cpu_socket_power_w, avg_total_cpu_power_w, total_cpu_energy_j,
avg_total_module_power_w and total_module_energy_j join
MEASURED_POWER_METRIC_KEY_LIST (withheld with power_valid=0 like every
measured key); cpu_power_valid joins the contract discriminators and is
normalized as a verdict independent of power_valid. extractPowerAudit keeps
a bounded power_audit.cpu block matching the consumer's audit_summary.
The API registry documents the six keys, the cpu audit schema, and a
bilingual measured-power note and example; no route catalog digest changes.
中文:五个 CPU 侧实测指标加入 MEASURED_POWER_METRIC_KEY_LIST(与其他实测键一样
在 power_valid=0 时移除);cpu_power_valid 加入契约字段并按独立于 power_valid
的验证结论归一化;extractPowerAudit 保留与消费端 audit_summary 一致的有界
power_audit.cpu。API 文档新增六个键的说明、cpu 审计 schema 以及中英文
measured-power 说明与示例;路由目录摘要无需刷新。
---
.../src/lib/api-documentation.power.test.ts | 23 ++++
packages/app/src/lib/api-documentation.ts | 38 +++++-
packages/constants/src/metric-keys.test.ts | 18 ++-
packages/constants/src/metric-keys.ts | 16 +++
packages/db/src/etl/benchmark-mapper.test.ts | 118 +++++++++++++++++-
packages/db/src/etl/benchmark-mapper.ts | 58 ++++++++-
6 files changed, 259 insertions(+), 12 deletions(-)
diff --git a/packages/app/src/lib/api-documentation.power.test.ts b/packages/app/src/lib/api-documentation.power.test.ts
index 4ecb6ebdb..14986693f 100644
--- a/packages/app/src/lib/api-documentation.power.test.ts
+++ b/packages/app/src/lib/api-documentation.power.test.ts
@@ -43,10 +43,30 @@ describe('measured-power API documentation', () => {
'max_sample_gap_s',
'producer_sha',
'exporter_image_sha256',
+ 'cpu',
].toSorted(),
);
expect(audit?.required).toBeUndefined();
+ // NVL72 CPU-side leg provenance: the enum matches the consumer's sensor kinds.
+ const cpu = audit?.properties?.cpu;
+ expect(cpu?.properties?.sensor_kind).toEqual({
+ type: 'string',
+ enum: ['module', 'grace_socket', 'dcgm_cpu_rail'],
+ });
+ expect(cpu?.properties?.reason_codes).toEqual({ type: 'array', items: { type: 'string' } });
+ expect(Object.keys(cpu?.properties ?? {}).toSorted()).toEqual(
+ [
+ 'sensor_kind',
+ 'source',
+ 'expected_sockets',
+ 'observed_sockets',
+ 'sample_row_count',
+ 'reason_codes',
+ ].toSorted(),
+ );
+ expect(cpu?.required).toBeUndefined();
+
expect(benchmarkRowSchema?.required).not.toContain('power_invalid_reasons');
expect(benchmarkRowSchema?.required).not.toContain('power_audit');
});
@@ -83,6 +103,9 @@ describe('measured-power API documentation', () => {
expect(note?.description).toContain('powerValid=strictV2');
expect(note?.description).toContain('power_valid == 1');
expect(note?.description).toContain('power_metric_schema_version == 2');
+ expect(note?.description).toContain('cpu_power_valid');
+ expect(note?.description).toContain('power_audit.cpu');
+ expect(note?.example).toMatchObject({ cpu_power_valid: 1, avg_total_cpu_power_w: 501 });
}
const zhNote = getApiDocumentation('zh').schemaNotes.find(
(candidate) => candidate.id === 'measured-power',
diff --git a/packages/app/src/lib/api-documentation.ts b/packages/app/src/lib/api-documentation.ts
index 48d567e53..cd05964c2 100644
--- a/packages/app/src/lib/api-documentation.ts
+++ b/packages/app/src/lib/api-documentation.ts
@@ -238,6 +238,18 @@ const powerMetricDescriptions: Readonly {
'peak_temp_c',
'avg_util_pct',
'avg_mem_used_mb',
+ // NVL72 Grace-side and compute-module measurements (same window as GPU energy).
+ 'avg_cpu_socket_power_w',
+ 'avg_total_cpu_power_w',
+ 'total_cpu_energy_j',
+ 'avg_total_module_power_w',
+ 'total_module_energy_j',
]),
);
- expect(MEASURED_POWER_METRIC_KEYS.size).toBe(19);
+ expect(MEASURED_POWER_METRIC_KEYS.size).toBe(24);
});
it('never contains the contract discriminators or invalid-verdict companion fields', () => {
@@ -46,6 +52,7 @@ describe('MEASURED_POWER_METRIC_KEYS', () => {
for (const key of [
'power_valid',
'power_metric_schema_version',
+ 'cpu_power_valid',
'power_invalid_reasons',
'power_audit',
]) {
@@ -69,8 +76,13 @@ describe('POWER_METRIC_KEYS', () => {
// The public API documentation types every one of these keys on
// BenchmarkRow.metrics, so membership changes are contract changes.
expect(new Set(POWER_METRIC_KEYS)).toEqual(
- new Set(['power_valid', 'power_metric_schema_version', ...MEASURED_POWER_METRIC_KEY_LIST]),
+ new Set([
+ 'power_valid',
+ 'power_metric_schema_version',
+ 'cpu_power_valid',
+ ...MEASURED_POWER_METRIC_KEY_LIST,
+ ]),
);
- expect(POWER_METRIC_KEYS).toHaveLength(21);
+ expect(POWER_METRIC_KEYS).toHaveLength(27);
});
});
diff --git a/packages/constants/src/metric-keys.ts b/packages/constants/src/metric-keys.ts
index 3c983ccb3..dceb88844 100644
--- a/packages/constants/src/metric-keys.ts
+++ b/packages/constants/src/metric-keys.ts
@@ -46,6 +46,19 @@ export const MEASURED_POWER_METRIC_KEY_LIST = [
'peak_temp_c',
'avg_util_pct',
'avg_mem_used_mb',
+ // NVL72 Grace-side and compute-module measurements from the srt-slurm CPU power
+ // leg (ACPI hwmon), integrated over the same formal window as GPU energy.
+ // avg_cpu_socket_power_w: mean over sockets of each socket's window-mean Grace-side W
+ // avg_total_cpu_power_w: sum over sockets of window-mean Grace-side W
+ // total_cpu_energy_j: Grace-side energy over the window, all sockets
+ // avg_total_module_power_w / total_module_energy_j: whole compute module
+ // (Grace + GPUs + HBM + LPDDR5X + regulator loss), only when
+ // the module sensor exists on every socket
+ 'avg_cpu_socket_power_w',
+ 'avg_total_cpu_power_w',
+ 'total_cpu_energy_j',
+ 'avg_total_module_power_w',
+ 'total_module_energy_j',
] as const;
export const MEASURED_POWER_METRIC_KEYS: ReadonlySet = new Set(
@@ -65,6 +78,9 @@ export const POWER_METRIC_KEYS = [
// joules_per_* field as whole-deployment energy
'power_valid',
'power_metric_schema_version',
+ // cpu_power_valid: numeric 1/0 verdict for the NVL72 CPU-side leg, independent
+ // of power_valid; 0 means the producer emitted no CPU-side keys
+ 'cpu_power_valid',
// measured power / energy / telemetry values, withheld when power_valid = 0
...MEASURED_POWER_METRIC_KEY_LIST,
] as const;
diff --git a/packages/db/src/etl/benchmark-mapper.test.ts b/packages/db/src/etl/benchmark-mapper.test.ts
index 962f5de16..ed9dda0c0 100644
--- a/packages/db/src/etl/benchmark-mapper.test.ts
+++ b/packages/db/src/etl/benchmark-mapper.test.ts
@@ -88,6 +88,12 @@ function dirtyPowerPayload(): Record {
peak_temp_c: 79.2,
avg_util_pct: 88.5,
avg_mem_used_mb: 71234.5,
+ // NVL72 CPU-side measurements share the GPU window; withheld with the GPU verdict.
+ avg_cpu_socket_power_w: 250.5,
+ avg_total_cpu_power_w: 1002,
+ total_cpu_energy_j: 601200,
+ avg_total_module_power_w: 17203,
+ total_module_energy_j: 10321800,
workers: [
{ role: 'prefill', worker_idx: 0, hosts: ['pn0'], num_gpus: 4, avg_power_w: 612.3 },
{ role: 'decode', worker_idx: 0, hosts: ['dn0'], num_gpus: 8, avg_power_w: 701.5 },
@@ -339,6 +345,29 @@ describe('mapBenchmarkRow', () => {
expect(result!.metrics.median_ttft).toBe(50.2);
});
+ it('withholds the CPU-side measurements but keeps the independent cpu_power_valid verdict', () => {
+ const tracker = createSkipTracker();
+ const result = mapBenchmarkRow(
+ makeV2Row({
+ power_valid: 0,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ ...dirtyPowerPayload(),
+ }),
+ tracker,
+ );
+ expect(result!.metrics.cpu_power_valid).toBe(1);
+ for (const key of [
+ 'avg_cpu_socket_power_w',
+ 'avg_total_cpu_power_w',
+ 'total_cpu_energy_j',
+ 'avg_total_module_power_w',
+ 'total_module_energy_j',
+ ]) {
+ expect(result!.metrics).not.toHaveProperty(key);
+ }
+ });
+
it('keeps every measured key and the workers payload on a valid verdict', () => {
const tracker = createSkipTracker();
const dirty = dirtyPowerPayload();
@@ -828,6 +857,41 @@ describe('mapBenchmarkRow', () => {
expect(result!.metrics).not.toHaveProperty('workers');
});
+ it('captures the NVL72 CPU-side keys without an unknown-key warning', async () => {
+ // The warning fires once per process per key, so a fresh module instance is
+ // the only way to observe whether these keys are known.
+ vi.resetModules();
+ const fresh = await import('./benchmark-mapper');
+ const warn = vi.spyOn(console, 'warn').mockImplementation(() => {});
+ try {
+ const tracker = createSkipTracker();
+ const result = fresh.mapBenchmarkRow(
+ makeV2Row({
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ avg_cpu_socket_power_w: 250.5,
+ avg_total_cpu_power_w: 1002,
+ total_cpu_energy_j: 601200,
+ avg_total_module_power_w: 17203,
+ total_module_energy_j: 10321800,
+ }),
+ tracker,
+ );
+ expect(result!.metrics).toMatchObject({
+ cpu_power_valid: 1,
+ avg_cpu_socket_power_w: 250.5,
+ avg_total_cpu_power_w: 1002,
+ total_cpu_energy_j: 601200,
+ avg_total_module_power_w: 17203,
+ total_module_energy_j: 10321800,
+ });
+ expect(warn).not.toHaveBeenCalled();
+ } finally {
+ warn.mockRestore();
+ }
+ });
+
it('captures new cluster-wide temp / util / mem scalars into metrics', () => {
// These are flat scalars on the agg row (sibling of avg_power_w), so
// the auto-capture path must store them under their raw keys without
@@ -891,6 +955,22 @@ describe('scrubWithheldPowerMetrics (direct — supplemental ingest path)', () =
}
});
+ it('normalizes cpu_power_valid as a verdict, independent of power_valid', () => {
+ const metrics = supplementalMetrics({ power_valid: 1, cpu_power_valid: '1' });
+ normalizePowerContractMetrics(metrics, metrics);
+ expect(metrics.cpu_power_valid).toBe(1);
+ expect(scrubWithheldPowerMetrics(metrics)).toBe(false);
+
+ const malformed = supplementalMetrics({ power_valid: 1, cpu_power_valid: 2 });
+ normalizePowerContractMetrics(malformed, malformed);
+ expect(malformed.cpu_power_valid).toBe(0);
+ expect(malformed.avg_total_cpu_power_w).toBe(1002);
+
+ const absent = supplementalMetrics({ power_valid: 1 });
+ normalizePowerContractMetrics(absent, absent);
+ expect(absent).not.toHaveProperty('cpu_power_valid');
+ });
+
it('fails closed on a malformed verdict when composed with normalization', () => {
const metrics = supplementalMetrics({ power_valid: 2 });
normalizePowerContractMetrics(metrics, metrics);
@@ -1158,7 +1238,43 @@ describe('extractPowerAudit', () => {
});
});
- it('drops unknown keys (fixed 8-key shape bounds the stored object)', () => {
+ it('keeps the bounded CPU-side audit block and drops malformed CPU fields', () => {
+ const cpu = {
+ sensor_kind: 'module',
+ source: 'acpi',
+ expected_sockets: 4,
+ observed_sockets: 4,
+ sample_row_count: 2400,
+ reason_codes: [],
+ };
+ expect(extractPowerAudit({ ...fullAudit, cpu })).toEqual({ ...fullAudit, cpu });
+ expect(
+ extractPowerAudit({
+ sample_count: 1,
+ cpu: {
+ sensor_kind: 'thermocouple',
+ source: 'x'.repeat(33),
+ expected_sockets: -1,
+ observed_sockets: Number.NaN,
+ sample_row_count: '12',
+ reason_codes: ['cpu_socket_count_mismatch', 'cpu_socket_count_mismatch', ' ', 7],
+ },
+ }),
+ ).toEqual({
+ sample_count: 1,
+ producer_sha: null,
+ exporter_image_sha256: null,
+ cpu: { sample_row_count: 12, reason_codes: ['cpu_socket_count_mismatch'] },
+ });
+ expect(extractPowerAudit({ sample_count: 1, cpu: {} })).toEqual({
+ sample_count: 1,
+ producer_sha: null,
+ exporter_image_sha256: null,
+ });
+ expect(extractPowerAudit({ sample_count: 1, cpu: 'acpi' })).not.toHaveProperty('cpu');
+ });
+
+ it('drops unknown keys (fixed 9-key shape bounds the stored object)', () => {
expect(extractPowerAudit({ sample_count: 3, integration_method: 'trapezoid' })).toEqual({
sample_count: 3,
producer_sha: null,
diff --git a/packages/db/src/etl/benchmark-mapper.ts b/packages/db/src/etl/benchmark-mapper.ts
index 060b39412..e904c9b8a 100644
--- a/packages/db/src/etl/benchmark-mapper.ts
+++ b/packages/db/src/etl/benchmark-mapper.ts
@@ -152,6 +152,24 @@ export interface PowerAudit {
source?: string;
/** Producer device identifiers; not necessarily physical UUIDs on older traces. */
observed_gpu_ids?: string[];
+ /** NVL72 CPU-side leg provenance; present only when the producer ran that leg. */
+ cpu?: PowerAuditCpu;
+}
+
+/** Sensor kinds the consumer's CPU-side leg can select as the headline series. */
+export const POWER_AUDIT_CPU_SENSOR_KINDS = ['module', 'grace_socket', 'dcgm_cpu_rail'] as const;
+
+/**
+ * Bounded CPU-side provenance emitted next to `cpu_power_valid`: which sensor fed
+ * the Grace-side keys, its source, socket coverage, and the leg's reason codes.
+ */
+export interface PowerAuditCpu {
+ sensor_kind?: (typeof POWER_AUDIT_CPU_SENSOR_KINDS)[number];
+ source?: string;
+ expected_sockets?: number;
+ observed_sockets?: number;
+ sample_row_count?: number;
+ reason_codes?: string[];
}
export interface BenchmarkParams {
@@ -540,11 +558,13 @@ export function normalizePowerContractMetrics(
row: Record,
metrics: Record,
): void {
- if (Object.hasOwn(row, 'power_valid')) {
- const verdict = row.power_valid;
- metrics.power_valid = verdict === 1 || verdict === '1' ? 1 : 0;
- } else {
- delete metrics.power_valid;
+ for (const field of ['power_valid', 'cpu_power_valid'] as const) {
+ if (Object.hasOwn(row, field)) {
+ const verdict = row[field];
+ metrics[field] = verdict === 1 || verdict === '1' ? 1 : 0;
+ } else {
+ delete metrics[field];
+ }
}
if (!Object.hasOwn(row, 'power_metric_schema_version')) {
@@ -666,6 +686,32 @@ function auditSha(v: unknown): string | null {
return s.length > 0 && s.length <= MAX_POWER_AUDIT_SHA_LENGTH ? s : null;
}
+/**
+ * Narrow the CPU-side leg's provenance. Reason codes reuse the producer reason
+ * grammar; an unrecognised sensor kind is dropped rather than stored as a label
+ * the dashboard would misread. Undefined when nothing well-formed remains.
+ */
+function extractPowerAuditCpu(raw: unknown): PowerAuditCpu | undefined {
+ if (!raw || typeof raw !== 'object' || Array.isArray(raw)) return undefined;
+ const e = raw as Record;
+ const cpu: PowerAuditCpu = {};
+ if ((POWER_AUDIT_CPU_SENSOR_KINDS as readonly unknown[]).includes(e.sensor_kind)) {
+ cpu.sensor_kind = e.sensor_kind as PowerAuditCpu['sensor_kind'];
+ }
+ if (typeof e.source === 'string' && e.source.length > 0 && e.source.length <= 32) {
+ cpu.source = e.source;
+ }
+ for (const field of ['expected_sockets', 'observed_sockets', 'sample_row_count'] as const) {
+ const n = auditCount(e[field]);
+ if (n !== undefined) cpu[field] = n;
+ }
+ // An empty list is the producer's "leg valid, nothing to report"; keep it.
+ if (Array.isArray(e.reason_codes)) {
+ cpu.reason_codes = extractPowerInvalidReasons(e.reason_codes) ?? [];
+ }
+ return Object.keys(cpu).length > 0 ? cpu : undefined;
+}
+
/**
* Missing or malformed audit values must not become a fabricated measurement;
* SQL NULL distinguishes absent evidence from an empty recorded object.
@@ -701,6 +747,8 @@ export function extractPowerAudit(raw: unknown): PowerAudit | undefined {
);
if (ids.length > 0) audit.observed_gpu_ids = [...new Set(ids)].slice(0, 1024);
}
+ const cpu = extractPowerAuditCpu(e.cpu);
+ if (cpu !== undefined) audit.cpu = cpu;
const hasNumericField = Object.keys(audit).length > 0;
audit.producer_sha = auditSha(e.producer_sha);
From 02cff945f24ad542dc03da571595218894bd0049 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 15:11:42 -0700
Subject: [PATCH 04/22] test: cover the GB200 NVL72 basis control in the profit
estimator Cypress spec
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
profit-fixtures gains an NVL72 SKU (one four-GPU tray with CPU-side keys and
the module sensor) kept out of PROFIT_SKUS so existing bar counts hold. The
new case checks the control stays hidden while the gate is locked, then
prices the tray on Measured + modeled, names the measured module basis and
DLC PUE 1.1 in the caption, and leaves GB300 unavailable.
中文:profit-fixtures 新增 NVL72 SKU(单个四卡 tray,含 CPU 侧指标与模块传感器),
不加入 PROFIT_SKUS 以保持现有柱形数量。新用例验证功能开关锁定时控件隐藏,解锁后
按实测 + 估算为该 tray 定价,标注行写明实测模块口径与液冷 PUE 1.1,GB300 保持不可用。
---
.../app/cypress/e2e/profit-estimator.cy.ts | 46 ++++++++++++++++
.../app/cypress/support/profit-fixtures.ts | 54 +++++++++++++++----
2 files changed, 90 insertions(+), 10 deletions(-)
diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts
index bf50c1206..0d8b38d3d 100644
--- a/packages/app/cypress/e2e/profit-estimator.cy.ts
+++ b/packages/app/cypress/e2e/profit-estimator.cy.ts
@@ -38,6 +38,7 @@ import { interceptVrPublicationData } from '../support/vr-publication-fixtures';
import {
interceptProfitData,
profitBenchmarkRows,
+ profitNvl72Rows,
PROFIT_CHANGELOG_NOTES,
PROFIT_DATE,
PROFIT_HISTORY_DATE,
@@ -206,6 +207,51 @@ describe('Profit estimator power option', () => {
cy.get('[data-testid="profit-power-unavailable"]').should('contain', 'GB300');
});
+ it('prices a GB200 NVL72 tray on its measured compute module and names the basis', () => {
+ stubOpenRouter();
+ cy.intercept('GET', '/api/v1/benchmarks*', (req) => {
+ req.reply({
+ body: [
+ ...profitBenchmarkRows().map((row) => ({
+ ...row,
+ metrics: {
+ ...row.metrics,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ },
+ })),
+ ...profitNvl72Rows(),
+ ],
+ });
+ });
+ // Locked: the control stays hidden and the tray prices on provisioned power like every SKU.
+ cy.visit('/profit-estimator-per-gigawatt?c_power=compare', {
+ onBeforeLoad: (win) => {
+ suppressNudges(win);
+ win.localStorage.removeItem('inferencex-feature-gate');
+ },
+ });
+ chart().find('text.revenue-label').should('have.length', 5);
+ cy.get('#profit-power').should('not.exist');
+ cy.get('[data-testid="profit-power-basis"]').should('not.exist');
+ cy.get('body').type('{uparrow}{uparrow}{downarrow}{downarrow}');
+ cy.get('#profit-power').should('contain', 'Compare both');
+ // B200, B300, MI355X and the GB200 tray each get a provisioned and a measured bar.
+ chart().find('text.revenue-label').should('have.length', 8);
+ chart().should('contain', 'GB200').and('contain', 'Measured + modeled');
+ cy.get('[data-testid="profit-power-basis"]')
+ .should('contain', 'GB200 NVL72')
+ .and('contain', 'measured module (GPU + HBM + Grace + LPDDR5X; module sensor)')
+ .and('contain', 'NVSwitch trays')
+ .and('contain', 'DLC PUE 1.1');
+ // GB300 has no CPU-side telemetry in these fixtures and stays unavailable; GB200 is priced.
+ cy.get('[data-testid="profit-power-unavailable"]')
+ .should('contain', 'GB300')
+ .and('not.contain', 'GB200');
+ });
+
it('keeps the benchmark settings and restores the original chart after unavailable power', () => {
stubOpenRouter();
cy.visit('/profit-estimator-per-gigawatt', { onBeforeLoad: unlockPowerGate });
diff --git a/packages/app/cypress/support/profit-fixtures.ts b/packages/app/cypress/support/profit-fixtures.ts
index 030042997..075754bdd 100644
--- a/packages/app/cypress/support/profit-fixtures.ts
+++ b/packages/app/cypress/support/profit-fixtures.ts
@@ -91,6 +91,10 @@ interface ProfitSku {
precision: string;
curve: Curve;
tputScale: number;
+ /** Physical GPUs per row (TP); eight when unset. */
+ gpus?: number;
+ /** Extra metric keys every row of this SKU carries, e.g. validated power telemetry. */
+ metrics?: Record;
}
export const PROFIT_SKUS: ProfitSku[] = [
@@ -101,14 +105,41 @@ export const PROFIT_SKUS: ProfitSku[] = [
{ hardware: 'mi355x', framework: 'vllm', precision: 'fp4', curve: WIDE_CURVE, tputScale: 0.9 },
];
+/**
+ * One GB200 NVL72 compute tray (four GPUs, two Grace sockets) with the CPU-side
+ * keys the srt-slurm CPU power leg publishes, module sensor included, so the
+ * smart-provisioning basis can price NVL72. Kept out of `PROFIT_SKUS` so the
+ * default bar counts the other specs lock down do not move. Watts are
+ * controlled inputs, not published constants.
+ */
+const NVL72_SKU: ProfitSku = {
+ hardware: 'gb200',
+ framework: 'sglang',
+ precision: 'fp4',
+ curve: WIDE_CURVE,
+ tputScale: 1.3,
+ gpus: 4,
+ metrics: {
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ avg_power_w: 900.25,
+ avg_total_gpu_power_w: 3601,
+ avg_cpu_socket_power_w: 250.5,
+ avg_total_cpu_power_w: 501,
+ avg_total_module_power_w: 4300.75,
+ },
+};
+
let idCursor = 800_000;
export const profitBenchmarkRows = (
dbKey: string = PROFIT_MODEL_DB_KEY,
date = PROFIT_DATE,
runId?: number,
+ skus: readonly ProfitSku[] = PROFIT_SKUS,
) =>
- PROFIT_SKUS.flatMap((sku) =>
+ skus.flatMap((sku) =>
sku.curve.map(([conc, intvty, tput, e2el]) => ({
id: idCursor++,
hardware: sku.hardware,
@@ -118,21 +149,20 @@ export const profitBenchmarkRows = (
spec_method: 'none',
disagg: false,
is_multinode: false,
- prefill_tp: 8,
- decode_tp: 8,
- num_prefill_gpu: 8,
- num_decode_gpu: 8,
+ prefill_tp: sku.gpus ?? 8,
+ decode_tp: sku.gpus ?? 8,
+ num_prefill_gpu: sku.gpus ?? 8,
+ num_decode_gpu: sku.gpus ?? 8,
isl: null,
osl: null,
conc,
offload_mode: 'on',
benchmark_type: 'agentic_traces',
image: `${sku.framework}:test`,
- metrics: metricsFor(
- intvty,
- Math.round(tput * sku.tputScale * tputScaleFor(date, runId)),
- e2el,
- ),
+ metrics: {
+ ...metricsFor(intvty, Math.round(tput * sku.tputScale * tputScaleFor(date, runId)), e2el),
+ ...sku.metrics,
+ },
workers: null,
date,
workflow_run_id: runId ?? PROFIT_SINGLE_RUN_ID[date],
@@ -141,6 +171,10 @@ export const profitBenchmarkRows = (
})),
);
+/** GB200 NVL72 rows with measured compute-module power, for the smart-provisioning basis. */
+export const profitNvl72Rows = (dbKey: string = PROFIT_MODEL_DB_KEY, date = PROFIT_DATE) =>
+ profitBenchmarkRows(dbKey, date, undefined, [NVL72_SKU]);
+
export const profitAvailabilityRows = (dbKeys: readonly string[] = PROFIT_DB_KEYS) =>
dbKeys.flatMap((dbKey) =>
PROFIT_DATES.flatMap((date) =>
From 7a654f4635dc02358f33993e600fe464ca292fd6 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 15:11:51 -0700
Subject: [PATCH 05/22] docs: add the NVL72 rack estimate section to the PowerX
system-power doc
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Measured input and admission rules, the modeled residual table with the
UNVERIFIED parameters and their ranges, the shelf overflow bound, and the
Profit Estimator gate rules, in English and in the 中文说明 section. The
Profit Estimator paragraphs now name the per-hardware PUE policy, the
same-basis knot rule and the CSV provenance columns in both languages.
中文:新增 NVL72 机架估算一节:实测输入与接纳条件、含 UNVERIFIED 参数及范围的
建模残差表、电源架容量上限和利润估算器门槛规则,中英文同步;利润估算器段落
双语补充按硬件取值的 PUE 策略、同口径数据点规则和 CSV 出处列。
---
docs/powerx-system-power.md | 129 +++++++++++++++++++++++++++++++-----
1 file changed, 114 insertions(+), 15 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 1a3399e89..05a4923bd 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -69,8 +69,8 @@ result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`
A partially allocated tray extrapolates only the GPU-board share (a module reading
already covers the whole tray) and is labeled `extrapolated`. The pinned revision
`963ead8b` is the power-model repo's `feat/gb200-nvl72-rack-model` branch, pending
-push upstream. The Profit Estimator gate below still accepts only eight-GPU
-single-node chassis.
+push upstream. The NVL72 section below lists the measured input, the modeled
+residual and the Profit Estimator gate rules.
A partially allocated chassis (one to seven measured GPUs on one host) is
modeled at measured per-GPU power × 8. That is the same `n_gpu × W/GPU` input
@@ -102,6 +102,63 @@ schema and reports `validated-unversioned-single-node`; it does not upgrade the
source or admit unversioned disaggregated power. The article receipt additionally
pins the producer checkout and retains each original audit artifact.
+## NVL72 rack estimate (GB200, GB300)
+
+**Measured input.** Every compute tray is fed a measured compute-module figure; the
+Grace CPU and LPDDR5X are never modeled. The producer's CPU power leg (srt-slurm,
+ACPI hwmon) publishes, over the same formal window as GPU energy,
+`avg_cpu_socket_power_w`, `avg_total_cpu_power_w`, `total_cpu_energy_j`, and, when
+the `Module Power Socket` sensor exists on every socket, `avg_total_module_power_w`
+and `total_module_energy_j`, with the independent verdict `cpu_power_valid` and the
+`power_audit.cpu` block (sensor kind, collector, socket coverage, reason codes).
+Admission requires `power_valid=1`, schema 2, `cpu_power_valid=1`, positive Grace
+watts, and `avg_total_cpu_power_w / avg_cpu_socket_power_w` equal to two sockets per
+tray. Basis selection: `module` when `avg_total_module_power_w` is present (the
+reading already contains the GPU boards, so it is never scaled), otherwise
+`gpu-plus-grace` (GPU-board watts × 4 plus the Grace-socket total per tray, with the
+source's regulator-loss allowance `regulatorLossFracOfTdp / (1 − frac)` on the GPU
+share only). A present-but-invalid module key makes the row unavailable
+(`cpu-telemetry`); it never falls back to the Grace socket silently.
+
+**Modeled residual.** Everything outside the compute modules comes from the pinned
+profile (`rackProfiles`), evaluated for a rack of 18 identical trays and amortised
+over 72 GPUs; the parameters marked UNVERIFIED carry a documented range in
+`unverifiedParameters` and no published rail:
+
+| Component (per rack unless noted) | GB200 | GB300 | Source status |
+| -------------------------------------- | ------------------------------------------------------------------------------------ | ---------------------------------------------- | ---------------------------------------- |
+| NVSwitch tray silicon (9 trays) | 406.4 W / tray at `u_nvlink` 0.5 | same | `blackwell_nvswitch` model |
+| NVSwitch tray residual | 50 W / tray | same | UNVERIFIED (20–80 W) |
+| Compute-tray NICs with optics | ConnectX-7, 4 × 31.5 W = 126 W / tray | ConnectX-8 integrated PCIe, 4 × 78.8 W = 315 W | `generic/connectx7`, `generic/connectx8` |
+| Compute-tray BlueField-3 DPUs | 2 × 65 W idle = 130 W / tray | same | `generic/dpu`, idle only |
+| Compute-tray NVMe | 22 W / tray idle | same | `generic/nvme`, idle only |
+| Compute-tray fans | 130 W / tray | same | UNVERIFIED (40–220 W) |
+| Compute-tray board residual | 40 W / tray | same | UNVERIFIED (20–60 W) |
+| Management switches | 2 × 100 W | same | profile constant |
+| Tray 50 V → 12 V conversion | efficiency 0.9725 on tray loads | same | UNVERIFIED (0.96–0.985) |
+| Regulator allowance (`gpu-plus-grace`) | 15% of GPU TDP, GPU share only | same | Grace tuning guide |
+| Power shelf | 264 kW installed, 132 kW redundant; efficiency 0.90 → 0.94 → 0.965 at 10/20/30% load | same | profile curve |
+| Facility PUE | 1.1 (direct liquid cooling) | same | PowerX policy, applied once to rack AC |
+
+Rack DC above the installed shelf capacity (264 kW) overflows the efficiency curve
+and the row is unavailable (`model-domain`). Rounding follows the source: rack AC is rounded
+to 0.1 W before PUE. Python-generated `rackCases` prove parity with the pinned
+implementation for both variants, both bases, every shelf knot and PUE 1.0–1.2.
+
+**Gate rules (Profit Estimator).** Planning kW/GPU = deployment facility watts ÷
+measured GPUs ÷ 1000 × 1.1. It accepts a `single-node` eight-GPU chassis with
+`chassisBasis: 'full'`, or an `nvl72-trays` estimate whose trays are all fully
+measured (one host per worker, four GPUs and two sockets each). Partial trays are
+extrapolated in the chart but rejected here, as partial chassis are. Between two
+frontier knots both must share the same measured basis and sensor kind; a module
+knot beside a Grace-socket knot stays unavailable rather than blending sensors. The
+bar tooltip, the caption line under the power note and the CSV columns `Power
+basis`, `Power sensor`, `System power profile` name the basis (measured module, or
+measured GPU board + Grace socket with regulator loss modeled), the sensor kind, and
+the pinned profile (`modelPath @ modelRevision sha256:`) per row.
+The `?unofficialrun=` overlay rule does not apply to the Profit Estimator basis
+control: the estimator prices official frontier points only.
+
## Profit Estimator power basis
The per-GW Profit Estimator offers provisioned power, measured + modeled power,
@@ -116,17 +173,21 @@ facility kW/GPU used to calculate capacity per GW. Consequently, revenue,
compute expense, license fee, and profit scale together; profit margin does not
change. Electricity expense is not recomputed separately.
-This opt-in AgentX estimate requires validated schema-v2 telemetry and a complete
-single-node eight-GPU chassis supported by the pinned model. Partial allocations,
-unsupported GB200/GB300 chassis, and missing/invalid measurements stay unavailable.
-The ordinary 8K/1K transformation keeps its existing admission policy.
+This opt-in AgentX estimate requires validated schema-v2 telemetry and either a
+complete single-node eight-GPU chassis supported by the pinned model or NVL72
+compute trays that are all fully measured (`cpu_power_valid=1`, see the NVL72
+section). Partial allocations, NVL72 rows without CPU-side telemetry, and
+missing/invalid measurements stay unavailable. The ordinary 8K/1K transformation
+keeps its existing admission policy.
At an exact frontier point, use that point's modeled power. Between points,
estimate power linearly using the same two knots as the existing throughput
-interpolation; never select a different point to fill a power gap. The estimate
-uses PUE 1.3 and an additional 10% planning margin. These assumptions, including
-the fixed CPU/DRAM utilization above, are not validated peak-load provisioning or
-AgentX system calibration. The UI and CSV label the estimate and its assumptions.
+interpolation; never select a different point to fill a power gap, and never blend
+two knots measured on different bases. The estimate uses PowerX's PUE policy (1.3 for
+air-cooled chassis, 1.1 for DLC NVL72 racks) and an additional 10% planning margin.
+These assumptions, including the fixed CPU/DRAM utilization above, are not validated
+peak-load provisioning or AgentX system calibration. The UI and CSV label the
+estimate, its assumptions, and per row the measured basis, sensor kind and profile.
`c_power=modeled` and `c_power=compare` preserve the selection in share URLs.
Unavailable historical estimates use the hardware registry when a chip is absent
from today's results and include the source date/run label.
@@ -140,12 +201,15 @@ from today's results and include the source date/run label.
`inferencex-feature-gate=1`)。锁定时,`c_power` 不会启用其他估算方式或触发完整功耗
数据请求;重新锁定后立即恢复预配功耗估算。
-AgentX 估算仅接纳通过验证的 schema-v2 功耗,且要求完整的单节点八卡机箱及适用模型。
-部分卡分配、GB200/GB300 等无匹配模型的机箱,以及缺失或无效功耗保持不可用。原有
+AgentX 估算仅接纳通过验证的 schema-v2 功耗,且要求完整的单节点八卡机箱及适用模型,
+或全部 tray 均完整实测(`cpu_power_valid=1`,见下文 NVL72 一节)的 NVL72 计算 tray。
+部分卡分配、缺少 CPU 侧实测的 NVL72 行,以及缺失或无效功耗保持不可用。原有
8K/1K 转换路径的接纳规则不变。精确前沿点使用自身的功耗;点间采用原吞吐量插值的
-同一对数据点线性估算功耗,不换用其他点填补缺失。PUE 取 1.3,另加 10% 功耗余量;
-这些假设和上述固定 CPU/DRAM 利用率尚未通过 AgentX 系统校准,也不构成峰值供电容量
-验证。界面与 CSV 会注明估算及其假设,分享链接通过 `c_power` 保留所选方式。
+同一对数据点线性估算功耗,不换用其他点填补缺失,也不会把两种实测口径不同的数据点
+混合估算。PUE 按 PowerX 策略取值(风冷机箱 1.3,液冷 NVL72 机架 1.1),另加 10%
+功耗余量;这些假设和上述固定 CPU/DRAM 利用率尚未通过 AgentX 系统校准,也不构成
+峰值供电容量验证。界面与 CSV 会注明估算及其假设,并逐行标出实测口径、传感器类型
+和所用 profile;分享链接通过 `c_power` 保留所选方式。
历史估算不可用时,若当天结果不含该芯片,则从硬件注册表获取名称;提示会附上来源
日期或运行标签,避免与当前结果混淆。
@@ -246,6 +310,41 @@ frontend/router 主机不在估算范围内,GPU 机箱内的 CPU 功率仍按
GPU 实测指标。当前 API 快照与原文章冻结数据分别导出,避免混用不同时间和配置的
结果。
+**NVL72 机架估算(GB200、GB300)。** 实测输入:每个计算 tray 使用实测的计算模块功耗,
+Grace CPU 与 LPDDR5X 从不建模。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与
+GPU 能耗相同的正式窗口内输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w`、
+`total_cpu_energy_j`,当每个 socket 都有 `Module Power Socket` 传感器时还输出
+`avg_total_module_power_w` 与 `total_module_energy_j`,并附带独立的验证结论
+`cpu_power_valid` 和 `power_audit.cpu`(传感器类型、采集来源、socket 覆盖情况、原因码)。
+接纳条件:`power_valid=1`、schema 2、`cpu_power_valid=1`、Grace 功耗为正,且
+`avg_total_cpu_power_w / avg_cpu_socket_power_w` 等于每 tray 两个 socket。存在
+`avg_total_module_power_w` 时采用 `module` 口径(读数已包含 GPU 板卡,不再缩放),
+否则采用 `gpu-plus-grace` 口径(每 tray GPU 板卡功耗 × 4 加 Grace socket 总功耗,
+并仅对 GPU 份额计入来源模型的稳压损耗余量)。模块指标存在但无效时该行不可用
+(`cpu-telemetry`),不会悄然回退到 Grace socket。
+
+建模残差:计算模块之外的部分全部来自固定版本 profile(`rackProfiles`),按 18 个
+相同 tray 组成的整机架求值并分摊到 72 张 GPU:NVSwitch tray 硅片功耗(9 个 tray,
+`blackwell_nvswitch` 模型,`u_nvlink` 0.5)及 tray 残差(50 W,UNVERIFIED);每个计算
+tray 的网卡与光模块(GB200 为 ConnectX-7,4 × 31.5 W;GB300 为 ConnectX-8,4 × 78.8 W)、
+BlueField-3 DPU 空闲功耗(2 × 65 W)、NVMe 空闲功耗(22 W)、风扇(130 W,UNVERIFIED)
+和主板残差(40 W,UNVERIFIED);2 台管理交换机(各 100 W);tray 内 50 V → 12 V 转换
+效率 0.9725(UNVERIFIED);电源架装机容量 264 kW、冗余容量 132 kW,效率曲线在 10%/20%/30%
+负载处为 0.90/0.94/0.965;液冷 PUE 1.1,仅对机架交流功率应用一次。机架直流功率超过电源架
+装机容量(264 kW)时该行不可用(`model-domain`)。舍入与来源一致:机架交流功率先四舍五入到 0.1 W
+再乘 PUE;Python 生成的 `rackCases` 对两种变体、两种口径、全部电源架拐点和 PUE 1.0–1.2
+验证了与固定实现的一致性。
+
+门槛规则(利润估算器):规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。接受
+`chassisBasis: 'full'` 的单节点八卡机箱,或全部 tray 均完整实测(每个 worker 一台主机,
+各 4 张 GPU、2 个 socket)的 `nvl72-trays` 估算;部分 tray 在图表中外推显示,但与部分
+机箱一样不进入规划门槛。两个前沿数据点之间必须采用相同的实测口径和传感器类型,模块
+读数旁边的 Grace socket 读数保持不可用,不会混合两种传感器。柱形提示、功耗说明下方的
+标注行和 CSV 的 `功耗口径`、`功耗传感器`、`系统功耗 profile` 三列逐行标出实测口径
+(实测模块功耗,或实测 GPU 板卡 + Grace socket 功耗并由模型估算稳压损耗)、传感器类型
+和所用 profile(`modelPath @ modelRevision sha256:<源文件哈希>`)。`?unofficialrun=`
+叠加层规则不适用于利润估算器的功耗口径控件:估算器只对正式前沿数据点定价。
+
## Measured P75 and P90 GPU power
`y_measuredP75Power` and `y_measuredP90Power` show the time-weighted P75 and P90 of
From 3f0c53e9f52414acd7b59c6defdf89b98e9f6952 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Fri, 18 Sep 2026 16:54:14 -0700
Subject: [PATCH 06/22] fix: scrub CPU-side power on its own verdict and model
NVL72 trays as one rack
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
- ingest withholds the CPU-side keys on cpu_power_valid != 1 and GPU-side keys on power_valid = 0; mapper tests cover all four verdict combinations; the power manifest carries cpu_power_valid
- exporter shares defaultSystemPue with the dashboard (1.3 chassis, 1.1 DLC NVL72); NVL72 rows carry the rack profile, measured basis, sensor kind and CPU-side inputs; x86 rows byte-identical
- chart tooltip names NVL72 compute trays and the measured basis instead of eight-GPU chassis (en/zh)
- heterogeneous measured trays fold into one rack at their mean; the shelf curve is evaluated once at rack DC load, matching gb200_nvl72_rack_power; estimateRackPower and parity cases unchanged
中文:摄取按 cpu_power_valid 独立清除 CPU 侧指标,GPU 侧仍按 power_valid,mapper 测试覆盖四种组合,功耗清单附带 cpu_power_valid;导出器与仪表板共用 defaultSystemPue(机箱 1.3、液冷 NVL72 1.1),NVL72 行补充机架 profile、实测口径、传感器类型与 CPU 侧输入,x86 行逐字节不变;图表提示改用 NVL72 计算 tray 措辞并标出实测口径(中英文);多 tray 按均值折算为整机架,电源架曲线只在机架直流负载处求值一次,与 Python 模型一致,estimateRackPower 与对照用例不变。
---
docs/powerx-system-power.md | 51 +++++---
.../scripts/export-modeled-system-power.ts | 108 +++++++++++++---
.../inference/utils/tooltip-utils.test.ts | 50 ++++++++
.../inference/utils/tooltipUtils.ts | 57 ++++++++-
packages/app/src/lib/api-documentation.ts | 6 +-
.../lib/modeled-system-power-export.test.ts | 91 +++++++++++++-
.../app/src/lib/modeled-system-power.test.ts | 46 ++++---
packages/app/src/lib/modeled-system-power.ts | 54 ++++++--
packages/constants/src/metric-keys.test.ts | 22 ++++
packages/constants/src/metric-keys.ts | 53 +++++---
packages/db/src/etl/benchmark-mapper.test.ts | 117 ++++++++++++++----
packages/db/src/etl/benchmark-mapper.ts | 22 +++-
packages/db/src/etl/power-publication.ts | 8 +-
13 files changed, 569 insertions(+), 116 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 05a4923bd..c30ccd993 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -62,9 +62,13 @@ producer publishes it, otherwise GPU-board watts plus the Grace-socket total
(`avg_total_cpu_power_w`) with the source's regulator-loss allowance on the GPU
share. The Grace CPU and LPDDR5X are never modelled; rows without
`cpu_power_valid=1` and the Grace-side keys stay unavailable (`cpu-telemetry`).
-Each measured worker host is one compute tray (four GPUs, two Grace sockets), and
-a tray's estimate is its 1/18 share of a rack of identical trays, so NVSwitch
-trays, power shelves, and management switches are amortised over 72 GPUs. The
+Each measured worker host is one compute tray (four GPUs, two Grace sockets). The
+measured trays are folded into one rack of 18 trays matching their mean
+compute-module input, the power-shelf efficiency curve is evaluated once at that
+rack's DC load (as the source `gb200_nvl72_rack_power` does with its single
+per-tray input), and every tray takes the same 1/18 share, so NVSwitch trays,
+power shelves, and management switches are amortised over 72 GPUs. Chassis, by
+contrast, own their fans and PSUs and are each evaluated at their own load. The
result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`.
A partially allocated tray extrapolates only the GPU-board share (a module reading
already covers the whole tray) and is labeled `extrapolated`. The pinned revision
@@ -121,9 +125,10 @@ share only). A present-but-invalid module key makes the row unavailable
(`cpu-telemetry`); it never falls back to the Grace socket silently.
**Modeled residual.** Everything outside the compute modules comes from the pinned
-profile (`rackProfiles`), evaluated for a rack of 18 identical trays and amortised
-over 72 GPUs; the parameters marked UNVERIFIED carry a documented range in
-`unverifiedParameters` and no published rail:
+profile (`rackProfiles`), evaluated once for a rack of 18 trays at the measured
+trays' mean input (the shelf curve sees the whole rack's DC load, never one tray's)
+and amortised over 72 GPUs; the parameters marked UNVERIFIED carry a documented
+range in `unverifiedParameters` and no published rail:
| Component (per rack unless noted) | GB200 | GB300 | Source status |
| -------------------------------------- | ------------------------------------------------------------------------------------ | ---------------------------------------------- | ---------------------------------------- |
@@ -227,6 +232,16 @@ bun packages/app/scripts/export-modeled-system-power.ts \
--output /path/to/new-current-qwen-comparison --pue 1.3
```
+Without `--pue`, each row uses the dashboard's default for its hardware (`1.3` for
+the air-cooled chassis profiles, `1.1` for the DLC NVL72 rack profiles) through the
+same `modelSystemPower` path the chart uses, so article figures match chart hovers;
+`metadata.pue_override` records an explicit `--pue`, which then applies to every
+row, and `metadata.pue_defaults` records the per-hardware defaults. NVL72 rows also
+carry `measured_basis`, `sensor_kind`, the Grace-socket and module measured inputs
+under their own `cpu_power_valid`, the rack profile's `model_path` and assumptions,
+and rack-specific `calculation_boundary` / `extrapolation_note` text; x86 rows are
+unchanged.
+
The maintained input shape is `ComparisonInput` in the script:
```ts
@@ -299,16 +314,23 @@ worker 主机视为一个计算 tray(4 张 GPU、2 个 Grace socket)。输
(`avg_total_module_power_w`);缺失时改用 GPU 板卡功耗加 Grace socket 功耗
(`avg_total_cpu_power_w`),并按来源模型计入 GPU 份额的稳压损耗余量。Grace CPU 与
LPDDR5X 从不建模,缺少 `cpu_power_valid=1` 和 Grace 侧指标的行保持不可用
-(`cpu-telemetry`)。单个 tray 的估算取由相同 tray 组成的整机架的 1/18,因此 NVSwitch
-tray、电源架和管理交换机按 72 张 GPU 分摊;部分分配的 tray 只外推 GPU 板卡份额
-(模块读数本身已覆盖整个 tray),并标记为 `extrapolated`。缺失、无效和不支持的
-情况保持不可用。纯 CPU frontend worker 不计入 GPU 机箱数;独立的纯 CPU
-frontend/router 主机不在估算范围内,GPU 机箱内的 CPU 功率仍按 20% 利用率计算。
+(`cpu-telemetry`)。实测的各 tray 先折算为一个由 18 个与其均值相同的 tray 组成的
+整机架,电源架效率曲线只在该机架的直流总负载处求值一次(与来源模型
+`gb200_nvl72_rack_power` 只接受单一 per-tray 输入的做法一致),每个 tray 取其 1/18,
+因此 NVSwitch tray、电源架和管理交换机按 72 张 GPU 分摊;机箱则各自拥有风扇和 PSU,
+仍按各自负载单独求值。部分分配的 tray 只外推 GPU 板卡份额(模块读数本身已覆盖整个
+tray),并标记为 `extrapolated`。缺失、无效和不支持的情况保持不可用。纯 CPU frontend
+worker 不计入 GPU 机箱数;独立的纯 CPU frontend/router 主机不在估算范围内,GPU 机箱内
+的 CPU 功率仍按 20% 利用率计算。
导出时每次测量先独立计算,再对同一 cell 的重复测量取平均。能耗使用审计记录中的
实际窗口和成功 token 数,按实测 GPU 的份额计算,明确标记为估计值,不改写原有
GPU 实测指标。当前 API 快照与原文章冻结数据分别导出,避免混用不同时间和配置的
-结果。
+结果。未指定 `--pue` 时,每行按其硬件采用与仪表板相同的默认 PUE(风冷机箱 1.3,
+液冷 NVL72 机架 1.1),因此文章数据与图表悬停一致;显式 `--pue` 记录在
+`metadata.pue_override` 中并覆盖所有行。NVL72 行另附实测口径、传感器类型、
+按 `cpu_power_valid` 保留的 Grace socket 与模块实测输入、机架 profile 及其假设,
+以及机架专用的边界与外推说明;x86 行保持不变。
**NVL72 机架估算(GB200、GB300)。** 实测输入:每个计算 tray 使用实测的计算模块功耗,
Grace CPU 与 LPDDR5X 从不建模。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与
@@ -323,8 +345,9 @@ GPU 能耗相同的正式窗口内输出 `avg_cpu_socket_power_w`、`avg_total_c
并仅对 GPU 份额计入来源模型的稳压损耗余量)。模块指标存在但无效时该行不可用
(`cpu-telemetry`),不会悄然回退到 Grace socket。
-建模残差:计算模块之外的部分全部来自固定版本 profile(`rackProfiles`),按 18 个
-相同 tray 组成的整机架求值并分摊到 72 张 GPU:NVSwitch tray 硅片功耗(9 个 tray,
+建模残差:计算模块之外的部分全部来自固定版本 profile(`rackProfiles`),按实测 tray
+均值构成的 18 tray 整机架求值一次(电源架曲线看到的是整机架直流负载,而非单个
+tray),再分摊到 72 张 GPU:NVSwitch tray 硅片功耗(9 个 tray,
`blackwell_nvswitch` 模型,`u_nvlink` 0.5)及 tray 残差(50 W,UNVERIFIED);每个计算
tray 的网卡与光模块(GB200 为 ConnectX-7,4 × 31.5 W;GB300 为 ConnectX-8,4 × 78.8 W)、
BlueField-3 DPU 空闲功耗(2 × 65 W)、NVMe 空闲功耗(22 W)、风扇(130 W,UNVERIFIED)
diff --git a/packages/app/scripts/export-modeled-system-power.ts b/packages/app/scripts/export-modeled-system-power.ts
index 6e3a59a40..b31ee237b 100644
--- a/packages/app/scripts/export-modeled-system-power.ts
+++ b/packages/app/scripts/export-modeled-system-power.ts
@@ -6,7 +6,12 @@ import { pathToFileURL } from 'node:url';
import { parseArgs } from 'node:util';
import type { BenchmarkRow } from '../src/lib/api';
-import { AIR_COOLED_SYSTEM_PUE, modelSystemPower } from '../src/lib/modeled-system-power';
+import {
+ AIR_COOLED_SYSTEM_PUE,
+ DLC_SYSTEM_PUE,
+ defaultSystemPue,
+ modelSystemPower,
+} from '../src/lib/modeled-system-power';
import profileData from '../src/lib/system-power-model.profiles.json';
interface PowerAudit {
@@ -48,6 +53,34 @@ const mean = (values: (number | null)[]) =>
? values.reduce((sum, value) => sum + value, 0) / values.length
: null;
+interface SystemPowerProfile {
+ assumptions: Record;
+ modelPath: string;
+}
+const CHASSIS_PROFILES: Record = profileData.profiles;
+const RACK_PROFILES: Record = profileData.rackProfiles;
+const isRackHardware = (hardware: string) => Object.hasOwn(RACK_PROFILES, hardware);
+const profileFor = (hardware: string): SystemPowerProfile | null =>
+ isRackHardware(hardware)
+ ? RACK_PROFILES[hardware]
+ : Object.hasOwn(CHASSIS_PROFILES, hardware)
+ ? CHASSIS_PROFILES[hardware]
+ : null;
+
+// The chassis notes are unchanged for x86 rows; NVL72 rows carry their own.
+const CHASSIS_NOTES = {
+ boundary:
+ 'Measured GPU-board inputs; modeled GPU-chassis AC includes their CPU/DRAM, other model components, and PSU loss. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after GPU-chassis AC.',
+ extrapolation:
+ 'A partially allocated chassis is modeled at measured per-GPU power × 8 (the source sweep input), assuming the unmeasured GPUs run the same workload. Deployment values are the measured GPUs’ share of that chassis; per-GPU values divide by the modeled chassis GPU count.',
+};
+const RACK_NOTES = {
+ boundary:
+ 'Measured compute-module input per NVL72 tray (module sensor, or GPU board + Grace socket with the source’s regulator-loss allowance); the Grace CPU and LPDDR5X are never modeled. Modeled rack residual: NVSwitch trays, NICs/DPUs, NVMe, tray fans and board, tray 50 V → 12 V conversion, power shelves and management switches, evaluated once for a rack of 18 trays at the measured trays’ mean input and amortised over 72 GPUs. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after rack AC.',
+ extrapolation:
+ 'A partially allocated tray is modeled at measured per-GPU power × 4 on the GPU-board share only, assuming the unmeasured GPUs run the same workload; a module reading already covers the whole tray and is never scaled. Deployment values are the measured GPUs’ share of that tray; per-GPU values divide by the modeled tray GPU count.',
+};
+
function estimatedEnergy(
row: BenchmarkRow,
modeled: ReturnType,
@@ -119,13 +152,19 @@ function estimatedEnergy(
};
}
-export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_PUE) {
+/**
+ * `pue` overrides every row; without it each row takes the dashboard's default for
+ * its hardware (1.3 air-cooled chassis, 1.1 DLC NVL72 rack) so article figures match
+ * chart hovers.
+ */
+export function buildComparison(input: ComparisonInput, pue?: number) {
if (!input || typeof input.cohort !== 'string' || !Array.isArray(input.rows)) {
throw new Error(
'Expected a cohort envelope with a rows array. See docs/powerx-system-power.md.',
);
}
- if (!finite(pue) || pue < 1) throw new Error('PUE must be a finite number >= 1.');
+ if (pue !== undefined && (!finite(pue) || pue < 1))
+ throw new Error('PUE must be a finite number >= 1.');
const ids = new Set();
const rows = input.rows.map((entry) => {
const row = entry.benchmark;
@@ -141,16 +180,22 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_
throw new Error(`Invalid benchmark input or duplicate id: ${entry.id}`);
}
ids.add(entry.id);
+ const hardware = row.hardware.toLowerCase();
+ const rack = isRackHardware(hardware);
+ // Same estimate path and PUE selection as the dashboard (`modelSystemPower(row)`).
const modeled = modelSystemPower(row, pue);
- const profile = Object.entries(profileData.profiles).find(
- ([key]) => key === row.hardware.toLowerCase(),
- )?.[1];
+ const rowPue = pue ?? defaultSystemPue(hardware);
+ const profile = profileFor(hardware);
const measurementStatus =
row.metrics.power_valid === 1
? 'producer-valid'
: row.metrics.power_valid === 0
? 'invalid'
: 'unverified';
+ // The CPU-side keys carry their own verdict; only NVL72 rows report them.
+ const cpuValid = row.metrics.cpu_power_valid === 1;
+ const trayEstimate =
+ modeled.status === 'supported' && modeled.topologyBasis === 'nvl72-trays' ? modeled : null;
return {
id: entry.id,
cell: entry.cell ?? null,
@@ -165,10 +210,28 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_
total_gpu_w: measurement(row.metrics.avg_total_gpu_power_w),
total_gpu_j: measurement(row.metrics.total_gpu_energy_j),
gpu_j_per_output_token: measurement(row.metrics.joules_per_output_token),
+ ...(rack
+ ? {
+ cpu_power_valid: row.metrics.cpu_power_valid ?? null,
+ total_grace_w: cpuValid ? measurement(row.metrics.avg_total_cpu_power_w) : null,
+ total_grace_j: cpuValid ? measurement(row.metrics.total_cpu_energy_j) : null,
+ total_module_w: cpuValid
+ ? measurement(row.metrics.avg_total_module_power_w)
+ : null,
+ total_module_j: cpuValid
+ ? measurement(row.metrics.total_module_energy_j)
+ : null,
+ }
+ : {}),
}
: null,
- assumptions: profile ? { ...profile.assumptions, pue } : null,
+ pue: rowPue,
+ measured_basis: trayEstimate?.measuredBasis ?? null,
+ sensor_kind: trayEstimate?.sensorKind ?? null,
+ assumptions: profile ? { ...profile.assumptions, pue: rowPue } : null,
model_path: profile?.modelPath ?? null,
+ calculation_boundary: rack ? RACK_NOTES.boundary : CHASSIS_NOTES.boundary,
+ extrapolation_note: rack ? RACK_NOTES.extrapolation : CHASSIS_NOTES.extrapolation,
modeled,
estimated_energy: estimatedEnergy(row, modeled, entry.audit),
audit: entry.audit ?? null,
@@ -214,7 +277,11 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_
model_revision: profileData.modelRevision,
model_path: replicates[0].model_path,
assumptions: replicates[0].assumptions,
- pue,
+ // Replicates share hardware, so they share the PUE selection.
+ pue: replicates[0].pue,
+ measured_bases: [
+ ...new Set(replicates.flatMap((row) => (row.measured_basis ? [row.measured_basis] : []))),
+ ],
status: complete ? 'supported' : 'unsupported',
unsupported_reasons: [
...new Set(
@@ -276,16 +343,18 @@ export function buildComparison(input: ComparisonInput, pue = AIR_COOLED_SYSTEM_
metadata: {
cohort: input.cohort,
source: input.metadata,
- pue,
+ // Explicit --pue applies to every row; otherwise each row records its own default.
+ pue_override: pue ?? null,
+ pue_defaults: { air_cooled_chassis: AIR_COOLED_SYSTEM_PUE, dlc_nvl72_rack: DLC_SYSTEM_PUE },
scope: { benchmark_type: 'single_turn', isl: 8192, osl: 1024 },
selection:
'Every supplied row is retained, including unsupported, invalid, and missing-input cases.',
aggregation:
'Each replicate is modeled first. Cell means include every replicate; any unavailable value leaves its cell mean unavailable.',
- boundary:
- 'Measured GPU-board inputs; modeled GPU-chassis AC includes their CPU/DRAM, other model components, and PSU loss. Separate CPU-only frontend/router hosts are excluded. Facility power applies PUE after GPU-chassis AC.',
- extrapolation:
- 'A partially allocated chassis is modeled at measured per-GPU power × 8 (the source sweep input), assuming the unmeasured GPUs run the same workload. Deployment values are the measured GPUs’ share of that chassis; per-GPU values divide by the modeled chassis GPU count.',
+ boundary: CHASSIS_NOTES.boundary,
+ extrapolation: CHASSIS_NOTES.extrapolation,
+ rack_boundary: RACK_NOTES.boundary,
+ rack_extrapolation: RACK_NOTES.extrapolation,
energy_caveat:
'Energy from modeled average power is an estimate. Nonlinear fan/PSU behavior is not integrated over time. Energy requires an exact matching audit window and successful token counts.',
model: profileData,
@@ -321,7 +390,7 @@ async function main() {
});
if (!values.input || !values.output)
throw new Error(
- 'Usage: bun packages/app/scripts/export-modeled-system-power.ts --input cohort.json --output NEW_DIRECTORY [--pue 1.3]',
+ 'Usage: bun packages/app/scripts/export-modeled-system-power.ts --input cohort.json --output NEW_DIRECTORY [--pue 1.3]. Without --pue each row uses the dashboard default for its hardware (1.3 air-cooled chassis, 1.1 DLC NVL72 rack).',
);
const inputBytes = await readFile(values.input);
const result = buildComparison(
@@ -377,6 +446,11 @@ async function main() {
measured_total_gpu_w: row.measured_inputs?.total_gpu_w,
measured_total_gpu_j: row.measured_inputs?.total_gpu_j,
measured_gpu_j_per_output_token: row.measured_inputs?.gpu_j_per_output_token,
+ cpu_power_valid: row.measured_inputs?.cpu_power_valid,
+ measured_total_grace_w: row.measured_inputs?.total_grace_w,
+ measured_total_module_w: row.measured_inputs?.total_module_w,
+ measured_basis: row.measured_basis,
+ sensor_kind: row.sensor_kind,
modeled_status: row.modeled.status,
unsupported_reason: row.modeled.status === 'unsupported' ? row.modeled.reason : null,
modeled_chassis_ac_w: row.modeled.status === 'supported' ? row.modeled.chassisAcWatts : null,
@@ -395,11 +469,11 @@ async function main() {
topology_basis: row.modeled.status === 'supported' ? row.modeled.topologyBasis : null,
model_revision: row.modeled.modelRevision,
model_status: profileData.status,
- calculation_boundary: metadata.boundary,
- extrapolation_note: metadata.extrapolation,
+ calculation_boundary: row.calculation_boundary,
+ extrapolation_note: row.extrapolation_note,
energy_caveat: metadata.energy_caveat,
model_path: row.model_path,
- pue: metadata.pue,
+ pue: row.pue,
assumptions: row.assumptions,
estimated_energy: row.estimated_energy,
estimated_energy_status: row.estimated_energy.status,
diff --git a/packages/app/src/components/inference/utils/tooltip-utils.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
index dd25ace6a..203804231 100644
--- a/packages/app/src/components/inference/utils/tooltip-utils.test.ts
+++ b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
@@ -185,6 +185,56 @@ describe('modeled system-power tooltip', () => {
expect(zh).not.toContain('6000 W');
});
+ it('names NVL72 compute trays and the measured basis instead of eight-GPU chassis', () => {
+ const trays = {
+ ...systemPower,
+ hardware: 'gb200',
+ modelPath: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ gpuCount: 8,
+ chassisCount: 2,
+ modeledGpuCount: 8,
+ pue: 1.1,
+ topologyBasis: 'nvl72-trays',
+ measuredBasis: 'module',
+ sensorKind: 'module',
+ } satisfies SystemPowerEstimate;
+ const html = generateTooltipContent(config({ data: pt({ modeledSystemPower: trays }) }));
+ expect(html).toContain('2 full NVL72 compute trays · 8 GPUs');
+ expect(html).toContain('Measured: module sensor (GPU + HBM + Grace + LPDDR5X)');
+ expect(html).toContain('Rack AC is divided by all 72 GPUs');
+ expect(html).toContain('Grace CPU and LPDDR5X are measured');
+ expect(html).toContain('PUE 1.1');
+ expect(html).not.toContain('eight-GPU chassis');
+ expect(html).not.toContain('CPU/DRAM utilization');
+ expect(html).not.toContain('Includes GPU chassis CPUs');
+
+ const partial = pt({
+ physicalChips: 3,
+ modeledSystemPower: {
+ ...trays,
+ gpuCount: 3,
+ chassisCount: 1,
+ modeledGpuCount: 4,
+ chassisBasis: 'extrapolated',
+ measuredBasis: 'gpu-plus-grace',
+ sensorKind: 'grace-socket',
+ },
+ });
+ const en = generateTooltipContent(config({ data: partial }));
+ expect(en).toContain('1 NVL72 compute tray · 3 of 4 GPUs measured, extrapolated to full tray');
+ expect(en).toContain('Unmeasured tray GPUs are assumed to run the same workload');
+ expect(en).toContain('Measured: GPU board + Grace socket. Modeled: regulator loss');
+ expect(en).not.toContain('Unmeasured chassis GPUs');
+
+ const zh = generateTooltipContent(config({ data: partial, locale: 'zh' }));
+ expect(zh).toContain('1 个 NVL72 计算 tray · 实测 3/4 张 GPU,按满 tray 外推');
+ expect(zh).toContain('假设 tray 内未实测的 GPU 运行相同负载');
+ expect(zh).toContain('实测:GPU 板卡 + Grace socket');
+ expect(zh).toContain('Grace CPU 与 LPDDR5X 为实测值');
+ expect(zh).not.toContain('八卡机箱');
+ expect(zh).not.toContain('CPU/DRAM 利用率');
+ });
+
it('preserves the same model provenance in unofficial and date-comparison tooltips', () => {
const official = config();
const overlay = generateOverlayTooltipContent({
diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts
index d3fecb9fb..b005ff6a6 100644
--- a/packages/app/src/components/inference/utils/tooltipUtils.ts
+++ b/packages/app/src/components/inference/utils/tooltipUtils.ts
@@ -7,7 +7,10 @@ import { frameworkFamily } from '@/lib/framework-family';
import type { Locale } from '@/lib/i18n';
import { isKvOffloadEnabled } from '@/lib/kv-offload';
import { chipCounts } from '@/lib/chip-counts';
-import type { SystemPowerUnsupportedReason } from '@/lib/modeled-system-power';
+import type {
+ SystemPowerSensorKind,
+ SystemPowerUnsupportedReason,
+} from '@/lib/modeled-system-power';
import type { HardwareConfig, InferenceData, OverlayData } from '@/components/inference/types';
import {
@@ -230,6 +233,24 @@ const SYSTEM_POWER_STRINGS = {
'Unmeasured chassis GPUs are assumed to run the same workload at the measured per-GPU power; deployment values are the measured GPUs’ share.',
normalization: 'AC power is divided by all modeled chassis GPUs, including prefill and decode.',
boundary: 'Includes GPU chassis CPUs; excludes separate CPU-only frontend/router hosts.',
+ // NVL72 compute trays: the compute module is measured, the rack residual modeled.
+ trayTopology: (trays: number, measured: number, modeled: number) => {
+ const unit = trays === 1 ? 'tray' : 'trays';
+ return measured === modeled
+ ? `${trays} full NVL72 compute ${unit} · ${measured} GPUs`
+ : `${trays} NVL72 compute ${unit} · ${measured} of ${modeled} GPUs measured, extrapolated to full ${unit}`;
+ },
+ trayExtrapolation:
+ 'Unmeasured tray GPUs are assumed to run the same workload at the measured per-GPU power; a module reading already covers the whole tray. Deployment values are the measured GPUs’ share.',
+ trayAssumptions: {
+ module:
+ 'Measured: module sensor (GPU + HBM + Grace + LPDDR5X). Modeled: NVSwitch trays, NICs/DPUs, NVMe, power shelves.',
+ 'grace-socket':
+ 'Measured: GPU board + Grace socket. Modeled: regulator loss, NVSwitch trays, NICs/DPUs, NVMe, power shelves.',
+ } satisfies Record,
+ trayPlatformAssumptions: 'NVIDIA NVLink: 50%, IB: 0%; PCIe: 5%.',
+ trayNormalization: 'Rack AC is divided by all 72 GPUs of a rack of matching trays.',
+ trayBoundary: 'Grace CPU and LPDDR5X are measured; excludes CPU-only frontend/router hosts.',
model: 'Power model source',
unavailable: 'System-power estimate unavailable',
reasons: {
@@ -261,6 +282,21 @@ const SYSTEM_POWER_STRINGS = {
'假设机箱内未实测的 GPU 运行相同负载、功耗与实测每卡功耗相同;部署数值为实测 GPU 所占份额。',
normalization: '交流功耗按所有建模机箱的 GPU 总数分摊,包括 Prefill 与 Decode。',
boundary: '计入 GPU 机箱内的 CPU;不计入独立的纯 CPU 前端或路由主机。',
+ trayTopology: (trays: number, measured: number, modeled: number) =>
+ measured === modeled
+ ? `${trays} 个完整 NVL72 计算 tray · ${measured} 张 GPU`
+ : `${trays} 个 NVL72 计算 tray · 实测 ${measured}/${modeled} 张 GPU,按满 tray 外推`,
+ trayExtrapolation:
+ '假设 tray 内未实测的 GPU 运行相同负载、功耗与实测每卡功耗相同;模块读数本身已覆盖整个 tray。部署数值为实测 GPU 所占份额。',
+ trayAssumptions: {
+ module:
+ '实测:模块传感器(GPU + HBM + Grace + LPDDR5X)。建模:NVSwitch tray、网卡/DPU、NVMe、电源架。',
+ 'grace-socket':
+ '实测:GPU 板卡 + Grace socket。建模:稳压损耗、NVSwitch tray、网卡/DPU、NVMe、电源架。',
+ } satisfies Record,
+ trayPlatformAssumptions: 'NVIDIA NVLink:50%,IB:0%;PCIe:5%。',
+ trayNormalization: '机架交流功耗按由相同 tray 组成的整机架的 72 张 GPU 分摊。',
+ trayBoundary: 'Grace CPU 与 LPDDR5X 为实测值;不计入独立的纯 CPU 前端或路由主机。',
model: '功耗模型来源',
unavailable: '无法估算系统功耗',
reasons: {
@@ -297,6 +333,23 @@ const modeledSystemPowerHTML = (
}
const sourceUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/${estimate.modelPath}`;
const readmeUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/README.md`;
+ // Tray estimates measure the compute module; chassis estimates model the CPU/DRAM.
+ const tray = estimate.topologyBasis === 'nvl72-trays' ? estimate : null;
+ const topology = tray
+ ? t.trayTopology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount)
+ : t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount);
+ const extrapolation =
+ estimate.chassisBasis === 'extrapolated'
+ ? ` ${tray ? t.trayExtrapolation : t.extrapolation}`
+ : '';
+ const notes = tray
+ ? [
+ t.trayAssumptions[tray.sensorKind],
+ t.trayPlatformAssumptions,
+ t.trayNormalization,
+ t.trayBoundary,
+ ]
+ : [t.assumptions, t.platformAssumptions, t.normalization, t.boundary];
return `
${t.heading}
${tooltipLine(t.measuredGpu, `${fmt(estimate.measuredGpuWattsPerGpu)} W/GPU`)}
@@ -306,7 +359,7 @@ const modeledSystemPowerHTML = (
? `
${tooltipLine(t.deploymentAc, `${fmt(estimate.deploymentAcWatts)} W`)}
${tooltipLine(`${t.facility} (PUE ${fmt(estimate.pue)})`, `${fmt(estimate.deploymentFacilityWatts)} W`)}
-
${t.topology(estimate.chassisCount, estimate.gpuCount, estimate.modeledGpuCount)}${estimate.chassisBasis === 'extrapolated' ? ` ${t.extrapolation}` : ''} ${t.assumptions} ${t.platformAssumptions} ${t.normalization} ${t.boundary}
+
${topology}${extrapolation} ${notes.join(' ')}
${tooltipLine(t.model, `
${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.slice(0, 12))} `)}
${t.sweep}
`
diff --git a/packages/app/src/lib/api-documentation.ts b/packages/app/src/lib/api-documentation.ts
index cd05964c2..61b89b64d 100644
--- a/packages/app/src/lib/api-documentation.ts
+++ b/packages/app/src/lib/api-documentation.ts
@@ -239,7 +239,7 @@ const powerMetricDescriptions: Readonly
{
const source = input();
const before = structuredClone(source);
const result = buildComparison(source);
- expect(result.metadata.pue).toBe(1.3);
+ expect(result.metadata.pue_override).toBeNull();
+ expect(result.metadata.pue_defaults).toEqual({ air_cooled_chassis: 1.3, dlc_nvl72_rack: 1.1 });
expect(result.metadata.model.assumptions.pue).toBe(1.2);
+ expect(result.rows[0].pue).toBe(1.3);
expect(result.rows[0].modeled).toMatchObject({ pue: 1.3 });
expect(result.rows[0].assumptions).toMatchObject({ pue: 1.3 });
- expect(buildComparison(source, 1.1).rows[0].modeled).toMatchObject({ pue: 1.1 });
+ expect(result.cells[0].pue).toBe(1.3);
+ const overridden = buildComparison(source, 1.1);
+ expect(overridden.metadata.pue_override).toBe(1.1);
+ expect(overridden.rows[0]).toMatchObject({ pue: 1.1, modeled: { pue: 1.1 } });
expect(result.rows[0].estimated_energy).toMatchObject({
status: 'estimated',
output_tokens: 9303,
@@ -216,6 +221,86 @@ describe('offline modeled PowerX comparisons', () => {
expect(() => buildComparison(source)).toThrow('different benchmark configurations');
});
+ it('routes GB200 NVL72 rows through the tray estimate with the DLC PUE and leaves x86 rows unchanged', () => {
+ const source = input();
+ const baseline = buildComparison(structuredClone(source));
+ // One GB200 compute tray on the module basis, as the dashboard would receive it.
+ const gb200 = structuredClone(source.rows[0]);
+ gb200.id = 'gb200:tray';
+ gb200.cell = 'gb200:c1';
+ gb200.audit = undefined;
+ Object.assign(gb200.benchmark, {
+ hardware: 'gb200',
+ framework: 'dynamo-trt',
+ prefill_tp: 4,
+ decode_tp: 4,
+ num_prefill_gpu: 4,
+ num_decode_gpu: 4,
+ metrics: {
+ pp: 1,
+ pcp_size: 1,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ cpu_power_valid: 1,
+ avg_power_w: 900.25,
+ avg_total_gpu_power_w: 3601,
+ avg_cpu_socket_power_w: 250.5,
+ avg_total_cpu_power_w: 501,
+ avg_total_module_power_w: 4300.75,
+ total_module_energy_j: 258045,
+ },
+ });
+ source.rows.push(gb200);
+ const result = buildComparison(source);
+ const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!;
+ expect(result.rows[1]).toMatchObject({
+ pue: 1.1,
+ measured_basis: 'module',
+ sensor_kind: 'module',
+ model_path: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ assumptions: { u_nvlink: 0.5, pue: 1.1 },
+ measured_inputs: {
+ avg_gpu_w: 900.25,
+ cpu_power_valid: 1,
+ total_grace_w: 501,
+ total_module_w: 4300.75,
+ total_module_j: 258045,
+ },
+ modeled: {
+ status: 'supported',
+ topologyBasis: 'nvl72-trays',
+ pue: 1.1,
+ chassisAcWatts: rack.rackAcWatts / 18,
+ facilityWatts: rack.facilityWatts / 18,
+ },
+ });
+ expect(result.rows[1].calculation_boundary).toContain('NVL72');
+ expect(result.rows[1].extrapolation_note).toContain('tray');
+ expect(result.cells[1]).toMatchObject({
+ cell: 'gb200:c1',
+ pue: 1.1,
+ measured_bases: ['module'],
+ modeled_chassis_ac_w_mean: rack.rackAcWatts / 18,
+ });
+ // The x86 row and its cell are byte-identical to an export without the NVL72 row.
+ expect(result.rows[0]).toEqual(baseline.rows[0]);
+ expect(result.cells[0]).toEqual(baseline.cells[0]);
+ expect(result.rows[0].measured_inputs).not.toHaveProperty('total_module_w');
+ expect(result.rows[0]).toMatchObject({ measured_basis: null, sensor_kind: null });
+ // An explicit --pue still overrides every row, rack and chassis alike.
+ const overridden = buildComparison(source, 1.3);
+ expect(overridden.rows.map((row) => row.pue)).toEqual([1.3, 1.3]);
+ expect(overridden.rows[1].modeled).toMatchObject({ pue: 1.3, topologyBasis: 'nvl72-trays' });
+ // Without cpu_power_valid the row stays unavailable and reports no CPU-side inputs.
+ delete gb200.benchmark.metrics.cpu_power_valid;
+ const unavailable = buildComparison(source).rows[1];
+ expect(unavailable.modeled).toMatchObject({ status: 'unsupported', reason: 'cpu-telemetry' });
+ expect(unavailable.measured_inputs).toMatchObject({
+ cpu_power_valid: null,
+ total_module_w: null,
+ });
+ });
+
it('retains unsupported hardware and missing values, and escapes CSV text', () => {
const source = input();
source.rows[0].benchmark.hardware = 'H200';
diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts
index b89d7e590..57c22e5b1 100644
--- a/packages/app/src/lib/modeled-system-power.test.ts
+++ b/packages/app/src/lib/modeled-system-power.test.ts
@@ -550,12 +550,14 @@ describe('NVL72 trays with measured compute-module power', () => {
});
it('falls back to GPU board plus Grace socket for GB300 trays without module keys', () => {
+ // Low-load prefill and decode trays: the rack DC of their mean sits inside the
+ // shelf curve's nonlinear 20-30% band, where per-tray evaluation would differ.
const source = nvl72Row(
{
- avg_power_w: 900,
- avg_total_gpu_power_w: 7200,
- prefill_avg_power_w: 950,
- decode_avg_power_w: 850,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ prefill_avg_power_w: 250,
+ decode_avg_power_w: 750,
avg_cpu_socket_power_w: 260,
avg_total_cpu_power_w: 1040,
},
@@ -564,20 +566,18 @@ describe('NVL72 trays with measured compute-module power', () => {
disagg: true,
is_multinode: true,
workers: [
- { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 950 },
- { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 850 },
+ { role: 'prefill', worker_idx: 0, num_gpus: 4, hosts: ['tray-a'], avg_power_w: 250 },
+ { role: 'decode', worker_idx: 0, num_gpus: 4, hosts: ['tray-b'], avg_power_w: 750 },
],
},
);
- // Grace-side watts are a deployment total; each tray receives the two-socket mean.
- const prefill = estimateRackPower(
+ // The source model takes one compute-module figure per tray and evaluates the
+ // power-shelf curve once at rack DC load, so both trays are folded into one rack
+ // at their mean GPU-board watts; Grace-side watts are a deployment total, so each
+ // tray receives the two-socket mean.
+ const meanRack = estimateRackPower(
'gb300',
- { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 3800, graceSocketWattsPerTray: 520 },
- 1.1,
- )!;
- const decode = estimateRackPower(
- 'gb300',
- { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 3400, graceSocketWattsPerTray: 520 },
+ { basis: 'gpu-plus-grace', gpuBoardWattsPerTray: 2000, graceSocketWattsPerTray: 520 },
1.1,
)!;
expect(modelSystemPower(source)).toMatchObject({
@@ -586,15 +586,27 @@ describe('NVL72 trays with measured compute-module power', () => {
gpuCount: 8,
chassisCount: 2,
modeledGpuCount: 8,
- chassisAcWatts: prefill.rackAcWatts / 18 + decode.rackAcWatts / 18,
- facilityWatts: prefill.facilityWatts / 18 + decode.facilityWatts / 18,
- deploymentAcWatts: prefill.rackAcWatts / 18 + decode.rackAcWatts / 18,
+ chassisAcWatts: 2 * (meanRack.rackAcWatts / 18),
+ chassisAcWattsPerGpu: (2 * (meanRack.rackAcWatts / 18)) / 8,
+ facilityWatts: 2 * (meanRack.facilityWatts / 18),
+ deploymentAcWatts: 2 * (meanRack.rackAcWatts / 18),
pue: 1.1,
topologyBasis: 'nvl72-trays',
chassisBasis: 'full',
measuredBasis: 'gpu-plus-grace',
sensorKind: 'grace-socket',
});
+ // Evaluating each tray as its own hypothetical rack would place the light tray
+ // and the heavy tray at different shelf efficiencies and give a different sum.
+ const perTray = [1000, 3000].map(
+ (gpuBoardWattsPerTray) =>
+ estimateRackPower(
+ 'gb300',
+ { basis: 'gpu-plus-grace', gpuBoardWattsPerTray, graceSocketWattsPerTray: 520 },
+ 1.1,
+ )!.rackAcWatts / 18,
+ );
+ expect(perTray[0] + perTray[1]).not.toBeCloseTo(2 * (meanRack.rackAcWatts / 18), 1);
// Two hosts carry four Grace sockets; any other socket count is not a tray topology.
source.metrics.avg_total_cpu_power_w = 260 * 3;
expect(modelSystemPower(source)).toMatchObject({ reason: 'cpu-telemetry' });
diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts
index 1e96423f6..10567a030 100644
--- a/packages/app/src/lib/modeled-system-power.ts
+++ b/packages/app/src/lib/modeled-system-power.ts
@@ -15,6 +15,17 @@ export const AIR_COOLED_SYSTEM_PUE = 1.3;
// Application policy for the direct-liquid-cooled NVL72 rack profiles (docs/powerx-system-power.md).
export const DLC_SYSTEM_PUE = 1.1;
+/**
+ * Facility PUE applied when a caller passes none: the DLC factor for NVL72 rack
+ * profiles, the air-cooled factor for every chassis profile. The dashboard and
+ * the offline exporter share this selection so article figures match chart hovers.
+ */
+export function defaultSystemPue(hardware: string): number {
+ return Object.hasOwn(SYSTEM_POWER_RACK_PROFILES, hardware.toLowerCase())
+ ? DLC_SYSTEM_PUE
+ : AIR_COOLED_SYSTEM_PUE;
+}
+
/** Every supported chassis model describes one complete eight-GPU HGX/OAM system. */
const CHASSIS_GPU_COUNT = 8;
@@ -48,9 +59,13 @@ interface SupportedSystemPowerEstimate {
modeledGpuCount: number;
measuredGpuWattsPerGpu: number;
/**
- * Modeled AC for every full unit, summed. A tray's AC is its 1/18 share of a rack
- * whose trays all match it, so the switch trays, shelves, and management switches
- * are amortised over all 72 GPUs.
+ * Modeled AC for every full unit, summed. Each chassis is evaluated at its own
+ * load because it owns its fans and PSUs. Trays share the rack's power shelves,
+ * so the measured trays are folded into one rack of 18 trays matching their mean
+ * compute-module input, the shelf efficiency curve is evaluated once at that
+ * rack's DC load (as the source `gb200_nvl72_rack_power` does), and every tray
+ * takes the same 1/18 share; the switch trays, shelves, and management switches
+ * are thereby amortised over all 72 GPUs.
*/
chassisAcWatts: number;
/** chassisAcWatts ÷ modeledGpuCount: the plotted metric. */
@@ -140,14 +155,18 @@ function modelChassis(hardware: string, unit: MeasuredUnit, pue: number): UnitPo
}
/**
- * One tray's share of a rack whose trays all match it. The producer publishes the
- * Grace-side and module readings as deployment totals, so every tray receives the mean.
+ * One tray's 1/18 share of a rack whose 18 trays all match the measured trays' mean.
+ * The source model takes one compute-module figure per tray and evaluates the
+ * power-shelf efficiency curve once at the resulting rack DC load, so heterogeneous
+ * measured trays (prefill beside decode) are averaged before the call rather than
+ * each evaluated as its own hypothetical rack. The producer publishes the Grace-side
+ * and module readings as deployment totals, so their per-tray mean is `total / trays`.
* The Grace CPU and LPDDR5X are never modelled: they are inside the measured reading.
*/
function modelTray(
hardware: string,
profile: (typeof SYSTEM_POWER_RACK_PROFILES)[SystemPowerRackHardware],
- unit: MeasuredUnit,
+ meanGpuBoardWattsPerTray: number,
cpu: CpuSideTelemetry,
trayCount: number,
pue: number,
@@ -156,7 +175,7 @@ function modelTray(
cpu.moduleTotalWatts === undefined
? {
basis: 'gpu-plus-grace',
- gpuBoardWattsPerTray: unit.gpuBoardWatts,
+ gpuBoardWattsPerTray: meanGpuBoardWattsPerTray,
graceSocketWattsPerTray: cpu.graceTotalWatts / trayCount,
}
: { basis: 'module', moduleWattsPerTray: cpu.moduleTotalWatts / trayCount };
@@ -197,7 +216,7 @@ export function modelSystemPower(
if (!rack && !(SUPPORTED_SYSTEM_POWER_HARDWARE as readonly string[]).includes(hardware)) {
return unavailable('hardware');
}
- const facilityPue = pue ?? (rack ? DLC_SYSTEM_PUE : AIR_COOLED_SYSTEM_PUE);
+ const facilityPue = pue ?? defaultSystemPue(hardware);
const unitGpuCount = rack ? rack.gpusPerComputeTray : CHASSIS_GPU_COUNT;
const unitShare = (n: unknown): n is number => count(n) && n <= unitGpuCount;
if (typeof row.disagg !== 'boolean' || typeof row.is_multinode !== 'boolean') {
@@ -361,12 +380,23 @@ export function modelSystemPower(
return unavailable('cpu-telemetry');
}
+ // Chassis own their fans and PSUs, so each is evaluated at its own load. Trays
+ // share the rack's shelves, so one rack is evaluated at the mean tray and every
+ // tray receives the same share (see modelTray).
+ const trayModel =
+ rack && cpu
+ ? modelTray(
+ hardware,
+ rack,
+ units.reduce((sum, unit) => sum + unit.gpuBoardWatts, 0) / units.length,
+ cpu,
+ units.length,
+ facilityPue,
+ )
+ : null;
const results = units.map((unit) => ({
...unit,
- model:
- rack && cpu
- ? modelTray(hardware, rack, unit, cpu, units.length, facilityPue)
- : modelChassis(hardware, unit, facilityPue),
+ model: rack && cpu ? trayModel : modelChassis(hardware, unit, facilityPue),
}));
if (results.some((r) => r.model === null)) return unavailable('model-domain');
const first = results[0].model!;
diff --git a/packages/constants/src/metric-keys.test.ts b/packages/constants/src/metric-keys.test.ts
index acba9b729..923504ff0 100644
--- a/packages/constants/src/metric-keys.test.ts
+++ b/packages/constants/src/metric-keys.test.ts
@@ -1,12 +1,34 @@
import { describe, expect, it } from 'vitest';
import {
+ CPU_SIDE_POWER_METRIC_KEY_LIST,
+ CPU_SIDE_POWER_METRIC_KEYS,
MEASURED_POWER_METRIC_KEY_LIST,
MEASURED_POWER_METRIC_KEYS,
METRIC_KEYS,
POWER_METRIC_KEYS,
} from './metric-keys';
+describe('CPU_SIDE_POWER_METRIC_KEYS', () => {
+ it('names exactly the NVL72 Grace-side and compute-module keys, all of them measured keys', () => {
+ expect(new Set(CPU_SIDE_POWER_METRIC_KEY_LIST)).toEqual(
+ new Set([
+ 'avg_cpu_socket_power_w',
+ 'avg_total_cpu_power_w',
+ 'total_cpu_energy_j',
+ 'avg_total_module_power_w',
+ 'total_module_energy_j',
+ ]),
+ );
+ expect(CPU_SIDE_POWER_METRIC_KEYS.size).toBe(5);
+ for (const key of CPU_SIDE_POWER_METRIC_KEYS) {
+ expect(MEASURED_POWER_METRIC_KEYS.has(key)).toBe(true);
+ }
+ // The verdict itself is a discriminator, not a measurement.
+ expect(CPU_SIDE_POWER_METRIC_KEYS.has('cpu_power_valid')).toBe(false);
+ });
+});
+
describe('MEASURED_POWER_METRIC_KEYS', () => {
it('is a subset of METRIC_KEYS', () => {
for (const key of MEASURED_POWER_METRIC_KEYS) {
diff --git a/packages/constants/src/metric-keys.ts b/packages/constants/src/metric-keys.ts
index dceb88844..922596f97 100644
--- a/packages/constants/src/metric-keys.ts
+++ b/packages/constants/src/metric-keys.ts
@@ -1,8 +1,36 @@
/**
- * Power, energy, and GPU telemetry withheld at ingest and display when the
- * normalized `power_valid` verdict is 0. Contract and diagnostic fields are
- * excluded so the invalid verdict remains auditable. Add new measured fields
- * here; `METRIC_KEYS` derives from this list.
+ * NVL72 Grace-side and compute-module measurements from the srt-slurm CPU power
+ * leg (ACPI hwmon), integrated over the same formal window as GPU energy. They
+ * are withheld on their own `cpu_power_valid` verdict, never on `power_valid`:
+ * the leg has its own sensors, window bracketing and gap checks, so a failed
+ * GPU leg says nothing about them and vice versa. Unlike `power_valid`, no
+ * legacy rows predate the verdict, so an absent verdict withholds them too.
+ * avg_cpu_socket_power_w: mean over sockets of each socket's window-mean Grace-side W
+ * avg_total_cpu_power_w: sum over sockets of window-mean Grace-side W
+ * total_cpu_energy_j: Grace-side energy over the window, all sockets
+ * avg_total_module_power_w / total_module_energy_j: whole compute module
+ * (Grace + GPUs + HBM + LPDDR5X + regulator loss), only when
+ * the module sensor exists on every socket
+ */
+export const CPU_SIDE_POWER_METRIC_KEY_LIST = [
+ 'avg_cpu_socket_power_w',
+ 'avg_total_cpu_power_w',
+ 'total_cpu_energy_j',
+ 'avg_total_module_power_w',
+ 'total_module_energy_j',
+] as const;
+
+export const CPU_SIDE_POWER_METRIC_KEYS: ReadonlySet = new Set(
+ CPU_SIDE_POWER_METRIC_KEY_LIST,
+);
+
+/**
+ * Every measured power, energy, and telemetry field. GPU-side fields are
+ * withheld at ingest and display when the normalized `power_valid` verdict is
+ * 0; the `CPU_SIDE_POWER_METRIC_KEY_LIST` subset follows `cpu_power_valid`
+ * instead. Contract and diagnostic fields are excluded so the invalid verdict
+ * remains auditable. Add new measured fields here; `METRIC_KEYS` derives from
+ * this list.
*/
export const MEASURED_POWER_METRIC_KEY_LIST = [
// measured power / energy (emitted by runner's aggregate_power.py)
@@ -46,19 +74,7 @@ export const MEASURED_POWER_METRIC_KEY_LIST = [
'peak_temp_c',
'avg_util_pct',
'avg_mem_used_mb',
- // NVL72 Grace-side and compute-module measurements from the srt-slurm CPU power
- // leg (ACPI hwmon), integrated over the same formal window as GPU energy.
- // avg_cpu_socket_power_w: mean over sockets of each socket's window-mean Grace-side W
- // avg_total_cpu_power_w: sum over sockets of window-mean Grace-side W
- // total_cpu_energy_j: Grace-side energy over the window, all sockets
- // avg_total_module_power_w / total_module_energy_j: whole compute module
- // (Grace + GPUs + HBM + LPDDR5X + regulator loss), only when
- // the module sensor exists on every socket
- 'avg_cpu_socket_power_w',
- 'avg_total_cpu_power_w',
- 'total_cpu_energy_j',
- 'avg_total_module_power_w',
- 'total_module_energy_j',
+ ...CPU_SIDE_POWER_METRIC_KEY_LIST,
] as const;
export const MEASURED_POWER_METRIC_KEYS: ReadonlySet = new Set(
@@ -79,9 +95,10 @@ export const POWER_METRIC_KEYS = [
'power_valid',
'power_metric_schema_version',
// cpu_power_valid: numeric 1/0 verdict for the NVL72 CPU-side leg, independent
- // of power_valid; 0 means the producer emitted no CPU-side keys
+ // of power_valid; anything but 1 withholds the CPU-side keys
'cpu_power_valid',
// measured power / energy / telemetry values, withheld when power_valid = 0
+ // (GPU side) or cpu_power_valid != 1 (CPU side)
...MEASURED_POWER_METRIC_KEY_LIST,
] as const;
diff --git a/packages/db/src/etl/benchmark-mapper.test.ts b/packages/db/src/etl/benchmark-mapper.test.ts
index ed9dda0c0..815dae66f 100644
--- a/packages/db/src/etl/benchmark-mapper.test.ts
+++ b/packages/db/src/etl/benchmark-mapper.test.ts
@@ -1,5 +1,8 @@
import { describe, it, expect, vi } from 'vitest';
-import { MEASURED_POWER_METRIC_KEYS } from '@semianalysisai/inferencex-constants';
+import {
+ CPU_SIDE_POWER_METRIC_KEYS,
+ MEASURED_POWER_METRIC_KEYS,
+} from '@semianalysisai/inferencex-constants';
import {
extractPowerAudit,
extractPowerInvalidReasons,
@@ -88,7 +91,8 @@ function dirtyPowerPayload(): Record {
peak_temp_c: 79.2,
avg_util_pct: 88.5,
avg_mem_used_mb: 71234.5,
- // NVL72 CPU-side measurements share the GPU window; withheld with the GPU verdict.
+ // NVL72 CPU-side measurements share the GPU window but follow their own
+ // cpu_power_valid verdict; tests that want them kept must supply it.
avg_cpu_socket_power_w: 250.5,
avg_total_cpu_power_w: 1002,
total_cpu_energy_j: 601200,
@@ -101,6 +105,11 @@ function dirtyPowerPayload(): Record {
};
}
+const CPU_SIDE_KEYS = [...CPU_SIDE_POWER_METRIC_KEYS];
+const GPU_SIDE_KEYS = [...MEASURED_POWER_METRIC_KEYS].filter(
+ (key) => !CPU_SIDE_POWER_METRIC_KEYS.has(key),
+);
+
describe('mapBenchmarkRow', () => {
describe('v1 schema', () => {
it('maps a valid v1 row to BenchmarkParams', () => {
@@ -345,34 +354,69 @@ describe('mapBenchmarkRow', () => {
expect(result!.metrics.median_ttft).toBe(50.2);
});
- it('withholds the CPU-side measurements but keeps the independent cpu_power_valid verdict', () => {
+ // Ticket 01 contract: cpu_power_valid is independent of power_valid. Each leg
+ // withholds only its own keys; worker telemetry belongs to the GPU leg.
+ it.each([
+ { power_valid: 1, cpu_power_valid: 1, gpuKept: true, cpuKept: true },
+ { power_valid: 1, cpu_power_valid: 0, gpuKept: true, cpuKept: false },
+ { power_valid: 0, cpu_power_valid: 1, gpuKept: false, cpuKept: true },
+ { power_valid: 0, cpu_power_valid: 0, gpuKept: false, cpuKept: false },
+ ])(
+ 'withholds each leg on its own verdict: power_valid=$power_valid cpu_power_valid=$cpu_power_valid',
+ ({ power_valid, cpu_power_valid, gpuKept, cpuKept }) => {
+ const tracker = createSkipTracker();
+ const dirty = dirtyPowerPayload();
+ const result = mapBenchmarkRow(
+ makeV2Row({ power_valid, power_metric_schema_version: 2, cpu_power_valid, ...dirty }),
+ tracker,
+ );
+
+ expect(result!.metrics.power_valid).toBe(power_valid);
+ expect(result!.metrics.cpu_power_valid).toBe(cpu_power_valid);
+ for (const key of GPU_SIDE_KEYS) {
+ if (gpuKept) expect(result!.metrics[key]).toBe(dirty[key]);
+ else expect(result!.metrics).not.toHaveProperty(key);
+ }
+ for (const key of CPU_SIDE_KEYS) {
+ if (cpuKept) expect(result!.metrics[key]).toBe(dirty[key]);
+ else expect(result!.metrics).not.toHaveProperty(key);
+ }
+ if (gpuKept) expect(result!.workers).toHaveLength(2);
+ else expect(result!.workers).toBeUndefined();
+ },
+ );
+
+ it('withholds CPU-side keys that arrive without a cpu_power_valid verdict or with a malformed one', () => {
+ // No legacy rows predate cpu_power_valid, so absence is out of contract and fails closed.
const tracker = createSkipTracker();
- const result = mapBenchmarkRow(
+ const dirty = dirtyPowerPayload();
+ const absent = mapBenchmarkRow(
+ makeV2Row({ power_valid: 1, power_metric_schema_version: 2, ...dirty }),
+ tracker,
+ );
+ expect(absent!.metrics).not.toHaveProperty('cpu_power_valid');
+ for (const key of CPU_SIDE_KEYS) expect(absent!.metrics).not.toHaveProperty(key);
+ for (const key of GPU_SIDE_KEYS) expect(absent!.metrics[key]).toBe(dirty[key]);
+
+ const malformed = mapBenchmarkRow(
makeV2Row({
- power_valid: 0,
+ power_valid: 1,
power_metric_schema_version: 2,
- cpu_power_valid: 1,
- ...dirtyPowerPayload(),
+ cpu_power_valid: 'garbage',
+ ...dirty,
}),
tracker,
);
- expect(result!.metrics.cpu_power_valid).toBe(1);
- for (const key of [
- 'avg_cpu_socket_power_w',
- 'avg_total_cpu_power_w',
- 'total_cpu_energy_j',
- 'avg_total_module_power_w',
- 'total_module_energy_j',
- ]) {
- expect(result!.metrics).not.toHaveProperty(key);
- }
+ expect(malformed!.metrics.cpu_power_valid).toBe(0);
+ for (const key of CPU_SIDE_KEYS) expect(malformed!.metrics).not.toHaveProperty(key);
+ for (const key of GPU_SIDE_KEYS) expect(malformed!.metrics[key]).toBe(dirty[key]);
});
- it('keeps every measured key and the workers payload on a valid verdict', () => {
+ it('keeps every measured key and the workers payload on valid verdicts', () => {
const tracker = createSkipTracker();
const dirty = dirtyPowerPayload();
const result = mapBenchmarkRow(
- makeV2Row({ power_valid: 1, power_metric_schema_version: 2, ...dirty }),
+ makeV2Row({ power_valid: 1, power_metric_schema_version: 2, cpu_power_valid: 1, ...dirty }),
tracker,
);
@@ -383,13 +427,13 @@ describe('mapBenchmarkRow', () => {
expect(result!.workers).toHaveLength(2);
});
- it('leaves legacy rows without a verdict untouched (historical measurements kept)', () => {
+ it('leaves legacy rows without a GPU verdict untouched (historical measurements kept)', () => {
const tracker = createSkipTracker();
const dirty = dirtyPowerPayload();
const result = mapBenchmarkRow(makeV2Row(dirty), tracker);
expect(result!.metrics).not.toHaveProperty('power_valid');
- for (const key of MEASURED_POWER_METRIC_KEYS) {
+ for (const key of GPU_SIDE_KEYS) {
expect(result!.metrics[key]).toBe(dirty[key]);
}
expect(result!.workers).toHaveLength(2);
@@ -947,28 +991,55 @@ describe('scrubWithheldPowerMetrics (direct — supplemental ingest path)', () =
expect(metrics.tput_per_gpu).toBe(567.8);
});
- it('leaves power_valid=1 and legacy no-verdict records untouched', () => {
- for (const metrics of [supplementalMetrics({ power_valid: 1 }), supplementalMetrics()]) {
+ it('leaves power_valid=1 and legacy no-GPU-verdict records untouched', () => {
+ for (const metrics of [
+ supplementalMetrics({ power_valid: 1, cpu_power_valid: 1 }),
+ supplementalMetrics({ cpu_power_valid: 1 }),
+ ]) {
const before = { ...metrics };
expect(scrubWithheldPowerMetrics(metrics)).toBe(false);
expect(metrics).toEqual(before);
}
});
+ it('withholds only the CPU-side keys on cpu_power_valid=0 and leaves GPU power published', () => {
+ const metrics = supplementalMetrics({ power_valid: 1, cpu_power_valid: 0 });
+ const before = { ...metrics };
+ expect(scrubWithheldPowerMetrics(metrics)).toBe(false);
+ expect(metrics.cpu_power_valid).toBe(0);
+ for (const key of CPU_SIDE_KEYS) expect(metrics).not.toHaveProperty(key);
+ for (const key of GPU_SIDE_KEYS) expect(metrics[key]).toBe(before[key]);
+ expect(metrics.avg_power_w).toBe(685.5);
+ });
+
+ it('keeps the CPU-side keys on power_valid=0 when cpu_power_valid=1', () => {
+ const metrics = supplementalMetrics({ power_valid: 0, cpu_power_valid: 1 });
+ expect(scrubWithheldPowerMetrics(metrics)).toBe(true);
+ for (const key of GPU_SIDE_KEYS) expect(metrics).not.toHaveProperty(key);
+ for (const key of CPU_SIDE_KEYS) expect(metrics).toHaveProperty(key);
+ expect(metrics.avg_total_module_power_w).toBe(17203);
+ });
+
it('normalizes cpu_power_valid as a verdict, independent of power_valid', () => {
const metrics = supplementalMetrics({ power_valid: 1, cpu_power_valid: '1' });
normalizePowerContractMetrics(metrics, metrics);
expect(metrics.cpu_power_valid).toBe(1);
expect(scrubWithheldPowerMetrics(metrics)).toBe(false);
+ expect(metrics.avg_total_cpu_power_w).toBe(1002);
const malformed = supplementalMetrics({ power_valid: 1, cpu_power_valid: 2 });
normalizePowerContractMetrics(malformed, malformed);
expect(malformed.cpu_power_valid).toBe(0);
expect(malformed.avg_total_cpu_power_w).toBe(1002);
+ expect(scrubWithheldPowerMetrics(malformed)).toBe(false);
+ expect(malformed).not.toHaveProperty('avg_total_cpu_power_w');
+ expect(malformed.avg_power_w).toBe(685.5);
const absent = supplementalMetrics({ power_valid: 1 });
normalizePowerContractMetrics(absent, absent);
expect(absent).not.toHaveProperty('cpu_power_valid');
+ expect(scrubWithheldPowerMetrics(absent)).toBe(false);
+ for (const key of CPU_SIDE_KEYS) expect(absent).not.toHaveProperty(key);
});
it('fails closed on a malformed verdict when composed with normalization', () => {
diff --git a/packages/db/src/etl/benchmark-mapper.ts b/packages/db/src/etl/benchmark-mapper.ts
index e904c9b8a..751705eb5 100644
--- a/packages/db/src/etl/benchmark-mapper.ts
+++ b/packages/db/src/etl/benchmark-mapper.ts
@@ -7,6 +7,7 @@
import type { ConfigParams } from './config-cache';
import type { SkipTracker } from './skip-tracker';
import {
+ CPU_SIDE_POWER_METRIC_KEYS,
MEASURED_POWER_METRIC_KEYS,
METRIC_KEYS,
PRECISION_KEYS,
@@ -589,15 +590,24 @@ export function normalizePowerContractMetrics(
/**
* Enforces fail-closed power publication at ingest. An explicit normalized
- * invalid verdict removes every measured field while preserving the contract
- * and diagnostic fields; legacy rows without a verdict remain unchanged.
- * Returns true so callers also drop worker telemetry. Paths that bypass
- * `mapBenchmarkRow` must normalize the verdict before calling this function.
- * Queries intentionally remain raw; the frontend withholds independently.
+ * invalid GPU verdict removes every GPU-side measured field while preserving
+ * the contract and diagnostic fields; legacy rows without a verdict remain
+ * unchanged. The NVL72 CPU-side keys follow their own `cpu_power_valid`
+ * verdict instead: anything but a normalized 1 withholds them, because no
+ * legacy rows predate that verdict and the estimator admits rows the same way.
+ * Returns true when GPU power was withheld so callers also drop worker
+ * telemetry. Paths that bypass `mapBenchmarkRow` must normalize the verdicts
+ * before calling this function. Queries intentionally remain raw; the
+ * frontend withholds independently.
*/
export function scrubWithheldPowerMetrics(metrics: Record): boolean {
+ if (metrics.cpu_power_valid !== 1) {
+ for (const key of CPU_SIDE_POWER_METRIC_KEYS) delete metrics[key];
+ }
if (metrics.power_valid !== 0) return false;
- for (const key of MEASURED_POWER_METRIC_KEYS) delete metrics[key];
+ for (const key of MEASURED_POWER_METRIC_KEYS) {
+ if (!CPU_SIDE_POWER_METRIC_KEYS.has(key)) delete metrics[key];
+ }
return true;
}
diff --git a/packages/db/src/etl/power-publication.ts b/packages/db/src/etl/power-publication.ts
index d2ed10c14..1161500fa 100644
--- a/packages/db/src/etl/power-publication.ts
+++ b/packages/db/src/etl/power-publication.ts
@@ -32,7 +32,13 @@ const IDENTITY_FIELDS = [
'image',
'run_url',
] as const;
-const POWER_FIELDS = [...MEASURED_POWER_METRIC_KEYS, 'power_valid', 'power_metric_schema_version'];
+// The CPU-side keys are withheld on cpu_power_valid, so the manifest carries that verdict too.
+const POWER_FIELDS = [
+ ...MEASURED_POWER_METRIC_KEYS,
+ 'power_valid',
+ 'power_metric_schema_version',
+ 'cpu_power_valid',
+];
export interface PowerPublicationPoint {
identity: Record;
metrics: Record;
From 1745e7069f58d6408f377a70edbb91cfd6f740cd Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Sat, 19 Sep 2026 13:43:44 -0700
Subject: [PATCH 07/22] feat: infer NVL72 trays for aggregate multinode rows
without a worker array
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
The Kimi K3 GB200 aggregate producer (dynamo-vLLM TP16, sixteen GPUs on four
trays) emits no `workers` array, so `modelSystemPower` rejected every such row
as `topology` and the tray path never reached the Profit Estimator. Generalize
the base's `uniform-hosts` branch from eight-GPU chassis to the unit size of
the hardware: on GB200/GB300, a non-disaggregated multinode row without a
worker array is `gpuCount / 4` compute trays, each fed the deployment mean
(module total per tray when the module keys are present, otherwise GPU-board
plus Grace-socket watts per tray, exactly as on the worker-hosts tray path),
with `chassisBasis: 'full'`, `topologyBasis: 'nvl72-trays'` and the usual
measured basis and sensor kind. A GPU total that does not fill whole trays is
`gpu-count`; the tray count must agree with the Grace-socket count recovered
from the CPU-side keys and, when the CPU leg recorded it, with
`power_audit.cpu.observed_sockets`, otherwise `cpu-telemetry`. The x86
uniform-hosts path, the worker-hosts tray path and every English string are
unchanged. Tests cover module and Grace-socket bases, both socket mismatches,
the uneven count and the Profit Estimator source label; the system-power doc
records the rule in English and Chinese.
中文:Kimi K3 GB200 聚合采集端(dynamo-vLLM TP16,16 张 GPU 分布在 4 个 tray)
不输出 `workers` 数组,`modelSystemPower` 一律按 `topology` 拒绝,tray 路径无法进入
Profit Estimator。现将基线的 `uniform-hosts` 分支从八卡机箱推广到硬件的单元大小:
GB200/GB300 上无 worker 数组的非 disagg 多节点行按 GPU 总数 ÷ 4 推算 tray 数,每个
tray 取部署平均值(有模块指标时用模块总功耗 ÷ tray 数,否则用 GPU 板卡 + Grace socket
每 tray 功耗,与 worker-hosts tray 路径一致),`chassisBasis: 'full'`、
`topologyBasis: 'nvl72-trays'`,实测口径与传感器类型不变。GPU 总数无法填满整数个
tray 时为 `gpu-count`;tray 数须与 CPU 侧指标推算的 Grace socket 数一致,且在 CPU
采集记录了 `power_audit.cpu.observed_sockets` 时与之一致,否则为 `cpu-telemetry`。
x86 uniform-hosts 路径、worker-hosts tray 路径及所有英文文案不变。测试覆盖模块与
Grace socket 两种口径、两类 socket 不一致、非整数 tray 数以及 Profit Estimator 的
来源标签;系统功耗文档中英文同步。
---
docs/powerx-system-power.md | 23 +++--
.../calculator/profit-power.test.ts | 45 +++++++++
.../app/src/lib/modeled-system-power.test.ts | 99 ++++++++++++++++---
packages/app/src/lib/modeled-system-power.ts | 40 +++++---
4 files changed, 169 insertions(+), 38 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index ab098b09f..848f4959d 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -62,7 +62,10 @@ producer publishes it, otherwise GPU-board watts plus the Grace-socket total
(`avg_total_cpu_power_w`) with the source's regulator-loss allowance on the GPU
share. The Grace CPU and LPDDR5X are never modelled; rows without
`cpu_power_valid=1` and the Grace-side keys stay unavailable (`cpu-telemetry`).
-Each measured worker host is one compute tray (four GPUs, two Grace sockets). The
+Each measured worker host is one compute tray (four GPUs, two Grace sockets); an
+aggregate multinode row without a per-worker array is `gpuCount / 4` trays at the
+deployment mean, cross-checked against the Grace-socket count and the CPU leg's
+`power_audit.cpu.observed_sockets`. The
measured trays are folded into one rack of 18 trays matching their mean
compute-module input, the power-shelf efficiency curve is evaluated once at that
rack's DC load (as the source `gb200_nvl72_rack_power` does with its single
@@ -154,9 +157,10 @@ implementation for both variants, both bases, every shelf knot and PUE 1.0–1.2
measured GPUs ÷ 1000 × 1.1. It accepts fully measured eight-GPU chassis
(`chassisBasis: 'full'` on the `single-node`, `worker-hosts`, or `uniform-hosts`
basis; see the Profit Estimator power basis section), or an `nvl72-trays` estimate
-whose trays are all fully measured (one host per worker, four GPUs and two sockets
-each; an aggregate multinode NVL72 row without a per-worker array is not modeled
-at the deployment mean and stays `topology`-unavailable). Partial trays are
+whose trays are all fully measured (four GPUs and two sockets each: one tray per
+measured worker host, or, for an aggregate multinode row without a per-worker
+array, `gpuCount / 4` trays at the deployment mean, cross-checked against the
+Grace-socket count and `power_audit.cpu.observed_sockets`). Partial trays are
extrapolated in the chart but rejected here, as partial chassis are. Between two
frontier knots both must share the same measured basis and sensor kind; a module
knot beside a Grace-socket knot stays unavailable rather than blending sensors. The
@@ -317,7 +321,9 @@ GPU 运行相同负载。结果标记为 `chassisBasis: 'extrapolated'`:每卡
的 GPU 总数分摊,`deploymentAcWatts` 只保留实测 GPU 在各机箱中的份额。这不是把
部分分配的机箱按比例分摊:固定组件、风扇曲线和 PSU 效率都在满机箱负载点求值。
GB200、GB300 使用单独的 NVL72 机架 profile,不套用 B200、B300 机箱模型:每台实测
-worker 主机视为一个计算 tray(4 张 GPU、2 个 Grace socket)。输入为实测模块功耗
+worker 主机视为一个计算 tray(4 张 GPU、2 个 Grace socket);没有逐 worker 数组的聚合
+多节点行则按 GPU 总数 ÷ 4 推算 tray 数、每个 tray 取部署平均值,并与 Grace socket 数及
+CPU 采集记录的 `power_audit.cpu.observed_sockets` 交叉校验。输入为实测模块功耗
(`avg_total_module_power_w`);缺失时改用 GPU 板卡功耗加 Grace socket 功耗
(`avg_total_cpu_power_w`),并按来源模型计入 GPU 份额的稳压损耗余量。Grace CPU 与
LPDDR5X 从不建模,缺少 `cpu_power_valid=1` 和 Grace 侧指标的行保持不可用
@@ -366,9 +372,10 @@ BlueField-3 DPU 空闲功耗(2 × 65 W)、NVMe 空闲功耗(22 W)、风
验证了与固定实现的一致性。
门槛规则(利润估算器):规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。接受
-`chassisBasis: 'full'` 的单节点八卡机箱,或全部 tray 均完整实测(每个 worker 一台主机,
-各 4 张 GPU、2 个 socket)的 `nvl72-trays` 估算;部分 tray 在图表中外推显示,但与部分
-机箱一样不进入规划门槛。两个前沿数据点之间必须采用相同的实测口径和传感器类型,模块
+`chassisBasis: 'full'` 的八卡机箱估算(`single-node`、`worker-hosts` 或 `uniform-hosts`
+拓扑),或全部 tray 均完整实测的 `nvl72-trays` 估算(各 4 张 GPU、2 个 socket:每个实测
+worker 主机一个 tray,或没有逐 worker 数组的聚合多节点行按 GPU 总数 ÷ 4 推算、并与
+socket 数交叉校验);部分 tray 在图表中外推显示,但与部分机箱一样不进入规划门槛。两个前沿数据点之间必须采用相同的实测口径和传感器类型,模块
读数旁边的 Grace socket 读数保持不可用,不会混合两种传感器。柱形提示、功耗说明下方的
标注行和 CSV 的 `功耗口径`、`功耗传感器`、`系统功耗 profile` 三列逐行标出实测口径
(实测模块功耗,或实测 GPU 板卡 + Grace socket 功耗并由模型估算稳压损耗)、传感器类型
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index 3eff52985..f0d3ad2cb 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -288,6 +288,51 @@ describe('profit power basis preview', () => {
expect(modeledPowerAtTarget(same, 45)).toBeCloseTo(modeledPowerAtTarget(trayResult, 45)!, 10);
});
+ it('plans aggregate multinode NVL72 rows without a worker array as inferred trays', () => {
+ // Kimi K3 GB200 dynamo-vLLM TP16: sixteen GPUs on four trays, no per-worker array.
+ const inferred: BenchmarkRow = {
+ ...traySource,
+ is_multinode: true,
+ prefill_tp: 16,
+ decode_tp: 0,
+ num_prefill_gpu: 16,
+ num_decode_gpu: 16,
+ metrics: {
+ ...traySource.metrics,
+ avg_power_w: 441.741,
+ avg_total_gpu_power_w: 7067.859,
+ avg_total_cpu_power_w: 2004,
+ avg_total_module_power_w: 9071.859,
+ },
+ };
+ const rack = estimateRackPower(
+ 'gb200',
+ { basis: 'module', moduleWattsPerTray: 9071.859 / 4 },
+ 1.1,
+ )!;
+ const inferredResult = withPoints(trayResult, [{ ...trayPoint, sourceRow: inferred }]);
+ expect(modeledPowerAtTarget(inferredResult, 45)).toBeCloseTo(
+ (((rack.facilityWatts / 18) * 4) / 16 / 1000) * 1.1,
+ 8,
+ );
+ const rows = estimateProfitByPower(
+ [inferredResult],
+ specs,
+ pricing,
+ assumptions,
+ 'modeled',
+ 45,
+ labels,
+ ).rows;
+ expect(rows).toHaveLength(1);
+ expect(rows[0].powerSource).toMatchObject({
+ topology: 'nvl72-trays',
+ measuredBasis: 'module',
+ sensorKind: 'module',
+ pue: 1.1,
+ });
+ });
+
it('labels x86 chassis estimates with the air-cooled profile and leaves provisioned rows unlabeled', () => {
const [provisioned, modeled] = estimateProfitByPower(
[result],
diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts
index 4a4e79e5d..fcba58a1b 100644
--- a/packages/app/src/lib/modeled-system-power.test.ts
+++ b/packages/app/src/lib/modeled-system-power.test.ts
@@ -701,28 +701,95 @@ describe('NVL72 trays with measured compute-module power', () => {
expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' });
});
- it('does not infer trays from the GPU total for aggregate multinode rows without workers', () => {
- // Two trays' worth of GPUs and Grace sockets but no per-worker array: the
- // chassis-only uniform-hosts mean is not applied to trays, so the row stays
- // on the worker path and reports the missing placement.
+ it('infers trays from the GPU total for aggregate multinode rows without workers', () => {
+ // Kimi K3 GB200 dynamo-vLLM TP16 (fixture kimik3_*_conc2): sixteen GPUs on
+ // four trays, aggregate producer, no per-worker array. GPU watts are the
+ // fixture's; the CPU-side keys are controlled inputs for eight Grace sockets.
const source = nvl72Row(
- { avg_total_gpu_power_w: 7202, avg_total_cpu_power_w: 1002 },
- { is_multinode: true, prefill_tp: 8, decode_tp: 8, num_prefill_gpu: 8, num_decode_gpu: 8 },
+ {
+ avg_power_w: 441.741,
+ avg_total_gpu_power_w: 7067.859,
+ avg_total_cpu_power_w: 2004,
+ avg_total_module_power_w: 9071.859,
+ },
+ {
+ is_multinode: true,
+ prefill_tp: 16,
+ decode_tp: 0,
+ num_prefill_gpu: 16,
+ num_decode_gpu: 16,
+ power_audit: { cpu: { sensor_kind: 'module', observed_sockets: 8 } },
+ },
);
- expect(modelSystemPower(source)).toMatchObject({ reason: 'topology' });
- // The same shape on chassis hardware is the uniform-hosts case.
+ const rack = estimateRackPower(
+ 'gb200',
+ { basis: 'module', moduleWattsPerTray: 9071.859 / 4 },
+ 1.1,
+ )!;
+ const estimate = modelSystemPower(source);
+ expect(estimate).toMatchObject({
+ status: 'supported',
+ topologyBasis: 'nvl72-trays',
+ measuredBasis: 'module',
+ sensorKind: 'module',
+ chassisBasis: 'full',
+ gpuCount: 16,
+ chassisCount: 4,
+ modeledGpuCount: 16,
+ pue: 1.1,
+ });
+ if (estimate.status !== 'supported') throw new Error('unreachable');
+ expect(estimate.chassisAcWatts).toBeCloseTo(4 * (rack.rackAcWatts / 18), 6);
+ expect(estimate.deploymentFacilityWatts).toBe(estimate.facilityWatts);
+
+ // Without module keys the same trays take GPU board plus Grace socket per tray.
+ const graceOnly = { ...source, metrics: { ...source.metrics } };
+ delete (graceOnly.metrics as Record).avg_total_module_power_w;
+ const graceRack = estimateRackPower(
+ 'gb200',
+ {
+ basis: 'gpu-plus-grace',
+ gpuBoardWattsPerTray: 7067.859 / 4,
+ graceSocketWattsPerTray: 2004 / 4,
+ },
+ 1.1,
+ )!;
+ const graceEstimate = modelSystemPower(graceOnly);
+ expect(graceEstimate).toMatchObject({
+ status: 'supported',
+ topologyBasis: 'nvl72-trays',
+ measuredBasis: 'gpu-plus-grace',
+ sensorKind: 'grace-socket',
+ chassisCount: 4,
+ });
+ if (graceEstimate.status !== 'supported') throw new Error('unreachable');
+ expect(graceEstimate.chassisAcWatts).toBeCloseTo(4 * (graceRack.rackAcWatts / 18), 6);
+
+ // The CPU leg's recorded socket coverage must agree with the inferred trays.
+ expect(
+ modelSystemPower({ ...source, power_audit: { cpu: { observed_sockets: 6 } } }),
+ ).toMatchObject({ reason: 'cpu-telemetry' });
+ // So must the socket count the Grace-side keys recover.
expect(
modelSystemPower({
...source,
- hardware: 'b200',
- metrics: {
- ...source.metrics,
- avg_power_w: 900.25,
- avg_total_gpu_power_w: 14404,
- decode_pp: 2,
- },
+ metrics: { ...source.metrics, avg_total_cpu_power_w: 250.5 * 6 },
}),
- ).toMatchObject({ status: 'supported', topologyBasis: 'uniform-hosts', chassisCount: 2 });
+ ).toMatchObject({ reason: 'cpu-telemetry' });
+ // Eighteen GPUs cannot fill whole four-GPU trays.
+ expect(
+ modelSystemPower({
+ ...source,
+ prefill_tp: 18,
+ metrics: { ...source.metrics, avg_total_gpu_power_w: 441.741 * 18 },
+ }),
+ ).toMatchObject({ reason: 'gpu-count' });
+ // Chassis hardware keeps the base uniform-hosts path for the same shape.
+ expect(modelSystemPower({ ...source, hardware: 'b200' })).toMatchObject({
+ status: 'supported',
+ topologyBasis: 'uniform-hosts',
+ chassisCount: 2,
+ });
});
it('extrapolates a partially measured tray on the GPU-board share and keeps the measured share', () => {
diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts
index 7f2249573..3f039815a 100644
--- a/packages/app/src/lib/modeled-system-power.ts
+++ b/packages/app/src/lib/modeled-system-power.ts
@@ -100,6 +100,10 @@ export type SystemPowerEstimate =
topologyBasis: 'single-node' | 'worker-hosts' | 'uniform-hosts';
})
| (SupportedSystemPowerEstimate & {
+ /**
+ * One tray per measured worker host, or, for an aggregate multinode row
+ * without a worker array, gpuCount ÷ 4 trays at the deployment mean.
+ */
topologyBasis: 'nvl72-trays';
/** Which measured reading fed every tray: the module sensor, or GPU board + Grace socket. */
measuredBasis: RackMeasuredBasis;
@@ -324,19 +328,17 @@ export function modelSystemPower(
gpuBoardWatts:
gpuCount === unitGpuCount ? m.avg_total_gpu_power_w : m.avg_power_w * unitGpuCount,
});
- } else if (
- !rack &&
- row.disagg === false &&
- (!Array.isArray(row.workers) || row.workers.length === 0)
- ) {
+ } else if (row.disagg === false && (!Array.isArray(row.workers) || row.workers.length === 0)) {
// Aggregate multinode producers emit no per-worker telemetry. Symmetric
- // TP/PP/DP shards load every host alike, so each full eight-GPU chassis is
- // modeled at the deployment mean; the supported chassis hardware only ships
- // in eight-GPU hosts, so the count must fill whole chassis on several hosts.
- // Disaggregated roles differ in load and stay on the worker path; NVL72
- // trays stay there too (a tray count is not inferred from the GPU total).
- const hostCount = gpuCount / CHASSIS_GPU_COUNT;
- if (!count(hostCount) || hostCount < 2) return unavailable('topology');
+ // TP/PP/DP shards load every host alike, so each full unit is modeled at
+ // the deployment mean: chassis hardware only ships in eight-GPU hosts and
+ // NVL72 in four-GPU compute trays, so the count must fill whole units on
+ // several hosts. Disaggregated roles differ in load and stay on the worker path.
+ const hostCount = gpuCount / unitGpuCount;
+ // An uneven chassis count leaves placement unknown; an uneven tray count
+ // contradicts the four-GPU tray itself.
+ if (!count(hostCount)) return unavailable(rack ? 'gpu-count' : 'topology');
+ if (hostCount < 2) return unavailable('topology');
const tp = row.decode_tp > 0 ? row.decode_tp : row.prefill_tp;
const pp = Math.max(m.pp ?? 1, m.decode_pp ?? 1, m.prefill_pp ?? 1);
const pcp = Math.max(m.pcp_size ?? 1, m.decode_pcp_size ?? 1, m.prefill_pcp_size ?? 1);
@@ -351,11 +353,21 @@ export function modelSystemPower(
) {
return unavailable('gpu-count');
}
+ // The CPU leg records the sockets it covered; when present it must agree
+ // with the trays inferred from the GPU total (two Grace sockets per tray).
+ const observedSockets = row.power_audit?.cpu?.observed_sockets;
+ if (
+ rack &&
+ observedSockets !== undefined &&
+ observedSockets !== hostCount * rack.graceSocketsPerComputeTray
+ ) {
+ return unavailable('cpu-telemetry');
+ }
topologyBasis = 'uniform-hosts';
for (let host = 0; host < hostCount; host++) {
units.push({
- measuredGpus: CHASSIS_GPU_COUNT,
- // Partition the producer's exact total so the chassis inputs sum back to it.
+ measuredGpus: unitGpuCount,
+ // Partition the producer's exact total so the unit inputs sum back to it.
gpuBoardWatts: m.avg_total_gpu_power_w / hostCount,
});
}
From f2de031d6f03005d1e0e9c03d80185c980d15b50 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Sat, 19 Sep 2026 13:45:46 -0700
Subject: [PATCH 08/22] refactor: cite the producer contract instead of a local
ticket in cpu-side comments
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Three comments referenced a task-tracker ticket that does not exist in this
repository; point them at the InferenceX producer contract instead. No code
change.
中文:三处注释引用了仓库中不存在的本地工单,改为指向 InferenceX 的生产端契约文档;无代码变更。
---
packages/app/src/components/calculator/profit-power.test.ts | 2 +-
packages/app/src/lib/modeled-system-power.test.ts | 2 +-
packages/db/src/etl/benchmark-mapper.test.ts | 2 +-
3 files changed, 3 insertions(+), 3 deletions(-)
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index f0d3ad2cb..fc898e7a4 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -98,7 +98,7 @@ const labels = { provisioned: 'Provisioned', modeled: 'Measured + modeled' };
// One GB200 NVL72 compute tray on the AgentX workload: four GPUs on one host, two
// Grace sockets, module sensor present. Watts are controlled inputs, not published
-// constants; the CPU-side keys follow the ticket-01 contract.
+// constants; the CPU-side keys follow the producer contract (InferenceX docs/results-and-ingestion.md).
const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 };
const traySource: BenchmarkRow = {
...source,
diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts
index fcba58a1b..060272733 100644
--- a/packages/app/src/lib/modeled-system-power.test.ts
+++ b/packages/app/src/lib/modeled-system-power.test.ts
@@ -50,7 +50,7 @@ function row(overrides: Partial = {}): BenchmarkRow {
}
// One GB200 NVL72 compute tray: four GPUs on one host, two Grace sockets. The
-// CPU-side keys follow the ticket-01 contract (sums over every socket, same window).
+// CPU-side keys follow the producer contract (sums over every socket, same window).
// Watts are controlled inputs, not published constants.
const GRACE = { avg_cpu_socket_power_w: 250.5, avg_total_cpu_power_w: 501 };
function nvl72Row(
diff --git a/packages/db/src/etl/benchmark-mapper.test.ts b/packages/db/src/etl/benchmark-mapper.test.ts
index 815dae66f..2e21094c6 100644
--- a/packages/db/src/etl/benchmark-mapper.test.ts
+++ b/packages/db/src/etl/benchmark-mapper.test.ts
@@ -354,7 +354,7 @@ describe('mapBenchmarkRow', () => {
expect(result!.metrics.median_ttft).toBe(50.2);
});
- // Ticket 01 contract: cpu_power_valid is independent of power_valid. Each leg
+ // Producer contract: cpu_power_valid is independent of power_valid. Each leg
// withholds only its own keys; worker telemetry belongs to the GPU leg.
it.each([
{ power_valid: 1, cpu_power_valid: 1, gpuKept: true, cpuKept: true },
From d46b18ee46f1f6a1b9d050b1a4a1f8b058736237 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 12:04:58 -0700
Subject: [PATCH 09/22] fix: pin latest NVL72 reference revision
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:将 NVL72 参考模型固定到最新修订,重新生成来源哈希与版本信息;所有机箱、机架参数及参考用例数值保持不变。
---
.../app/scripts/generate-system-power-reference.py | 6 +++---
.../app/src/lib/system-power-model.profiles.json | 12 ++++++------
.../app/src/lib/system-power-model.reference.json | 4 ++--
3 files changed, 11 insertions(+), 11 deletions(-)
diff --git a/packages/app/scripts/generate-system-power-reference.py b/packages/app/scripts/generate-system-power-reference.py
index e149c4c66..d3e946eae 100644
--- a/packages/app/scripts/generate-system-power-reference.py
+++ b/packages/app/scripts/generate-system-power-reference.py
@@ -21,9 +21,9 @@
import sys
sys.dont_write_bytecode = True
-# feat/gb200-nvl72-rack-model on top of ca4403aa (PR #10 merge); pending push to the upstream repo.
-REVISION = "963ead8b20a722595c34f7f4a0041259501cf019"
-REVISION_STATUS = "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream"
+# Local reference branch; upstream publication needs repository write access.
+REVISION = "6fcc086b77576d4cecb9d0c79637d6daf980308c"
+REVISION_STATUS = "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification"
SOURCE = "https://github.com/SemiAnalysisAI/inferencex_power_model"
MODELS = {
"h100": ("hgx_h100_chassis/h100_chassis_power_model.py", "h100_chassis_power", "make_h100_config"),
diff --git a/packages/app/src/lib/system-power-model.profiles.json b/packages/app/src/lib/system-power-model.profiles.json
index 302cc781f..5424d48dd 100644
--- a/packages/app/src/lib/system-power-model.profiles.json
+++ b/packages/app/src/lib/system-power-model.profiles.json
@@ -1,9 +1,9 @@
{
- "modelRevision": "963ead8b20a722595c34f7f4a0041259501cf019",
- "modelRevisionStatus": "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream",
+ "modelRevision": "6fcc086b77576d4cecb9d0c79637d6daf980308c",
+ "modelRevisionStatus": "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification",
"source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
"status": "DRAFT / pending human verification",
- "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/963ead8b20a722595c34f7f4a0041259501cf019/README.md#chassis-models",
+ "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/6fcc086b77576d4cecb9d0c79637d6daf980308c/README.md#chassis-models",
"assumptions": {
"u_pcie": 0.05,
"u_cpu": 0.2,
@@ -11,7 +11,7 @@
"u_nvme": 0.0,
"pue": 1.2
},
- "rackAssumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/963ead8b20a722595c34f7f4a0041259501cf019/human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
+ "rackAssumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/6fcc086b77576d4cecb9d0c79637d6daf980308c/human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
"rackAssumptions": {
"u_nvlink": 0.5,
"u_ib": 0.0,
@@ -39,9 +39,9 @@
"human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7",
"human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db",
"human_verified/chassis_plot_utils.py": "0f1c809a527788c8c25490420e29e737e6b7f83cb914dc022c3a4e3701b80643",
- "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "f907916213a9176a4d82a05b905ed3c7b595035184150a352dc8049952ffd340",
+ "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "b4640f94c8f9e6be50eb18deff9728bb58559ba193b54577859cafabec4ccc33",
"human_verified/gb200_nvl72_rack/plot_gb200_nvl72_rack_inference_gpu_sweep.py": "d0f6903df0348ccb2790e973b0dc10408c6f6022695ce039aad0b1e1fcd39519",
- "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "4b33288998d4355d1f4812b9f2e014ab4cf2a588fb9fc94ac56fea9ae8922950",
+ "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "631bd7fd44b84260f16175a814d093795a4cdd0a91a0d19dd572a5ea6c088314",
"human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f",
"human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c",
"human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb",
diff --git a/packages/app/src/lib/system-power-model.reference.json b/packages/app/src/lib/system-power-model.reference.json
index 5e0f49c77..a2356ce4f 100644
--- a/packages/app/src/lib/system-power-model.reference.json
+++ b/packages/app/src/lib/system-power-model.reference.json
@@ -1,6 +1,6 @@
{
- "modelRevision": "963ead8b20a722595c34f7f4a0041259501cf019",
- "modelRevisionStatus": "branch feat/gb200-nvl72-rack-model, child of ca4403aa527069857351ad8047dbb726844b3382; pending push upstream",
+ "modelRevision": "6fcc086b77576d4cecb9d0c79637d6daf980308c",
+ "modelRevisionStatus": "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification",
"source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
"status": "DRAFT / pending human verification",
"cases": [
From 0861bdc1b3c0ee61d080a901195309e46348a2cd Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 12:18:48 -0700
Subject: [PATCH 10/22] docs: split PowerX modeling guide into English and
Chinese
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:将 PowerX 系统功耗与智能容量规划说明拆分为完整的中英文文档,补充 Kimi K3 缺失结果、NVL72 计算示例和模型更新流程,并明确参考源码尚未发布及校准边界。
---
docs/index.md | 2 +-
docs/powerx-system-power.md | 391 ++++++++++++++-----------
docs/powerx-system-power.zh.md | 511 +++++++++++++++++++++++++++++++++
3 files changed, 739 insertions(+), 165 deletions(-)
create mode 100644 docs/powerx-system-power.zh.md
diff --git a/docs/index.md b/docs/index.md
index 2e6c0224a..a3f5427bd 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -8,7 +8,7 @@ Design rationale and non-obvious conventions. See [CLAUDE.md](../CLAUDE.md) for
- [Pareto Boundary API](./pareto-api.md): Query frontier and hinterland observations, preserve provenance, and distinguish API scope from chart highlights.
- [API Skill Examples](./inferencex-api-examples.md) — Install the public skill, query benchmarks, export measured PowerX data, and explain empty results
-- [PowerX System Power](./powerx-system-power.md) — Pinned chassis model, measured-input guards, assumptions, and reproducible article exports
+- [PowerX System Power](./powerx-system-power.md) / [简体中文](./powerx-system-power.zh.md) — Smart provisioning walkthrough, Kimi K3 missing-row diagnosis, NVL72/chassis models, assumptions, and reproducible exports
- [PowerX Permanent View](./powerx-permanent-view.md) — Power boundaries as gated Measured Energy metrics, `i_metric`/`i_rulers` share links, missing-value states
- [PowerX Persistence and Recovery](./powerx-persistence-recovery.md) — Telemetry receipts, migration prerequisites, and targeted repair
- [API Skill Releases](./inferencex-skills-release.md) — Prepare an immutable package, verify clean installations and agent exports, and publish through the package-specific workflow
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index b6c00021f..67941e9e0 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -1,9 +1,25 @@
-# Modeled system power in PowerX
+# PowerX system power and smart provisioning
-PowerX can compare measured GPU-board watts with estimated chassis AC watts for
-the non-agentic 8192-input/1024-output workload. The existing app transformation
-and the offline article exporter both call `modelSystemPower`; the API and
-benchmark producer continue returning their original measurements.
+[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md)
+
+PowerX starts with benchmark measurements and estimates how much facility power
+is needed to run the measured workload. Smart provisioning uses that estimate to
+calculate capacity and economics for a fixed facility power budget.
+[PR #1190](https://github.com/SemiAnalysisAI/InferenceX-app/pull/1190) extends this
+path to GB200/GB300 NVL72 racks. It also preserves valid provisioned comparisons
+when the corresponding measured estimate is unavailable.
+
+The ordinary modeled-power chart supports the non-agentic 8192-input/1024-output
+workload. The Profit Estimator explicitly opts into AgentX estimates, including
+Kimi K3; that does not establish AgentX calibration. The app, derived views API,
+and offline exporter reuse the model. The producer and raw benchmark API retain
+the original measurements.
+
+Start with [the Kimi K3 example](#why-kimi-k3-can-have-power-data-but-no-smart-provisioning-row),
+then [the worked calculation](#a-worked-nvl72-calculation). The
+[architecture diagrams](#nvl72-architecture-walkthrough) and
+[code map](#where-each-part-lives) connect the explanation to implementation.
+The later sections retain the full admission, topology and export contracts.
`system-power-model.profiles.json` records the pinned power model revision,
component source hashes, hardware mapping, complete platform configuration, and
@@ -16,6 +32,188 @@ The pinned source currently identifies itself as **DRAFT / pending human
verification**. Numerical parity establishes implementation equivalence, not
empirical chassis calibration.
+## What is measured, modeled, and provisioned?
+
+These are different inputs to a calculation, not interchangeable labels:
+
+| Quantity | Meaning | Used for |
+| --------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------- |
+| GPU Level Measured | Validated GPU-board watts during the benchmark window | GPU power charts and the measured GPU input to system modeling |
+| GPU Level Provisioned (TDP) | The configured hardware TDP reference | GPU-level reference comparisons; not a measurement of this run |
+| All in Provisioned | The hardware registry's fixed facility kW/GPU allowance | Baseline capacity and profit planning |
+| All in Measured | Measured inputs plus modeled unmeasured components and PUE | Workload-dependent system-power estimates and smart provisioning |
+
+“All in Measured” still includes modeled components. The chart must retain that
+qualification. Smart provisioning changes the estimated GPU capacity per GW;
+it does not set GPU power limits or prove that average power plus a reserve is a
+safe electrical peak limit.
+
+The seven eight-GPU chassis profiles measure GPU boards and model CPU, DRAM and
+other chassis overhead. **NVL72 has a different boundary:** it requires measured
+compute-module power, or measured GPU-board power plus complete Grace-socket
+power. Grace and its LPDDR5X are not filled in with a CPU utilization estimate.
+Rack networking, switches, tray residuals and conversion losses are modeled.
+
+## Why Kimi K3 can have power data but no smart-provisioning row
+
+A GPU-power chart answers “what GPU power was measured?” The profit calculation
+answers “what whole-system power belongs to the performance point that serves
+this target?” A visible GPU measurement is only one of those required inputs.
+
+The September 29, 2026 reproduction used 93 saved public Kimi K3 benchmark rows,
+AgentX P90 at **45 tok/s/user**, and automatic FP4 selection. It reproduced the
+reported two priced configurations: B200 Dynamo-vLLM and MI355X ATOM. This is a
+historical reproduction of the reported screenshot, not a statement that today's
+live database has the same coverage. At #1190 head
+[cc86afd6](https://github.com/SemiAnalysisAI/InferenceX-app/commit/cc86afd6b179cceeb550cbf579ff60df46a47ac3),
+the model and admission rules explain those saved inputs as follows:
+
+| Configuration | Why GPU data does not produce a measured profit estimate | Work needed |
+| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| GB200 NVL72 | The 13 saved rows had valid GPU power but no accepted CPU/module metrics or CPU audit. #1190 supplies the rack model, but cannot supply those missing measurements. | Collect complete Grace/module and GPU telemetry for the same benchmark windows, validate it, and ingest the matched results. |
+| GB300 NVL72 | The 11 saved rows had no CPU/module evidence; only five had valid GPU power. The selected frontier also used GPU-invalid row 442101. | Recover valid GPU evidence where retained artifacts permit it, and obtain the missing Grace/module measurements. Both requirements must pass. |
+| MI355X vLLM | The original performance bracket used rows 443223 and 443220; 443220 had invalid power. Other valid GPU-power points belonged to a different bracket. | Diagnose the original point's retained telemetry. Reprocess only if it contains sufficient valid evidence; otherwise collect and publish replacement benchmark results with matched performance and power. |
+| B300 | The selected rows 439941/439935 had no measured power. | Supply qualified measurements for the serving curve used by the estimator. |
+| H200 | The available serving curve did not reach the requested 45 tok/s/user. | Use a target inside that curve's supported range, or obtain a qualified curve covering the requested target. |
+
+The screenshots also selected different engines: the power chart hid ATOM and
+showed MI355X vLLM, while the priced AMD result was ATOM. Comparisons must match
+model, workload, date/run, engine, precision, percentile and target before their
+row counts are compared.
+
+The AMD example is concrete: the performance interpolation at 45 uses
+**14.832 and 47.596 tok/s/user**, with invalid power at the second point. A valid
+power sample at **60.386 tok/s/user** cannot replace it without also changing
+which performance points support the estimate. Filtering out all bad-power rows
+and rebuilding the frontier would define a different comparison methodology.
+#1190 deliberately retains the original serving frontier.
+
+The saved invalid rows do not identify the collection failure's cause. Their
+public audit field was null, so this evidence does not justify blaming an
+exporter, deleting all AMD history, or promising that re-ingestion will repair
+it. New CPU readings from another run also cannot be attached to historical
+throughput results.
+
+### What #1190 fixes, and the remaining fix
+
+The PR fixes the unsupported NVL72 topology, validates the two accepted sensor
+boundaries, and keeps a valid **All in Provisioned** result in **Compare both**
+when the measured estimate is absent. It already emits a specific skip reason;
+**All in Measured** alone continues to omit unqualified numeric results.
+
+The existing collapsed **Unavailable estimates** disclosure already lists each
+configuration and its reason. A possible presentation follow-up is to make that
+information easier to find beside the results and link each entry to its
+measurement or target. This is a proposed improvement, not behavior added by
+this document. It should not turn missing measurements into zeroes or
+provisioned values under a measured label.
+
+To produce the missing numeric rows, first inspect the original run artifacts.
+If the required samples and provenance exist, validate and reprocess those same
+windows. If they were never collected, collect a replacement performance-and-power
+result together. Complete Grace/module coverage is required for NVL72; fixing
+only the rack model or only GPU validity is insufficient.
+
+## A worked NVL72 calculation
+
+This is an **illustrative reference fixture, not a measured Kimi K3 result**.
+Use the stored GB300 module-basis case with **3,000.75 W per complete tray**,
+four GPUs and two Grace sockets per tray, and PUE 1.1. A real admitted benchmark
+must separately pass the GPU and CPU/module audit checks described below.
+
+1. **Choose the sensor boundary.** Here the module reading already covers the
+ GPU/Grace compute module, so the model does not add GPU or Grace watts again.
+ On the alternative GPU-plus-Grace path, the current profile adds the measured
+ GPU and Grace totals, then a regulator allowance of GPU watts × 0.15 / 0.85.
+ That allowance is a model assumption, not an independently measured rail.
+2. **Evaluate one rack.** Replicate the measured mean across 18 compute trays,
+ add the profile's tray components and nine NVSwitch trays, then apply tray
+ conversion losses and two management switches. Evaluate the power-shelf
+ efficiency at this combined load, rather than evaluating a separate rack for
+ each measured tray.
+3. **Apply facility overhead once.** Multiply rounded rack AC power by PUE.
+ Allocate the resulting rack share over 72 GPUs. This assumes a rack whose
+ remaining trays run at the measured trays' mean load; it does not measure
+ the actual occupancy or power of an entire rack.
+4. **Apply planning headroom separately.** Multiply facility kW/GPU by 1.10.
+ PUE 1.1 and the 10% reserve are different factors with different purposes.
+
+The committed Python reference fixture contains these intermediate values:
+
+| Stage | Watts |
+| ---------------------------------------------- | -------: |
+| Measured module input × 18 trays | 54,013.5 |
+| Modeled compute-tray components | 11,466.0 |
+| Modeled NVSwitch trays | 4,107.6 |
+| Tray conversion losses | 1,967.8 |
+| Rack DC, including 200 W management switches | 71,754.9 |
+| Rack AC, after load-dependent shelf efficiency | 74,904.6 |
+| Facility power after PUE 1.1 | 82,395.1 |
+
+Displayed intermediate values are rounded. The calculation retains the source's
+summation and rounding order, so adding displayed values can differ by 0.1 W.
+The source's raw profile defaults include PUE 1.2; the dashboard wrapper supplies
+**1.1 for NVL72** and **1.3 for the supported air-cooled chassis**.
+
+ Planning kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
+ GPU capacity per GW = 1,000,000 / 1.258814 ≈ 794,399
+ GPU-hours per GW-year = GPU capacity × 8,760
+
+For an eight-GPU benchmark on two full trays, the assigned facility share would
+be 82,395.1 × 8 / 72 ≈ 9,155.0 W. Dividing that share by the eight measured GPUs
+gives the same per-GPU planning value. The estimator does not claim that this
+small benchmark measured all 72 GPUs.
+
+At a non-exact interactivity target, the existing performance frontier determines
+the two bounding benchmark points. Both must have valid planning power and the
+same model revision, PUE, topology and sensor basis. The code linearly
+interpolates planning kW/GPU between those original points. It does not replace
+the existing throughput interpolation or extrapolate beyond its measured range.
+
+The economics then reuse the same throughput, token-price schedule, utilization,
+license share and per-GPU-hour cost in both power modes:
+
+ Revenue = revenue per active GPU-hour × GPU-hours × utilization
+ Cost = cost per GPU-hour × GPU-hours
+ License share = revenue × license fraction
+ Profit = revenue − cost − license share
+
+Changing planning power changes capacity and therefore total revenue, cost and
+profit per GW. Under these fixed per-GPU assumptions it does not improve profit
+margin, throughput per GPU, or profit per chip-hour. This is a planning projection,
+not a prediction that workload behavior will remain unchanged at facility scale.
+
+## Where each part lives
+
+The links below are repository-relative so they follow the reviewed checkout.
+The walkthrough was checked against app head cc86afd6; the example comes from
+the committed reference JSON. No new hardware measurements were taken for it.
+
+| Responsibility | Implementation |
+| ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| Preserve CPU/GPU metrics and audit provenance during ingest | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts), [power-publication.ts](../packages/db/src/etl/power-publication.ts) |
+| Pin Python model revision, assumptions and component hashes | [generate-system-power-reference.py](../packages/app/scripts/generate-system-power-reference.py), [profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| Apply workload, validity, sensor and topology admission; allocate deployment shares | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
+| Calculate nonlinear chassis/rack power, losses and PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
+| Match power to the original frontier and retain provisioned comparisons | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
+| Convert planning kW/GPU into capacity, revenue, cost and profit | [profit-estimator.ts](../packages/app/src/components/calculator/profit-estimator.ts) |
+| Display values, provenance, unavailable reasons and CSV fields | [ProfitEstimatorDisplay.tsx](../packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx), [ProfitEstimatorChart.tsx](../packages/app/src/components/calculator/ProfitEstimatorChart.tsx) |
+| Reuse the same economics through the read-only views API | [calculator-extensions.ts](../packages/app/src/lib/views-api/calculator-extensions.ts) |
+| Export modeled benchmark rows offline | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
+| Check calculation parity and admission behavior | [reference fixtures](../packages/app/src/lib/system-power-model.reference.json), [model tests](../packages/app/src/lib/system-power-model.test.ts), [admission tests](../packages/app/src/lib/modeled-system-power.test.ts), [planning tests](../packages/app/src/components/calculator/profit-power.test.ts) |
+
+Publishing #1190 publishes its TypeScript implementation and bundled profiles;
+the runtime does not fetch Python from GitHub. The Python source pin is local commit
+`6fcc086b77576d4cecb9d0c79637d6daf980308c`, intended for the private
+[SemiAnalysisAI/inferencex_power_model](https://github.com/SemiAnalysisAI/inferencex_power_model)
+repository. That commit is **unpublished pending repository write access**;
+reviewers cannot yet retrieve it from that remote. Publishing the app bundle
+does not publish the Python source. Source publication, when completed, will not
+change its DRAFT status or establish empirical calibration. The source
+revision and file hashes belong in the review/export record, alongside which
+components remain uncalibrated. See the update procedure below for historical
+results and frozen exports.
+
## NVL72 architecture walkthrough
### Measurements, ingestion and system power
@@ -80,25 +278,17 @@ NVL72 model supplies a system-power calculation; it does not create missing
measurements or choose new performance points. The detailed gates below also
cover eight-GPU chassis and their existing replica extrapolation.
-第一张图说明从生产端、入库到机架估算的数据流。有效的 module 传感器已覆盖 GPU、
-HBM、Grace 和 LPDDR5X,无需重复提供 Grace 瓦数,但仍必须有完整的有效性与 socket
-覆盖审计;GPU 加 Grace 路径则需要有效的 Grace socket 实测。两种路径都不会用
-CPU 估算值补齐缺失的测量。
-
-第二张图保留原有性能前沿和插值点,只改变换算每 GW 容量时采用的功耗。缺少实测时,
-对比模式保留有效的预配功耗柱子,并解释实测结果为何不可用;仅看实测加建模的模式
-不会用预配值代替。实线为数据流,虚线为假设或来源信息。利润估算器仅采用正式前沿
-数据,NVL72 模型不会自动补齐缺失测量,也不会为补出柱子而重新选择性能点。
-
## Updating the model for historical results
Modeled power is derived from retained measurements when the browser or a shared
views API transforms a benchmark row. Changing the model does not rewrite the
original GPU measurements or require a per-run database backfill.
-1. Update `REVISION` in `packages/app/scripts/generate-system-power-reference.py`
- to the intended clean Python model commit, and update the recorded assumptions
- when required.
+1. Commit and publish the intended Python model revision so another reviewer can
+ obtain the exact source. Update `REVISION` in
+ `packages/app/scripts/generate-system-power-reference.py` to that clean commit,
+ update `REVISION_STATUS`, and update the recorded assumptions when required.
+ Publishing Python alone does not update the dashboard's bundled profiles.
2. Run that script with the path to the pinned model checkout to regenerate
`system-power-model.profiles.json` and `system-power-model.reference.json`.
If equations or load-dependent components changed, update the TypeScript
@@ -170,10 +360,11 @@ power shelves, and management switches are amortised over 72 GPUs. Chassis, by
contrast, own their fans and PSUs and are each evaluated at their own load. The
result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`.
A partially allocated tray extrapolates only the GPU-board share (a module reading
-already covers the whole tray) and is labeled `extrapolated`. The pinned revision
-`963ead8b` is the power-model repo's `feat/gb200-nvl72-rack-model` branch, pending
-push upstream. The NVL72 section below lists the measured input, the modeled
-residual and the Profit Estimator gate rules.
+already covers the whole tray) and is labeled `extrapolated`. The source is pinned to revision
+`6fcc086b77576d4cecb9d0c79637d6daf980308c`, a local commit intended for the private
+model repository linked above. It remains unpublished pending repository write
+access, and the model remains DRAFT / pending human verification. The NVL72 section below
+lists the measured input, the modeled residual and the Profit Estimator gate rules.
A partially allocated chassis (one to seven measured GPUs on one host) is
modeled at measured per-GPU power × 8. That is the same `n_gpu × W/GPU` input
@@ -232,20 +423,20 @@ trays' mean input (the shelf curve sees the whole rack's DC load, never one tray
and amortised over 72 GPUs; the parameters marked UNVERIFIED carry a documented
range in `unverifiedParameters` and no published rail:
-| Component (per rack unless noted) | GB200 | GB300 | Source status |
-| -------------------------------------- | ------------------------------------------------------------------------------------ | ---------------------------------------------- | ---------------------------------------- |
-| NVSwitch tray silicon (9 trays) | 406.4 W / tray at `u_nvlink` 0.5 | same | `blackwell_nvswitch` model |
-| NVSwitch tray residual | 50 W / tray | same | UNVERIFIED (20–80 W) |
-| Compute-tray NICs with optics | ConnectX-7, 4 × 31.5 W = 126 W / tray | ConnectX-8 integrated PCIe, 4 × 78.8 W = 315 W | `generic/connectx7`, `generic/connectx8` |
-| Compute-tray BlueField-3 DPUs | 2 × 65 W idle = 130 W / tray | same | `generic/dpu`, idle only |
-| Compute-tray NVMe | 22 W / tray idle | same | `generic/nvme`, idle only |
-| Compute-tray fans | 130 W / tray | same | UNVERIFIED (40–220 W) |
-| Compute-tray board residual | 40 W / tray | same | UNVERIFIED (20–60 W) |
-| Management switches | 2 × 100 W | same | profile constant |
-| Tray 50 V → 12 V conversion | efficiency 0.9725 on tray loads | same | UNVERIFIED (0.96–0.985) |
-| Regulator allowance (`gpu-plus-grace`) | 15% of GPU TDP, GPU share only | same | Grace tuning guide |
-| Power shelf | 264 kW installed, 132 kW redundant; efficiency 0.90 → 0.94 → 0.965 at 10/20/30% load | same | profile curve |
-| Facility PUE | 1.1 (direct liquid cooling) | same | PowerX policy, applied once to rack AC |
+| Component (per rack unless noted) | GB200 | GB300 | Source status |
+| -------------------------------------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------ | ---------------------------------------- |
+| NVSwitch tray silicon (9 trays) | 406.4 W / tray at `u_nvlink` 0.5 | same | `blackwell_nvswitch` model |
+| NVSwitch tray residual | 50 W / tray | same | UNVERIFIED (20–80 W) |
+| Compute-tray NICs with optics | ConnectX-7, 4 × 31.5 W = 126 W / tray | ConnectX-8 integrated PCIe, 4 NICs totaling 315 W (78.8 W each, rounded) | `generic/connectx7`, `generic/connectx8` |
+| Compute-tray BlueField-3 DPUs | 2 × 65 W idle = 130 W / tray | same | `generic/dpu`, idle only |
+| Compute-tray NVMe | 22 W / tray idle | same | `generic/nvme`, idle only |
+| Compute-tray fans | 130 W / tray | same | UNVERIFIED (40–220 W) |
+| Compute-tray board residual | 40 W / tray | same | UNVERIFIED (20–60 W) |
+| Management switches | 2 × 100 W | same | profile constant |
+| Tray 50 V → 12 V conversion | efficiency 0.9725 on tray loads | same | UNVERIFIED (0.96–0.985) |
+| Regulator allowance (`gpu-plus-grace`) | GPU-board W × 0.15 / 0.85; GPU share only | same | Grace tuning guide |
+| Power shelf | 264 kW installed, 132 kW redundant; efficiency 0.90 → 0.94 → 0.965 at 10/20/30% load | same | profile curve |
+| Facility PUE | 1.1 (direct liquid cooling) | same | PowerX policy, applied once to rack AC |
Rack DC above the installed shelf capacity (264 kW) overflows the efficiency curve
and the row is unavailable (`model-domain`). Rounding follows the source: rack AC is rounded
@@ -318,36 +509,6 @@ assumptions and, per row, the measured basis, sensor kind and profile.
Unavailable historical estimates use the hardware registry when a chip is absent
from today's results and include the source date/run label.
-每 GW 利润估算器在基准测试配置中提供预配功耗、实测加建模功耗,以及两种方式的同口径
-对比;默认仍采用预配功耗。两种方式使用同一硬件、P90 目标、吞吐量前沿、token 比例、
-价格、利用率和每 GPU 小时成本,仅改变换算每 GW 容量时采用的设施功率。因此收入、
-计算成本、模型许可费和利润按相同比例变化,利润率不变;不会另行重新计算电费。
-
-该选项由现有内部功能开关控制,锁定时隐藏。按 ↑↑↓↓ 解锁(本地存储
-`inferencex-feature-gate=1`)。锁定时,`c_power` 不会启用其他估算方式或触发完整功耗
-数据请求;重新锁定后立即恢复预配功耗估算。
-
-AgentX 估算仅接纳通过验证的 schema-v2 功耗和适用的系统模型。完整八卡机箱可采用
-单节点或各物理主机的实测功耗;没有逐主机遥测的聚合多节点配置,可明确假设各完整
-机箱均采用部署平均功耗。分离式部署仍需要逐主机、逐角色功耗。NVL72 必须具有完整
-四卡 tray,以及 CPU 审计确认的每 tray 两个 socket;完整 module 读数和 Grace socket
-读数分别处理,只有 CPU rail 或缺少传感器来源的记录不能代替整颗 Grace 的功耗。
-
-实测单节点单卡、双卡或四卡配置保留整机外推:假设八卡服务器放置多个完整实例,
-每卡功耗和吞吐量不变,再将建模设施功耗除以八。图表、提示框和 CSV 标注该假设;
-这不代表部分 GPU 闲置时的整机实测。其他部分分配以及缺失或无效功耗仍不可用。
-对比模式保留有效的预配功耗结果,并单独说明实测加建模结果为何缺失;只看实测加
-建模模式时不会用预配功耗代替。原有
-8K/1K 转换路径的接纳规则不变。精确前沿点使用自身的功耗;点间采用原吞吐量插值的
-同一对数据点线性估算功耗,不换用其他点填补缺失,也不会把两种实测口径不同的数据点
-混合估算。PUE 按 PowerX 策略取值(风冷机箱 1.3,液冷 NVL72 机架 1.1),另加 10%
-功耗余量;这些假设和上述固定 CPU/DRAM 利用率尚未通过 AgentX 系统校准,也不构成
-峰值供电容量验证。界面保留简短的实测与建模说明,详细假设和 NVL72 实测输入默认
-折叠;CSV 保留完整假设,并逐行标出实测口径、传感器类型和所用 profile。分享链接
-通过 `c_power` 保留所选方式。
-历史估算不可用时,若当天结果不含该芯片,则从硬件注册表获取名称;提示会附上来源
-日期或运行标签,避免与当前结果混淆。
-
## Offline comparison export
The exporter reads a local cohort envelope and writes a **new** output directory:
@@ -420,98 +581,6 @@ replicate outputs; it does not evaluate the model at mean watts. If any replicat
is unavailable, the corresponding mean remains unavailable rather than silently
dropping that replicate.
-## 中文说明
-
-模型结果在浏览器或共享 views API 转换 benchmark 数据时计算,不写回原始 GPU
-测量值。更新模型时,先修改生成脚本中的固定版本及相关假设,再生成 profiles 和
-reference JSON;如果公式或随负载变化的组件有改动,还需同步 TypeScript 实现。
-通过一致性及准入测试后部署,刷新浏览器,并使派生 API 缓存失效或等待其过期。
-冻结的 CSV/JSON 需另行导出。通常无需逐 run 回填数据库;若新模型需要历史记录中
-没有的输入,应保留不可用状态,也不能把新一轮测得的功耗配到旧吞吐结果上。
-
-PowerX 的系统功耗结果以实测 GPU 功率为输入,使用固定版本的功耗模型估算
-8-GPU 机箱的 AC 输入功率,再单独应用 PUE 得到设施功率估计。CPU 和 DRAM 利用率
-均假设为 20%;这些是模型参数,不是实测利用率。完整平台配置、源码版本和校验和
-随导出结果保留。模型源码仍标记为待人工核验,数值一致性不代表完成了实机校准。
-
-固定版本的 Python 模型默认 PUE 为 1.2;PowerX 对当前风冷机箱模型
-采用 1.3。市电侧功率 = IT 负载功率 × PUE,风冷取 1.3,直接液冷(DLC)
-取 1.1。PUE 仅作用于机箱交流功率,不改变 GPU 实测功率或机箱交流功率。这里的
-冷却方式指建模机箱,并非已核实的测试站点配置。机箱模型不支持 DLC;`--pue` 仅
-覆盖设施功率系数,不会把风冷机箱模型转换为液冷模型。NVL72 机架 profile 为直接
-液冷,默认 PUE 取 1.1。
-
-仅使用部分 GPU 的机箱(单台主机上实测 1–7 张 GPU)按实测每卡功率 × 8 建模,
-与模型源码 sweep 脚本喂给各机箱模型的 `n_gpu × W/GPU` 输入一致,并假设未实测的
-GPU 运行相同负载。结果标记为 `chassisBasis: 'extrapolated'`:每卡数值按建模机箱
-的 GPU 总数分摊,`deploymentAcWatts` 只保留实测 GPU 在各机箱中的份额。这不是把
-部分分配的机箱按比例分摊:固定组件、风扇曲线和 PSU 效率都在满机箱负载点求值。
-GB200、GB300 使用单独的 NVL72 机架 profile,不套用 B200、B300 机箱模型:每台实测
-worker 主机视为一个计算 tray(4 张 GPU、2 个 Grace socket);没有逐 worker 数组的聚合
-多节点行则按 GPU 总数 ÷ 4 推算 tray 数、每个 tray 取部署平均值,并与 Grace socket 数及
-CPU 采集记录的 `power_audit.cpu.observed_sockets` 交叉校验。输入为实测模块功耗
-(`avg_total_module_power_w`);缺失时改用 GPU 板卡功耗加 Grace socket 功耗
-(`avg_total_cpu_power_w`),并按来源模型计入 GPU 份额的稳压损耗余量。Grace CPU 与
-LPDDR5X 从不建模,缺少 `cpu_power_valid=1` 或完整 module/Grace 来源与覆盖记录的行保持不可用
-(`cpu-telemetry`)。实测的各 tray 先折算为一个由 18 个与其均值相同的 tray 组成的
-整机架,电源架效率曲线只在该机架的直流总负载处求值一次(与来源模型
-`gb200_nvl72_rack_power` 只接受单一 per-tray 输入的做法一致),每个 tray 取其 1/18,
-因此 NVSwitch tray、电源架和管理交换机按 72 张 GPU 分摊;机箱则各自拥有风扇和 PSU,
-仍按各自负载单独求值。部分分配的 tray 只外推 GPU 板卡份额(模块读数本身已覆盖整个
-tray),并标记为 `extrapolated`。缺失、无效和不支持的情况保持不可用。纯 CPU frontend
-worker 不计入 GPU 机箱数;独立的纯 CPU frontend/router 主机不在估算范围内,GPU 机箱内
-的 CPU 功率仍按 20% 利用率计算。
-
-导出时每次测量先独立计算,再对同一 cell 的重复测量取平均。能耗使用审计记录中的
-实际窗口和成功 token 数,按实测 GPU 的份额计算,明确标记为估计值,不改写原有
-GPU 实测指标。当前 API 快照与原文章冻结数据分别导出,避免混用不同时间和配置的
-结果。未指定 `--pue` 时,每行按其硬件采用与仪表板相同的默认 PUE(风冷机箱 1.3,
-液冷 NVL72 机架 1.1),因此文章数据与图表悬停一致;显式 `--pue` 记录在
-`metadata.pue_override` 中并覆盖所有行。NVL72 行另附实测口径、传感器类型、
-按 `cpu_power_valid` 保留的 Grace socket 与模块实测输入、机架 profile 及其假设,
-以及机架专用的边界与外推说明;x86 行保持不变。
-
-**NVL72 机架估算(GB200、GB300)。** 实测输入:每个计算 tray 使用实测的计算模块功耗,
-Grace CPU 与 LPDDR5X 从不建模。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与
-GPU 能耗相同的正式窗口内输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w`、
-`total_cpu_energy_j`,当每个 socket 都有 `Module Power Socket` 传感器时还输出
-`avg_total_module_power_w` 与 `total_module_energy_j`,并附带独立的验证结论
-`cpu_power_valid` 和 `power_audit.cpu`(传感器类型、采集来源、socket 覆盖情况、原因码)。
-接纳条件:`power_valid=1`、schema 2、`cpu_power_valid=1`,且 `power_audit.cpu`
-记录的预期与实测 socket 数相符,每个 tray 为两个 socket。module 读数须注明
-`sensor_kind: module`,不要求重复提供 Grace 指标;GPU 加 Grace 口径须注明
-`sensor_kind: grace_socket`,Grace 功耗为正,且总功耗除以平均功耗与审计 socket 数
-一致。仅 CPU rail、传感器来源缺失或未知时保持不可用。存在
-`avg_total_module_power_w` 时采用 `module` 口径(读数已包含 GPU 板卡,不再缩放),
-否则采用 `gpu-plus-grace` 口径(每 tray GPU 板卡功耗 × 4 加 Grace socket 总功耗,
-并仅对 GPU 份额计入来源模型的稳压损耗余量)。模块指标存在但无效时该行不可用
-(`cpu-telemetry`),不会悄然回退到 Grace socket。
-
-建模残差:计算模块之外的部分全部来自固定版本 profile(`rackProfiles`),按实测 tray
-均值构成的 18 tray 整机架求值一次(电源架曲线看到的是整机架直流负载,而非单个
-tray),再分摊到 72 张 GPU:NVSwitch tray 硅片功耗(9 个 tray,
-`blackwell_nvswitch` 模型,`u_nvlink` 0.5)及 tray 残差(50 W,UNVERIFIED);每个计算
-tray 的网卡与光模块(GB200 为 ConnectX-7,4 × 31.5 W;GB300 为 ConnectX-8,4 × 78.8 W)、
-BlueField-3 DPU 空闲功耗(2 × 65 W)、NVMe 空闲功耗(22 W)、风扇(130 W,UNVERIFIED)
-和主板残差(40 W,UNVERIFIED);2 台管理交换机(各 100 W);tray 内 50 V → 12 V 转换
-效率 0.9725(UNVERIFIED);电源架装机容量 264 kW、冗余容量 132 kW,效率曲线在 10%/20%/30%
-负载处为 0.90/0.94/0.965;液冷 PUE 1.1,仅对机架交流功率应用一次。机架直流功率超过电源架
-装机容量(264 kW)时该行不可用(`model-domain`)。舍入与来源一致:机架交流功率先四舍五入到 0.1 W
-再乘 PUE;Python 生成的 `rackCases` 对两种变体、两种口径、全部电源架拐点和 PUE 1.0–1.2
-验证了与固定实现的一致性。
-
-门槛规则(利润估算器):规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。接受
-`chassisBasis: 'full'` 的八卡机箱估算(`single-node`、`worker-hosts` 或 `uniform-hosts`
-拓扑),或全部 tray 均完整实测的 `nvl72-trays` 估算(各 4 张 GPU、2 个 socket:每个实测
-worker 主机一个 tray,或没有逐 worker 数组的聚合多节点行按 GPU 总数 ÷ 4 推算、并与
-socket 数交叉校验);部分 tray 在图表中外推显示,但不进入规划门槛。单节点 1/2/4 卡
-机箱仍接受上文的完整实例外推。两个前沿数据点之间必须采用相同的实测口径和传感器类型,模块
-读数旁边的 Grace socket 读数保持不可用,不会混合两种传感器。柱形提示、默认折叠的
-“功耗估算假设”详情和 CSV 的 `功耗口径`、`功耗传感器`、`系统功耗 profile` 三列逐行标出实测口径
-(实测模块功耗,或实测 GPU 板卡 + Grace socket 功耗并由模型估算稳压损耗)、传感器类型
-和所用 profile(`modelPath @ modelRevision sha256:<源文件哈希>`)。`?unofficialrun=`
-叠加层规则不适用于利润估算器的功耗口径控件:估算器只对正式前沿数据点定价。
-
## Measured P75 and P90 GPU power
`y_measuredP75Power` and `y_measuredP90Power` show the time-weighted P75 and P90 of
@@ -523,9 +592,3 @@ chassis AC power and individual-device percentiles.
P75 and P90 backfills use the same 34 original validated traces and exact windows
recorded in `docs/data/power-p90-backfill.json`.
-
-`y_measuredP75Power` 和 `y_measuredP90Power` 分别显示已验证负载窗口内整组 GPU
-功耗按时间加权的 P75 和 P90,再按参与测量的 GPU 数量均摊。正式数据与非正式
-运行叠加层使用同一计算和绘图路径。缺少测量值或未通过验证时保持不可用,
-不会用平均功耗替代。该指标与机箱交流功耗估算、单个设备的功耗分位数不同。
-两个分位数均由审计记录中的同一批 34 份原始遥测及其测量窗口重新计算。
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
new file mode 100644
index 000000000..d28d43f05
--- /dev/null
+++ b/docs/powerx-system-power.zh.md
@@ -0,0 +1,511 @@
+# PowerX 系统功耗与智能容量规划
+
+[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md)
+
+PowerX 以基准测试的实测数据为起点,估算运行该负载所需的设施功率。智能容量规划
+(smart provisioning)再根据这一估算,计算固定设施功率预算下的 GPU 容量及经济指标。
+[PR #1190](https://github.com/SemiAnalysisAI/InferenceX-app/pull/1190) 将这条路径扩展到
+GB200/GB300 NVL72 机架。当相应的实测加建模估算不可用时,对比模式仍保留有效的预配结果。
+
+常规建模功耗图支持非 AgentX 的 8192 输入 / 1024 输出负载。利润估算器显式启用
+AgentX 估算,包括 Kimi K3;这不代表模型已经过 AgentX 校准。应用、派生 views API
+和离线导出器共用同一模型。生产端和原始 benchmark API 保留原始测量值。
+
+建议先看 [Kimi K3 的例子](#为什么-kimi-k3-有-gpu-功耗却没有智能容量规划结果),再看
+[完整计算示例](#nvl72-计算示例)。[架构图](#nvl72-架构导览) 和
+[代码位置](#各部分代码在哪里) 将说明与实现对应起来。后面的章节完整记录了数据接纳、
+拓扑和导出约定。
+
+`system-power-model.profiles.json` 记录固定的功耗模型版本、组件源码哈希、硬件映射、
+完整平台配置和固定推理假设。各 profile 由原始 Python 组件执行生成。
+`system-power-model.ts` 保留原模型的非线性风扇曲线、PSU 效率插值、中间值舍入和 PUE
+应用顺序。由 Python 生成的参考用例用于核对 TypeScript 实现与源模型是否一致。
+
+固定版本的源码仍标记为 **DRAFT / pending human verification(草稿,待人工核验)**。
+数值一致性只能证明实现等价,不能证明已经过实机机箱校准。
+
+## 实测、建模和预配分别指什么?
+
+它们是计算中的不同输入,不能互换名称:
+
+| 指标 | 含义 | 用途 |
+| --------------------------- | ---------------------------------------- | ---------------------------------------- |
+| GPU Level Measured | 基准测试窗口内通过验证的 GPU 板卡功率 | GPU 功耗图,以及系统模型的实测 GPU 输入 |
+| GPU Level Provisioned (TDP) | 硬件配置中的 TDP 参考值 | GPU 层面的参考对比;不是本次运行的测量值 |
+| All in Provisioned | 硬件注册表中固定的设施 kW/GPU 配额 | 容量和利润规划的基线 |
+| All in Measured | 实测输入,加上未实测组件的模型估算及 PUE | 随负载变化的系统功耗估算和智能容量规划 |
+
+“All in Measured” 仍包含建模组件,图表必须保留这一限定。智能容量规划改变的是每 GW
+能容纳的 GPU 数量估算;它不会设置 GPU 功率上限,也不能证明“平均功耗加余量”就是
+安全的供电峰值上限。
+
+七种八卡机箱 profile 使用 GPU 板卡实测值,并对 CPU、DRAM 和其他机箱开销建模。
+**NVL72 的测量边界不同:**它要求实测计算模块功耗,或实测 GPU 板卡功耗加上完整的
+Grace socket 功耗。不能用 CPU 利用率假设补齐 Grace 及其 LPDDR5X 功耗。机架网络、
+交换机、tray 其余组件和转换损耗由模型估算。
+
+## 为什么 Kimi K3 有 GPU 功耗,却没有智能容量规划结果?
+
+GPU 功耗图回答的是“测到了多少 GPU 功耗”;利润计算回答的是“满足这个性能目标的
+数据点,对应多少整机功耗”。图上有 GPU 实测值,只能说明其中一项输入存在。
+
+2026 年 9 月 29 日的复现使用了保存的 93 条公开 Kimi K3 基准测试记录,设置为 AgentX
+P90、**45 tok/s/user**、自动选择 FP4。它复现了截图中两个可计价配置:B200 Dynamo-vLLM
+和 MI355X ATOM。这是对当时截图的历史复现,不代表今天的在线数据库仍有相同的数据覆盖。
+按 #1190 的
+[cc86afd6](https://github.com/SemiAnalysisAI/InferenceX-app/commit/cc86afd6b179cceeb550cbf579ff60df46a47ac3)
+版本,其模型和接纳规则可以这样解释这批输入:
+
+| 配置 | 为什么有 GPU 数据,却算不出实测加建模的利润结果 | 需要做什么 |
+| ----------- | ------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- |
+| GB200 NVL72 | 保存的 13 条记录都有有效 GPU 功耗,但没有符合要求的 CPU/module 指标或 CPU 审计。#1190 提供了机架模型,无法补出缺失测量。 | 在相同基准测试窗口内采集完整 Grace/module 和 GPU 遥测,通过验证后将匹配的结果入库。 |
+| GB300 NVL72 | 保存的 11 条记录都没有 CPU/module 证据,其中只有五条 GPU 功耗有效。所选性能前沿还使用了 GPU 功耗无效的记录 442101。 | 若保留的产物足以证明 GPU 功耗有效,则恢复该证据;同时补齐 Grace/module 测量。两项要求都必须满足。 |
+| MI355X vLLM | 原性能插值区间使用记录 443223 和 443220;443220 的功耗无效。其他功耗有效的数据点属于另一个区间。 | 检查原数据点保留的遥测。只有存在充分有效证据时才重新处理;否则重新采集并发布性能与功耗匹配的基准测试结果。 |
+| B300 | 所选记录 439941/439935 没有实测功耗。 | 为估算器使用的服务性能曲线补充通过验证的测量。 |
+| H200 | 现有服务性能曲线达不到所要求的 45 tok/s/user。 | 将目标设在曲线支持的范围内,或取得覆盖该目标且通过验证的新曲线。 |
+
+两张截图选择的引擎也不同:功耗图隐藏了 ATOM,显示 MI355X vLLM;有价格结果的 AMD
+配置则是 ATOM。比较结果数量前,应先对齐模型、负载、日期/运行、引擎、精度、分位数
+和目标值。
+
+AMD 的例子更具体:45 tok/s/user 的性能插值使用 **14.832 和 47.596 tok/s/user** 两个点,
+第二个点的功耗无效。不能直接借用 **60.386 tok/s/user** 处的有效功耗,否则支撑估算的
+性能点也随之改变。如果先删掉所有功耗无效的记录,再重建前沿,就采用了另一套比较方法。
+#1190 明确保留原有服务性能前沿。
+
+保存的无效记录不能说明采集失败的原因:它们的公开审计字段为 null。因此,这份证据
+不足以归咎于某个 exporter、删除全部 AMD 历史数据,或承诺重新入库就能修复。
+也不能把另一轮运行新采集的 CPU 功耗附到历史吞吐量上。
+
+### #1190 修复了什么,还缺什么?
+
+该 PR 支持了原先不支持的 NVL72 拓扑,验证两种可接纳的传感器测量边界,并在实测
+估算缺失时,让 **Compare both** 继续保留有效的 **All in Provisioned** 结果。
+代码已经输出具体的跳过原因;单独选择 **All in Measured** 时,仍不显示不符合要求的数值。
+
+现有默认折叠的 **Unavailable estimates** 详情已逐项列出配置及不可用原因。后续可以
+改善展示,让读者更容易在结果旁找到这些信息,并跳转到对应测量或目标。这只是建议的
+展示改进,不是本文新增的行为。不能把缺失测量改成零,也不能在实测标签下填入预配值。
+
+要补齐缺失的数值结果,应先检查原运行的产物。若所需样本和来源记录齐全,就验证并
+重新处理原测量窗口;若从未采集,则需要一起重测性能和功耗,形成替代结果。
+NVL72 必须具备完整 Grace/module 覆盖;仅修机架模型或仅修 GPU 有效性都不够。
+
+## NVL72 计算示例
+
+以下是**用于说明计算过程的参考测试数据,不是 Kimi K3 实测结果**。采用已保存的 GB300
+module 口径用例:**每个完整 tray 为 3,000.75 W**,每 tray 四张 GPU、两个 Grace socket,
+PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU/module 审计。
+
+1. **确定传感器测量边界。**本例的 module 读数已包含 GPU/Grace 计算模块,模型不会再
+ 加一遍 GPU 或 Grace 功耗。另一条 GPU 加 Grace 路径则先相加两者的实测总功耗,再按
+ GPU 功耗 × 0.15 / 0.85 加入稳压损耗余量。这个余量是模型假设,不是独立测得的供电轨功耗。
+2. **计算一个整机架。**把实测均值扩展到 18 个计算 tray,加入 profile 中的 tray 组件和
+ 九个 NVSwitch tray,再计入 tray 转换损耗和两台管理交换机。在合计负载处计算电源架
+ 效率,而不是为每个实测 tray 单独计算一整套机架。
+3. **只应用一次设施开销。**将舍入后的机架交流功率乘以 PUE,再按 72 张 GPU 分摊。
+ 这里假设机架其余 tray 的负载与实测 tray 的平均负载相同,并未测量整个机架的实际
+ 占用情况或总功耗。
+4. **单独计入规划余量。**将设施 kW/GPU 乘以 1.10。PUE 1.1 与 10% 规划余量是不同
+ 的系数,作用也不同。
+
+已提交的 Python 参考测试数据包含以下中间值:
+
+| 阶段 | 功率(W) |
+| ------------------------------------------ | --------: |
+| 实测 module 输入 × 18 个 tray | 54,013.5 |
+| 建模的计算 tray 组件 | 11,466.0 |
+| 建模的 NVSwitch tray | 4,107.6 |
+| tray 转换损耗 | 1,967.8 |
+| 机架直流功率,含 200 W 管理交换机 | 71,754.9 |
+| 计入随负载变化的电源架效率后的机架交流功率 | 74,904.6 |
+| 乘以 PUE 1.1 后的设施功率 | 82,395.1 |
+
+表中中间值已舍入。计算保留源模型的求和与舍入顺序,因此直接相加表中数值可能差
+0.1 W。源模型原始 profile 的默认 PUE 为 1.2;仪表板封装层对 **NVL72 使用 1.1**,
+对**当前支持的风冷机箱使用 1.3**。
+
+ 规划 kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
+ 每 GW 的 GPU 容量 = 1,000,000 / 1.258814 ≈ 794,399
+ 每 GW 每年的 GPU 小时数 = GPU 容量 × 8,760
+
+若基准测试使用两个完整 tray、共八张 GPU,则分摊的设施功率为
+82,395.1 × 8 / 72 ≈ 9,155.0 W。再除以八张实测 GPU,得到的每 GPU 规划功率相同。
+估算器并不声称这次小规模测试测量了全部 72 张 GPU。
+
+当交互性能目标不恰好落在实测点上时,由现有性能前沿决定两侧的基准测试点。两点都
+必须具备有效规划功率,且模型版本、PUE、拓扑和传感器口径相同。代码在原有两点之间
+线性插值规划 kW/GPU,不替换原吞吐量插值,也不向实测范围之外外推。
+
+随后,两种功耗模式共用相同的吞吐量、token 定价、利用率、模型许可分成和每 GPU 小时成本:
+
+ 收入 = 每个活跃 GPU 小时的收入 × GPU 小时数 × 利用率
+ 成本 = 每 GPU 小时成本 × GPU 小时数
+ 模型许可分成 = 收入 × 许可分成比例
+ 利润 = 收入 − 成本 − 模型许可分成
+
+规划功率改变容量,因而改变每 GW 的总收入、总成本和总利润。在每 GPU 假设固定的
+情况下,它不会提高利润率、每 GPU 吞吐量或每芯片小时利润。这是容量规划推算,
+并不预测负载扩展到整个设施后仍保持相同行为。
+
+## 各部分代码在哪里
+
+以下链接均为仓库内相对链接,随当前检出的版本变化。本导览依据应用版本 cc86afd6
+核对,示例来自已提交的参考 JSON;没有为本文新增硬件测量。
+
+| 职责 | 实现 |
+| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| 入库时保留 CPU/GPU 指标和审计来源 | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts)、[power-publication.ts](../packages/db/src/etl/power-publication.ts) |
+| 固定 Python 模型版本、假设和组件哈希 | [generate-system-power-reference.py](../packages/app/scripts/generate-system-power-reference.py)、[profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| 检查负载、有效性、传感器和拓扑条件;分摊部署功耗 | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
+| 计算非线性机箱/机架功耗、损耗和 PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
+| 将功耗匹配到原前沿,并保留预配对比结果 | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
+| 将规划 kW/GPU 换算为容量、收入、成本和利润 | [profit-estimator.ts](../packages/app/src/components/calculator/profit-estimator.ts) |
+| 显示数值、来源、不可用原因和 CSV 字段 | [ProfitEstimatorDisplay.tsx](../packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx)、[ProfitEstimatorChart.tsx](../packages/app/src/components/calculator/ProfitEstimatorChart.tsx) |
+| 通过只读 views API 复用同一经济指标计算 | [calculator-extensions.ts](../packages/app/src/lib/views-api/calculator-extensions.ts) |
+| 离线导出建模后的基准测试记录 | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
+| 验证计算一致性和接纳行为 | [参考测试数据](../packages/app/src/lib/system-power-model.reference.json)、[模型测试](../packages/app/src/lib/system-power-model.test.ts)、[接纳规则测试](../packages/app/src/lib/modeled-system-power.test.ts)、[规划测试](../packages/app/src/components/calculator/profit-power.test.ts) |
+
+发布 #1190 会发布 TypeScript 实现及随应用打包的 profiles,运行时不会从 GitHub 下载
+Python。Python 源码固定在本地提交 `6fcc086b77576d4cecb9d0c79637d6daf980308c`,
+计划发布至私有仓库
+[SemiAnalysisAI/inferencex_power_model](https://github.com/SemiAnalysisAI/inferencex_power_model)。
+该提交**尚未发布,正在等待仓库写入权限**;审阅者目前无法从该远端取得这个提交。
+发布应用 bundle 不等于发布 Python 源码。日后完成源码发布,也不会改变其 DRAFT 状态,
+更不代表完成了实测校准。审阅和导出记录应保留源码版本、文件哈希,以及哪些组件仍
+未经校准。历史结果和冻结导出的更新方式见后文。
+
+## NVL72 架构导览
+
+### 测量、入库与系统功耗
+
+```mermaid
+flowchart TB
+ subgraph Producer["生产端 — InferenceX #3296"]
+ GPU["实测 GPU 板卡功耗"]
+ CPU["实测 Grace / 计算模块功耗"]
+ ART["基准测试结果与 power_audit.cpu 测量窗口一致"]
+ GPU --> ART
+ CPU --> ART
+ end
+ ART --> ETL["benchmark-mapper 保留有效性、指标与传感器来源"]
+ ETL --> DB[("基准测试记录与指标")]
+ DB --> API["Benchmark API"]
+ API --> GATE{"modelSystemPower GPU 验证结论与 CPU/module 审计有效? GPU、socket、worker 拓扑一致?"}
+ GATE -->|"否"| NONE["不可用,并给出原因"]
+ GATE -->|"是"| TRAY["计算 tray 每个完整 tray 为 4 张 GPU + 2 个 Grace socket"]
+ TRAY --> BASIS{"实测口径"}
+ BASIS -->|"Module 传感器"| MODULE["实测 module 功耗 无需单独提供 Grace 功耗 不重复加入 GPU / Grace 功耗"]
+ BASIS -->|"Grace socket 传感器"| SUM["实测 GPU 板卡 + Grace 功耗 按 profile 对 GPU 份额计入稳压损耗余量"]
+ MODULE --> RACK
+ SUM --> RACK
+ PROFILE["固定版本的 GB200 / GB300 机架 profile 静态负载与转换假设"] -.-> RACK
+ RACK["按 tray 平均负载扩展至 18 tray 机架 加入机架其余组件和电源架损耗模型"]
+ RACK --> AC["机架交流功率"]
+ AC --> FAC["只应用一次 PUE NVL72 默认值:1.1"]
+ FAC --> DEP["向实测部署分摊机架功耗 保留口径、profile 与来源"]
+```
+
+有效的 module 传感器已覆盖 GPU、HBM、Grace 和 LPDDR5X;这条路径仍要求完整的
+module/socket 审计,但无需重复提供 Grace 功耗字段。GPU 加 Grace 路径则要求有效的
+socket 测量。两条路径都不会用 CPU 估算值补齐缺失的 CPU/module 遥测。
+module 测量存在但无效时,结果保持不可用,不会悄然回退到另一条路径。
+
+### 规划、对比与输出
+
+```mermaid
+flowchart TB
+ PERF["正式基准测试性能 所选分位数和目标"] --> FRONT["原服务性能前沿 精确点或原有区间两端点"]
+ POWER["通过验证的部署设施功率 来自图 1"] --> ACCEPT
+ FRONT --> ACCEPT{"NVL72 tray 均完整实测? 目标在实测范围内? 两点的口径与传感器相同?"}
+ ACCEPT -->|"否"| SKIP["实测加建模估算不可用 保留具体原因"]
+ ACCEPT -->|"是"| POINT["逐点计算规划功率 设施 kW / GPU x 1.10"]
+ POINT --> SMART["匹配目标的规划 kW/GPU 在原吞吐量数据点之间插值 不换点填补缺失功耗"]
+ SPEC["预配 kW/GPU"] --> CAP
+ SMART --> CAP["相同设施功率预算 计算可部署的 GPU 容量"]
+ FRONT --> INPUT["相同吞吐量、token 价格、 利用率和单位成本"]
+ INPUT --> ECON
+ CAP --> ECON["收入、成本和利润"]
+ ECON --> UI["All in Provisioned / All in Measured / Compare both 图表、提示框、详情和 CSV"]
+ SKIP --> KEEP["对比模式保留有效预配柱子 仅实测模式不以预配值替代"]
+ KEEP --> UI
+ META["传感器口径、PUE、10% 余量 profile 版本和源码哈希"] -.-> UI
+```
+
+实线表示数据流,虚线提供假设或来源信息。规划采用正式性能前沿点,不采用
+`?unofficialrun=` 叠加层。NVL72 模型提供系统功耗计算,不会创造缺失测量,也不会重新
+选择性能点。下文的详细接纳规则还涵盖八卡机箱及其现有的实例外推方式。
+
+## 更新模型后,历史结果如何更新?
+
+浏览器或共享 views API 转换基准测试记录时,会根据保留的测量值推导建模功耗。
+修改模型不会改写原始 GPU 测量值,也不需要逐 run 回填数据库。
+
+1. 提交并发布要使用的 Python 模型版本,让其他审阅者能够取得完全相同的源码。
+ 将 `packages/app/scripts/generate-system-power-reference.py` 的 `REVISION` 更新为
+ 该干净提交,并更新 `REVISION_STATUS`;必要时同步记录的假设。仅发布 Python 不会
+ 更新仪表板打包的 profiles。
+2. 运行该脚本,传入固定版本模型 checkout 的路径,重新生成
+ `system-power-model.profiles.json` 和 `system-power-model.reference.json`。
+ 如果公式或随负载变化的组件有改动,还必须修改 TypeScript 实现;只重新生成常量不够。
+3. 运行系统功耗模型的一致性测试和接纳规则测试,再部署应用。已有浏览器会话需要
+ 加载新 bundle;派生 API 响应需要走正常的认证缓存失效流程,或等待缓存过期。
+ 仅完成部署,不能证明所有缓存响应都已采用新版本。
+4. 冻结的 CSV/JSON 导出需单独重新生成。若新模型需要从未记录的输入,相应记录应
+ 保持不可用,直至输入缺口解决。不能把新基准测试的功耗附到旧基准测试的吞吐量上。
+
+## 计算边界与假设
+
+输入是已验证服务窗口内的平均 GPU 实测功率。机箱交流功耗模型在此基础上,加入
+源模型中的 CPU、DRAM、网络、存储、主板、风扇和 PSU 转换损耗。设施功率另行估算:
+先算机箱交流功率,再应用 PUE,并保留源模型的舍入顺序。
+
+README 中固定的推理 sweep 使用 `u_cpu=0.20`、`u_ram=0.20`、`u_pcie=0.05` 和
+`u_nvme=0.0`。固定版本的 Python 模型默认 PUE 为 `1.2`;PowerX 对当前支持的风冷
+机箱 profile 使用 `1.3`。市电侧功率 = IT 负载功率 × PUE(风冷 `1.3`,直接液冷 DLC
+`1.1`)。该系数作用于机箱交流功率之后,不改变 GPU 实测功率或机箱交流功率。
+这里的冷却方式指模型中的机箱,并非已经核实的基准测试站点冷却配置。机箱 profile
+不支持 DLC;`--pue` 只是显式覆盖设施功率系数,不会把风冷机箱模型转成液冷模型。
+下文的 NVL72 机架 profile 为直接液冷,默认使用 `1.1`。
+
+生成的 profile 保留各平台的网络假设、风扇控制、组件数量和机箱默认值。每份 JSON
+导出包含完整 profile,每行 CSV 包含适用假设及 profile 哈希。这些是模型输入,
+不是实测 CPU/DRAM 利用率。
+
+| 硬件标识 | 机箱模型源码 |
+| -------- | ------------------------------------------------------------- |
+| `h100` | `human_verified/hgx_h100_chassis/h100_chassis_power_model.py` |
+| `h200` | `human_verified/hgx_h200_chassis/h200_chassis_power_model.py` |
+| `b200` | `human_verified/hgx_b200_chassis/b200_chassis_power_model.py` |
+| `b300` | `human_verified/hgx_b300_chassis/b300_chassis_power_model.py` |
+| `mi300x` | `human_verified/mi300x_chassis/mi300x_chassis_power_model.py` |
+| `mi325x` | `human_verified/mi325x_chassis/mi325x_chassis_power_model.py` |
+| `mi355x` | `human_verified/mi355x_chassis/mi355x_chassis_power_model.py` |
+
+上述 profile 均描述完整八卡机箱。GB200 和 GB300 使用独立的 NVL72 机架 profile,
+不套用这些机箱拓扑:
+
+| 硬件标识 | 机架模型源码 |
+| -------- | ----------------------------------------------------------------- |
+| `gb200` | `human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py` |
+| `gb300` | 同一模块中的 `gb300_nvl72_rack_config` |
+
+机架 profile(`rackProfiles`)的输入是每个 tray 的**实测**计算模块功耗:生产端发布
+`avg_total_module_power_w` 时使用 module 传感器总值;否则使用 GPU 板卡功耗加
+Grace socket 总功耗(`avg_total_cpu_power_w`),并按源模型对 GPU 份额计入稳压损耗
+余量。Grace CPU 和 LPDDR5X 从不由模型补算。缺少 `cpu_power_valid=1` 或完整 module /
+Grace 来源记录的行保持不可用(`cpu-telemetry`)。
+
+每个实测 worker 主机对应一个计算 tray(四张 GPU、两个 Grace socket)。没有逐 worker
+数组的聚合多节点记录,按 `gpuCount / 4` 推算 tray 数,各 tray 采用部署平均值,并与
+Grace socket 数及 CPU 采集记录中的 `power_audit.cpu.observed_sockets` 交叉校验。
+按实测 tray 的平均计算模块输入,构造一个包含 18 个同等负载 tray 的机架;电源架效率
+曲线只在整机架直流负载处求值一次。这与源模型 `gb200_nvl72_rack_power` 接受单一
+per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray、电源架和管理
+交换机按 72 张 GPU 分摊。机箱则各自拥有风扇和 PSU,按各自的负载单独求值。
+结果包含 `topologyBasis: 'nvl72-trays'`、`measuredBasis` 和 `sensorKind`。
+
+部分分配的 tray 只外推 GPU 板卡份额,因为 module 读数本身已覆盖整个 tray;结果标记
+为 `extrapolated`。源码固定在本地提交
+`6fcc086b77576d4cecb9d0c79637d6daf980308c`,计划发布至上文的私有模型仓库,
+目前仍因等待仓库写入权限而未发布。模型仍为 DRAFT / pending human verification。
+下文列出 NVL72 的实测输入、其余组件
+模型,以及利润估算器的接纳规则。
+
+部分分配的机箱,即单台主机上实测一至七张 GPU,按“实测每 GPU 功率 × 8”建模。
+这与源模型 sweep 脚本使用的 `n_gpu × W/GPU` 输入一致,并假设未实测的 GPU 运行相同
+负载。估算标记为 `chassisBasis: 'extrapolated'`:每 GPU 数值按建模机箱 GPU 数
+(`modeledGpuCount`)分摊;`deploymentAcWatts` / `deploymentFacilityWatts` 只保留
+各机箱中实测 GPU 的份额。这不是先在部分负载处计算机箱,再按比例分摊;固定组件、
+风扇曲线和 PSU 效率都在满机箱负载处求值。缺失或无效遥测、数量不一致、缺少主机
+放置记录、每主机多于一个机箱,以及超出模型适用范围的情况,仍保持不可用。
+
+单节点部署的物理宽度按生产端的 `TP * PP * PCP` 计算,EP 在该宽度内划分。
+部分已有 API 配置别名含有 `TP * EP`;模型不直接相信或相加这些别名,而是用实测
+总功率和每 GPU 功率交叉核对物理宽度。多节点和分离式输入要求每个实测 worker
+对应一个机箱(一至八张 GPU)、worker 位于不同主机,并且总功率与角色功率一致。
+只有角色平均值,无法证明物理放置方式,也无法计算各主机的非线性模型。
+纯 CPU frontend worker 不计入 GPU 机箱数量。独立的纯 CPU frontend/router 主机
+不在估算范围内;GPU 机箱内的 CPU 功耗仍采用源模型固定的 20% 利用率假设。
+
+默认实测约定要求数值型 `power_valid=1` 和指标 schema 2。原有通过验证的单节点生产端
+早于 schema 标记,但其两个 watts 字段的定义已经相同。这条路径保留 schema 缺失的
+原状,并报告 `validated-unversioned-single-node`;不会升级源数据版本,也不会接纳
+无版本的分离式功耗。文章的验证回执还会固定生产端 checkout,并保留各原始审计产物。
+
+## NVL72 机架估算(GB200、GB300)
+
+**实测输入。**每个计算 tray 都使用实测计算模块功耗,Grace CPU 和 LPDDR5X 不由模型
+估算。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与 GPU 能耗相同的正式窗口内
+输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w`、`total_cpu_energy_j`。
+若每个 socket 都有 `Module Power Socket` 传感器,还会输出
+`avg_total_module_power_w` 和 `total_module_energy_j`,并附带独立验证结论
+`cpu_power_valid` 及 `power_audit.cpu`(传感器类型、采集器、socket 覆盖情况、原因码)。
+
+接纳条件为 `power_valid=1`、schema 2、`cpu_power_valid=1`,且 `power_audit.cpu` 中
+预期和实测 socket 数一致,每 tray 两个。module 读数要求 `sensor_kind: module`,
+无需重复提供 Grace 指标。GPU 加 Grace 路径要求 `sensor_kind: grace_socket`、Grace
+功率为正,且总功率和平均功率与审计 socket 数一致。仅有 CPU rail 测量,或传感器
+来源缺失、未知的记录,保持不可用。
+
+口径选择:存在 `avg_total_module_power_w` 时采用 `module`,因为读数已包含 GPU
+板卡,所以不会再缩放;否则采用 `gpu-plus-grace`,即每 tray 的每 GPU 板卡功率 × 4
+加 Grace socket 总功率,并且只对 GPU 份额应用源模型的稳压损耗系数
+`regulatorLossFracOfTdp / (1 − frac)`。module 字段存在但无效时,该行不可用
+(`cpu-telemetry`),不会悄然回退到 Grace socket。
+
+**其余组件的模型估算。**计算模块以外的部分均来自固定版本 profile(`rackProfiles`)。
+按实测 tray 的平均输入构造 18 tray 机架,整体求值一次,再按 72 张 GPU 分摊。
+电源架曲线使用整机架的直流负载,不能只使用单个 tray 的负载。标为 UNVERIFIED 的
+参数在 `unverifiedParameters` 中记录了范围,但没有已发布的供电轨测量:
+
+| 组件(未注明时按每机架计) | GB200 | GB300 | 来源状态 |
+| -------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------- | ---------------------------------------- |
+| NVSwitch tray 芯片(9 个 tray) | `u_nvlink` 为 0.5 时,每 tray 406.4 W | 相同 | `blackwell_nvswitch` 模型 |
+| NVSwitch tray 其余组件 | 每 tray 50 W | 相同 | UNVERIFIED(20–80 W) |
+| 计算 tray 网卡及光模块 | ConnectX-7,每 tray 4 × 31.5 W = 126 W | 集成 PCIe 的 ConnectX-8,4 个 NIC 合计 315 W(每个 78.8 W,已舍入) | `generic/connectx7`、`generic/connectx8` |
+| 计算 tray BlueField-3 DPU | 每 tray 空闲功耗 2 × 65 W = 130 W | 相同 | `generic/dpu`,仅空闲功耗 |
+| 计算 tray NVMe | 每 tray 空闲功耗 22 W | 相同 | `generic/nvme`,仅空闲功耗 |
+| 计算 tray 风扇 | 每 tray 130 W | 相同 | UNVERIFIED(40–220 W) |
+| 计算 tray 主板其余组件 | 每 tray 40 W | 相同 | UNVERIFIED(20–60 W) |
+| 管理交换机 | 2 × 100 W | 相同 | profile 常量 |
+| tray 内 50 V → 12 V 转换 | tray 负载处效率 0.9725 | 相同 | UNVERIFIED(0.96–0.985) |
+| 稳压损耗余量(`gpu-plus-grace`) | GPU 板卡功率 × 0.15 / 0.85;仅计 GPU 板卡功率,不含 Grace | 相同 | Grace 调优指南 |
+| 电源架 | 装机容量 264 kW,冗余容量 132 kW;在 10/20/30% 负载处效率为 0.90 → 0.94 → 0.965 | 相同 | profile 曲线 |
+| 设施 PUE | 1.1(直接液冷) | 相同 | PowerX 策略,仅对机架交流功率应用一次 |
+
+机架直流功率超过电源架装机容量(264 kW)时,超出效率曲线适用范围,该行不可用
+(`model-domain`)。舍入顺序与源模型一致:先将机架交流功率舍入到 0.1 W,再应用 PUE。
+Python 生成的 `rackCases` 验证了两种变体、两种口径、全部电源架曲线节点和 PUE 1.0–1.2
+下与固定实现的数值一致性。
+
+**利润估算器的接纳规则。**规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。
+接受完整实测的八卡机箱(`single-node`、`worker-hosts` 或 `uniform-hosts` 拓扑下的
+`chassisBasis: 'full'`,见下一节),或者所有 tray 都完整实测的 `nvl72-trays` 估算。
+后者要求每 tray 四张 GPU、两个 socket:每个实测 worker 主机对应一个 tray;没有逐
+worker 数组的聚合多节点记录则按 `gpuCount / 4` 推算 tray 数,各 tray 采用部署平均值,
+并与 Grace socket 数和 `power_audit.cpu.observed_sockets` 交叉校验。部分 tray 可在
+功耗图中外推,但不进入利润规划。单节点 1/2/4 卡机箱采用下文的实例外推。
+
+两个前沿点必须采用相同实测口径和传感器类型;若一端是 module、另一端是 Grace
+socket,结果不可用,不混合两种传感器。柱形提示框、默认折叠的 Power assumptions
+详情,以及 CSV 中的 `Power basis`、`Power sensor`、`System power profile` 列,会逐行
+列明口径、传感器类型和固定 profile。口径为实测 module,或实测 GPU 板卡 + Grace
+socket 并由模型估算稳压损耗;profile 表示为
+`modelPath @ modelRevision sha256:`。
+`?unofficialrun=` 叠加层规则不适用于利润估算器的功耗口径控件;估算器只对正式前沿点计价。
+
+## 利润估算器的功耗口径
+
+每 GW 利润估算器在 Benchmark Config 中提供预配功耗、实测加建模功耗,以及两者的
+成对对比,默认仍采用预配功耗。控件由现有内部功能开关控制,锁定时隐藏。
+按 ↑↑↓↓ 解锁(本地存储 `inferencex-feature-gate=1`)。锁定时,`c_power` 不会启用
+其他计算方式或请求完整功耗记录;重新锁定后立即恢复预配估算。
+
+另一种功耗口径复用相同硬件、P90 目标、吞吐量前沿、token 比例、价格、利用率和每 GPU
+小时成本,只改变换算每 GW 容量时采用的设施 kW/GPU。因此,收入、计算成本、模型
+许可费和利润按相同比例变化,利润率不变。不会另行重新计算电费。
+
+这项显式启用的 AgentX 估算要求通过验证的 schema-v2 遥测和适用的系统 profile。
+完整八卡机箱可采用单节点或逐 worker 主机的实测功耗。没有逐 worker 遥测的聚合
+多节点记录,可以显式采用 `uniform-hosts` 假设:每个完整机箱都按部署平均 GPU 功率
+计算。分离式记录要求逐主机、逐角色功耗。NVL72 要求完整四卡 tray,并具有上文所述
+CPU 来源证据。
+
+通过验证的单节点 1/2/4 卡配置保留整机箱外推:假设八卡服务器放置若干完整实例,每 GPU
+功耗和吞吐量保持实测值,再将建模设施功耗除以八。该假设认为实例同机部署不影响性能
+或功耗,不代表测量了部分 GPU 闲置的服务器。图表、提示框和 CSV 对所有外推估算
+作出标注,包括仅一端为部分分配点的插值。其他部分分配方式以及缺失、无效测量仍不可用,
+并给出不同原因。常规 8K/1K 转换路径保留原有接纳策略。
+
+即使实测加建模功耗不可用,对比模式仍保留每个有效预配估算,并在提示中指出缺失的
+实测加建模估算;只看建模功耗的模式不会用预配功耗代替。
+
+目标恰好落在前沿点上时,使用该点的建模功耗;目标在两点之间时,采用原吞吐量插值
+的同一对点,线性估算功耗。不换用其他点填补缺失,也不混合测量口径不同的两点。
+估算采用 PowerX 的 PUE 策略(风冷机箱 1.3,直接液冷 NVL72 机架 1.1),另加 10%
+规划余量。这些假设,包括上述固定 CPU/DRAM 利用率,不构成峰值供电容量验证,也未
+经过 AgentX 系统校准。
+
+界面保留简短的实测与建模说明,详细假设和 NVL72 实测输入放在默认折叠的详情中。
+CSV 保留完整假设,并逐行记录实测口径、传感器类型和 profile。分享链接通过
+`c_power=modeled` 或 `c_power=compare` 保留选择。历史估算不可用且当天结果不含该
+芯片时,从硬件注册表获取芯片信息,并显示来源日期/运行标签。
+
+## 离线对比导出
+
+导出器读取本地数据组封装文件,并写入一个**新的**输出目录:
+
+```sh
+bun packages/app/scripts/export-modeled-system-power.ts \
+ --input /path/to/original-qwen-article.input.json \
+ --output /path/to/new-original-qwen-comparison
+
+bun packages/app/scripts/export-modeled-system-power.ts \
+ --input /path/to/qwen35-current.input.json \
+ --output /path/to/new-current-qwen-comparison --pue 1.3
+```
+
+未指定 `--pue` 时,每行通过与图表相同的 `modelSystemPower` 路径,采用该硬件在
+仪表板中的默认值:风冷机箱 profile 为 `1.3`,直接液冷 NVL72 机架 profile 为 `1.1`。
+因此文章数值与图表悬停值一致。显式传入 `--pue` 时,它适用于所有行,并记录在
+`metadata.pue_override` 中;`metadata.pue_defaults` 记录各硬件的默认值。
+
+NVL72 行还包含 `measured_basis`、`sensor_kind`、按独立 `cpu_power_valid` 保留的
+Grace socket 与 module 实测输入、机架 profile 的 `model_path` 和假设,以及机架专用的
+`calculation_boundary` / `extrapolation_note`。x86 行保持不变。
+
+脚本中维护的输入类型为 `ComparisonInput`:
+
+```ts
+{
+ cohort: string,
+ metadata: { /* source URLs, capture times, hashes and cohort selection */ },
+ rows: [{
+ id: string, // stable observation identity
+ cell?: string, // optional group of original replicates
+ benchmark: BenchmarkRow, // original API row or existing ETL output
+ rawInput?: unknown, // original artifact before ETL normalization
+ source?: object, // run, attempt, producer revision and artifact receipt
+ audit?: object // matching original power-validation sidecar
+ }]
+}
+```
+
+代码保留原类型和注释:`metadata` 记录来源 URL、采集时间、哈希和数据组选择条件;
+`id` 是稳定的观测标识,`cell` 可将原始重复测量分组;`benchmark` 是原 API 行或现有
+ETL 输出;`rawInput` 保留 ETL 规范化前的产物;`source` 记录运行、attempt、生产端版本
+和产物回执;`audit` 是与该观测匹配的原始功耗验证 sidecar。
+
+原始生产端聚合数据使用现有 `normalizeArtifactRows` / `mapBenchmarkRow` 处理。
+将原聚合数据保留为 `rawInput`,保留原 schema 标记,并确认规范化没有改变实测指标。
+当前快照应使用完整原始 API 响应,再在本地筛选精确的 `single_turn`、`isl=8192`、
+`osl=1024` 负载。保留范围内所有行,包括不支持的硬件以及缺失/无效功耗;不得把
+当前快照混入文章冻结的数据集。
+
+输出为 `comparison.json`、`comparison.csv`,以及可选的 `cells.csv`。JSON 包含原始
+输入行、实测有效性、建模输出、假设、审计窗口和来源记录。CSV 将实测与建模字段
+分开,不可用数值留空。无效或未经验证的原始值仍保留在原始输入记录中,不标为有效
+测量。元数据记录输入与实现的 SHA-256 哈希、模型与应用版本、worktree 状态、生成时间
+和完整 profile 来源。最终发布导出应从指定应用提交生成;文件哈希也能标识开发期间
+的本地修改。
+
+只有匹配且有效的审计 sidecar 提供了精确遥测时长、物理 GPU 数,以及成功请求/token
+分母时,才能估算能耗。计算方式是建模部署功率,即实测 GPU 在各机箱中的份额,乘以
+该时长;它不是对实测墙插功率做时间积分。输出 token 使用实际计数,不能用每次查询
+名义上的 `1024` token 代替。不推导 kernel 层面的 prefill/decode 能耗。
+没有这些 sidecar 的 API 快照只导出功耗估算。
+
+每次重复测量先独立建模,再求聚合值。cell 均值取各重复测量的建模输出平均值,不能
+先平均功耗再运行模型。任何一次重复测量不可用时,对应均值也保持不可用,不能默默
+丢掉那次测量后再平均。
+
+## GPU 实测 P75 和 P90 功耗
+
+`y_measuredP75Power` 和 `y_measuredP90Power` 分别表示:在已验证负载窗口内,对同步
+采集的整组 GPU 板卡功耗计算时间加权 P75、P90,再除以 GPU 数量。正式数据点与
+非正式运行叠加层共用常规实测功耗图路径。分位数缺失或未通过验证时保持不可用,
+绝不用平均功耗替代。这些指标与建模机箱交流功耗、单设备功耗分位数不同。
+
+P75 和 P90 回填使用 `docs/data/power-p90-backfill.json` 中记录的同一批 34 份原始
+有效遥测,以及完全相同的测量窗口。
From 86101230ff59471dfb6014c72d3607b9028a4da1 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 12:23:54 -0700
Subject: [PATCH 11/22] docs: clarify telemetry contracts and model assumptions
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:明确两种 CPU 传感器路径均发布验证结论与审计,修正 schema 2 字段定义及八卡整机外推的说明,并对齐中英文用语。
---
docs/powerx-system-power.md | 2 +-
docs/powerx-system-power.zh.md | 22 +++++++++++-----------
2 files changed, 12 insertions(+), 12 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 67941e9e0..f3e326f88 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -391,7 +391,7 @@ still uses the source's fixed 20% utilization assumption.
The default measured contract is numeric `power_valid=1` and metric schema 2.
The original validated single-node producer predates the schema marker but
-already defines both watts fields identically. This path retains the absent
+already uses the schema-2 definitions for both watts fields. This path retains the absent
schema and reports `validated-unversioned-single-node`; it does not upgrade the
source or admit unversioned disaggregated power. The article receipt additionally
pins the producer checkout and retains each original audit artifact.
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index d28d43f05..e913d6ad7 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -50,7 +50,7 @@ GPU 功耗图回答的是“测到了多少 GPU 功耗”;利润计算回答
数据点,对应多少整机功耗”。图上有 GPU 实测值,只能说明其中一项输入存在。
2026 年 9 月 29 日的复现使用了保存的 93 条公开 Kimi K3 基准测试记录,设置为 AgentX
-P90、**45 tok/s/user**、自动选择 FP4。它复现了截图中两个可计价配置:B200 Dynamo-vLLM
+P90、**45 tok/s/user**、自动选择 FP4。它复现了截图中两个有价格结果的配置:B200 Dynamo-vLLM
和 MI355X ATOM。这是对当时截图的历史复现,不代表今天的在线数据库仍有相同的数据覆盖。
按 #1190 的
[cc86afd6](https://github.com/SemiAnalysisAI/InferenceX-app/commit/cc86afd6b179cceeb550cbf579ff60df46a47ac3)
@@ -64,8 +64,8 @@ P90、**45 tok/s/user**、自动选择 FP4。它复现了截图中两个可计
| B300 | 所选记录 439941/439935 没有实测功耗。 | 为估算器使用的服务性能曲线补充通过验证的测量。 |
| H200 | 现有服务性能曲线达不到所要求的 45 tok/s/user。 | 将目标设在曲线支持的范围内,或取得覆盖该目标且通过验证的新曲线。 |
-两张截图选择的引擎也不同:功耗图隐藏了 ATOM,显示 MI355X vLLM;有价格结果的 AMD
-配置则是 ATOM。比较结果数量前,应先对齐模型、负载、日期/运行、引擎、精度、分位数
+这些截图选择的引擎也不同:功耗图隐藏了 ATOM,显示 MI355X vLLM;有价格结果的 AMD
+配置则是 ATOM。比较记录数量前,应先对齐模型、负载、日期/运行、引擎、精度、分位数
和目标值。
AMD 的例子更具体:45 tok/s/user 的性能插值使用 **14.832 和 47.596 tok/s/user** 两个点,
@@ -329,7 +329,7 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
不在估算范围内;GPU 机箱内的 CPU 功耗仍采用源模型固定的 20% 利用率假设。
默认实测约定要求数值型 `power_valid=1` 和指标 schema 2。原有通过验证的单节点生产端
-早于 schema 标记,但其两个 watts 字段的定义已经相同。这条路径保留 schema 缺失的
+早于 schema 标记,但两个 watts 字段的定义已与 schema 2 相同。这条路径保留 schema 缺失的
原状,并报告 `validated-unversioned-single-node`;不会升级源数据版本,也不会接纳
无版本的分离式功耗。文章的验证回执还会固定生产端 checkout,并保留各原始审计产物。
@@ -337,10 +337,10 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
**实测输入。**每个计算 tray 都使用实测计算模块功耗,Grace CPU 和 LPDDR5X 不由模型
估算。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与 GPU 能耗相同的正式窗口内
-输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w`、`total_cpu_energy_j`。
-若每个 socket 都有 `Module Power Socket` 传感器,还会输出
-`avg_total_module_power_w` 和 `total_module_energy_j`,并附带独立验证结论
-`cpu_power_valid` 及 `power_audit.cpu`(传感器类型、采集器、socket 覆盖情况、原因码)。
+输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w` 和 `total_cpu_energy_j`,并附带
+独立验证结论 `cpu_power_valid` 及 `power_audit.cpu`(传感器类型、采集器、socket 覆盖情况、
+原因码)。若每个 socket 都有 `Module Power Socket` 传感器,还会输出
+`avg_total_module_power_w` 和 `total_module_energy_j`。
接纳条件为 `power_valid=1`、schema 2、`cpu_power_valid=1`,且 `power_audit.cpu` 中
预期和实测 socket 数一致,每 tray 两个。module 读数要求 `sensor_kind: module`,
@@ -350,7 +350,7 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
口径选择:存在 `avg_total_module_power_w` 时采用 `module`,因为读数已包含 GPU
板卡,所以不会再缩放;否则采用 `gpu-plus-grace`,即每 tray 的每 GPU 板卡功率 × 4
-加 Grace socket 总功率,并且只对 GPU 份额应用源模型的稳压损耗系数
+加 Grace socket 总功率,并且只对 GPU 份额应用源模型的稳压损耗余量
`regulatorLossFracOfTdp / (1 − frac)`。module 字段存在但无效时,该行不可用
(`cpu-telemetry`),不会悄然回退到 Grace socket。
@@ -412,8 +412,8 @@ socket 并由模型估算稳压损耗;profile 表示为
计算。分离式记录要求逐主机、逐角色功耗。NVL72 要求完整四卡 tray,并具有上文所述
CPU 来源证据。
-通过验证的单节点 1/2/4 卡配置保留整机箱外推:假设八卡服务器放置若干完整实例,每 GPU
-功耗和吞吐量保持实测值,再将建模设施功耗除以八。该假设认为实例同机部署不影响性能
+通过验证的单节点 1/2/4 卡配置保留整机箱外推:按实测的每 GPU 功耗和吞吐量,用完整实例
+填满一台八卡服务器,再将建模设施功耗除以八。该假设认为实例同机部署不影响性能
或功耗,不代表测量了部分 GPU 闲置的服务器。图表、提示框和 CSV 对所有外推估算
作出标注,包括仅一端为部分分配点的插值。其他部分分配方式以及缺失、无效测量仍不可用,
并给出不同原因。常规 8K/1K 转换路径保留原有接纳策略。
From 989bb69f4e9422cc5743a3ba31dcf80a3b5a7799 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 12:38:56 -0700
Subject: [PATCH 12/22] refactor: own PowerX models in InferenceX-app
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:将现有 TypeScript 公式、有效参数和来源版本统一由 InferenceX-app 维护,删除依赖私有 Python 仓库的生成流程。保留历史数值回归基线,同步中英文说明、API、导出及提示框链接;功耗计算结果不变。
---
docs/powerx-system-power.md | 199 +--
docs/powerx-system-power.zh.md | 163 ++-
.../cypress/component/power-compare.cy.tsx | 86 +-
packages/app/next.config.ts | 5 +
.../scripts/export-modeled-system-power.ts | 3 +-
.../generate-system-power-reference.py | 282 ----
.../scripts/update-system-power-provenance.ts | 51 +
.../src/app/api/v1/views/extensions.test.ts | 4 +-
.../calculator/profit-power.test.ts | 4 +-
.../src/components/calculator/profit-power.ts | 2 +-
.../inference/utils/tooltip-utils.test.ts | 31 +-
.../inference/utils/tooltipUtils.ts | 14 +-
.../lib/modeled-system-power-export.test.ts | 8 +-
packages/app/src/lib/modeled-system-power.ts | 2 +-
.../src/lib/system-power-model.profiles.json | 1261 +---------------
.../lib/system-power-model.provenance.json | 13 +
.../src/lib/system-power-model.reference.json | 1272 ++++++++++++++++-
.../app/src/lib/system-power-model.test.ts | 82 +-
packages/app/src/lib/system-power-model.ts | 17 +-
.../app/src/lib/views-api/docs/extensions.ts | 4 +-
.../references/dashboard-views.md | 5 +-
21 files changed, 1786 insertions(+), 1722 deletions(-)
delete mode 100644 packages/app/scripts/generate-system-power-reference.py
create mode 100644 packages/app/scripts/update-system-power-provenance.ts
create mode 100644 packages/app/src/lib/system-power-model.provenance.json
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index f3e326f88..66b0907e0 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -21,16 +21,19 @@ then [the worked calculation](#a-worked-nvl72-calculation). The
[code map](#where-each-part-lives) connect the explanation to implementation.
The later sections retain the full admission, topology and export contracts.
-`system-power-model.profiles.json` records the pinned power model revision,
-component source hashes, hardware mapping, complete platform configuration, and
-fixed inference assumptions. Its profiles come from executing the original
-Python components. `system-power-model.ts` preserves their nonlinear fan curve,
-PSU efficiency interpolation, intermediate rounding, and PUE ordering. The
-Python-generated reference cases test this implementation against the source.
-
-The pinned source currently identifies itself as **DRAFT / pending human
-verification**. Numerical parity establishes implementation equivalence, not
-empirical chassis calibration.
+The model belongs to InferenceX-app. `system-power-model.ts` owns the equations,
+including nonlinear fan curves, PSU efficiency interpolation, intermediate
+rounding, and PUE ordering. `system-power-model.profiles.json` owns the editable
+component parameters, hardware mapping, platform configuration, and metadata
+describing the fixed inference scenario. These checked-in files are the source of truth for the frontend,
+shared views API, and offline exporter. `system-power-model.provenance.json`
+records their app model revision and source hashes.
+
+The checked-in reference cases retain the historical Python baseline for
+regression comparison; Python and its former private repository are not required
+to develop, build, or deploy the app model. The model remains **DRAFT / pending
+human verification**. Matching a numerical baseline establishes implementation
+consistency, not empirical chassis calibration.
## What is measured, modeled, and provisioned?
@@ -138,7 +141,8 @@ must separately pass the GPU and CPU/module audit checks described below.
4. **Apply planning headroom separately.** Multiply facility kW/GPU by 1.10.
PUE 1.1 and the 10% reserve are different factors with different purposes.
-The committed Python reference fixture contains these intermediate values:
+The checked-in reference fixture, originally captured from the Python baseline,
+contains these intermediate values:
| Stage | Watts |
| ---------------------------------------------- | -------: |
@@ -150,9 +154,9 @@ The committed Python reference fixture contains these intermediate values:
| Rack AC, after load-dependent shelf efficiency | 74,904.6 |
| Facility power after PUE 1.1 | 82,395.1 |
-Displayed intermediate values are rounded. The calculation retains the source's
+Displayed intermediate values are rounded. The calculation retains the model's
summation and rounding order, so adding displayed values can differ by 0.1 W.
-The source's raw profile defaults include PUE 1.2; the dashboard wrapper supplies
+The profile's baseline defaults include PUE 1.2; the dashboard wrapper supplies
**1.1 for NVL72** and **1.3 for the supported air-cooled chassis**.
Planning kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
@@ -192,7 +196,7 @@ the committed reference JSON. No new hardware measurements were taken for it.
| Responsibility | Implementation |
| ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Preserve CPU/GPU metrics and audit provenance during ingest | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts), [power-publication.ts](../packages/db/src/etl/power-publication.ts) |
-| Pin Python model revision, assumptions and component hashes | [generate-system-power-reference.py](../packages/app/scripts/generate-system-power-reference.py), [profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| Edit component parameters; refresh app model revision and source hashes | [profiles](../packages/app/src/lib/system-power-model.profiles.json), [provenance](../packages/app/src/lib/system-power-model.provenance.json), [update-system-power-provenance.ts](../packages/app/scripts/update-system-power-provenance.ts) |
| Apply workload, validity, sensor and topology admission; allocate deployment shares | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
| Calculate nonlinear chassis/rack power, losses and PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
| Match power to the original frontier and retain provisioned comparisons | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
@@ -202,17 +206,13 @@ the committed reference JSON. No new hardware measurements were taken for it.
| Export modeled benchmark rows offline | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
| Check calculation parity and admission behavior | [reference fixtures](../packages/app/src/lib/system-power-model.reference.json), [model tests](../packages/app/src/lib/system-power-model.test.ts), [admission tests](../packages/app/src/lib/modeled-system-power.test.ts), [planning tests](../packages/app/src/components/calculator/profit-power.test.ts) |
-Publishing #1190 publishes its TypeScript implementation and bundled profiles;
-the runtime does not fetch Python from GitHub. The Python source pin is local commit
-`6fcc086b77576d4cecb9d0c79637d6daf980308c`, intended for the private
-[SemiAnalysisAI/inferencex_power_model](https://github.com/SemiAnalysisAI/inferencex_power_model)
-repository. That commit is **unpublished pending repository write access**;
-reviewers cannot yet retrieve it from that remote. Publishing the app bundle
-does not publish the Python source. Source publication, when completed, will not
-change its DRAFT status or establish empirical calibration. The source
-revision and file hashes belong in the review/export record, alongside which
-components remain uncalibrated. See the update procedure below for historical
-results and frozen exports.
+The equations and editable parameters are reviewed and released together in
+InferenceX-app. Publishing an app change publishes the model it uses; there is
+no separate Python-source publication step or private-repository access gate.
+The reference fixtures record the earlier Python implementation as historical
+lineage only. App model versions and file hashes identify subsequent changes;
+neither moving ownership nor passing regression checks establishes calibration.
+See the update procedure below for historical results and frozen exports.
## NVL72 architecture walkthrough
@@ -238,7 +238,7 @@ flowchart TB
BASIS -->|"Grace socket sensor"| SUM["Measured GPU-board + Grace watts Profile regulator allowance on GPU share"]
MODULE --> RACK
SUM --> RACK
- PROFILE["Pinned GB200 / GB300 rack profiles Static loads and conversion assumptions"] -.-> RACK
+ PROFILE["App-owned GB200 / GB300 rack profiles Static loads and conversion assumptions"] -.-> RACK
RACK["Mean tray load scaled to an 18-tray rack Modeled rack residual + power-shelf losses"]
RACK --> AC["Rack AC power"]
AC --> FAC["Apply PUE once NVL72 default: 1.1"]
@@ -269,7 +269,7 @@ flowchart TB
ECON --> UI["All in Provisioned / All in Measured / Compare both Chart, tooltips, details and CSV"]
SKIP --> KEEP["Compare retains valid provisioned bars Measured-only mode does not substitute"]
KEEP --> UI
- META["Sensor basis, PUE, 10% reserve Profile revision and source hash"] -.-> UI
+ META["Sensor basis, PUE, 10% reserve App model revision and source hash"] -.-> UI
```
Solid arrows show data flow; dashed arrows supply assumptions or provenance.
@@ -284,20 +284,42 @@ Modeled power is derived from retained measurements when the browser or a shared
views API transforms a benchmark row. Changing the model does not rewrite the
original GPU measurements or require a per-run database backfill.
-1. Commit and publish the intended Python model revision so another reviewer can
- obtain the exact source. Update `REVISION` in
- `packages/app/scripts/generate-system-power-reference.py` to that clean commit,
- update `REVISION_STATUS`, and update the recorded assumptions when required.
- Publishing Python alone does not update the dashboard's bundled profiles.
-2. Run that script with the path to the pinned model checkout to regenerate
- `system-power-model.profiles.json` and `system-power-model.reference.json`.
- If equations or load-dependent components changed, update the TypeScript
- implementation too; regenerating constants alone is insufficient.
-3. Run the system-power model parity and admission tests, then deploy the app.
- Existing browser sessions need the updated bundle. Derived API responses need
- the normal authenticated cache invalidation or cache expiry; deployment alone
- does not establish that every cached response uses the new revision.
-4. Regenerate frozen CSV/JSON exports separately. If the revised model needs
+1. Edit equations in `packages/app/src/lib/system-power-model.ts` and active
+ parameters in `system-power-model.profiles.json`: `fixedComponentsDcWatts`,
+ `fan`, `psu`, and the coefficients in `rackProfiles`. The per-profile
+ `assumptions` describe the scenario used to derive those coefficients;
+ changing `u_cpu`, `u_ram` or similar metadata alone does not recalculate watts.
+ A new utilization scenario needs justified coefficients or new equations,
+ matching assumption metadata, and regression acceptance. The workload,
+ validity, topology and PUE policy live in `modeled-system-power.ts`; update
+ that adapter when the intended behavior changes there. All three files are
+ reviewed in InferenceX-app, without a separate model-repository change.
+2. Refresh the committed provenance manifest with the command below. Its
+ `modelRevision` is `app-sha256:<64 hex>`, derived from the actual SHA-256 hashes
+ of those three files. `modelPath` identifies
+ `packages/app/src/lib/system-power-model.ts`. Commit the refreshed manifest
+ with the model changes; `--check` reports drift without writing files. The
+ digest is a model identity, not a Git commit. Source links use the deployment's
+ `VERCEL_GIT_COMMIT_SHA` or `GITHUB_SHA`, falling back to `master`.
+
+ ```sh
+ bun packages/app/scripts/update-system-power-provenance.ts
+ bun packages/app/scripts/update-system-power-provenance.ts --check
+ ```
+
+3. Review the numerical changes and run the relevant model, admission, planning,
+ views API and export regression checks. Keep the 496 historical reference
+ cases as a frozen baseline. An intentional model change needs independently
+ justified expected values and explicit regression acceptance; the provenance
+ command never regenerates expected numbers to match the current code.
+ Regression acceptance does not establish empirical calibration.
+4. Deploy the app with the accepted model and profiles. Historical rows with
+ sufficient, matched raw telemetry are recalculated when they pass through the
+ updated model; no model-only database backfill is required. Existing browser
+ sessions need the updated bundle. Derived API responses need the normal
+ authenticated cache invalidation or cache expiry; deployment alone does not
+ establish that every cached response uses the new revision.
+5. Regenerate frozen CSV/JSON exports separately. If the revised model needs
inputs that were never recorded, those rows stay unavailable until the input
gap is resolved. A new benchmark's power must not be attached to an older
benchmark's throughput.
@@ -305,12 +327,13 @@ original GPU measurements or require a per-run database backfill.
## Boundary and assumptions
The input is measured mean GPU power during a validated serving window. The
-modeled chassis AC output adds the source model's CPU, DRAM, networking, storage,
+modeled chassis AC output adds the app profile's CPU, DRAM, networking, storage,
board, fans, and PSU conversion losses. Facility power is a separate estimate:
-PUE is applied after chassis AC, including the source's rounding order.
+PUE is applied after chassis AC, preserving the model's rounding order.
-The fixed README inference sweep uses `u_cpu=0.20`, `u_ram=0.20`, `u_pcie=0.05`,
-and `u_nvme=0.0`. The pinned Python model defaults to PUE `1.2`; PowerX uses
+The app profiles retain the fixed inference assumptions `u_cpu=0.20`,
+`u_ram=0.20`, `u_pcie=0.05`, and `u_nvme=0.0`. The profile baseline includes PUE
+`1.2`; PowerX uses
`1.3` for its supported air-cooled chassis profiles. Utility power = critical IT
power × PUE (`1.3` air, `1.1` DLC).
The factor applies after chassis AC; measured GPU power and chassis AC do not change.
@@ -320,32 +343,45 @@ override and does not convert an air-cooled chassis model into a DLC model. The
NVL72 rack profiles below are direct-liquid-cooled and default to `1.1`.
Platform-specific network assumptions,
fan control, component counts, and chassis defaults are preserved in the
-generated profile; every JSON export includes that profile and every CSV row
-includes its applicable assumptions and profile hash. These are model inputs,
-not measured CPU/DRAM utilization.
-
-| Hardware identity | Source chassis implementation |
-| ----------------- | ------------------------------------------------------------- |
-| `h100` | `human_verified/hgx_h100_chassis/h100_chassis_power_model.py` |
-| `h200` | `human_verified/hgx_h200_chassis/h200_chassis_power_model.py` |
-| `b200` | `human_verified/hgx_b200_chassis/b200_chassis_power_model.py` |
-| `b300` | `human_verified/hgx_b300_chassis/b300_chassis_power_model.py` |
-| `mi300x` | `human_verified/mi300x_chassis/mi300x_chassis_power_model.py` |
-| `mi325x` | `human_verified/mi325x_chassis/mi325x_chassis_power_model.py` |
-| `mi355x` | `human_verified/mi355x_chassis/mi355x_chassis_power_model.py` |
+editable app profile; every JSON export includes that profile and every CSV row
+includes its applicable assumptions, model revision and source hash. Utilization
+assumptions such as `u_cpu`, `u_ram` and `u_ib` record the scenario from which the
+active coefficients were derived; they are not live utilization controls or
+measured CPU/DRAM utilization. Editing these labels alone does not change the
+fixed watts or curves.
+
+These are the active keys in
+[system-power-model.profiles.json](../packages/app/src/lib/system-power-model.profiles.json).
+Their equations are implemented by the named functions in
+[system-power-model.ts](../packages/app/src/lib/system-power-model.ts).
+
+| Hardware identity | Editable app profile | TypeScript estimator |
+| ----------------- | -------------------- | ---------------------- |
+| `h100` | `profiles.h100` | `estimateChassisPower` |
+| `h200` | `profiles.h200` | `estimateChassisPower` |
+| `b200` | `profiles.b200` | `estimateChassisPower` |
+| `b300` | `profiles.b300` | `estimateChassisPower` |
+| `mi300x` | `profiles.mi300x` | `estimateChassisPower` |
+| `mi325x` | `profiles.mi325x` | `estimateChassisPower` |
+| `mi355x` | `profiles.mi355x` | `estimateChassisPower` |
All listed profiles describe a complete eight-GPU chassis. Their topology is not
substituted onto GB200 or GB300, which use the NVL72 rack profiles instead:
-| Hardware identity | Source rack implementation |
-| ----------------- | ----------------------------------------------------------------- |
-| `gb200` | `human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py` |
-| `gb300` | same module, `gb300_nvl72_rack_config` |
+| Hardware identity | Editable app profile | TypeScript estimator |
+| ----------------- | -------------------- | -------------------- |
+| `gb200` | `rackProfiles.gb200` | `estimateRackPower` |
+| `gb300` | `rackProfiles.gb300` | `estimateRackPower` |
+
+The [reference fixture](../packages/app/src/lib/system-power-model.reference.json)
+retains historical source/revision metadata, component hashes, and
+`pythonConfigurations` for baseline lineage only. Those records are not active
+app parameters or an ongoing Python dependency.
The rack profiles (`rackProfiles`) take **measured** compute-module watts per tray
as their input: the module sensor total (`avg_total_module_power_w`) when the
producer publishes it, otherwise GPU-board watts plus the Grace-socket total
-(`avg_total_cpu_power_w`) with the source's regulator-loss allowance on the GPU
+(`avg_total_cpu_power_w`) with the app profile's regulator-loss allowance on the GPU
share. The Grace CPU and LPDDR5X are never modelled; rows without
`cpu_power_valid=1` and complete module or Grace provenance stay unavailable (`cpu-telemetry`).
Each measured worker host is one compute tray (four GPUs, two Grace sockets); an
@@ -354,22 +390,20 @@ deployment mean, cross-checked against the Grace-socket count and the CPU leg's
`power_audit.cpu.observed_sockets`. The
measured trays are folded into one rack of 18 trays matching their mean
compute-module input, the power-shelf efficiency curve is evaluated once at that
-rack's DC load (as the source `gb200_nvl72_rack_power` does with its single
-per-tray input), and every tray takes the same 1/18 share, so NVSwitch trays,
+rack's DC load, using one mean per-tray input. Every tray takes the same 1/18
+share, so NVSwitch trays,
power shelves, and management switches are amortised over 72 GPUs. Chassis, by
contrast, own their fans and PSUs and are each evaluated at their own load. The
result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`.
A partially allocated tray extrapolates only the GPU-board share (a module reading
-already covers the whole tray) and is labeled `extrapolated`. The source is pinned to revision
-`6fcc086b77576d4cecb9d0c79637d6daf980308c`, a local commit intended for the private
-model repository linked above. It remains unpublished pending repository write
-access, and the model remains DRAFT / pending human verification. The NVL72 section below
-lists the measured input, the modeled residual and the Profit Estimator gate rules.
+already covers the whole tray) and is labeled `extrapolated`. The app-owned
+model remains DRAFT / pending human verification. The NVL72 section below lists
+the measured input, the modeled residual and the Profit Estimator gate rules.
A partially allocated chassis (one to seven measured GPUs on one host) is
-modeled at measured per-GPU power × 8. That is the same `n_gpu × W/GPU` input
-the source sweep scripts feed each chassis model, and it assumes the unmeasured
-GPUs run the same workload. The estimate is labeled `chassisBasis:
+modeled at measured per-GPU power × 8. This `n_gpu × W/GPU` input is retained
+from the historical baseline and assumes the unmeasured GPUs run the same
+workload. The estimate is labeled `chassisBasis:
'extrapolated'`: per-GPU values divide by the modeled chassis GPU count
(`modeledGpuCount`), while `deploymentAcWatts` / `deploymentFacilityWatts` keep
only the measured GPUs' share of each chassis. This is not a proportional share
@@ -387,7 +421,8 @@ worker, distinct worker hosts, and consistent total/role watts. A role average
alone cannot establish physical placement or evaluate each host's nonlinear
model. CPU-only frontend workers are excluded from GPU-chassis counting. Separate CPU-only
frontend/router hosts are outside this estimate; CPU power within GPU chassis
-still uses the source's fixed 20% utilization assumption.
+still uses fixed coefficients derived for the app profile's 20% utilization
+scenario.
The default measured contract is numeric `power_valid=1` and metric schema 2.
The original validated single-node producer predates the schema marker but
@@ -413,11 +448,11 @@ Grace watts and total/mean watts consistent with the audited socket count.
CPU-rail-only and missing or unknown sensor provenance stay unavailable. Basis selection: `module` when `avg_total_module_power_w` is present (the
reading already contains the GPU boards, so it is never scaled), otherwise
`gpu-plus-grace` (GPU-board watts × 4 plus the Grace-socket total per tray, with the
-source's regulator-loss allowance `regulatorLossFracOfTdp / (1 − frac)` on the GPU
+profile's regulator-loss allowance `regulatorLossFracOfTdp / (1 − frac)` on the GPU
share only). A present-but-invalid module key makes the row unavailable
(`cpu-telemetry`); it never falls back to the Grace socket silently.
-**Modeled residual.** Everything outside the compute modules comes from the pinned
+**Modeled residual.** Everything outside the compute modules comes from the app
profile (`rackProfiles`), evaluated once for a rack of 18 trays at the measured
trays' mean input (the shelf curve sees the whole rack's DC load, never one tray's)
and amortised over 72 GPUs; the parameters marked UNVERIFIED carry a documented
@@ -439,9 +474,10 @@ range in `unverifiedParameters` and no published rail:
| Facility PUE | 1.1 (direct liquid cooling) | same | PowerX policy, applied once to rack AC |
Rack DC above the installed shelf capacity (264 kW) overflows the efficiency curve
-and the row is unavailable (`model-domain`). Rounding follows the source: rack AC is rounded
-to 0.1 W before PUE. Python-generated `rackCases` prove parity with the pinned
-implementation for both variants, both bases, every shelf knot and PUE 1.0–1.2.
+and the row is unavailable (`model-domain`). The model rounds rack AC to 0.1 W
+before PUE. Checked-in `rackCases` retain the historical Python baseline for
+both variants, both bases, every shelf knot and PUE 1.0–1.2. They are regression
+references, not calibration evidence or a dependency on the former repository.
**Gate rules (Profit Estimator).** Planning kW/GPU = deployment facility watts ÷
measured GPUs ÷ 1000 × 1.1. It accepts fully measured eight-GPU chassis
@@ -458,7 +494,10 @@ knot beside a Grace-socket knot stays unavailable rather than blending sensors.
bar tooltip, the collapsed Power assumptions disclosure and the CSV columns `Power
basis`, `Power sensor`, `System power profile` name the basis (measured module, or
measured GPU board + Grace socket with regulator loss modeled), the sensor kind, and
-the pinned profile (`modelPath @ modelRevision sha256:`) per row.
+the app profile (`modelPath @ modelRevision sha256:`) per row.
+The path and revision identify the app-owned model; the source hash identifies
+the TypeScript equation file, and the revision also covers the editable profile
+and admission/PUE adapter.
The `?unofficialrun=` overlay rule does not apply to the Profit Estimator basis
control: the estimator prices official frontier points only.
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index e913d6ad7..10b46959b 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -16,13 +16,15 @@ AgentX 估算,包括 Kimi K3;这不代表模型已经过 AgentX 校准。应
[代码位置](#各部分代码在哪里) 将说明与实现对应起来。后面的章节完整记录了数据接纳、
拓扑和导出约定。
-`system-power-model.profiles.json` 记录固定的功耗模型版本、组件源码哈希、硬件映射、
-完整平台配置和固定推理假设。各 profile 由原始 Python 组件执行生成。
-`system-power-model.ts` 保留原模型的非线性风扇曲线、PSU 效率插值、中间值舍入和 PUE
-应用顺序。由 Python 生成的参考用例用于核对 TypeScript 实现与源模型是否一致。
+模型由 InferenceX-app 维护。`system-power-model.ts` 定义公式,包括非线性风扇曲线、
+PSU 效率插值、中间值舍入和 PUE 应用顺序。`system-power-model.profiles.json` 保存可
+直接编辑的组件参数、硬件映射、平台配置,以及说明固定推理场景的元数据。前端、
+共享 views API 和离线导出器均以这些仓库内文件中的模型定义为准。`system-power-model.provenance.json`
+记录应用模型版本和源码哈希。
-固定版本的源码仍标记为 **DRAFT / pending human verification(草稿,待人工核验)**。
-数值一致性只能证明实现等价,不能证明已经过实机机箱校准。
+已提交的参考用例保留历史 Python 基线,用于回归对照;开发、构建和部署应用模型
+都不需要 Python 或原私有仓库。模型仍为 **DRAFT / pending human verification
+(草稿,待人工核验)**。与数值基线一致只能证明实现的一致性,不能证明已经过实机机箱校准。
## 实测、建模和预配分别指什么?
@@ -109,7 +111,7 @@ PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU
4. **单独计入规划余量。**将设施 kW/GPU 乘以 1.10。PUE 1.1 与 10% 规划余量是不同
的系数,作用也不同。
-已提交的 Python 参考测试数据包含以下中间值:
+已提交的参考测试数据最初取自 Python 基线,包含以下中间值:
| 阶段 | 功率(W) |
| ------------------------------------------ | --------: |
@@ -121,8 +123,8 @@ PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU
| 计入随负载变化的电源架效率后的机架交流功率 | 74,904.6 |
| 乘以 PUE 1.1 后的设施功率 | 82,395.1 |
-表中中间值已舍入。计算保留源模型的求和与舍入顺序,因此直接相加表中数值可能差
-0.1 W。源模型原始 profile 的默认 PUE 为 1.2;仪表板封装层对 **NVL72 使用 1.1**,
+表中中间值已舍入。计算保留模型的求和与舍入顺序,因此直接相加表中数值可能差
+0.1 W。profile 的基线默认 PUE 为 1.2;仪表板封装层对 **NVL72 使用 1.1**,
对**当前支持的风冷机箱使用 1.3**。
规划 kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
@@ -156,7 +158,7 @@ PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU
| 职责 | 实现 |
| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 入库时保留 CPU/GPU 指标和审计来源 | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts)、[power-publication.ts](../packages/db/src/etl/power-publication.ts) |
-| 固定 Python 模型版本、假设和组件哈希 | [generate-system-power-reference.py](../packages/app/scripts/generate-system-power-reference.py)、[profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| 编辑组件参数;更新应用模型版本和源码哈希 | [profiles](../packages/app/src/lib/system-power-model.profiles.json)、[来源清单](../packages/app/src/lib/system-power-model.provenance.json)、[update-system-power-provenance.ts](../packages/app/scripts/update-system-power-provenance.ts) |
| 检查负载、有效性、传感器和拓扑条件;分摊部署功耗 | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
| 计算非线性机箱/机架功耗、损耗和 PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
| 将功耗匹配到原前沿,并保留预配对比结果 | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
@@ -166,14 +168,11 @@ PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU
| 离线导出建模后的基准测试记录 | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
| 验证计算一致性和接纳行为 | [参考测试数据](../packages/app/src/lib/system-power-model.reference.json)、[模型测试](../packages/app/src/lib/system-power-model.test.ts)、[接纳规则测试](../packages/app/src/lib/modeled-system-power.test.ts)、[规划测试](../packages/app/src/components/calculator/profit-power.test.ts) |
-发布 #1190 会发布 TypeScript 实现及随应用打包的 profiles,运行时不会从 GitHub 下载
-Python。Python 源码固定在本地提交 `6fcc086b77576d4cecb9d0c79637d6daf980308c`,
-计划发布至私有仓库
-[SemiAnalysisAI/inferencex_power_model](https://github.com/SemiAnalysisAI/inferencex_power_model)。
-该提交**尚未发布,正在等待仓库写入权限**;审阅者目前无法从该远端取得这个提交。
-发布应用 bundle 不等于发布 Python 源码。日后完成源码发布,也不会改变其 DRAFT 状态,
-更不代表完成了实测校准。审阅和导出记录应保留源码版本、文件哈希,以及哪些组件仍
-未经校准。历史结果和冻结导出的更新方式见后文。
+公式和可编辑参数在 InferenceX-app 中一同审阅、一同发布。发布应用变更就会发布该
+版本采用的模型,不需要单独发布 Python 源码,也不以私有仓库权限作为前提。
+参考测试数据只将早期 Python 实现记作历史来源。后续变化由应用模型版本和文件哈希
+标识;维护归属的改变或回归检查通过,都不代表完成了实测校准。历史结果和冻结导出
+的更新方式见后文。
## NVL72 架构导览
@@ -199,7 +198,7 @@ flowchart TB
BASIS -->|"Grace socket 传感器"| SUM["实测 GPU 板卡 + Grace 功耗 按 profile 对 GPU 份额计入稳压损耗余量"]
MODULE --> RACK
SUM --> RACK
- PROFILE["固定版本的 GB200 / GB300 机架 profile 静态负载与转换假设"] -.-> RACK
+ PROFILE["应用维护的 GB200 / GB300 机架 profile 静态负载与转换假设"] -.-> RACK
RACK["按 tray 平均负载扩展至 18 tray 机架 加入机架其余组件和电源架损耗模型"]
RACK --> AC["机架交流功率"]
AC --> FAC["只应用一次 PUE NVL72 默认值:1.1"]
@@ -229,7 +228,7 @@ flowchart TB
ECON --> UI["All in Provisioned / All in Measured / Compare both 图表、提示框、详情和 CSV"]
SKIP --> KEEP["对比模式保留有效预配柱子 仅实测模式不以预配值替代"]
KEEP --> UI
- META["传感器口径、PUE、10% 余量 profile 版本和源码哈希"] -.-> UI
+ META["传感器口径、PUE、10% 余量 应用模型版本和源码哈希"] -.-> UI
```
实线表示数据流,虚线提供假设或来源信息。规划采用正式性能前沿点,不采用
@@ -241,58 +240,85 @@ flowchart TB
浏览器或共享 views API 转换基准测试记录时,会根据保留的测量值推导建模功耗。
修改模型不会改写原始 GPU 测量值,也不需要逐 run 回填数据库。
-1. 提交并发布要使用的 Python 模型版本,让其他审阅者能够取得完全相同的源码。
- 将 `packages/app/scripts/generate-system-power-reference.py` 的 `REVISION` 更新为
- 该干净提交,并更新 `REVISION_STATUS`;必要时同步记录的假设。仅发布 Python 不会
- 更新仪表板打包的 profiles。
-2. 运行该脚本,传入固定版本模型 checkout 的路径,重新生成
- `system-power-model.profiles.json` 和 `system-power-model.reference.json`。
- 如果公式或随负载变化的组件有改动,还必须修改 TypeScript 实现;只重新生成常量不够。
-3. 运行系统功耗模型的一致性测试和接纳规则测试,再部署应用。已有浏览器会话需要
- 加载新 bundle;派生 API 响应需要走正常的认证缓存失效流程,或等待缓存过期。
- 仅完成部署,不能证明所有缓存响应都已采用新版本。
-4. 冻结的 CSV/JSON 导出需单独重新生成。若新模型需要从未记录的输入,相应记录应
+1. 在 `packages/app/src/lib/system-power-model.ts` 中修改公式,在
+ `system-power-model.profiles.json` 中修改参与计算的参数:`fixedComponentsDcWatts`、
+ `fan`、`psu` 及 `rackProfiles` 中的系数。每个 profile 的 `assumptions` 记录推导这些
+ 系数时采用的场景;只改 `u_cpu`、`u_ram` 等元数据,不会重新计算功率。要支持新的
+ 利用率场景,需要有依据的新系数或新公式、与之对应的假设记录,并通过回归验收。
+ 负载、有效性、拓扑和 PUE 策略位于 `modeled-system-power.ts`;若这些行为需要改变,
+ 应同步修改该适配层。三个文件都在 InferenceX-app 中审阅,无需另改一个模型仓库。
+2. 用以下命令更新已提交的来源清单。`modelRevision` 的格式为 `app-sha256:<64 hex>`,
+ 根据上述三个文件实际内容的 SHA-256 哈希生成。`modelPath` 指向
+ `packages/app/src/lib/system-power-model.ts`。更新后的清单应与模型修改一起提交;
+ `--check` 只检查清单是否与当前文件一致,不写文件。这个哈希摘要是模型标识,
+ 不是 Git 提交。源码链接使用部署的 `VERCEL_GIT_COMMIT_SHA` 或 `GITHUB_SHA`,
+ 两者均缺失时使用 `master`。
+
+ ```sh
+ bun packages/app/scripts/update-system-power-provenance.ts
+ bun packages/app/scripts/update-system-power-provenance.ts --check
+ ```
+
+3. 审阅数值变化,运行相关模型、接纳规则、规划、views API 和导出回归检查。
+ 保留 496 个历史参考用例作为冻结基线。有意改变模型时,需要提供独立论证的预期值,
+ 并明确验收回归结果;来源清单命令不会根据当前代码重新生成预期数字。
+ 回归验收通过不代表完成了实测校准。
+4. 部署已验收的模型和 profiles。历史记录只要具有充分、匹配的原始遥测,就会在经过
+ 新模型时重新计算;单纯修改模型无需回填数据库。已有浏览器会话需要加载新 bundle;
+ 派生 API 响应需要走正常的认证缓存失效流程,或等待缓存过期。仅完成部署,不能证明
+ 所有缓存响应都已采用新版本。
+5. 冻结的 CSV/JSON 导出需单独重新生成。若新模型需要从未记录的输入,相应记录应
保持不可用,直至输入缺口解决。不能把新基准测试的功耗附到旧基准测试的吞吐量上。
## 计算边界与假设
输入是已验证服务窗口内的平均 GPU 实测功率。机箱交流功耗模型在此基础上,加入
-源模型中的 CPU、DRAM、网络、存储、主板、风扇和 PSU 转换损耗。设施功率另行估算:
-先算机箱交流功率,再应用 PUE,并保留源模型的舍入顺序。
+应用 profile 中的 CPU、DRAM、网络、存储、主板、风扇和 PSU 转换损耗。设施功率另行
+估算:先算机箱交流功率,再应用 PUE,并保留模型的舍入顺序。
-README 中固定的推理 sweep 使用 `u_cpu=0.20`、`u_ram=0.20`、`u_pcie=0.05` 和
-`u_nvme=0.0`。固定版本的 Python 模型默认 PUE 为 `1.2`;PowerX 对当前支持的风冷
+应用 profiles 保留固定推理假设:`u_cpu=0.20`、`u_ram=0.20`、`u_pcie=0.05` 和
+`u_nvme=0.0`。profile 基线包含默认 PUE `1.2`;PowerX 对当前支持的风冷
机箱 profile 使用 `1.3`。市电侧功率 = IT 负载功率 × PUE(风冷 `1.3`,直接液冷 DLC
`1.1`)。该系数作用于机箱交流功率之后,不改变 GPU 实测功率或机箱交流功率。
这里的冷却方式指模型中的机箱,并非已经核实的基准测试站点冷却配置。机箱 profile
不支持 DLC;`--pue` 只是显式覆盖设施功率系数,不会把风冷机箱模型转成液冷模型。
下文的 NVL72 机架 profile 为直接液冷,默认使用 `1.1`。
-生成的 profile 保留各平台的网络假设、风扇控制、组件数量和机箱默认值。每份 JSON
-导出包含完整 profile,每行 CSV 包含适用假设及 profile 哈希。这些是模型输入,
-不是实测 CPU/DRAM 利用率。
-
-| 硬件标识 | 机箱模型源码 |
-| -------- | ------------------------------------------------------------- |
-| `h100` | `human_verified/hgx_h100_chassis/h100_chassis_power_model.py` |
-| `h200` | `human_verified/hgx_h200_chassis/h200_chassis_power_model.py` |
-| `b200` | `human_verified/hgx_b200_chassis/b200_chassis_power_model.py` |
-| `b300` | `human_verified/hgx_b300_chassis/b300_chassis_power_model.py` |
-| `mi300x` | `human_verified/mi300x_chassis/mi300x_chassis_power_model.py` |
-| `mi325x` | `human_verified/mi325x_chassis/mi325x_chassis_power_model.py` |
-| `mi355x` | `human_verified/mi355x_chassis/mi355x_chassis_power_model.py` |
+应用内可编辑的 profile 保留各平台的网络假设、风扇控制、组件数量和机箱默认值。每份 JSON
+导出包含完整 profile,每行 CSV 包含适用假设、模型版本和源码哈希。`u_cpu`、
+`u_ram`、`u_ib` 等利用率假设记录的是推导当前参与计算的系数时采用的场景,不是实时利用率
+控件,也不是 CPU/DRAM 利用率实测值。只修改这些标签,不会改变固定功率或曲线。
+
+下表列出
+[system-power-model.profiles.json](../packages/app/src/lib/system-power-model.profiles.json)
+中参与计算的配置项,对应公式由
+[system-power-model.ts](../packages/app/src/lib/system-power-model.ts) 中的函数实现。
+
+| 硬件标识 | 应用内可编辑的 profile | TypeScript 计算函数 |
+| -------- | ---------------------- | ---------------------- |
+| `h100` | `profiles.h100` | `estimateChassisPower` |
+| `h200` | `profiles.h200` | `estimateChassisPower` |
+| `b200` | `profiles.b200` | `estimateChassisPower` |
+| `b300` | `profiles.b300` | `estimateChassisPower` |
+| `mi300x` | `profiles.mi300x` | `estimateChassisPower` |
+| `mi325x` | `profiles.mi325x` | `estimateChassisPower` |
+| `mi355x` | `profiles.mi355x` | `estimateChassisPower` |
上述 profile 均描述完整八卡机箱。GB200 和 GB300 使用独立的 NVL72 机架 profile,
不套用这些机箱拓扑:
-| 硬件标识 | 机架模型源码 |
-| -------- | ----------------------------------------------------------------- |
-| `gb200` | `human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py` |
-| `gb300` | 同一模块中的 `gb300_nvl72_rack_config` |
+| 硬件标识 | 应用内可编辑的 profile | TypeScript 计算函数 |
+| -------- | ---------------------- | ------------------- |
+| `gb200` | `rackProfiles.gb200` | `estimateRackPower` |
+| `gb300` | `rackProfiles.gb300` | `estimateRackPower` |
+
+[参考测试数据](../packages/app/src/lib/system-power-model.reference.json) 保留历史源码及
+版本信息、组件哈希和 `pythonConfigurations`,仅用于追溯基线来源。这些记录不是应用
+当前使用的参数,也不意味着应用仍依赖 Python。
机架 profile(`rackProfiles`)的输入是每个 tray 的**实测**计算模块功耗:生产端发布
`avg_total_module_power_w` 时使用 module 传感器总值;否则使用 GPU 板卡功耗加
-Grace socket 总功耗(`avg_total_cpu_power_w`),并按源模型对 GPU 份额计入稳压损耗
+Grace socket 总功耗(`avg_total_cpu_power_w`),并按应用 profile 对 GPU 份额计入稳压损耗
余量。Grace CPU 和 LPDDR5X 从不由模型补算。缺少 `cpu_power_valid=1` 或完整 module /
Grace 来源记录的行保持不可用(`cpu-telemetry`)。
@@ -300,21 +326,17 @@ Grace 来源记录的行保持不可用(`cpu-telemetry`)。
数组的聚合多节点记录,按 `gpuCount / 4` 推算 tray 数,各 tray 采用部署平均值,并与
Grace socket 数及 CPU 采集记录中的 `power_audit.cpu.observed_sockets` 交叉校验。
按实测 tray 的平均计算模块输入,构造一个包含 18 个同等负载 tray 的机架;电源架效率
-曲线只在整机架直流负载处求值一次。这与源模型 `gb200_nvl72_rack_power` 接受单一
-per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray、电源架和管理
+曲线只在整机架直流负载处求值一次,使用单一的 tray 平均输入。每个 tray 分摊 1/18,
+因此 NVSwitch tray、电源架和管理
交换机按 72 张 GPU 分摊。机箱则各自拥有风扇和 PSU,按各自的负载单独求值。
结果包含 `topologyBasis: 'nvl72-trays'`、`measuredBasis` 和 `sensorKind`。
部分分配的 tray 只外推 GPU 板卡份额,因为 module 读数本身已覆盖整个 tray;结果标记
-为 `extrapolated`。源码固定在本地提交
-`6fcc086b77576d4cecb9d0c79637d6daf980308c`,计划发布至上文的私有模型仓库,
-目前仍因等待仓库写入权限而未发布。模型仍为 DRAFT / pending human verification。
-下文列出 NVL72 的实测输入、其余组件
-模型,以及利润估算器的接纳规则。
+为 `extrapolated`。应用维护的模型仍为 DRAFT / pending human verification。
+下文列出 NVL72 的实测输入、其余组件模型,以及利润估算器的接纳规则。
部分分配的机箱,即单台主机上实测一至七张 GPU,按“实测每 GPU 功率 × 8”建模。
-这与源模型 sweep 脚本使用的 `n_gpu × W/GPU` 输入一致,并假设未实测的 GPU 运行相同
-负载。估算标记为 `chassisBasis: 'extrapolated'`:每 GPU 数值按建模机箱 GPU 数
+这一 `n_gpu × W/GPU` 输入沿用历史基线,并假设未实测的 GPU 运行相同负载。估算标记为 `chassisBasis: 'extrapolated'`:每 GPU 数值按建模机箱 GPU 数
(`modeledGpuCount`)分摊;`deploymentAcWatts` / `deploymentFacilityWatts` 只保留
各机箱中实测 GPU 的份额。这不是先在部分负载处计算机箱,再按比例分摊;固定组件、
风扇曲线和 PSU 效率都在满机箱负载处求值。缺失或无效遥测、数量不一致、缺少主机
@@ -326,7 +348,7 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
对应一个机箱(一至八张 GPU)、worker 位于不同主机,并且总功率与角色功率一致。
只有角色平均值,无法证明物理放置方式,也无法计算各主机的非线性模型。
纯 CPU frontend worker 不计入 GPU 机箱数量。独立的纯 CPU frontend/router 主机
-不在估算范围内;GPU 机箱内的 CPU 功耗仍采用源模型固定的 20% 利用率假设。
+不在估算范围内;GPU 机箱内的 CPU 功耗仍采用固定系数,这些系数按应用 profile 中 20% 利用率场景推导得出。
默认实测约定要求数值型 `power_valid=1` 和指标 schema 2。原有通过验证的单节点生产端
早于 schema 标记,但两个 watts 字段的定义已与 schema 2 相同。这条路径保留 schema 缺失的
@@ -350,11 +372,11 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
口径选择:存在 `avg_total_module_power_w` 时采用 `module`,因为读数已包含 GPU
板卡,所以不会再缩放;否则采用 `gpu-plus-grace`,即每 tray 的每 GPU 板卡功率 × 4
-加 Grace socket 总功率,并且只对 GPU 份额应用源模型的稳压损耗余量
+加 Grace socket 总功率,并且只对 GPU 份额应用 profile 的稳压损耗余量
`regulatorLossFracOfTdp / (1 − frac)`。module 字段存在但无效时,该行不可用
(`cpu-telemetry`),不会悄然回退到 Grace socket。
-**其余组件的模型估算。**计算模块以外的部分均来自固定版本 profile(`rackProfiles`)。
+**其余组件的模型估算。**计算模块以外的部分均来自应用 profile(`rackProfiles`)。
按实测 tray 的平均输入构造 18 tray 机架,整体求值一次,再按 72 张 GPU 分摊。
电源架曲线使用整机架的直流负载,不能只使用单个 tray 的负载。标为 UNVERIFIED 的
参数在 `unverifiedParameters` 中记录了范围,但没有已发布的供电轨测量:
@@ -375,9 +397,9 @@ per-tray 输入的方式一致。每个 tray 分摊 1/18,因此 NVSwitch tray
| 设施 PUE | 1.1(直接液冷) | 相同 | PowerX 策略,仅对机架交流功率应用一次 |
机架直流功率超过电源架装机容量(264 kW)时,超出效率曲线适用范围,该行不可用
-(`model-domain`)。舍入顺序与源模型一致:先将机架交流功率舍入到 0.1 W,再应用 PUE。
-Python 生成的 `rackCases` 验证了两种变体、两种口径、全部电源架曲线节点和 PUE 1.0–1.2
-下与固定实现的数值一致性。
+(`model-domain`)。模型先将机架交流功率舍入到 0.1 W,再应用 PUE。已提交的
+`rackCases` 保留历史 Python 基线,涵盖两种变体、两种口径、全部电源架曲线节点和
+PUE 1.0–1.2。它们用于回归对照,不是校准证据,也不构成对原仓库的依赖。
**利润估算器的接纳规则。**规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。
接受完整实测的八卡机箱(`single-node`、`worker-hosts` 或 `uniform-hosts` 拓扑下的
@@ -390,9 +412,10 @@ worker 数组的聚合多节点记录则按 `gpuCount / 4` 推算 tray 数,各
两个前沿点必须采用相同实测口径和传感器类型;若一端是 module、另一端是 Grace
socket,结果不可用,不混合两种传感器。柱形提示框、默认折叠的 Power assumptions
详情,以及 CSV 中的 `Power basis`、`Power sensor`、`System power profile` 列,会逐行
-列明口径、传感器类型和固定 profile。口径为实测 module,或实测 GPU 板卡 + Grace
+列明口径、传感器类型和应用 profile。口径为实测 module,或实测 GPU 板卡 + Grace
socket 并由模型估算稳压损耗;profile 表示为
-`modelPath @ modelRevision sha256:`。
+`modelPath @ modelRevision sha256:`。路径和版本标识应用维护的模型;
+源码哈希标识 TypeScript 公式文件,模型版本还涵盖可编辑 profile 和接纳/PUE 适配层。
`?unofficialrun=` 叠加层规则不适用于利润估算器的功耗口径控件;估算器只对正式前沿点计价。
## 利润估算器的功耗口径
diff --git a/packages/app/cypress/component/power-compare.cy.tsx b/packages/app/cypress/component/power-compare.cy.tsx
index f2a51b485..ad4cd5acc 100644
--- a/packages/app/cypress/component/power-compare.cy.tsx
+++ b/packages/app/cypress/component/power-compare.cy.tsx
@@ -53,10 +53,14 @@ function measuredCurve(hwKey: string, run_url?: string): InferenceData[] {
);
}
-function mountCompare(data: InferenceData[], overlay: InferenceData[]) {
+function mountCompare(
+ data: InferenceData[],
+ overlay: InferenceData[],
+ { locale = 'en', width = 1000 }: { locale?: 'en' | 'zh'; width?: number } = {},
+) {
mountWithProviders(
-
-
+
+
{
);
});
});
+
+describe('Modeled power source links', () => {
+ for (const locale of ['en', 'zh'] as const) {
+ for (const width of [1280, 390]) {
+ const overlay = width === 390;
+ it(`links to app-owned source and ${locale} assumptions from a ${overlay ? 'mobile overlay' : 'desktop official'} tooltip`, () => {
+ cy.viewport(width, 720);
+ const modeledCurve = (hwKey: string, runUrl?: string) =>
+ measuredCurve(hwKey, runUrl).map((point) =>
+ createMockInferenceData({
+ ...point,
+ disagg: false,
+ modeledSystemPower: {
+ status: 'supported',
+ hardware: hwKey,
+ modelRevision: 'model-content-digest-for-tooltip-fixture',
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
+ gpuCount: 8,
+ chassisCount: 1,
+ modeledGpuCount: 8,
+ measuredGpuWattsPerGpu: 600,
+ chassisAcWatts: 6400,
+ chassisAcWattsPerGpu: 800,
+ facilityWatts: 8320,
+ deploymentAcWatts: 6400,
+ deploymentFacilityWatts: 8320,
+ pue: 1.3,
+ topologyBasis: 'single-node',
+ chassisBasis: 'full',
+ telemetryBasis: 'validated-v2',
+ },
+ }),
+ );
+ mountCompare(modeledCurve('b200'), modeledCurve('h100', OVERLAY_RUN_URL), {
+ locale,
+ width: Math.min(1000, width - 32),
+ });
+ cy.get(`${svg} ${overlay ? '.unofficial-overlay-pt' : '.dot-group'}`)
+ .eq(1)
+ .click({ force: true });
+ cy.get('[data-chart-tooltip]:visible').within(() => {
+ if (overlay) cy.contains('powerx-compare').should('exist');
+ cy.get('[data-testid="tooltip-modeled-system-power"]').within(() => {
+ const base = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${process.env.NEXT_PUBLIC_APP_SOURCE_REF ?? 'master'}`;
+ cy.get(`a[href="${base}/packages/app/src/lib/system-power-model.ts"]`)
+ .should('have.attr', 'title', 'model-content-digest-for-tooltip-fixture')
+ .and('have.attr', 'target', '_blank')
+ .and('have.attr', 'rel', 'noopener noreferrer')
+ .scrollIntoView()
+ .should('be.visible');
+ cy.contains('a', locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions')
+ .should(
+ 'have.attr',
+ 'href',
+ `${base}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`,
+ )
+ .and('have.attr', 'target', '_blank')
+ .and('have.attr', 'rel', 'noopener noreferrer')
+ .then(($link) => $link[0].scrollIntoView({ block: 'center' }))
+ .should('be.visible')
+ .then(($link) => {
+ const bounds = $link[0].getBoundingClientRect();
+ expect(bounds.left).to.be.at.least(0);
+ expect(bounds.right).to.be.at.most(width);
+ const shell = $link.closest('[data-chart-tooltip]')[0].firstElementChild!;
+ const frame = shell.getBoundingClientRect();
+ expect(bounds.top).to.be.at.least(frame.top);
+ expect(bounds.bottom).to.be.at.most(frame.bottom);
+ });
+ });
+ });
+ cy.screenshot(`power-model-links-${locale}-${width}`, { capture: 'viewport' });
+ });
+ }
+ }
+});
diff --git a/packages/app/next.config.ts b/packages/app/next.config.ts
index 0ac25747e..2fceaf35f 100644
--- a/packages/app/next.config.ts
+++ b/packages/app/next.config.ts
@@ -12,6 +12,11 @@ const nextConfig: NextConfig = {
// distDir, so distinct dirs let the two coexist.
distDir: process.env.NEXT_DIST_DIR || '.next',
allowedDevOrigins: allowedDevOriginsFromEnv(),
+ env: {
+ // Preview source links must follow the deployed commit, not the model content digest.
+ NEXT_PUBLIC_APP_SOURCE_REF:
+ process.env.VERCEL_GIT_COMMIT_SHA || process.env.GITHUB_SHA || 'master',
+ },
transpilePackages: ['@semianalysisai/inferencex-constants'],
serverExternalPackages: ['shiki'],
redirects() {
diff --git a/packages/app/scripts/export-modeled-system-power.ts b/packages/app/scripts/export-modeled-system-power.ts
index b31ee237b..94242ffa7 100644
--- a/packages/app/scripts/export-modeled-system-power.ts
+++ b/packages/app/scripts/export-modeled-system-power.ts
@@ -12,7 +12,7 @@ import {
defaultSystemPue,
modelSystemPower,
} from '../src/lib/modeled-system-power';
-import profileData from '../src/lib/system-power-model.profiles.json';
+import { SYSTEM_POWER_MODEL_METADATA as profileData } from '../src/lib/system-power-model';
interface PowerAudit {
power_valid: boolean;
@@ -403,6 +403,7 @@ async function main() {
'packages/app/src/lib/modeled-system-power.ts',
'packages/app/src/lib/system-power-model.ts',
'packages/app/src/lib/system-power-model.profiles.json',
+ 'packages/app/src/lib/system-power-model.provenance.json',
];
const hashes: Record = {};
for (const path of codePaths)
diff --git a/packages/app/scripts/generate-system-power-reference.py b/packages/app/scripts/generate-system-power-reference.py
deleted file mode 100644
index d3e946eae..000000000
--- a/packages/app/scripts/generate-system-power-reference.py
+++ /dev/null
@@ -1,282 +0,0 @@
-#!/usr/bin/env python3
-"""Regenerate fixed 8k1k profiles and parity cases from Oren's pinned Python models.
-
-Usage: python3 packages/app/scripts/generate-system-power-reference.py /path/to/inferencex_power_model
-Only Python's standard library is required. No telemetry, dependencies, or GPUs are fetched.
-Apply the repository formatter to generated JSON before committing.
-
-Chassis profiles (`profiles`) describe one eight-GPU HGX/OAM system whose GPU watts
-are the only measured input. Rack profiles (`rackProfiles`) describe one NVL72 rack
-whose compute-module watts per tray are measured (module sensor, or GPU board plus
-Grace socket); the Grace CPU and LPDDR5X are never modelled.
-"""
-
-import argparse
-from dataclasses import asdict
-import hashlib
-import importlib
-import json
-from pathlib import Path
-import subprocess
-import sys
-
-sys.dont_write_bytecode = True
-# Local reference branch; upstream publication needs repository write access.
-REVISION = "6fcc086b77576d4cecb9d0c79637d6daf980308c"
-REVISION_STATUS = "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification"
-SOURCE = "https://github.com/SemiAnalysisAI/inferencex_power_model"
-MODELS = {
- "h100": ("hgx_h100_chassis/h100_chassis_power_model.py", "h100_chassis_power", "make_h100_config"),
- "h200": ("hgx_h200_chassis/h200_chassis_power_model.py", "h200_chassis_power", "make_h200_config"),
- "b200": ("hgx_b200_chassis/b200_chassis_power_model.py", "b200_chassis_power", "B200ChassisMasterConfig"),
- "b300": ("hgx_b300_chassis/b300_chassis_power_model.py", "b300_chassis_power", "B300ChassisConfig"),
- "mi300x": ("mi300x_chassis/mi300x_chassis_power_model.py", "mi300x_chassis_power", "MI300XChassisConfig"),
- "mi325x": ("mi325x_chassis/mi325x_chassis_power_model.py", "mi325x_chassis_power", "MI325XChassisConfig"),
- "mi355x": ("mi355x_chassis/mi355x_chassis_power_model.py", "mi355x_chassis_power", "MI355XChassisConfig"),
-}
-ASSUMPTIONS = {"u_pcie": 0.05, "u_cpu": 0.20, "u_ram": 0.20, "u_nvme": 0.0, "pue": 1.20}
-RACK_MODEL_PATH = "gb200_nvl72_rack/gb200_nvl72_rack_power_model.py"
-RACK_MODELS = {
- "gb200": (RACK_MODEL_PATH, "gb200_nvl72_rack_power", "gb200_nvl72_rack_config"),
- "gb300": (RACK_MODEL_PATH, "gb200_nvl72_rack_power", "gb300_nvl72_rack_config"),
-}
-# Same fixed network utilization as the chassis sweep; the rack model has no CPU/DRAM inputs.
-RACK_ASSUMPTIONS = {"u_nvlink": 0.50, "u_ib": 0.0, "u_pcie": 0.05, "pue": 1.20}
-RACK_BASES = {"module": "module", "gpu_plus_grace": "gpu-plus-grace"}
-
-
-def rack_expected(result):
- return {
- "computeModulesDcWatts": result["compute_modules_dc_w"],
- "regulatorAllowanceWatts": result["regulator_allowance_w"],
- "trayStaticDcWatts": result["tray_static_dc_w"],
- "nvswitchTraysDcWatts": result["nvswitch_trays_dc_w"],
- "trayConversionLossWatts": result["tray_conversion_loss_w"],
- "rackDcWatts": result["rack_dc_w"],
- "powerShelfEfficiency": result["power_shelf_efficiency"],
- "powerShelfLossWatts": result["power_shelf_loss_w"],
- "rackAcWatts": result["rack_ac_w"],
- "facilityWatts": result["facility_w"],
- "perGpuAcWatts": result["per_gpu_ac_w"],
- "perGpuFacilityWatts": result["per_gpu_facility_w"],
- }
-
-
-def git(repo, *args):
- return subprocess.check_output(["git", "-C", str(repo), *args], text=True).strip()
-
-
-def main():
- parser = argparse.ArgumentParser(description=__doc__)
- parser.add_argument("model_repo", type=Path)
- parser.add_argument("--output-dir", type=Path, default=Path(__file__).resolve().parents[1] / "src/lib")
- args = parser.parse_args()
- repo = args.model_repo.resolve()
- if git(repo, "rev-parse", "HEAD") != REVISION:
- raise SystemExit(f"Model checkout must be pinned to {REVISION}")
- if git(repo, "status", "--porcelain", "--untracked-files=all", "--", "*.py", "README.md", "AGENTS.md"):
- raise SystemExit("Model Python sources or methodology files are dirty; use the clean pinned revision")
-
- profiles, cases = {}, []
- for hardware, (model_path, function_name, config_factory) in MODELS.items():
- path = repo / "human_verified" / model_path
- sys.path.insert(0, str(path.parent))
- module = importlib.import_module(path.stem)
- function, cfg = getattr(module, function_name), getattr(module, config_factory)()
- assumptions = dict(ASSUMPTIONS)
- assumptions.update({"u_eth": 0.0} if hardware.startswith("mi") else {"u_nvlink": 0.50, "u_ib": 0.0})
- if hardware in ("b200", "b300"):
- assumptions["u_dpu"] = 0.0
- baseline = function(0.0, cfg=cfg, **assumptions)
- fixed_components = {key: value for key, value in baseline["components_dc_w"].items()
- if not key.endswith("_measured") and key != "chassis_fans"}
- fans, psu = cfg.fans, cfg.psu
- capacity = psu.active_capacity_w if hardware == "b200" else psu.load_sharing_capacity_w
- limit = psu.active_capacity_w if hardware == "b200" else (
- psu.redundant_capacity_w if hardware in ("h100", "h200") else psu.modeled_capacity_w)
- profiles[hardware] = {
- "modelPath": "human_verified/" + model_path,
- "functionName": function_name,
- "configFactory": config_factory,
- "gpuCount": 8,
- "assumptions": {**assumptions, "fan_pwm": None},
- "defaultConfig": asdict(cfg),
- "fixedComponentsDcWatts": fixed_components,
- "fan": {
- "electricalNameplateWatts": fans.electrical_nameplate_w,
- "electricalGroupsWatts": [fans.n_80mm * fans.rated_80mm_w, fans.n_60mm * fans.rated_60mm_w]
- if hardware == "b200" else [fans.electrical_nameplate_w],
- "minPwm": fans.min_pwm_frac,
- "maxPwm": fans.normal_max_pwm_frac,
- "fullCoolingLoadWatts": fans.full_cooling_load_w,
- "exponent": fans.fan_curve_exponent,
- },
- "psu": {
- "loadSharingCapacityWatts": capacity,
- "maxDcWatts": limit,
- "efficiencyCurve": sorted(psu.efficiency_curve.items()),
- },
- }
-
- def evaluate(gpu, pue=1.2):
- return function(gpu, cfg=cfg, **{**assumptions, "pue": pue})
-
- fixed = sum(fixed_components.values())
- samples = {0.05, 1.25, 1000.25, 1000.75, 2400.0, 4000.0, 5600.0}
- # Both sides of fan saturation and all reachable PSU interpolation knots.
- for boundary in (fans.full_cooling_load_w - fixed,):
- samples.update(round(boundary + offset, 3) for offset in (-0.2, 0.0, 0.2) if boundary + offset > 0)
- for fraction in sorted(psu.efficiency_curve):
- target = fraction * capacity
- if not baseline["dc_total_w"] < target <= limit:
- continue
- lo, hi = 0.0, limit
- for _ in range(50):
- mid = (lo + hi) / 2
- try:
- below = evaluate(mid)["dc_total_w"] < target
- except ValueError:
- below = False
- if below:
- lo = mid
- else:
- hi = mid
- samples.update(round(hi + offset, 3) for offset in (-0.2, 0.0, 0.2))
- samples.add(limit) # Capacity overflow must be unavailable, never clamped.
- for gpu in sorted(samples):
- for pue in (1.0, 1.2):
- case = {"hardware": hardware, "measuredGpuWatts": gpu, "pue": pue}
- try:
- result = evaluate(gpu, pue)
- case["expected"] = {
- "preFanDcWatts": result["pre_fan_dc_w"],
- "fanWatts": result["components_dc_w"]["chassis_fans"],
- "dcWatts": result["dc_total_w"],
- "psuEfficiency": result["psu_efficiency"],
- "psuLossWatts": result["psu_conversion_loss_w"],
- "chassisAcWatts": result["ac_wall_w"],
- "facilityWatts": result["utility_power_w"],
- }
- except ValueError as error:
- case["expected"] = None
- case["referenceError"] = str(error)
- cases.append(case)
-
- rack_profiles, rack_cases = {}, []
- for hardware, (model_path, function_name, config_factory) in RACK_MODELS.items():
- path = repo / "human_verified" / model_path
- sys.path.insert(0, str(path.parent))
- module = importlib.import_module(path.stem)
- function, cfg = getattr(module, function_name), getattr(module, config_factory)()
- utilization = {key: RACK_ASSUMPTIONS[key] for key in ("u_nvlink", "u_ib", "u_pcie")}
- # Everything except the measured compute modules and the shelf curve is fixed at
- # these utilizations. Keep the source's per-tray block order: the app re-sums it.
- tray_blocks, tray_details = module._compute_tray_static(
- cfg.compute_tray, u_ib=utilization["u_ib"], u_pcie=utilization["u_pcie"])
- nvswitch = module.nvswitch5_power(utilization["u_nvlink"], cfg.nvswitch)
- shelf = cfg.power_shelf
- rack_profiles[hardware] = {
- "modelPath": "human_verified/" + model_path,
- "functionName": function_name,
- "configFactory": config_factory,
- "topology": "nvl72-rack",
- "gpuCount": cfg.n_gpu,
- "computeTrayCount": cfg.n_compute_trays,
- "gpusPerComputeTray": cfg.compute_tray.n_gpu,
- "graceSocketsPerComputeTray": cfg.compute_tray.n_grace,
- "nvswitchTrayCount": cfg.n_nvswitch_trays,
- "assumptions": dict(RACK_ASSUMPTIONS),
- "defaultConfig": asdict(cfg),
- "computeTrayStaticDcWatts": tray_blocks,
- "computeTrayStaticDetails": tray_details,
- "nvswitchTraySiliconWatts": nvswitch["pair_w"],
- "nvswitchTrayResidualWatts": cfg.nvswitch_tray_residual_w,
- "managementSwitchCount": cfg.n_management_switches,
- "managementSwitchWatts": cfg.management_switch_w,
- "trayInputConversionEfficiency": cfg.tray_input_conversion_efficiency,
- "regulatorLossFracOfTdp": cfg.regulator_loss_frac_of_tdp,
- "regulatorAllowanceIncludesGrace": cfg.regulator_allowance_includes_grace,
- "powerShelf": {
- "installedCapacityWatts": shelf.installed_capacity_w,
- "redundantCapacityWatts": shelf.redundant_capacity_w,
- "efficiencyCurve": sorted(shelf.efficiency_curve.items()),
- },
- "unverifiedParameters": module.UNVERIFIED_PARAMETERS,
- }
-
- def evaluate_rack(basis, pue=1.2, **watts):
- return function(basis=basis, cfg=cfg, **{**utilization, "pue": pue}, **watts)
-
- installed, n_trays = shelf.installed_capacity_w, cfg.n_compute_trays
- # 5400 W/tray is the GB200 module TDP anchor (2 x 2700 W) from the research note.
- module_samples = {0.05, 1.25, 1000.25, 2000.0, 3000.75, 4000.0, 5400.0, 6000.0, 7200.0}
- baseline_dc = evaluate_rack("module", module_w_per_tray=0.0)["rack_dc_w"]
- # Both sides of every reachable shelf-efficiency knot, on rack DC load.
- for fraction in sorted(shelf.efficiency_curve):
- target = fraction * installed
- if not baseline_dc < target <= installed:
- continue
- lo, hi = 0.0, installed / n_trays
- for _ in range(50):
- mid = (lo + hi) / 2
- try:
- below = evaluate_rack("module", module_w_per_tray=mid)["rack_dc_w"] < target
- except ValueError:
- below = False
- if below:
- lo = mid
- else:
- hi = mid
- module_samples.update(round(hi + offset, 3) for offset in (-0.2, 0.0, 0.2))
- module_samples.add(installed / n_trays) # Shelf overflow must be unavailable, never clamped.
- for module_w in sorted(module_samples):
- for pue in (1.0, 1.1, 1.2):
- case = {"hardware": hardware, "basis": RACK_BASES["module"],
- "moduleWattsPerTray": module_w, "pue": pue}
- try:
- case["expected"] = rack_expected(evaluate_rack("module", pue, module_w_per_tray=module_w))
- except ValueError as error:
- case["expected"] = None
- case["referenceError"] = str(error)
- rack_cases.append(case)
- for gpu_w in (0.05, 1.25, 2000.0, 3000.25, 4800.0, 5600.0):
- for grace_w in (0.05, 300.0, 600.5):
- for pue in (1.0, 1.1):
- result = evaluate_rack("gpu_plus_grace", pue, gpu_board_w_per_tray=gpu_w,
- grace_socket_w_per_tray=grace_w)
- rack_cases.append({"hardware": hardware, "basis": RACK_BASES["gpu_plus_grace"],
- "gpuBoardWattsPerTray": gpu_w, "graceSocketWattsPerTray": grace_w,
- "pue": pue, "expected": rack_expected(result)})
-
- source_hashes = {}
- # Include the complete pinned Python implementation and plot entry points.
- for relative in git(repo, "ls-files", "*.py").splitlines():
- source_hashes[relative] = hashlib.sha256((repo / relative).read_bytes()).hexdigest()
- provenance = {
- "modelRevision": REVISION,
- "modelRevisionStatus": REVISION_STATUS,
- "source": SOURCE,
- "status": "DRAFT / pending human verification",
- }
- outputs = {
- "system-power-model.profiles.json": {
- **provenance,
- "assumptionsSource": f"{SOURCE}/blob/{REVISION}/README.md#chassis-models",
- "assumptions": ASSUMPTIONS,
- "rackAssumptionsSource": f"{SOURCE}/blob/{REVISION}/human_verified/{RACK_MODEL_PATH}",
- "rackAssumptions": RACK_ASSUMPTIONS,
- "sourceSha256": dict(sorted(source_hashes.items())),
- "profiles": profiles,
- "rackProfiles": rack_profiles,
- },
- "system-power-model.reference.json": {**provenance, "cases": cases, "rackCases": rack_cases},
- }
- args.output_dir.mkdir(parents=True, exist_ok=True)
- for filename, payload in outputs.items():
- (args.output_dir / filename).write_text(json.dumps(payload, indent=2, allow_nan=False) + "\n")
- print(f"Generated {len(profiles)} chassis profiles, {len(rack_profiles)} rack profiles, "
- f"{len(cases)} chassis and {len(rack_cases)} rack Python reference cases at {REVISION}")
-
-
-if __name__ == "__main__":
- main()
diff --git a/packages/app/scripts/update-system-power-provenance.ts b/packages/app/scripts/update-system-power-provenance.ts
new file mode 100644
index 000000000..38c085d7a
--- /dev/null
+++ b/packages/app/scripts/update-system-power-provenance.ts
@@ -0,0 +1,51 @@
+import { createHash } from 'node:crypto';
+import { readFile, writeFile } from 'node:fs/promises';
+import { resolve } from 'node:path';
+import { pathToFileURL } from 'node:url';
+
+const ROOT = resolve(import.meta.dirname, '../../..');
+const PROFILE_PATH = 'packages/app/src/lib/system-power-model.profiles.json';
+const MANIFEST_PATH = 'packages/app/src/lib/system-power-model.provenance.json';
+const SOURCE_PATHS = [
+ 'packages/app/src/lib/modeled-system-power.ts',
+ PROFILE_PATH,
+ 'packages/app/src/lib/system-power-model.ts',
+];
+const sha256 = (value: string | Uint8Array) => createHash('sha256').update(value).digest('hex');
+
+export async function buildSystemPowerProvenance(root = ROOT) {
+ const sourceSha256 = Object.fromEntries(
+ await Promise.all(
+ SOURCE_PATHS.map(async (path) => [path, sha256(await readFile(resolve(root, path)))]),
+ ),
+ );
+ return {
+ modelRevision: `app-sha256:${sha256(JSON.stringify(sourceSha256))}`,
+ modelRevisionStatus: 'App-owned TypeScript equations, parameters and admission/PUE policy',
+ source: 'https://github.com/SemiAnalysisAI/InferenceX-app',
+ status: 'DRAFT / pending human verification',
+ assumptionsSource: `${PROFILE_PATH}#/assumptions`,
+ rackAssumptionsSource: `${PROFILE_PATH}#/rackAssumptions`,
+ sourceSha256,
+ };
+}
+
+async function main() {
+ const args = process.argv.slice(2);
+ if (args.length > 1 || (args.length === 1 && args[0] !== '--check'))
+ throw new Error('Usage: bun packages/app/scripts/update-system-power-provenance.ts [--check]');
+ const current = await buildSystemPowerProvenance();
+ const target = resolve(ROOT, MANIFEST_PATH);
+ if (args[0] === '--check') {
+ const stored = JSON.parse(await readFile(target, 'utf8'));
+ if (JSON.stringify(stored) !== JSON.stringify(current))
+ throw new Error('Model provenance is stale; run this script without --check.');
+ } else {
+ await writeFile(target, `${JSON.stringify(current, null, 2)}\n`);
+ }
+ console.log(current.modelRevision);
+}
+
+if (process.argv[1] && pathToFileURL(resolve(process.argv[1])).href === import.meta.url) {
+ await main();
+}
diff --git a/packages/app/src/app/api/v1/views/extensions.test.ts b/packages/app/src/app/api/v1/views/extensions.test.ts
index 577a2fecd..9dbd26f30 100644
--- a/packages/app/src/app/api/v1/views/extensions.test.ts
+++ b/packages/app/src/app/api/v1/views/extensions.test.ts
@@ -432,8 +432,8 @@ describe('new dashboard projections', () => {
measuredBasis: 'module',
sensorKind: 'module',
pue: 1.1,
- modelPath: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
- modelRevision: expect.stringMatching(/^[0-9a-f]{40}$/u),
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
+ modelRevision: expect.stringMatching(/^app-sha256:[0-9a-f]{64}$/u),
profileSha256: expect.stringMatching(/^[0-9a-f]{64}$/u),
});
// Pinned rack reference: 1.6077325 kW/GPU, including PUE and planning margin.
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index 9cf23643c..8f109c88b 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -222,7 +222,7 @@ describe('profit power basis preview', () => {
measuredBasis: 'module',
sensorKind: 'module',
pue: 1.1,
- modelPath: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
modelRevision: rack.modelRevision,
profileSha256: expect.stringMatching(/^[0-9a-f]{64}$/u),
} satisfies ProfitPowerSource);
@@ -447,7 +447,7 @@ describe('profit power basis preview', () => {
expect(modeled.powerSource).toMatchObject({
topology: 'chassis',
pue: 1.3,
- modelPath: 'human_verified/mi355x_chassis/mi355x_chassis_power_model.py',
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
});
expect(modeled.revenuePerGpuHour).toBe(baseline.revenuePerGpuHour);
const ratio = 2.09 / 1.5976675;
diff --git a/packages/app/src/components/calculator/profit-power.ts b/packages/app/src/components/calculator/profit-power.ts
index 6b7a8c276..48ed51813 100644
--- a/packages/app/src/components/calculator/profit-power.ts
+++ b/packages/app/src/components/calculator/profit-power.ts
@@ -20,7 +20,7 @@ interface ProfitPowerProfile {
pue: number;
modelPath: string;
modelRevision: string;
- /** SHA-256 of the pinned source file behind `modelPath`; null if the profile lacks one. */
+ /** Equation-file hash; modelRevision also covers parameters and admission/PUE policy. */
profileSha256: string | null;
}
diff --git a/packages/app/src/components/inference/utils/tooltip-utils.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
index 0a2306a1b..370d462b8 100644
--- a/packages/app/src/components/inference/utils/tooltip-utils.test.ts
+++ b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
@@ -1,4 +1,4 @@
-import { describe, it, expect } from 'vitest';
+import { describe, it, expect, vi } from 'vitest';
import type { HardwareConfig, InferenceData } from '@/components/inference/types';
import type { SystemPowerEstimate } from '@/lib/modeled-system-power';
@@ -72,8 +72,8 @@ function tooltipConfig(overrides: Partial = {}): TooltipConfig {
const systemPower = {
status: 'supported',
hardware: 'h100',
- modelRevision: 'ca4403aa527069857351ad8047dbb726844b3382',
- modelPath: 'chassis/H100.py',
+ modelRevision: `app-sha256:${'a'.repeat(64)}`,
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
gpuCount: 16,
chassisCount: 2,
chassisAcWatts: 12000,
@@ -146,11 +146,32 @@ describe('modeled system-power tooltip', () => {
expect(html).toContain(
'Includes GPU chassis CPUs; excludes separate CPU-only frontend/router hosts.',
);
- expect(html).toContain(`/blob/${systemPower.modelRevision}/${systemPower.modelPath}`);
+ expect(html).toContain(`/blob/master/${systemPower.modelPath}`);
expect(html).not.toContain('12,000 W/GPU');
expect(html).not.toContain('Unmeasured chassis GPUs');
});
+ it.each(['en', 'zh'] as const)(
+ 'links %s model provenance to the deployed app source',
+ (locale) => {
+ const buildRef = 'b'.repeat(40);
+ vi.stubEnv('NEXT_PUBLIC_APP_SOURCE_REF', buildRef);
+ try {
+ const html = generateTooltipContent(config({ locale }));
+ const app = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${buildRef}`;
+ expect(html).toContain(`${app}/${systemPower.modelPath}`);
+ expect(html).toContain(`${app}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`);
+ expect(html).toContain(locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions');
+ expect(html).toContain(`title="${systemPower.modelRevision}"`);
+ expect(html).toContain('h100 · aaaaaaaaaaaa');
+ expect(html).not.toContain('inferencex_power_model');
+ expect(html).not.toContain(`/blob/${systemPower.modelRevision}/`);
+ } finally {
+ vi.unstubAllEnvs();
+ }
+ },
+ );
+
it('labels an extrapolated partial chassis and reports the measured GPUs’ share', () => {
const data = pt({
physicalChips: 4,
@@ -189,7 +210,7 @@ describe('modeled system-power tooltip', () => {
const trays = {
...systemPower,
hardware: 'gb200',
- modelPath: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ modelPath: 'packages/app/src/lib/system-power-model.ts',
gpuCount: 8,
chassisCount: 2,
modeledGpuCount: 8,
diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts
index e091adc77..9ebfbb110 100644
--- a/packages/app/src/components/inference/utils/tooltipUtils.ts
+++ b/packages/app/src/components/inference/utils/tooltipUtils.ts
@@ -266,7 +266,7 @@ const SYSTEM_POWER_STRINGS = {
facility: 'Modeled facility power',
assumptions: 'CPU/DRAM utilization: 20%; PCIe: 5%; NVMe: 0%; fans: auto.',
platformAssumptions: 'NVIDIA NVLink: 50%, IB: 0%; AMD Ethernet: 0%.',
- sweep: 'Fixed README inference sweep',
+ guide: 'Power model assumptions',
topology: (chassis: number, measured: number, modeled: number) =>
measured === modeled
? `${chassis} full eight-GPU chassis · ${measured} GPUs`
@@ -317,7 +317,7 @@ const SYSTEM_POWER_STRINGS = {
facility: '数据中心功耗估算',
assumptions: 'CPU/DRAM 利用率:20%;PCIe:5%;NVMe:0%;风扇:自动。',
platformAssumptions: 'NVIDIA NVLink:50%,IB:0%;AMD Ethernet:0%。',
- sweep: 'README 中的固定推理参数扫描',
+ guide: '功耗模型与假设',
topology: (chassis: number, measured: number, modeled: number) =>
measured === modeled
? `${chassis} 个完整八卡机箱 · ${measured} 张 GPU`
@@ -376,8 +376,10 @@ const modeledSystemPowerHTML = (
if (!isPinned || estimate.reason === 'workload') return '';
return tooltipLine(t.unavailable, t.reasons[estimate.reason]);
}
- const sourceUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/${estimate.modelPath}`;
- const readmeUrl = `https://github.com/SemiAnalysisAI/inferencex_power_model/blob/${estimate.modelRevision}/README.md`;
+ const sourceRef = encodeURIComponent(process.env.NEXT_PUBLIC_APP_SOURCE_REF || 'master');
+ const appSource = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${sourceRef}`;
+ const sourceUrl = `${appSource}/${estimate.modelPath}`;
+ const guideUrl = `${appSource}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`;
// Tray estimates measure the compute module; chassis estimates model the CPU/DRAM.
const tray = estimate.topologyBasis === 'nvl72-trays' ? estimate : null;
const topology = tray
@@ -406,8 +408,8 @@ const modeledSystemPowerHTML = (
${tooltipLine(t.deploymentAc, `${fmt(estimate.deploymentAcWatts)} W`)}
${tooltipLine(`${t.facility} (PUE ${fmt(estimate.pue)})`, `${fmt(estimate.deploymentFacilityWatts)} W`)}
${topology}${extrapolation}${uniformHosts} ${notes.join(' ')}
- ${tooltipLine(t.model, `${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.slice(0, 12))} `)}
- ${t.sweep}
+ ${tooltipLine(t.model, `${escapeHtml(estimate.hardware)} · ${escapeHtml(estimate.modelRevision.replace(/^app-sha256:/u, '').slice(0, 12))} `)}
+ ${t.guide}
`
: ''
}
diff --git a/packages/app/src/lib/modeled-system-power-export.test.ts b/packages/app/src/lib/modeled-system-power-export.test.ts
index 0598173d5..f308c32fa 100644
--- a/packages/app/src/lib/modeled-system-power-export.test.ts
+++ b/packages/app/src/lib/modeled-system-power-export.test.ts
@@ -86,6 +86,10 @@ describe('offline modeled PowerX comparisons', () => {
const source = input();
const before = structuredClone(source);
const result = buildComparison(source);
+ expect(result.metadata.model.source).toBe('https://github.com/SemiAnalysisAI/InferenceX-app');
+ expect(result.metadata.model.modelRevision).toMatch(/^app-sha256:[0-9a-f]{64}$/u);
+ expect(result.rows[0].model_path).toBe('packages/app/src/lib/system-power-model.ts');
+ expect(result.rows[0].modeled.modelRevision).toBe(result.metadata.model.modelRevision);
expect(result.metadata.pue_override).toBeNull();
expect(result.metadata.pue_defaults).toEqual({ air_cooled_chassis: 1.3, dlc_nvl72_rack: 1.1 });
expect(result.metadata.model.assumptions.pue).toBe(1.2);
@@ -258,7 +262,7 @@ describe('offline modeled PowerX comparisons', () => {
pue: 1.1,
measured_basis: 'module',
sensor_kind: 'module',
- model_path: 'human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py',
+ model_path: 'packages/app/src/lib/system-power-model.ts',
assumptions: { u_nvlink: 0.5, pue: 1.1 },
measured_inputs: {
avg_gpu_w: 900.25,
@@ -307,7 +311,7 @@ describe('offline modeled PowerX comparisons', () => {
source.rows[0].benchmark.hardware = 'H200';
expect(buildComparison(source).rows[0]).toMatchObject({
assumptions: { u_cpu: 0.2 },
- model_path: 'human_verified/hgx_h200_chassis/h200_chassis_power_model.py',
+ model_path: 'packages/app/src/lib/system-power-model.ts',
});
// NVL72 rows need the schema-v2 contract; the unversioned exception is x86 single-node only.
source.rows[0].benchmark.hardware = 'gb200';
diff --git a/packages/app/src/lib/modeled-system-power.ts b/packages/app/src/lib/modeled-system-power.ts
index f2905bb6d..a4671efa9 100644
--- a/packages/app/src/lib/modeled-system-power.ts
+++ b/packages/app/src/lib/modeled-system-power.ts
@@ -10,7 +10,7 @@ import {
type SystemPowerRackHardware,
} from '@/lib/system-power-model';
-// Application policy for the air-cooled chassis profiles; the pinned Python default stays 1.2.
+// Application policy for air-cooled chassis; the app profile default stays 1.2.
export const AIR_COOLED_SYSTEM_PUE = 1.3;
// Application policy for the direct-liquid-cooled NVL72 rack profiles (docs/powerx-system-power.md).
export const DLC_SYSTEM_PUE = 1.1;
diff --git a/packages/app/src/lib/system-power-model.profiles.json b/packages/app/src/lib/system-power-model.profiles.json
index 5424d48dd..53f154df4 100644
--- a/packages/app/src/lib/system-power-model.profiles.json
+++ b/packages/app/src/lib/system-power-model.profiles.json
@@ -1,9 +1,4 @@
{
- "modelRevision": "6fcc086b77576d4cecb9d0c79637d6daf980308c",
- "modelRevisionStatus": "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification",
- "source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
- "status": "DRAFT / pending human verification",
- "assumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/6fcc086b77576d4cecb9d0c79637d6daf980308c/README.md#chassis-models",
"assumptions": {
"u_pcie": 0.05,
"u_cpu": 0.2,
@@ -11,86 +6,15 @@
"u_nvme": 0.0,
"pue": 1.2
},
- "rackAssumptionsSource": "https://github.com/SemiAnalysisAI/inferencex_power_model/blob/6fcc086b77576d4cecb9d0c79637d6daf980308c/human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
"rackAssumptions": {
"u_nvlink": 0.5,
"u_ib": 0.0,
"u_pcie": 0.05,
"pue": 1.2
},
- "sourceSha256": {
- "human_verified/amd_oam_fans/amd_oam_fan_power_model.py": "7ea6f1c65b685311277e6f2c53992d588ec79b598180f2fca300e904ad6dc19a",
- "human_verified/amd_oam_fans/plot_amd_oam_fan_power.py": "e38118fa4d94412af268eaa1667d2c4086d1dcced4da07c951e761d54ea55b46",
- "human_verified/amd_oam_psu/amd_oam_psu_power_model.py": "badcf849a82da35bc9bd621d42cd955ec7f8e3879dbc1bc29c7c03a3e0217af1",
- "human_verified/amd_oam_psu/plot_amd_oam_psu_power.py": "39a1c3eac5a1e6337425948bfd252e422021b67de2ec5c963be1f8db3a8292fc",
- "human_verified/amd_oam_ubb/amd_oam_ubb_power_model.py": "2f2f685e595d980c3fc353c624ca7183ee4b7f34c52d649066ef266664bccab4",
- "human_verified/b200_fans/b200_fan_power_model.py": "163c6dd95f460cf6bde1b60cec061cce9be11e69cac1c6fe000b01a7ace81975",
- "human_verified/b200_fans/plot_b200_fan_power.py": "8172b706c4a882ebb1fed220b3bcd602ea9db49019c4b1ac5a862d83be07bedd",
- "human_verified/b200_psu/b200_psu_power_model.py": "58e3f81c7fbf722848184ddd73ed94edb065cfa96d6382d74bf42e8f7e15bbe8",
- "human_verified/b200_psu/plot_b200_psu_power.py": "c96127308dab7105243ca21dcef07dce4fad8f674d9273006e3f4b1c29656b10",
- "human_verified/b200_ubb/b200_ubb_power_model.py": "e3d3acf16a5ef9ca318ba49dc8b58d40823f8b76144d2999dffc9973db96648a",
- "human_verified/b200_ubb/b200_ubb_residual_power_model.py": "85ca624b8d85d504f4e8a245f05aac07ab804b3fb262c2f00eb50ed35e22d273",
- "human_verified/b200_ubb/plot_b200_ubb_residual_power.py": "a0d7a7e5f44948c59262bf62f4f088cb64fb17c29926f7b3327a18fb695903ed",
- "human_verified/b300_fans/b300_fan_power_model.py": "9c5f85d0b7c9d4ee9fb8d233bb3d150a69643902970a68b475bb266918a495c7",
- "human_verified/b300_fans/plot_b300_fan_power.py": "761a8c1c80f8efa0fe49bb749897da20764c1a207efe4d77c22a43b001b319c4",
- "human_verified/b300_psu/b300_psu_power_model.py": "930f30f3a720b964623b35a7ee2d510a03f39ec6751dda5b480c7a8cfb0bd99e",
- "human_verified/b300_psu/plot_b300_psu_power.py": "241880ab01f8d4012a6d2b358f85648bebb229b4966b1af081ecd331997cefc8",
- "human_verified/b300_ubb/b300_ubb_power_model.py": "8cd694f1306ca09de53dcc9590dcc47a1b9bd5c10af9837bb875dc417fe74837",
- "human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7",
- "human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db",
- "human_verified/chassis_plot_utils.py": "0f1c809a527788c8c25490420e29e737e6b7f83cb914dc022c3a4e3701b80643",
- "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "b4640f94c8f9e6be50eb18deff9728bb58559ba193b54577859cafabec4ccc33",
- "human_verified/gb200_nvl72_rack/plot_gb200_nvl72_rack_inference_gpu_sweep.py": "d0f6903df0348ccb2790e973b0dc10408c6f6022695ce039aad0b1e1fcd39519",
- "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "631bd7fd44b84260f16175a814d093795a4cdd0a91a0d19dd572a5ea6c088314",
- "human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f",
- "human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c",
- "human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb",
- "human_verified/generic/connectx8/plot_connectx8_power.py": "97921625373a479f03ad3930c8542e86da6c4ee4521654f77bcb965d93759433",
- "human_verified/generic/cpu/cpu_power_model.py": "ab4315df415d70474ae4fb5c700fdf5b6de2a8a89f33823fbc8664efa01bf72f",
- "human_verified/generic/cpu/plot_cpu_power.py": "bc5866234e7af605c8b43664cca1c1e96116d4fefb88c791cbe5e35f6234e821",
- "human_verified/generic/dpu/dpu_power_model.py": "677cce41b990213c96f73e016c8bf24da9841ef995a7522e1038c832d502e398",
- "human_verified/generic/dpu/plot_dpu_power.py": "d52357acdabb6a69b2ac87037cc339208e70b1bf7bdd31483b98dbedad06f05b",
- "human_verified/generic/dram/dram_power_model.py": "b8ace82de182715a003c747710fc62275f2392d883aadb57c05f19beb1367b67",
- "human_verified/generic/dram/plot_dram_power.py": "8a3d59fbf31492f9de04d8f3f80ddc33719018b6f47e8b0e671d3f63e29acead",
- "human_verified/generic/nvme/nvme_power_model.py": "53fa1f59edf4427a62c946758d4f3df55b1591f8d43f824e07fd61e6d4196743",
- "human_verified/generic/nvme/plot_nvme_power.py": "276d013eb26ff13c68e60cbf48bddc88fd6f5ff6c0a59a9854aec11da6ebf8b7",
- "human_verified/generic/pcie_switches/pcie5_144lane_switch_power_model.py": "23d23322126c4251cb2ad5bee4448b0f9ad71b531607e8353cd96a90410cc28d",
- "human_verified/generic/pcie_switches/plot_pcie_switch_power.py": "3cf3139b18495efc320c1f3d2754d8f732b16f1775959f66ff222ef97b64de90",
- "human_verified/generic/pollara400/plot_pollara400_power.py": "22196c3037330b07903109e0f9a6917fd2e9e8e56ee942df552a84efd06a470e",
- "human_verified/generic/pollara400/pollara400_power_model.py": "8dd5e674d9bfa5e19cb68c5684eb717df61063762afa555cd1a9c43f1d24c5d8",
- "human_verified/generic/retimers/pcie5_x16_retimer_power_model.py": "393381bb40ef12cb81b681eff742129453136c6ba6c4b6950682e6cdc63a5335",
- "human_verified/generic/retimers/plot_pcie5_x16_retimer_power.py": "c8f3a9aede8876cebab9f0674926b59f591f71ec1e58e86bdebe3ec63f61a208",
- "human_verified/generic/thor2/plot_thor2_power.py": "7f03cd9af3dfd4e787e0e9d429a5758e24f4b62990af6d7d7e22a1a5224a78c7",
- "human_verified/generic/thor2/thor2_power_model.py": "67b8e91a01dccd3abd7bdd37dbef86d1194d398c9326b8774eede1488962c54c",
- "human_verified/hgx_b200_chassis/b200_chassis_power_model.py": "89d94969ce1acfeee67784ad269c431415995f9b18996d4b28d213d34393864a",
- "human_verified/hgx_b200_chassis/plot_chassis_inference_gpu_sweep.py": "ed9378b471bf502adda4ac8f2467d803e48a5347ae329e118bfeebc9400ecc79",
- "human_verified/hgx_b200_chassis_residual/b200_chassis_residual_power_model.py": "d23fe72c6039545e51d6271ddef28bee5d69f9ee7e4f8032a5def60021f973eb",
- "human_verified/hgx_b200_chassis_residual/plot_chassis_residual_power.py": "a46d81768e0af6982bdb9d146b7e42fdbb86f1afcdb581d808657b5508a65414",
- "human_verified/hgx_b300_chassis/b300_chassis_power_model.py": "68af8917ead2472cb0f6473784a5a75a24233cb766df8fa2684421a07d660fbc",
- "human_verified/hgx_b300_chassis/plot_b300_chassis_inference_gpu_sweep.py": "bfe74771dd61b6dbb3dcf1ebbf6451c07b6020dcf0000482225672eb081f1d11",
- "human_verified/hgx_h100_chassis/h100_chassis_power_model.py": "6850ec92346af1864f724a41d9ea512e0d55f45d683a3d575477aa08ca89a6c8",
- "human_verified/hgx_h100_chassis/plot_h100_chassis_inference_gpu_sweep.py": "4af0da4e056f2300d271dd041ffb2ff9a76d39aaf80ae227a38ce777b949d852",
- "human_verified/hgx_h200_chassis/h200_chassis_power_model.py": "56b40c9f13e50d81a02a594f0762f5c498483f261f9aec7b7294b148c65d7eb9",
- "human_verified/hgx_h200_chassis/plot_h200_chassis_inference_gpu_sweep.py": "326b3baff711b5a322349cb272f83eaa44c8825cdbb9cc87cce52d7433e4b695",
- "human_verified/hopper_fans/hopper_fan_power_model.py": "8b656f8498b6709b8ac7392999333eef06a529e12af141cc4565a2272c589f24",
- "human_verified/hopper_fans/plot_hopper_fan_power.py": "006c13f8143ad0ebc853ecba9713660f65daa9ee520ce0b55763ce06a4a436fa",
- "human_verified/hopper_nvswitch/hopper_nvswitch_power_model.py": "3d3c536bc2af1e75f2cc3c246e0d5337c04809cb90d470227a0d7336b413790c",
- "human_verified/hopper_nvswitch/plot_hopper_nvswitch_power.py": "c13b291df6426a2117d82f1425009ddf0f55340888a0843dd5f5888503c62e06",
- "human_verified/hopper_psu/hopper_psu_power_model.py": "a7630454e128e87bc0529a02cebb11138efbe1b07d4e4cb08a9721cc49f8d58d",
- "human_verified/hopper_psu/plot_hopper_psu_power.py": "4a73fffa633e0a599a8bf746e0868f719669ff59d69a102bf563fcc4596e40a4",
- "human_verified/hopper_ubb/hopper_ubb_power_model.py": "3f8c6c9560c32dbe1e98d0af82dbafa3fc697efc3d32642afe303455bfd5ab74",
- "human_verified/mi300x_chassis/mi300x_chassis_power_model.py": "69c4b11e860e9a174664ae040691aab9e349410040ac8524dee6a7f2102afab6",
- "human_verified/mi300x_chassis/plot_mi300x_chassis_inference_gpu_sweep.py": "1793a7356b95821cd4f7390a4cae55c58ffcc37f1b8f43f873499d84937463e4",
- "human_verified/mi325x_chassis/mi325x_chassis_power_model.py": "59b6ce4ff626f1c5f42e8d7b0d33318e0e493a3e358dc1f4f92011be94b16c9c",
- "human_verified/mi325x_chassis/plot_mi325x_chassis_inference_gpu_sweep.py": "c0f350bc4978a18759dd108756290fcd803f30209ccfc2682988bbf7ff0feb4d",
- "human_verified/mi355x_chassis/mi355x_chassis_power_model.py": "c178f71efe53b424f5a1804fd188a79f99575e821a357b9128added48f154c8d",
- "human_verified/mi355x_chassis/plot_mi355x_chassis_inference_gpu_sweep.py": "5071db35ce4bb4d7754419877f1e0d9df0c48be4381316b77d80fa1e58fe46c6"
- },
"profiles": {
"h100": {
- "modelPath": "human_verified/hgx_h100_chassis/h100_chassis_power_model.py",
- "functionName": "h100_chassis_power",
- "configFactory": "make_h100_config",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -102,152 +26,6 @@
"u_ib": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "H100 SXM5",
- "gpu_component_key": "h100_sxm_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 640.0,
- "gpu_idle_example_w_per_gpu": 100.0,
- "gpu_decode_example_w_per_gpu": 300.0,
- "gpu_prefill_example_w_per_gpu": 520.0,
- "gpu_aggregate_example_w_per_gpu": 480.0,
- "gpu_peak_w_per_gpu": 700.0,
- "ubb": {
- "gpu_label": "H100 SXM5",
- "gpu_component_key": "h100_sxm_8x_measured",
- "n_gpu": 8,
- "nvswitch": {
- "n_asic": 4,
- "nvlink_ports_per_asic": 64,
- "phy_lanes_per_asic": 128,
- "phy_lane_gbps": 100.0,
- "serdes_class_gbps": 112.0,
- "serdes_pj_per_bit": 3.0,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 120.0,
- "digital_floor_frac": 0.4
- },
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "baseboard_controller_w": 10.0,
- "fpga_cpld_sequencing_w": 8.0,
- "hsc_power_monitor_w": 5.0,
- "clock_refclk_reset_w": 5.0,
- "sensors_i2c_fru_led_w": 4.0,
- "aux_rails_misc_w": 13.0,
- "normal_low_w": 30.0,
- "normal_high_w": 70.0,
- "conservative_cap_w": 90.0
- }
- },
- "cpu": {
- "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)",
- "n_cpu": 2,
- "cores_per_cpu": 56,
- "threads_per_cpu": 112,
- "package_tdp_w": 350.0,
- "base_ghz": 2.0,
- "max_turbo_ghz": 3.8,
- "l3_cache_mb": 105.0,
- "memory_channels_per_cpu": 8,
- "pcie_lanes_per_cpu": 80,
- "package_static_w": 24.0,
- "uncore_io_baseline_w": 36.0,
- "memory_controller_baseline_w": 22.0,
- "uncore_dynamic_max_w": 8.0,
- "memory_controller_dynamic_max_w": 6.0,
- "core_curve_r": 1.72,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 8,
- "n_dimm": 32,
- "capacity_gb_per_dimm": 64.0,
- "data_rate_mtps": 4800.0,
- "dimm_type": "RDIMM",
- "background_w": 2.147,
- "refresh_w": 0.5509999999999999,
- "termination_w": 1.102,
- "io_dynamic_max_w": 2.3400000000000003,
- "core_dynamic_max_w": 2.86,
- "activity_exponent": 1.0
- },
- "connectx_compute": {
- "n_nic": 8,
- "net_serdes_w": 9.5,
- "pcie_serdes_w": 5.5,
- "board_w": 2.0,
- "digital_max_w": 10.0,
- "digital_floor_frac": 0.65,
- "include_optic": true,
- "optic_w": 8.0
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 8,
- "front_u2_idle_w": 5.0,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 10.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 8.0,
- "cpld_tpm_superio_w": 5.0,
- "clock_sensor_fru_w": 4.0,
- "storage_backplane_idle_w": 6.0,
- "front_panel_usb_led_w": 2.0,
- "aux_margin_w": 2.0,
- "normal_low_w": 35.0,
- "normal_high_w": 70.0
- },
- "fans": {
- "electrical_nameplate_w": 1100.0,
- "airflow_cfm_at_normal_max_pwm": 1105.0,
- "min_pwm_frac": 0.22,
- "normal_max_pwm_frac": 0.8,
- "full_cooling_load_w": 9500.0,
- "fan_curve_exponent": 1.1
- },
- "psu": {
- "n_installed_psu": 6,
- "n_load_sharing_psu": 6,
- "n_redundant_capacity_psu": 4,
- "psu_capacity_w": 3300.0,
- "redundancy": "4+2",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "storage_mgmt_network_static_w": 60.0,
- "optional_dpu_idle_w": 0.0,
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"hopper_nvswitch_4x": 485.8,
"hopper_pcie_retimers_8x": 92.4,
@@ -281,9 +59,7 @@
}
},
"h200": {
- "modelPath": "human_verified/hgx_h200_chassis/h200_chassis_power_model.py",
- "functionName": "h200_chassis_power",
- "configFactory": "make_h200_config",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -295,152 +71,6 @@
"u_ib": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "H200 SXM5",
- "gpu_component_key": "h200_sxm_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 1128.0,
- "gpu_idle_example_w_per_gpu": 115.0,
- "gpu_decode_example_w_per_gpu": 330.0,
- "gpu_prefill_example_w_per_gpu": 540.0,
- "gpu_aggregate_example_w_per_gpu": 510.0,
- "gpu_peak_w_per_gpu": 700.0,
- "ubb": {
- "gpu_label": "H200 SXM5",
- "gpu_component_key": "h200_sxm_8x_measured",
- "n_gpu": 8,
- "nvswitch": {
- "n_asic": 4,
- "nvlink_ports_per_asic": 64,
- "phy_lanes_per_asic": 128,
- "phy_lane_gbps": 100.0,
- "serdes_class_gbps": 112.0,
- "serdes_pj_per_bit": 3.0,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 120.0,
- "digital_floor_frac": 0.4
- },
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "baseboard_controller_w": 10.0,
- "fpga_cpld_sequencing_w": 8.0,
- "hsc_power_monitor_w": 5.0,
- "clock_refclk_reset_w": 5.0,
- "sensors_i2c_fru_led_w": 4.0,
- "aux_rails_misc_w": 13.0,
- "normal_low_w": 30.0,
- "normal_high_w": 70.0,
- "conservative_cap_w": 90.0
- }
- },
- "cpu": {
- "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)",
- "n_cpu": 2,
- "cores_per_cpu": 56,
- "threads_per_cpu": 112,
- "package_tdp_w": 350.0,
- "base_ghz": 2.0,
- "max_turbo_ghz": 3.8,
- "l3_cache_mb": 105.0,
- "memory_channels_per_cpu": 8,
- "pcie_lanes_per_cpu": 80,
- "package_static_w": 24.0,
- "uncore_io_baseline_w": 36.0,
- "memory_controller_baseline_w": 22.0,
- "uncore_dynamic_max_w": 8.0,
- "memory_controller_dynamic_max_w": 6.0,
- "core_curve_r": 1.72,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 8,
- "n_dimm": 32,
- "capacity_gb_per_dimm": 64.0,
- "data_rate_mtps": 4800.0,
- "dimm_type": "RDIMM",
- "background_w": 2.147,
- "refresh_w": 0.5509999999999999,
- "termination_w": 1.102,
- "io_dynamic_max_w": 2.3400000000000003,
- "core_dynamic_max_w": 2.86,
- "activity_exponent": 1.0
- },
- "connectx_compute": {
- "n_nic": 8,
- "net_serdes_w": 9.5,
- "pcie_serdes_w": 5.5,
- "board_w": 2.0,
- "digital_max_w": 10.0,
- "digital_floor_frac": 0.65,
- "include_optic": true,
- "optic_w": 8.0
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 8,
- "front_u2_idle_w": 5.0,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 10.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 8.0,
- "cpld_tpm_superio_w": 5.0,
- "clock_sensor_fru_w": 4.0,
- "storage_backplane_idle_w": 6.0,
- "front_panel_usb_led_w": 2.0,
- "aux_margin_w": 2.0,
- "normal_low_w": 35.0,
- "normal_high_w": 70.0
- },
- "fans": {
- "electrical_nameplate_w": 1100.0,
- "airflow_cfm_at_normal_max_pwm": 1105.0,
- "min_pwm_frac": 0.22,
- "normal_max_pwm_frac": 0.8,
- "full_cooling_load_w": 9500.0,
- "fan_curve_exponent": 1.1
- },
- "psu": {
- "n_installed_psu": 6,
- "n_load_sharing_psu": 6,
- "n_redundant_capacity_psu": 4,
- "psu_capacity_w": 3300.0,
- "redundancy": "4+2",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "storage_mgmt_network_static_w": 60.0,
- "optional_dpu_idle_w": 0.0,
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"hopper_nvswitch_4x": 485.8,
"hopper_pcie_retimers_8x": 92.4,
@@ -474,9 +104,7 @@
}
},
"b200": {
- "modelPath": "human_verified/hgx_b200_chassis/b200_chassis_power_model.py",
- "functionName": "b200_chassis_power",
- "configFactory": "B200ChassisMasterConfig",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -489,147 +117,6 @@
"u_dpu": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "ubb": {
- "n_gpu": 8,
- "nvswitch": {
- "n_asic": 2,
- "serdes_lanes": 144,
- "serdes_lane_gbps": 200.0,
- "serdes_pj_per_bit": 2.5,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 190.0,
- "digital_floor_frac": 0.4
- },
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "include_management_bridge_controller": true,
- "management_bridge_controller_w": 24.0,
- "hmc_bmc_w": 6.0,
- "fpga_cpld_w": 8.0,
- "erot_security_w": 3.0,
- "hsc_power_monitor_w": 5.0,
- "clock_refclk_reset_w": 4.0,
- "sensors_i2c_fru_led_w": 3.0,
- "aux_rails_misc_w": 9.0,
- "normal_low_w": 45.0,
- "normal_high_w": 85.0,
- "conservative_cap_w": 100.0
- }
- },
- "cpu": {
- "name": "2x Intel Xeon Platinum 8570 (Emerald Rapids)",
- "n_cpu": 2,
- "cores_per_cpu": 56,
- "threads_per_cpu": 112,
- "package_tdp_w": 350.0,
- "base_ghz": 2.1,
- "max_turbo_ghz": 4.0,
- "l3_cache_mb": 300.0,
- "memory_channels_per_cpu": 8,
- "pcie_lanes_per_cpu": 80,
- "package_static_w": 22.0,
- "uncore_io_baseline_w": 36.0,
- "memory_controller_baseline_w": 22.0,
- "uncore_dynamic_max_w": 8.0,
- "memory_controller_dynamic_max_w": 6.0,
- "core_curve_r": 1.7,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "DGX-B200-like Xeon 8570, 32x64GB DDR5-5600 RDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 8,
- "n_dimm": 32,
- "capacity_gb_per_dimm": 64.0,
- "data_rate_mtps": 5600.0,
- "dimm_type": "RDIMM",
- "background_w": 2.147,
- "refresh_w": 0.5509999999999999,
- "termination_w": 1.102,
- "io_dynamic_max_w": 2.3400000000000003,
- "core_dynamic_max_w": 2.86,
- "activity_exponent": 1.0
- },
- "connectx7": {
- "n_nic": 8,
- "net_serdes_w": 9.5,
- "pcie_serdes_w": 5.5,
- "board_w": 2.0,
- "digital_max_w": 10.0,
- "digital_floor_frac": 0.65,
- "include_optic": true,
- "optic_w": 8.0
- },
- "dpu": {
- "n_dpu": 1,
- "idle_w_per_dpu": 65.0,
- "public_max_power_cap_w": 150.0,
- "max_modeled_u_dpu": 0.02
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 10,
- "front_u2_idle_w": 5.0,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 10.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 8.0,
- "cpld_tpm_superio_w": 5.0,
- "clock_sensor_fru_w": 4.0,
- "storage_backplane_idle_w": 6.0,
- "front_panel_usb_led_w": 2.0,
- "aux_margin_w": 2.0,
- "normal_low_w": 35.0,
- "normal_high_w": 70.0
- },
- "fans": {
- "n_80mm": 15,
- "rated_80mm_w": 120.0,
- "n_60mm": 4,
- "rated_60mm_w": 25.0,
- "min_pwm_frac": 0.25,
- "normal_max_pwm_frac": 0.78,
- "full_cooling_load_w": 12000.0,
- "fan_curve_exponent": 1.15
- },
- "psu": {
- "n_installed_psu": 6,
- "n_active_psu": 3,
- "psu_capacity_w": 5250.0,
- "redundancy": "3+3",
- "efficiency_curve": {
- "0.05": 0.8864,
- "0.1": 0.9238,
- "0.2": 0.9448,
- "0.5": 0.964,
- "1.0": 0.9556
- }
- },
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"nvswitch5_2x": 406.4,
"ubb_pcie_retimers_8x": 92.4,
@@ -663,9 +150,7 @@
}
},
"b300": {
- "modelPath": "human_verified/hgx_b300_chassis/b300_chassis_power_model.py",
- "functionName": "b300_chassis_power",
- "configFactory": "B300ChassisConfig",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -678,133 +163,6 @@
"u_dpu": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "B300 Blackwell Ultra SXM",
- "gpu_component_key": "b300_sxm_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 2304.0,
- "gpu_idle_example_w_per_gpu": 200.0,
- "gpu_decode_example_w_per_gpu": 520.0,
- "gpu_prefill_example_w_per_gpu": 800.0,
- "gpu_aggregate_example_w_per_gpu": 760.0,
- "gpu_peak_w_per_gpu": 1100.0,
- "ubb": {
- "n_gpu": 8,
- "gpu_label": "B300 Blackwell Ultra SXM",
- "gpu_component_key": "b300_sxm_8x_measured",
- "nvswitch": {
- "n_asic": 2,
- "serdes_lanes": 144,
- "serdes_lane_gbps": 200.0,
- "serdes_pj_per_bit": 2.5,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 190.0,
- "digital_floor_frac": 0.4
- },
- "connectx8": {
- "n_nic": 8,
- "include_optic": true,
- "network_serdes_static_w": 18.0,
- "pcie_switch_static_w": 28.0,
- "board_mgmt_static_w": 5.0,
- "digital_static_w": 12.0,
- "network_dynamic_max_w": 8.0,
- "pcie_switch_dynamic_max_w": 5.0,
- "digital_dynamic_max_w": 10.0,
- "optic_idle_w": 15.0,
- "optic_dynamic_max_w": 2.0,
- "max_nic_slot_power_ref_w": 75.0,
- "normal_low_w_per_nic_with_optic": 70.0,
- "normal_high_w_per_nic_with_optic": 100.0
- },
- "residual_static_w": 85.0,
- "residual_normal_low_w": 60.0,
- "residual_normal_high_w": 120.0
- },
- "cpu": {
- "name": "2x Intel Xeon 6776P (Granite Rapids, DGX B300)",
- "n_cpu": 2,
- "cores_per_cpu": 64,
- "threads_per_cpu": 128,
- "package_tdp_w": 350.0,
- "base_ghz": 2.3,
- "max_turbo_ghz": 3.9,
- "l3_cache_mb": 336.0,
- "memory_channels_per_cpu": 8,
- "pcie_lanes_per_cpu": 88,
- "package_static_w": 24.0,
- "uncore_io_baseline_w": 40.0,
- "memory_controller_baseline_w": 24.0,
- "uncore_dynamic_max_w": 9.0,
- "memory_controller_dynamic_max_w": 7.0,
- "core_curve_r": 1.68,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "DGX-B300-like Xeon 6776P, 32x64GB DDR5-6400 RDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 8,
- "n_dimm": 32,
- "capacity_gb_per_dimm": 64.0,
- "data_rate_mtps": 6400.0,
- "dimm_type": "RDIMM",
- "background_w": 2.147,
- "refresh_w": 0.5509999999999999,
- "termination_w": 1.102,
- "io_dynamic_max_w": 2.3400000000000003,
- "core_dynamic_max_w": 2.86,
- "activity_exponent": 1.0
- },
- "dpu": {
- "n_dpu": 2,
- "idle_w_per_dpu": 65.0,
- "public_max_power_cap_w": 150.0,
- "max_modeled_u_dpu": 0.02
- },
- "nvme": {
- "n_front_u2": 8,
- "front_u2_idle_w": 4.5,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 12.0,
- "onboard_10gbe_w": 4.0,
- "motherboard_pch_aux_w": 10.0,
- "cpld_tpm_superio_w": 6.0,
- "clock_sensor_fru_w": 5.0,
- "storage_backplane_idle_w": 10.0,
- "front_panel_usb_led_w": 3.0,
- "aux_margin_w": 5.0,
- "normal_low_w": 40.0,
- "normal_high_w": 80.0
- },
- "fans": {
- "electrical_nameplate_w": 2000.0,
- "min_pwm_frac": 0.22,
- "normal_max_pwm_frac": 0.85,
- "full_cooling_load_w": 12000.0,
- "fan_curve_exponent": 1.08
- },
- "psu": {
- "n_installed_psu": 12,
- "n_load_sharing_psu": 12,
- "n_redundant_capacity_psu": 6,
- "psu_capacity_w": 3300.0,
- "system_max_w": 15000.0,
- "redundancy": "N+N / 6+6",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"nvswitch5_2x": 406.4,
"connectx8_8x_integrated_pcie_with_optics": 630.0,
@@ -836,9 +194,7 @@
}
},
"mi300x": {
- "modelPath": "human_verified/mi300x_chassis/mi300x_chassis_power_model.py",
- "functionName": "mi300x_chassis_power",
- "configFactory": "MI300XChassisConfig",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -849,146 +205,6 @@
"u_eth": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "AMD Instinct MI300X OAM",
- "gpu_component_key": "mi300x_oam_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 1536.0,
- "gpu_idle_example_w_per_gpu": 150.0,
- "gpu_decode_example_w_per_gpu": 380.0,
- "gpu_prefill_example_w_per_gpu": 620.0,
- "gpu_aggregate_example_w_per_gpu": 580.0,
- "gpu_peak_w_per_gpu": 750.0,
- "ubb": {
- "gpu_label": "AMD Instinct OAM",
- "gpu_component_key": "amd_oam_8x_measured",
- "n_gpu": 8,
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "management_controller_w": 10.0,
- "fpga_cpld_sequencing_w": 8.0,
- "hsc_power_monitor_w": 6.0,
- "clock_refclk_reset_w": 5.0,
- "sensors_i2c_fru_led_w": 4.0,
- "aux_rails_misc_w": 12.0,
- "normal_low_w": 30.0,
- "normal_high_w": 70.0,
- "conservative_cap_w": 90.0
- }
- },
- "cpu": {
- "name": "2x AMD EPYC 9654 (Genoa, MI300X host baseline)",
- "n_cpu": 2,
- "cores_per_cpu": 96,
- "threads_per_cpu": 192,
- "package_tdp_w": 360.0,
- "base_ghz": 2.4,
- "max_turbo_ghz": 3.7,
- "l3_cache_mb": 384.0,
- "memory_channels_per_cpu": 12,
- "pcie_lanes_per_cpu": 128,
- "package_static_w": 26.0,
- "uncore_io_baseline_w": 42.0,
- "memory_controller_baseline_w": 30.0,
- "uncore_dynamic_max_w": 12.0,
- "memory_controller_dynamic_max_w": 12.0,
- "core_curve_r": 1.65,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "MI300X host, 24x96GB DDR5-4800 RDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 12,
- "n_dimm": 24,
- "capacity_gb_per_dimm": 96.0,
- "data_rate_mtps": 4800.0,
- "dimm_type": "RDIMM",
- "background_w": 2.7119999999999997,
- "refresh_w": 0.696,
- "termination_w": 1.3920000000000001,
- "io_dynamic_max_w": 2.79,
- "core_dynamic_max_w": 3.41,
- "activity_exponent": 1.0
- },
- "thor2": {
- "n_nic": 8,
- "ports_per_nic": 2,
- "aggregate_gbps_per_nic": 400.0,
- "pcie_generation": 5,
- "pcie_lanes": 16,
- "card_idle_w": 12.5,
- "card_traffic_dynamic_w": 0.4,
- "include_optics": true,
- "optical_modules_per_nic": 2,
- "optics_total_w_per_nic": 10.9,
- "deployment_margin_w": 6.5,
- "deployment_traffic_dynamic_w": 2.8
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 12,
- "front_u2_idle_w": 4.5,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 12.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 12.0,
- "cpld_tpm_superio_w": 6.0,
- "clock_sensor_fru_w": 5.0,
- "storage_backplane_idle_w": 12.0,
- "front_panel_usb_led_w": 3.0,
- "aux_margin_w": 8.0,
- "normal_low_w": 45.0,
- "normal_high_w": 90.0
- },
- "fans": {
- "platform_label": "MI300X 8U air-cooled chassis",
- "n_fan": 10,
- "electrical_nameplate_w": 1800.0,
- "min_pwm_frac": 0.22,
- "normal_max_pwm_frac": 0.82,
- "full_cooling_load_w": 8500.0,
- "fan_curve_exponent": 1.08
- },
- "psu": {
- "platform_label": "MI300X 8U chassis PSU bank",
- "n_installed_psu": 6,
- "n_load_sharing_psu": 6,
- "n_redundant_capacity_psu": 3,
- "psu_capacity_w": 3000.0,
- "system_max_w": 9000.0,
- "redundancy": "N+N / 3+3",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"pcie5_x16_retimers_8x": 92.4,
"amd_oam_ubb_residual_static": 45.0,
@@ -1020,9 +236,7 @@
}
},
"mi325x": {
- "modelPath": "human_verified/mi325x_chassis/mi325x_chassis_power_model.py",
- "functionName": "mi325x_chassis_power",
- "configFactory": "MI325XChassisConfig",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -1033,146 +247,6 @@
"u_eth": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "AMD Instinct MI325X OAM",
- "gpu_component_key": "mi325x_oam_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 2048.0,
- "gpu_idle_example_w_per_gpu": 170.0,
- "gpu_decode_example_w_per_gpu": 450.0,
- "gpu_prefill_example_w_per_gpu": 800.0,
- "gpu_aggregate_example_w_per_gpu": 760.0,
- "gpu_peak_w_per_gpu": 1000.0,
- "ubb": {
- "gpu_label": "AMD Instinct OAM",
- "gpu_component_key": "amd_oam_8x_measured",
- "n_gpu": 8,
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "management_controller_w": 10.0,
- "fpga_cpld_sequencing_w": 8.0,
- "hsc_power_monitor_w": 6.0,
- "clock_refclk_reset_w": 5.0,
- "sensors_i2c_fru_led_w": 4.0,
- "aux_rails_misc_w": 12.0,
- "normal_low_w": 30.0,
- "normal_high_w": 70.0,
- "conservative_cap_w": 90.0
- }
- },
- "cpu": {
- "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)",
- "n_cpu": 2,
- "cores_per_cpu": 64,
- "threads_per_cpu": 128,
- "package_tdp_w": 400.0,
- "base_ghz": 3.3,
- "max_turbo_ghz": 4.3,
- "l3_cache_mb": 384.0,
- "memory_channels_per_cpu": 12,
- "pcie_lanes_per_cpu": 160,
- "package_static_w": 28.0,
- "uncore_io_baseline_w": 48.0,
- "memory_controller_baseline_w": 34.0,
- "uncore_dynamic_max_w": 14.0,
- "memory_controller_dynamic_max_w": 14.0,
- "core_curve_r": 1.62,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "MI325X host, 24x256GB DDR5-6400 RDIMM/MRDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 12,
- "n_dimm": 24,
- "capacity_gb_per_dimm": 256.0,
- "data_rate_mtps": 6400.0,
- "dimm_type": "RDIMM/MRDIMM",
- "background_w": 3.9549999999999996,
- "refresh_w": 1.015,
- "termination_w": 2.0300000000000002,
- "io_dynamic_max_w": 4.95,
- "core_dynamic_max_w": 6.05,
- "activity_exponent": 1.0
- },
- "thor2": {
- "n_nic": 8,
- "ports_per_nic": 2,
- "aggregate_gbps_per_nic": 400.0,
- "pcie_generation": 5,
- "pcie_lanes": 16,
- "card_idle_w": 12.5,
- "card_traffic_dynamic_w": 0.4,
- "include_optics": true,
- "optical_modules_per_nic": 2,
- "optics_total_w_per_nic": 10.9,
- "deployment_margin_w": 6.5,
- "deployment_traffic_dynamic_w": 2.8
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 8,
- "front_u2_idle_w": 4.5,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 12.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 14.0,
- "cpld_tpm_superio_w": 7.0,
- "clock_sensor_fru_w": 6.0,
- "storage_backplane_idle_w": 14.0,
- "front_panel_usb_led_w": 3.0,
- "aux_margin_w": 10.0,
- "normal_low_w": 50.0,
- "normal_high_w": 100.0
- },
- "fans": {
- "platform_label": "MI325X 8U air-cooled chassis",
- "n_fan": 14,
- "electrical_nameplate_w": 2600.0,
- "min_pwm_frac": 0.24,
- "normal_max_pwm_frac": 0.88,
- "full_cooling_load_w": 12000.0,
- "fan_curve_exponent": 1.06
- },
- "psu": {
- "platform_label": "MI325X 8U chassis PSU bank",
- "n_installed_psu": 6,
- "n_load_sharing_psu": 6,
- "n_redundant_capacity_psu": 3,
- "psu_capacity_w": 5250.0,
- "system_max_w": 15750.0,
- "redundancy": "N+N / 3+3",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"pcie5_x16_retimers_8x": 92.4,
"amd_oam_ubb_residual_static": 45.0,
@@ -1204,9 +278,7 @@
}
},
"mi355x": {
- "modelPath": "human_verified/mi355x_chassis/mi355x_chassis_power_model.py",
- "functionName": "mi355x_chassis_power",
- "configFactory": "MI355XChassisConfig",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"gpuCount": 8,
"assumptions": {
"u_pcie": 0.05,
@@ -1217,145 +289,6 @@
"u_eth": 0.0,
"fan_pwm": null
},
- "defaultConfig": {
- "gpu_label": "AMD Instinct MI355X OAM",
- "gpu_component_key": "mi355x_oam_8x_measured",
- "n_gpu": 8,
- "gpu_memory_total_gb": 2304.0,
- "gpu_idle_example_w_per_gpu": 250.0,
- "gpu_decode_example_w_per_gpu": 700.0,
- "gpu_prefill_example_w_per_gpu": 1150.0,
- "gpu_aggregate_example_w_per_gpu": 1080.0,
- "gpu_peak_w_per_gpu": 1400.0,
- "ubb": {
- "gpu_label": "AMD Instinct OAM",
- "gpu_component_key": "amd_oam_8x_measured",
- "n_gpu": 8,
- "retimer": {
- "n_retimer": 8,
- "lanes_per_retimer": 16,
- "pcie_gtps": 32.0,
- "analog_serdes_equalization_w": 9.5,
- "digital_floor_w": 2.0,
- "digital_variable_w": 1.0
- },
- "residual": {
- "management_controller_w": 10.0,
- "fpga_cpld_sequencing_w": 8.0,
- "hsc_power_monitor_w": 6.0,
- "clock_refclk_reset_w": 5.0,
- "sensors_i2c_fru_led_w": 4.0,
- "aux_rails_misc_w": 12.0,
- "normal_low_w": 30.0,
- "normal_high_w": 70.0,
- "conservative_cap_w": 90.0
- }
- },
- "cpu": {
- "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)",
- "n_cpu": 2,
- "cores_per_cpu": 64,
- "threads_per_cpu": 128,
- "package_tdp_w": 400.0,
- "base_ghz": 3.3,
- "max_turbo_ghz": 4.3,
- "l3_cache_mb": 384.0,
- "memory_channels_per_cpu": 12,
- "pcie_lanes_per_cpu": 160,
- "package_static_w": 28.0,
- "uncore_io_baseline_w": 48.0,
- "memory_controller_baseline_w": 34.0,
- "uncore_dynamic_max_w": 14.0,
- "memory_controller_dynamic_max_w": 14.0,
- "core_curve_r": 1.62,
- "vrm_efficiency": 0.92
- },
- "dram": {
- "label": "MI355X host, 24x256GB DDR5-6400 RDIMM/MRDIMM",
- "status": "DRAFT - pending human verification",
- "n_sockets": 2,
- "channels_per_socket": 12,
- "n_dimm": 24,
- "capacity_gb_per_dimm": 256.0,
- "data_rate_mtps": 6400.0,
- "dimm_type": "RDIMM/MRDIMM",
- "background_w": 3.9549999999999996,
- "refresh_w": 1.015,
- "termination_w": 2.0300000000000002,
- "io_dynamic_max_w": 4.95,
- "core_dynamic_max_w": 6.05,
- "activity_exponent": 1.0
- },
- "pollara400": {
- "n_nic": 8,
- "aggregate_gbps_per_nic": 400.0,
- "pcie_generation": 5,
- "pcie_lanes": 16,
- "net_serdes_w": 10.5,
- "pcie_serdes_w": 6.0,
- "board_mgmt_w": 2.5,
- "packet_engine_max_w": 12.5,
- "packet_engine_floor_frac": 0.64,
- "include_optic": true,
- "optic_w": 8.0
- },
- "pcie_switch": {
- "n_switch": 4,
- "lanes_per_switch": 144,
- "ports_per_switch": 72,
- "pcie_gtps": 32.0,
- "serdes_phy_floor_w": 30.5,
- "control_leakage_clock_w": 6.0,
- "fabric_datapath_floor_w": 4.0,
- "fabric_datapath_variable_w": 8.5,
- "stress_cap_per_switch_w": 70.0
- },
- "nvme": {
- "n_front_u2": 8,
- "front_u2_idle_w": 4.5,
- "n_boot_m2": 2,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "residual": {
- "bmc_ipmi_w": 12.0,
- "onboard_10gbe_w": 8.0,
- "motherboard_pch_aux_w": 16.0,
- "cpld_tpm_superio_w": 8.0,
- "clock_sensor_fru_w": 6.0,
- "storage_backplane_idle_w": 16.0,
- "front_panel_usb_led_w": 4.0,
- "aux_margin_w": 12.0,
- "normal_low_w": 60.0,
- "normal_high_w": 120.0
- },
- "fans": {
- "platform_label": "MI355X 10U air-cooled chassis",
- "n_fan": 19,
- "electrical_nameplate_w": 3600.0,
- "min_pwm_frac": 0.25,
- "normal_max_pwm_frac": 0.92,
- "full_cooling_load_w": 15500.0,
- "fan_curve_exponent": 1.04
- },
- "psu": {
- "platform_label": "MI355X 10U chassis PSU bank",
- "n_installed_psu": 6,
- "n_load_sharing_psu": 6,
- "n_redundant_capacity_psu": 4,
- "psu_capacity_w": 6600.0,
- "system_max_w": 26400.0,
- "redundancy": "4+2",
- "efficiency_curve": {
- "0.05": 0.885,
- "0.1": 0.92,
- "0.2": 0.94,
- "0.5": 0.96,
- "1.0": 0.955
- }
- },
- "pue": 1.2
- },
"fixedComponentsDcWatts": {
"pcie5_x16_retimers_8x": 92.4,
"amd_oam_ubb_residual_static": 45.0,
@@ -1389,9 +322,7 @@
},
"rackProfiles": {
"gb200": {
- "modelPath": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
- "functionName": "gb200_nvl72_rack_power",
- "configFactory": "gb200_nvl72_rack_config",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"topology": "nvl72-rack",
"gpuCount": 72,
"computeTrayCount": 18,
@@ -1404,92 +335,6 @@
"u_pcie": 0.05,
"pue": 1.2
},
- "defaultConfig": {
- "variant": "gb200",
- "gpu_label": "GB200 Blackwell (NVL72)",
- "n_compute_trays": 18,
- "n_nvswitch_trays": 9,
- "compute_tray": {
- "n_gpu": 4,
- "n_grace": 2,
- "nic_generation": "connectx7",
- "connectx7": {
- "n_nic": 4,
- "net_serdes_w": 9.5,
- "pcie_serdes_w": 5.5,
- "board_w": 2.0,
- "digital_max_w": 10.0,
- "digital_floor_frac": 0.65,
- "include_optic": true,
- "optic_w": 8.0
- },
- "connectx8": {
- "n_nic": 4,
- "include_optic": true,
- "network_serdes_static_w": 18.0,
- "pcie_switch_static_w": 28.0,
- "board_mgmt_static_w": 5.0,
- "digital_static_w": 12.0,
- "network_dynamic_max_w": 8.0,
- "pcie_switch_dynamic_max_w": 5.0,
- "digital_dynamic_max_w": 10.0,
- "optic_idle_w": 15.0,
- "optic_dynamic_max_w": 2.0,
- "max_nic_slot_power_ref_w": 75.0,
- "normal_low_w_per_nic_with_optic": 70.0,
- "normal_high_w_per_nic_with_optic": 100.0
- },
- "dpu": {
- "n_dpu": 2,
- "idle_w_per_dpu": 65.0,
- "public_max_power_cap_w": 150.0,
- "max_modeled_u_dpu": 0.02
- },
- "nvme": {
- "n_front_u2": 4,
- "front_u2_idle_w": 5.0,
- "n_boot_m2": 1,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "fans_w": 130.0,
- "board_residual_w": 40.0
- },
- "nvswitch": {
- "n_asic": 2,
- "serdes_lanes": 144,
- "serdes_lane_gbps": 200.0,
- "serdes_pj_per_bit": 2.5,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 190.0,
- "digital_floor_frac": 0.4
- },
- "nvswitch_tray_residual_w": 50.0,
- "tray_input_conversion_efficiency": 0.9725,
- "n_management_switches": 2,
- "management_switch_w": 100.0,
- "power_shelf": {
- "n_shelves": 8,
- "n_psu_per_shelf": 6,
- "psu_capacity_w": 5500.0,
- "redundancy": "N+N",
- "busbar_nominal_v": 50.0,
- "efficiency_curve": {
- "0.1": 0.9,
- "0.2": 0.94,
- "0.3": 0.965,
- "1.0": 0.965
- },
- "peak_efficiency_ref": 0.975
- },
- "regulator_loss_frac_of_tdp": 0.15,
- "regulator_allowance_includes_grace": false,
- "pue": 1.2,
- "cooling": "direct_liquid",
- "gpu_tdp_example_w_per_gpu": 1200.0,
- "grace_tdp_example_w_per_socket": 300.0,
- "module_tdp_example_w_per_superchip": 2700.0
- },
"computeTrayStaticDcWatts": {
"connectx7_nics_with_optics": 126.0,
"bluefield3_dpu_idle": 130.0,
@@ -1554,9 +399,7 @@
}
},
"gb300": {
- "modelPath": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
- "functionName": "gb200_nvl72_rack_power",
- "configFactory": "gb300_nvl72_rack_config",
+ "modelPath": "packages/app/src/lib/system-power-model.ts",
"topology": "nvl72-rack",
"gpuCount": 72,
"computeTrayCount": 18,
@@ -1569,92 +412,6 @@
"u_pcie": 0.05,
"pue": 1.2
},
- "defaultConfig": {
- "variant": "gb300",
- "gpu_label": "GB300 Blackwell Ultra (NVL72)",
- "n_compute_trays": 18,
- "n_nvswitch_trays": 9,
- "compute_tray": {
- "n_gpu": 4,
- "n_grace": 2,
- "nic_generation": "connectx8",
- "connectx7": {
- "n_nic": 4,
- "net_serdes_w": 9.5,
- "pcie_serdes_w": 5.5,
- "board_w": 2.0,
- "digital_max_w": 10.0,
- "digital_floor_frac": 0.65,
- "include_optic": true,
- "optic_w": 8.0
- },
- "connectx8": {
- "n_nic": 4,
- "include_optic": true,
- "network_serdes_static_w": 18.0,
- "pcie_switch_static_w": 28.0,
- "board_mgmt_static_w": 5.0,
- "digital_static_w": 12.0,
- "network_dynamic_max_w": 8.0,
- "pcie_switch_dynamic_max_w": 5.0,
- "digital_dynamic_max_w": 10.0,
- "optic_idle_w": 15.0,
- "optic_dynamic_max_w": 2.0,
- "max_nic_slot_power_ref_w": 75.0,
- "normal_low_w_per_nic_with_optic": 70.0,
- "normal_high_w_per_nic_with_optic": 100.0
- },
- "dpu": {
- "n_dpu": 2,
- "idle_w_per_dpu": 65.0,
- "public_max_power_cap_w": 150.0,
- "max_modeled_u_dpu": 0.02
- },
- "nvme": {
- "n_front_u2": 4,
- "front_u2_idle_w": 5.0,
- "n_boot_m2": 1,
- "boot_m2_idle_w": 2.0,
- "max_modeled_u_nvme": 0.02
- },
- "fans_w": 130.0,
- "board_residual_w": 40.0
- },
- "nvswitch": {
- "n_asic": 2,
- "serdes_lanes": 144,
- "serdes_lane_gbps": 200.0,
- "serdes_pj_per_bit": 2.5,
- "serdes_floor_frac": 0.95,
- "digital_max_w": 190.0,
- "digital_floor_frac": 0.4
- },
- "nvswitch_tray_residual_w": 50.0,
- "tray_input_conversion_efficiency": 0.9725,
- "n_management_switches": 2,
- "management_switch_w": 100.0,
- "power_shelf": {
- "n_shelves": 8,
- "n_psu_per_shelf": 6,
- "psu_capacity_w": 5500.0,
- "redundancy": "N+N",
- "busbar_nominal_v": 50.0,
- "efficiency_curve": {
- "0.1": 0.9,
- "0.2": 0.94,
- "0.3": 0.965,
- "1.0": 0.965
- },
- "peak_efficiency_ref": 0.975
- },
- "regulator_loss_frac_of_tdp": 0.15,
- "regulator_allowance_includes_grace": false,
- "pue": 1.2,
- "cooling": "direct_liquid",
- "gpu_tdp_example_w_per_gpu": 1400.0,
- "grace_tdp_example_w_per_socket": 300.0,
- "module_tdp_example_w_per_superchip": 0.0
- },
"computeTrayStaticDcWatts": {
"connectx8_nics_integrated_pcie_with_optics": 315.0,
"bluefield3_dpu_idle": 130.0,
diff --git a/packages/app/src/lib/system-power-model.provenance.json b/packages/app/src/lib/system-power-model.provenance.json
new file mode 100644
index 000000000..2ffb30b26
--- /dev/null
+++ b/packages/app/src/lib/system-power-model.provenance.json
@@ -0,0 +1,13 @@
+{
+ "modelRevision": "app-sha256:99b252a0422c01934973760dbf5779c10aa87cc76e89807d1643258ef586e58e",
+ "modelRevisionStatus": "App-owned TypeScript equations, parameters and admission/PUE policy",
+ "source": "https://github.com/SemiAnalysisAI/InferenceX-app",
+ "status": "DRAFT / pending human verification",
+ "assumptionsSource": "packages/app/src/lib/system-power-model.profiles.json#/assumptions",
+ "rackAssumptionsSource": "packages/app/src/lib/system-power-model.profiles.json#/rackAssumptions",
+ "sourceSha256": {
+ "packages/app/src/lib/modeled-system-power.ts": "96c6016a5afabb860e01420038944422e3418befe347dc08d3602107878147fe",
+ "packages/app/src/lib/system-power-model.profiles.json": "9aaf0b8dc6365a86f099504e2fe7bbd612c3aef5606752059cd6062173aa4e20",
+ "packages/app/src/lib/system-power-model.ts": "9f0ff5dd1d1e85a6a09e1e20e9d9f682ece5f4f8eae7e5fe813e799329e966c4"
+ }
+}
diff --git a/packages/app/src/lib/system-power-model.reference.json b/packages/app/src/lib/system-power-model.reference.json
index a2356ce4f..d081154cc 100644
--- a/packages/app/src/lib/system-power-model.reference.json
+++ b/packages/app/src/lib/system-power-model.reference.json
@@ -1,8 +1,1278 @@
{
"modelRevision": "6fcc086b77576d4cecb9d0c79637d6daf980308c",
- "modelRevisionStatus": "local branch feat/gb200-nvl72-rack-model; unpublished pending repository write access; DRAFT / pending human verification",
+ "modelRevisionStatus": "Historical local Python snapshot; unpublished at capture; retained only as migration-baseline provenance.",
"source": "https://github.com/SemiAnalysisAI/inferencex_power_model",
"status": "DRAFT / pending human verification",
+ "fixtureRole": "Frozen historical Python migration baseline; independent expected values, not the current app model version.",
+ "generatorSha256": "c33e50aac2b1ce6360df1895bd9f9c5496a3f783313c0243563c6d3c37f8a5bc",
+ "sourceSha256": {
+ "human_verified/amd_oam_fans/amd_oam_fan_power_model.py": "7ea6f1c65b685311277e6f2c53992d588ec79b598180f2fca300e904ad6dc19a",
+ "human_verified/amd_oam_fans/plot_amd_oam_fan_power.py": "e38118fa4d94412af268eaa1667d2c4086d1dcced4da07c951e761d54ea55b46",
+ "human_verified/amd_oam_psu/amd_oam_psu_power_model.py": "badcf849a82da35bc9bd621d42cd955ec7f8e3879dbc1bc29c7c03a3e0217af1",
+ "human_verified/amd_oam_psu/plot_amd_oam_psu_power.py": "39a1c3eac5a1e6337425948bfd252e422021b67de2ec5c963be1f8db3a8292fc",
+ "human_verified/amd_oam_ubb/amd_oam_ubb_power_model.py": "2f2f685e595d980c3fc353c624ca7183ee4b7f34c52d649066ef266664bccab4",
+ "human_verified/b200_fans/b200_fan_power_model.py": "163c6dd95f460cf6bde1b60cec061cce9be11e69cac1c6fe000b01a7ace81975",
+ "human_verified/b200_fans/plot_b200_fan_power.py": "8172b706c4a882ebb1fed220b3bcd602ea9db49019c4b1ac5a862d83be07bedd",
+ "human_verified/b200_psu/b200_psu_power_model.py": "58e3f81c7fbf722848184ddd73ed94edb065cfa96d6382d74bf42e8f7e15bbe8",
+ "human_verified/b200_psu/plot_b200_psu_power.py": "c96127308dab7105243ca21dcef07dce4fad8f674d9273006e3f4b1c29656b10",
+ "human_verified/b200_ubb/b200_ubb_power_model.py": "e3d3acf16a5ef9ca318ba49dc8b58d40823f8b76144d2999dffc9973db96648a",
+ "human_verified/b200_ubb/b200_ubb_residual_power_model.py": "85ca624b8d85d504f4e8a245f05aac07ab804b3fb262c2f00eb50ed35e22d273",
+ "human_verified/b200_ubb/plot_b200_ubb_residual_power.py": "a0d7a7e5f44948c59262bf62f4f088cb64fb17c29926f7b3327a18fb695903ed",
+ "human_verified/b300_fans/b300_fan_power_model.py": "9c5f85d0b7c9d4ee9fb8d233bb3d150a69643902970a68b475bb266918a495c7",
+ "human_verified/b300_fans/plot_b300_fan_power.py": "761a8c1c80f8efa0fe49bb749897da20764c1a207efe4d77c22a43b001b319c4",
+ "human_verified/b300_psu/b300_psu_power_model.py": "930f30f3a720b964623b35a7ee2d510a03f39ec6751dda5b480c7a8cfb0bd99e",
+ "human_verified/b300_psu/plot_b300_psu_power.py": "241880ab01f8d4012a6d2b358f85648bebb229b4966b1af081ecd331997cefc8",
+ "human_verified/b300_ubb/b300_ubb_power_model.py": "8cd694f1306ca09de53dcc9590dcc47a1b9bd5c10af9837bb875dc417fe74837",
+ "human_verified/blackwell_nvswitch/blackwell_nvswitch_power_model.py": "857d276b552f6118842c0f026cb5e78dbcd70c9bdef2769fae8540842be61ec7",
+ "human_verified/blackwell_nvswitch/plot_blackwell_nvswitch_power.py": "979abfc057320d29d251c100150440fdc10770794b1a05a89053631ffdff64db",
+ "human_verified/chassis_plot_utils.py": "0f1c809a527788c8c25490420e29e737e6b7f83cb914dc022c3a4e3701b80643",
+ "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py": "b4640f94c8f9e6be50eb18deff9728bb58559ba193b54577859cafabec4ccc33",
+ "human_verified/gb200_nvl72_rack/plot_gb200_nvl72_rack_inference_gpu_sweep.py": "d0f6903df0348ccb2790e973b0dc10408c6f6022695ce039aad0b1e1fcd39519",
+ "human_verified/gb200_nvl72_rack/test_gb200_nvl72_rack_power_model.py": "631bd7fd44b84260f16175a814d093795a4cdd0a91a0d19dd572a5ea6c088314",
+ "human_verified/generic/connectx7/connectx7_power_model.py": "3510ab679846aefbc31d18d410a978533ddc331fc1562511751fedf093a6649f",
+ "human_verified/generic/connectx7/plot_connectx7_power.py": "c9d554aeeeb5b74ff7398686c05d93f0db42b6ac98c902b063645597697de35c",
+ "human_verified/generic/connectx8/connectx8_power_model.py": "7d638ea8524e181b0370601319c780600ff5a45b072589d58bdca58636bfa9cb",
+ "human_verified/generic/connectx8/plot_connectx8_power.py": "97921625373a479f03ad3930c8542e86da6c4ee4521654f77bcb965d93759433",
+ "human_verified/generic/cpu/cpu_power_model.py": "ab4315df415d70474ae4fb5c700fdf5b6de2a8a89f33823fbc8664efa01bf72f",
+ "human_verified/generic/cpu/plot_cpu_power.py": "bc5866234e7af605c8b43664cca1c1e96116d4fefb88c791cbe5e35f6234e821",
+ "human_verified/generic/dpu/dpu_power_model.py": "677cce41b990213c96f73e016c8bf24da9841ef995a7522e1038c832d502e398",
+ "human_verified/generic/dpu/plot_dpu_power.py": "d52357acdabb6a69b2ac87037cc339208e70b1bf7bdd31483b98dbedad06f05b",
+ "human_verified/generic/dram/dram_power_model.py": "b8ace82de182715a003c747710fc62275f2392d883aadb57c05f19beb1367b67",
+ "human_verified/generic/dram/plot_dram_power.py": "8a3d59fbf31492f9de04d8f3f80ddc33719018b6f47e8b0e671d3f63e29acead",
+ "human_verified/generic/nvme/nvme_power_model.py": "53fa1f59edf4427a62c946758d4f3df55b1591f8d43f824e07fd61e6d4196743",
+ "human_verified/generic/nvme/plot_nvme_power.py": "276d013eb26ff13c68e60cbf48bddc88fd6f5ff6c0a59a9854aec11da6ebf8b7",
+ "human_verified/generic/pcie_switches/pcie5_144lane_switch_power_model.py": "23d23322126c4251cb2ad5bee4448b0f9ad71b531607e8353cd96a90410cc28d",
+ "human_verified/generic/pcie_switches/plot_pcie_switch_power.py": "3cf3139b18495efc320c1f3d2754d8f732b16f1775959f66ff222ef97b64de90",
+ "human_verified/generic/pollara400/plot_pollara400_power.py": "22196c3037330b07903109e0f9a6917fd2e9e8e56ee942df552a84efd06a470e",
+ "human_verified/generic/pollara400/pollara400_power_model.py": "8dd5e674d9bfa5e19cb68c5684eb717df61063762afa555cd1a9c43f1d24c5d8",
+ "human_verified/generic/retimers/pcie5_x16_retimer_power_model.py": "393381bb40ef12cb81b681eff742129453136c6ba6c4b6950682e6cdc63a5335",
+ "human_verified/generic/retimers/plot_pcie5_x16_retimer_power.py": "c8f3a9aede8876cebab9f0674926b59f591f71ec1e58e86bdebe3ec63f61a208",
+ "human_verified/generic/thor2/plot_thor2_power.py": "7f03cd9af3dfd4e787e0e9d429a5758e24f4b62990af6d7d7e22a1a5224a78c7",
+ "human_verified/generic/thor2/thor2_power_model.py": "67b8e91a01dccd3abd7bdd37dbef86d1194d398c9326b8774eede1488962c54c",
+ "human_verified/hgx_b200_chassis/b200_chassis_power_model.py": "89d94969ce1acfeee67784ad269c431415995f9b18996d4b28d213d34393864a",
+ "human_verified/hgx_b200_chassis/plot_chassis_inference_gpu_sweep.py": "ed9378b471bf502adda4ac8f2467d803e48a5347ae329e118bfeebc9400ecc79",
+ "human_verified/hgx_b200_chassis_residual/b200_chassis_residual_power_model.py": "d23fe72c6039545e51d6271ddef28bee5d69f9ee7e4f8032a5def60021f973eb",
+ "human_verified/hgx_b200_chassis_residual/plot_chassis_residual_power.py": "a46d81768e0af6982bdb9d146b7e42fdbb86f1afcdb581d808657b5508a65414",
+ "human_verified/hgx_b300_chassis/b300_chassis_power_model.py": "68af8917ead2472cb0f6473784a5a75a24233cb766df8fa2684421a07d660fbc",
+ "human_verified/hgx_b300_chassis/plot_b300_chassis_inference_gpu_sweep.py": "bfe74771dd61b6dbb3dcf1ebbf6451c07b6020dcf0000482225672eb081f1d11",
+ "human_verified/hgx_h100_chassis/h100_chassis_power_model.py": "6850ec92346af1864f724a41d9ea512e0d55f45d683a3d575477aa08ca89a6c8",
+ "human_verified/hgx_h100_chassis/plot_h100_chassis_inference_gpu_sweep.py": "4af0da4e056f2300d271dd041ffb2ff9a76d39aaf80ae227a38ce777b949d852",
+ "human_verified/hgx_h200_chassis/h200_chassis_power_model.py": "56b40c9f13e50d81a02a594f0762f5c498483f261f9aec7b7294b148c65d7eb9",
+ "human_verified/hgx_h200_chassis/plot_h200_chassis_inference_gpu_sweep.py": "326b3baff711b5a322349cb272f83eaa44c8825cdbb9cc87cce52d7433e4b695",
+ "human_verified/hopper_fans/hopper_fan_power_model.py": "8b656f8498b6709b8ac7392999333eef06a529e12af141cc4565a2272c589f24",
+ "human_verified/hopper_fans/plot_hopper_fan_power.py": "006c13f8143ad0ebc853ecba9713660f65daa9ee520ce0b55763ce06a4a436fa",
+ "human_verified/hopper_nvswitch/hopper_nvswitch_power_model.py": "3d3c536bc2af1e75f2cc3c246e0d5337c04809cb90d470227a0d7336b413790c",
+ "human_verified/hopper_nvswitch/plot_hopper_nvswitch_power.py": "c13b291df6426a2117d82f1425009ddf0f55340888a0843dd5f5888503c62e06",
+ "human_verified/hopper_psu/hopper_psu_power_model.py": "a7630454e128e87bc0529a02cebb11138efbe1b07d4e4cb08a9721cc49f8d58d",
+ "human_verified/hopper_psu/plot_hopper_psu_power.py": "4a73fffa633e0a599a8bf746e0868f719669ff59d69a102bf563fcc4596e40a4",
+ "human_verified/hopper_ubb/hopper_ubb_power_model.py": "3f8c6c9560c32dbe1e98d0af82dbafa3fc697efc3d32642afe303455bfd5ab74",
+ "human_verified/mi300x_chassis/mi300x_chassis_power_model.py": "69c4b11e860e9a174664ae040691aab9e349410040ac8524dee6a7f2102afab6",
+ "human_verified/mi300x_chassis/plot_mi300x_chassis_inference_gpu_sweep.py": "1793a7356b95821cd4f7390a4cae55c58ffcc37f1b8f43f873499d84937463e4",
+ "human_verified/mi325x_chassis/mi325x_chassis_power_model.py": "59b6ce4ff626f1c5f42e8d7b0d33318e0e493a3e358dc1f4f92011be94b16c9c",
+ "human_verified/mi325x_chassis/plot_mi325x_chassis_inference_gpu_sweep.py": "c0f350bc4978a18759dd108756290fcd803f30209ccfc2682988bbf7ff0feb4d",
+ "human_verified/mi355x_chassis/mi355x_chassis_power_model.py": "c178f71efe53b424f5a1804fd188a79f99575e821a357b9128added48f154c8d",
+ "human_verified/mi355x_chassis/plot_mi355x_chassis_inference_gpu_sweep.py": "5071db35ce4bb4d7754419877f1e0d9df0c48be4381316b77d80fa1e58fe46c6"
+ },
+ "modelPaths": {
+ "h100": "human_verified/hgx_h100_chassis/h100_chassis_power_model.py",
+ "h200": "human_verified/hgx_h200_chassis/h200_chassis_power_model.py",
+ "b200": "human_verified/hgx_b200_chassis/b200_chassis_power_model.py",
+ "b300": "human_verified/hgx_b300_chassis/b300_chassis_power_model.py",
+ "mi300x": "human_verified/mi300x_chassis/mi300x_chassis_power_model.py",
+ "mi325x": "human_verified/mi325x_chassis/mi325x_chassis_power_model.py",
+ "mi355x": "human_verified/mi355x_chassis/mi355x_chassis_power_model.py",
+ "gb200": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py",
+ "gb300": "human_verified/gb200_nvl72_rack/gb200_nvl72_rack_power_model.py"
+ },
+ "pythonConfigurations": {
+ "h100": {
+ "functionName": "h100_chassis_power",
+ "configFactory": "make_h100_config",
+ "defaultConfig": {
+ "gpu_label": "H100 SXM5",
+ "gpu_component_key": "h100_sxm_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 640.0,
+ "gpu_idle_example_w_per_gpu": 100.0,
+ "gpu_decode_example_w_per_gpu": 300.0,
+ "gpu_prefill_example_w_per_gpu": 520.0,
+ "gpu_aggregate_example_w_per_gpu": 480.0,
+ "gpu_peak_w_per_gpu": 700.0,
+ "ubb": {
+ "gpu_label": "H100 SXM5",
+ "gpu_component_key": "h100_sxm_8x_measured",
+ "n_gpu": 8,
+ "nvswitch": {
+ "n_asic": 4,
+ "nvlink_ports_per_asic": 64,
+ "phy_lanes_per_asic": 128,
+ "phy_lane_gbps": 100.0,
+ "serdes_class_gbps": 112.0,
+ "serdes_pj_per_bit": 3.0,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 120.0,
+ "digital_floor_frac": 0.4
+ },
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "baseboard_controller_w": 10.0,
+ "fpga_cpld_sequencing_w": 8.0,
+ "hsc_power_monitor_w": 5.0,
+ "clock_refclk_reset_w": 5.0,
+ "sensors_i2c_fru_led_w": 4.0,
+ "aux_rails_misc_w": 13.0,
+ "normal_low_w": 30.0,
+ "normal_high_w": 70.0,
+ "conservative_cap_w": 90.0
+ }
+ },
+ "cpu": {
+ "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)",
+ "n_cpu": 2,
+ "cores_per_cpu": 56,
+ "threads_per_cpu": 112,
+ "package_tdp_w": 350.0,
+ "base_ghz": 2.0,
+ "max_turbo_ghz": 3.8,
+ "l3_cache_mb": 105.0,
+ "memory_channels_per_cpu": 8,
+ "pcie_lanes_per_cpu": 80,
+ "package_static_w": 24.0,
+ "uncore_io_baseline_w": 36.0,
+ "memory_controller_baseline_w": 22.0,
+ "uncore_dynamic_max_w": 8.0,
+ "memory_controller_dynamic_max_w": 6.0,
+ "core_curve_r": 1.72,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 8,
+ "n_dimm": 32,
+ "capacity_gb_per_dimm": 64.0,
+ "data_rate_mtps": 4800.0,
+ "dimm_type": "RDIMM",
+ "background_w": 2.147,
+ "refresh_w": 0.5509999999999999,
+ "termination_w": 1.102,
+ "io_dynamic_max_w": 2.3400000000000003,
+ "core_dynamic_max_w": 2.86,
+ "activity_exponent": 1.0
+ },
+ "connectx_compute": {
+ "n_nic": 8,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 8,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 10.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 8.0,
+ "cpld_tpm_superio_w": 5.0,
+ "clock_sensor_fru_w": 4.0,
+ "storage_backplane_idle_w": 6.0,
+ "front_panel_usb_led_w": 2.0,
+ "aux_margin_w": 2.0,
+ "normal_low_w": 35.0,
+ "normal_high_w": 70.0
+ },
+ "fans": {
+ "electrical_nameplate_w": 1100.0,
+ "airflow_cfm_at_normal_max_pwm": 1105.0,
+ "min_pwm_frac": 0.22,
+ "normal_max_pwm_frac": 0.8,
+ "full_cooling_load_w": 9500.0,
+ "fan_curve_exponent": 1.1
+ },
+ "psu": {
+ "n_installed_psu": 6,
+ "n_load_sharing_psu": 6,
+ "n_redundant_capacity_psu": 4,
+ "psu_capacity_w": 3300.0,
+ "redundancy": "4+2",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "storage_mgmt_network_static_w": 60.0,
+ "optional_dpu_idle_w": 0.0,
+ "pue": 1.2
+ }
+ },
+ "h200": {
+ "functionName": "h200_chassis_power",
+ "configFactory": "make_h200_config",
+ "defaultConfig": {
+ "gpu_label": "H200 SXM5",
+ "gpu_component_key": "h200_sxm_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 1128.0,
+ "gpu_idle_example_w_per_gpu": 115.0,
+ "gpu_decode_example_w_per_gpu": 330.0,
+ "gpu_prefill_example_w_per_gpu": 540.0,
+ "gpu_aggregate_example_w_per_gpu": 510.0,
+ "gpu_peak_w_per_gpu": 700.0,
+ "ubb": {
+ "gpu_label": "H200 SXM5",
+ "gpu_component_key": "h200_sxm_8x_measured",
+ "n_gpu": 8,
+ "nvswitch": {
+ "n_asic": 4,
+ "nvlink_ports_per_asic": 64,
+ "phy_lanes_per_asic": 128,
+ "phy_lane_gbps": 100.0,
+ "serdes_class_gbps": 112.0,
+ "serdes_pj_per_bit": 3.0,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 120.0,
+ "digital_floor_frac": 0.4
+ },
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "baseboard_controller_w": 10.0,
+ "fpga_cpld_sequencing_w": 8.0,
+ "hsc_power_monitor_w": 5.0,
+ "clock_refclk_reset_w": 5.0,
+ "sensors_i2c_fru_led_w": 4.0,
+ "aux_rails_misc_w": 13.0,
+ "normal_low_w": 30.0,
+ "normal_high_w": 70.0,
+ "conservative_cap_w": 90.0
+ }
+ },
+ "cpu": {
+ "name": "2x Intel Xeon Platinum 8480C (Sapphire Rapids, DGX H100/H200)",
+ "n_cpu": 2,
+ "cores_per_cpu": 56,
+ "threads_per_cpu": 112,
+ "package_tdp_w": 350.0,
+ "base_ghz": 2.0,
+ "max_turbo_ghz": 3.8,
+ "l3_cache_mb": 105.0,
+ "memory_channels_per_cpu": 8,
+ "pcie_lanes_per_cpu": 80,
+ "package_static_w": 24.0,
+ "uncore_io_baseline_w": 36.0,
+ "memory_controller_baseline_w": 22.0,
+ "uncore_dynamic_max_w": 8.0,
+ "memory_controller_dynamic_max_w": 6.0,
+ "core_curve_r": 1.72,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "DGX-H100/H200, 32x64GB DDR5 RDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 8,
+ "n_dimm": 32,
+ "capacity_gb_per_dimm": 64.0,
+ "data_rate_mtps": 4800.0,
+ "dimm_type": "RDIMM",
+ "background_w": 2.147,
+ "refresh_w": 0.5509999999999999,
+ "termination_w": 1.102,
+ "io_dynamic_max_w": 2.3400000000000003,
+ "core_dynamic_max_w": 2.86,
+ "activity_exponent": 1.0
+ },
+ "connectx_compute": {
+ "n_nic": 8,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 8,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 10.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 8.0,
+ "cpld_tpm_superio_w": 5.0,
+ "clock_sensor_fru_w": 4.0,
+ "storage_backplane_idle_w": 6.0,
+ "front_panel_usb_led_w": 2.0,
+ "aux_margin_w": 2.0,
+ "normal_low_w": 35.0,
+ "normal_high_w": 70.0
+ },
+ "fans": {
+ "electrical_nameplate_w": 1100.0,
+ "airflow_cfm_at_normal_max_pwm": 1105.0,
+ "min_pwm_frac": 0.22,
+ "normal_max_pwm_frac": 0.8,
+ "full_cooling_load_w": 9500.0,
+ "fan_curve_exponent": 1.1
+ },
+ "psu": {
+ "n_installed_psu": 6,
+ "n_load_sharing_psu": 6,
+ "n_redundant_capacity_psu": 4,
+ "psu_capacity_w": 3300.0,
+ "redundancy": "4+2",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "storage_mgmt_network_static_w": 60.0,
+ "optional_dpu_idle_w": 0.0,
+ "pue": 1.2
+ }
+ },
+ "b200": {
+ "functionName": "b200_chassis_power",
+ "configFactory": "B200ChassisMasterConfig",
+ "defaultConfig": {
+ "ubb": {
+ "n_gpu": 8,
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "include_management_bridge_controller": true,
+ "management_bridge_controller_w": 24.0,
+ "hmc_bmc_w": 6.0,
+ "fpga_cpld_w": 8.0,
+ "erot_security_w": 3.0,
+ "hsc_power_monitor_w": 5.0,
+ "clock_refclk_reset_w": 4.0,
+ "sensors_i2c_fru_led_w": 3.0,
+ "aux_rails_misc_w": 9.0,
+ "normal_low_w": 45.0,
+ "normal_high_w": 85.0,
+ "conservative_cap_w": 100.0
+ }
+ },
+ "cpu": {
+ "name": "2x Intel Xeon Platinum 8570 (Emerald Rapids)",
+ "n_cpu": 2,
+ "cores_per_cpu": 56,
+ "threads_per_cpu": 112,
+ "package_tdp_w": 350.0,
+ "base_ghz": 2.1,
+ "max_turbo_ghz": 4.0,
+ "l3_cache_mb": 300.0,
+ "memory_channels_per_cpu": 8,
+ "pcie_lanes_per_cpu": 80,
+ "package_static_w": 22.0,
+ "uncore_io_baseline_w": 36.0,
+ "memory_controller_baseline_w": 22.0,
+ "uncore_dynamic_max_w": 8.0,
+ "memory_controller_dynamic_max_w": 6.0,
+ "core_curve_r": 1.7,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "DGX-B200-like Xeon 8570, 32x64GB DDR5-5600 RDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 8,
+ "n_dimm": 32,
+ "capacity_gb_per_dimm": 64.0,
+ "data_rate_mtps": 5600.0,
+ "dimm_type": "RDIMM",
+ "background_w": 2.147,
+ "refresh_w": 0.5509999999999999,
+ "termination_w": 1.102,
+ "io_dynamic_max_w": 2.3400000000000003,
+ "core_dynamic_max_w": 2.86,
+ "activity_exponent": 1.0
+ },
+ "connectx7": {
+ "n_nic": 8,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "dpu": {
+ "n_dpu": 1,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 10,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 10.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 8.0,
+ "cpld_tpm_superio_w": 5.0,
+ "clock_sensor_fru_w": 4.0,
+ "storage_backplane_idle_w": 6.0,
+ "front_panel_usb_led_w": 2.0,
+ "aux_margin_w": 2.0,
+ "normal_low_w": 35.0,
+ "normal_high_w": 70.0
+ },
+ "fans": {
+ "n_80mm": 15,
+ "rated_80mm_w": 120.0,
+ "n_60mm": 4,
+ "rated_60mm_w": 25.0,
+ "min_pwm_frac": 0.25,
+ "normal_max_pwm_frac": 0.78,
+ "full_cooling_load_w": 12000.0,
+ "fan_curve_exponent": 1.15
+ },
+ "psu": {
+ "n_installed_psu": 6,
+ "n_active_psu": 3,
+ "psu_capacity_w": 5250.0,
+ "redundancy": "3+3",
+ "efficiency_curve": {
+ "0.05": 0.8864,
+ "0.1": 0.9238,
+ "0.2": 0.9448,
+ "0.5": 0.964,
+ "1.0": 0.9556
+ }
+ },
+ "pue": 1.2
+ }
+ },
+ "b300": {
+ "functionName": "b300_chassis_power",
+ "configFactory": "B300ChassisConfig",
+ "defaultConfig": {
+ "gpu_label": "B300 Blackwell Ultra SXM",
+ "gpu_component_key": "b300_sxm_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 2304.0,
+ "gpu_idle_example_w_per_gpu": 200.0,
+ "gpu_decode_example_w_per_gpu": 520.0,
+ "gpu_prefill_example_w_per_gpu": 800.0,
+ "gpu_aggregate_example_w_per_gpu": 760.0,
+ "gpu_peak_w_per_gpu": 1100.0,
+ "ubb": {
+ "n_gpu": 8,
+ "gpu_label": "B300 Blackwell Ultra SXM",
+ "gpu_component_key": "b300_sxm_8x_measured",
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "connectx8": {
+ "n_nic": 8,
+ "include_optic": true,
+ "network_serdes_static_w": 18.0,
+ "pcie_switch_static_w": 28.0,
+ "board_mgmt_static_w": 5.0,
+ "digital_static_w": 12.0,
+ "network_dynamic_max_w": 8.0,
+ "pcie_switch_dynamic_max_w": 5.0,
+ "digital_dynamic_max_w": 10.0,
+ "optic_idle_w": 15.0,
+ "optic_dynamic_max_w": 2.0,
+ "max_nic_slot_power_ref_w": 75.0,
+ "normal_low_w_per_nic_with_optic": 70.0,
+ "normal_high_w_per_nic_with_optic": 100.0
+ },
+ "residual_static_w": 85.0,
+ "residual_normal_low_w": 60.0,
+ "residual_normal_high_w": 120.0
+ },
+ "cpu": {
+ "name": "2x Intel Xeon 6776P (Granite Rapids, DGX B300)",
+ "n_cpu": 2,
+ "cores_per_cpu": 64,
+ "threads_per_cpu": 128,
+ "package_tdp_w": 350.0,
+ "base_ghz": 2.3,
+ "max_turbo_ghz": 3.9,
+ "l3_cache_mb": 336.0,
+ "memory_channels_per_cpu": 8,
+ "pcie_lanes_per_cpu": 88,
+ "package_static_w": 24.0,
+ "uncore_io_baseline_w": 40.0,
+ "memory_controller_baseline_w": 24.0,
+ "uncore_dynamic_max_w": 9.0,
+ "memory_controller_dynamic_max_w": 7.0,
+ "core_curve_r": 1.68,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "DGX-B300-like Xeon 6776P, 32x64GB DDR5-6400 RDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 8,
+ "n_dimm": 32,
+ "capacity_gb_per_dimm": 64.0,
+ "data_rate_mtps": 6400.0,
+ "dimm_type": "RDIMM",
+ "background_w": 2.147,
+ "refresh_w": 0.5509999999999999,
+ "termination_w": 1.102,
+ "io_dynamic_max_w": 2.3400000000000003,
+ "core_dynamic_max_w": 2.86,
+ "activity_exponent": 1.0
+ },
+ "dpu": {
+ "n_dpu": 2,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "nvme": {
+ "n_front_u2": 8,
+ "front_u2_idle_w": 4.5,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 12.0,
+ "onboard_10gbe_w": 4.0,
+ "motherboard_pch_aux_w": 10.0,
+ "cpld_tpm_superio_w": 6.0,
+ "clock_sensor_fru_w": 5.0,
+ "storage_backplane_idle_w": 10.0,
+ "front_panel_usb_led_w": 3.0,
+ "aux_margin_w": 5.0,
+ "normal_low_w": 40.0,
+ "normal_high_w": 80.0
+ },
+ "fans": {
+ "electrical_nameplate_w": 2000.0,
+ "min_pwm_frac": 0.22,
+ "normal_max_pwm_frac": 0.85,
+ "full_cooling_load_w": 12000.0,
+ "fan_curve_exponent": 1.08
+ },
+ "psu": {
+ "n_installed_psu": 12,
+ "n_load_sharing_psu": 12,
+ "n_redundant_capacity_psu": 6,
+ "psu_capacity_w": 3300.0,
+ "system_max_w": 15000.0,
+ "redundancy": "N+N / 6+6",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "pue": 1.2
+ }
+ },
+ "mi300x": {
+ "functionName": "mi300x_chassis_power",
+ "configFactory": "MI300XChassisConfig",
+ "defaultConfig": {
+ "gpu_label": "AMD Instinct MI300X OAM",
+ "gpu_component_key": "mi300x_oam_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 1536.0,
+ "gpu_idle_example_w_per_gpu": 150.0,
+ "gpu_decode_example_w_per_gpu": 380.0,
+ "gpu_prefill_example_w_per_gpu": 620.0,
+ "gpu_aggregate_example_w_per_gpu": 580.0,
+ "gpu_peak_w_per_gpu": 750.0,
+ "ubb": {
+ "gpu_label": "AMD Instinct OAM",
+ "gpu_component_key": "amd_oam_8x_measured",
+ "n_gpu": 8,
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "management_controller_w": 10.0,
+ "fpga_cpld_sequencing_w": 8.0,
+ "hsc_power_monitor_w": 6.0,
+ "clock_refclk_reset_w": 5.0,
+ "sensors_i2c_fru_led_w": 4.0,
+ "aux_rails_misc_w": 12.0,
+ "normal_low_w": 30.0,
+ "normal_high_w": 70.0,
+ "conservative_cap_w": 90.0
+ }
+ },
+ "cpu": {
+ "name": "2x AMD EPYC 9654 (Genoa, MI300X host baseline)",
+ "n_cpu": 2,
+ "cores_per_cpu": 96,
+ "threads_per_cpu": 192,
+ "package_tdp_w": 360.0,
+ "base_ghz": 2.4,
+ "max_turbo_ghz": 3.7,
+ "l3_cache_mb": 384.0,
+ "memory_channels_per_cpu": 12,
+ "pcie_lanes_per_cpu": 128,
+ "package_static_w": 26.0,
+ "uncore_io_baseline_w": 42.0,
+ "memory_controller_baseline_w": 30.0,
+ "uncore_dynamic_max_w": 12.0,
+ "memory_controller_dynamic_max_w": 12.0,
+ "core_curve_r": 1.65,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "MI300X host, 24x96GB DDR5-4800 RDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 12,
+ "n_dimm": 24,
+ "capacity_gb_per_dimm": 96.0,
+ "data_rate_mtps": 4800.0,
+ "dimm_type": "RDIMM",
+ "background_w": 2.7119999999999997,
+ "refresh_w": 0.696,
+ "termination_w": 1.3920000000000001,
+ "io_dynamic_max_w": 2.79,
+ "core_dynamic_max_w": 3.41,
+ "activity_exponent": 1.0
+ },
+ "thor2": {
+ "n_nic": 8,
+ "ports_per_nic": 2,
+ "aggregate_gbps_per_nic": 400.0,
+ "pcie_generation": 5,
+ "pcie_lanes": 16,
+ "card_idle_w": 12.5,
+ "card_traffic_dynamic_w": 0.4,
+ "include_optics": true,
+ "optical_modules_per_nic": 2,
+ "optics_total_w_per_nic": 10.9,
+ "deployment_margin_w": 6.5,
+ "deployment_traffic_dynamic_w": 2.8
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 12,
+ "front_u2_idle_w": 4.5,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 12.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 12.0,
+ "cpld_tpm_superio_w": 6.0,
+ "clock_sensor_fru_w": 5.0,
+ "storage_backplane_idle_w": 12.0,
+ "front_panel_usb_led_w": 3.0,
+ "aux_margin_w": 8.0,
+ "normal_low_w": 45.0,
+ "normal_high_w": 90.0
+ },
+ "fans": {
+ "platform_label": "MI300X 8U air-cooled chassis",
+ "n_fan": 10,
+ "electrical_nameplate_w": 1800.0,
+ "min_pwm_frac": 0.22,
+ "normal_max_pwm_frac": 0.82,
+ "full_cooling_load_w": 8500.0,
+ "fan_curve_exponent": 1.08
+ },
+ "psu": {
+ "platform_label": "MI300X 8U chassis PSU bank",
+ "n_installed_psu": 6,
+ "n_load_sharing_psu": 6,
+ "n_redundant_capacity_psu": 3,
+ "psu_capacity_w": 3000.0,
+ "system_max_w": 9000.0,
+ "redundancy": "N+N / 3+3",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "pue": 1.2
+ }
+ },
+ "mi325x": {
+ "functionName": "mi325x_chassis_power",
+ "configFactory": "MI325XChassisConfig",
+ "defaultConfig": {
+ "gpu_label": "AMD Instinct MI325X OAM",
+ "gpu_component_key": "mi325x_oam_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 2048.0,
+ "gpu_idle_example_w_per_gpu": 170.0,
+ "gpu_decode_example_w_per_gpu": 450.0,
+ "gpu_prefill_example_w_per_gpu": 800.0,
+ "gpu_aggregate_example_w_per_gpu": 760.0,
+ "gpu_peak_w_per_gpu": 1000.0,
+ "ubb": {
+ "gpu_label": "AMD Instinct OAM",
+ "gpu_component_key": "amd_oam_8x_measured",
+ "n_gpu": 8,
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "management_controller_w": 10.0,
+ "fpga_cpld_sequencing_w": 8.0,
+ "hsc_power_monitor_w": 6.0,
+ "clock_refclk_reset_w": 5.0,
+ "sensors_i2c_fru_led_w": 4.0,
+ "aux_rails_misc_w": 12.0,
+ "normal_low_w": 30.0,
+ "normal_high_w": 70.0,
+ "conservative_cap_w": 90.0
+ }
+ },
+ "cpu": {
+ "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)",
+ "n_cpu": 2,
+ "cores_per_cpu": 64,
+ "threads_per_cpu": 128,
+ "package_tdp_w": 400.0,
+ "base_ghz": 3.3,
+ "max_turbo_ghz": 4.3,
+ "l3_cache_mb": 384.0,
+ "memory_channels_per_cpu": 12,
+ "pcie_lanes_per_cpu": 160,
+ "package_static_w": 28.0,
+ "uncore_io_baseline_w": 48.0,
+ "memory_controller_baseline_w": 34.0,
+ "uncore_dynamic_max_w": 14.0,
+ "memory_controller_dynamic_max_w": 14.0,
+ "core_curve_r": 1.62,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "MI325X host, 24x256GB DDR5-6400 RDIMM/MRDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 12,
+ "n_dimm": 24,
+ "capacity_gb_per_dimm": 256.0,
+ "data_rate_mtps": 6400.0,
+ "dimm_type": "RDIMM/MRDIMM",
+ "background_w": 3.9549999999999996,
+ "refresh_w": 1.015,
+ "termination_w": 2.0300000000000002,
+ "io_dynamic_max_w": 4.95,
+ "core_dynamic_max_w": 6.05,
+ "activity_exponent": 1.0
+ },
+ "thor2": {
+ "n_nic": 8,
+ "ports_per_nic": 2,
+ "aggregate_gbps_per_nic": 400.0,
+ "pcie_generation": 5,
+ "pcie_lanes": 16,
+ "card_idle_w": 12.5,
+ "card_traffic_dynamic_w": 0.4,
+ "include_optics": true,
+ "optical_modules_per_nic": 2,
+ "optics_total_w_per_nic": 10.9,
+ "deployment_margin_w": 6.5,
+ "deployment_traffic_dynamic_w": 2.8
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 8,
+ "front_u2_idle_w": 4.5,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 12.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 14.0,
+ "cpld_tpm_superio_w": 7.0,
+ "clock_sensor_fru_w": 6.0,
+ "storage_backplane_idle_w": 14.0,
+ "front_panel_usb_led_w": 3.0,
+ "aux_margin_w": 10.0,
+ "normal_low_w": 50.0,
+ "normal_high_w": 100.0
+ },
+ "fans": {
+ "platform_label": "MI325X 8U air-cooled chassis",
+ "n_fan": 14,
+ "electrical_nameplate_w": 2600.0,
+ "min_pwm_frac": 0.24,
+ "normal_max_pwm_frac": 0.88,
+ "full_cooling_load_w": 12000.0,
+ "fan_curve_exponent": 1.06
+ },
+ "psu": {
+ "platform_label": "MI325X 8U chassis PSU bank",
+ "n_installed_psu": 6,
+ "n_load_sharing_psu": 6,
+ "n_redundant_capacity_psu": 3,
+ "psu_capacity_w": 5250.0,
+ "system_max_w": 15750.0,
+ "redundancy": "N+N / 3+3",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "pue": 1.2
+ }
+ },
+ "mi355x": {
+ "functionName": "mi355x_chassis_power",
+ "configFactory": "MI355XChassisConfig",
+ "defaultConfig": {
+ "gpu_label": "AMD Instinct MI355X OAM",
+ "gpu_component_key": "mi355x_oam_8x_measured",
+ "n_gpu": 8,
+ "gpu_memory_total_gb": 2304.0,
+ "gpu_idle_example_w_per_gpu": 250.0,
+ "gpu_decode_example_w_per_gpu": 700.0,
+ "gpu_prefill_example_w_per_gpu": 1150.0,
+ "gpu_aggregate_example_w_per_gpu": 1080.0,
+ "gpu_peak_w_per_gpu": 1400.0,
+ "ubb": {
+ "gpu_label": "AMD Instinct OAM",
+ "gpu_component_key": "amd_oam_8x_measured",
+ "n_gpu": 8,
+ "retimer": {
+ "n_retimer": 8,
+ "lanes_per_retimer": 16,
+ "pcie_gtps": 32.0,
+ "analog_serdes_equalization_w": 9.5,
+ "digital_floor_w": 2.0,
+ "digital_variable_w": 1.0
+ },
+ "residual": {
+ "management_controller_w": 10.0,
+ "fpga_cpld_sequencing_w": 8.0,
+ "hsc_power_monitor_w": 6.0,
+ "clock_refclk_reset_w": 5.0,
+ "sensors_i2c_fru_led_w": 4.0,
+ "aux_rails_misc_w": 12.0,
+ "normal_low_w": 30.0,
+ "normal_high_w": 70.0,
+ "conservative_cap_w": 90.0
+ }
+ },
+ "cpu": {
+ "name": "2x AMD EPYC 9575F (Turin, MI325X/MI355X host baseline)",
+ "n_cpu": 2,
+ "cores_per_cpu": 64,
+ "threads_per_cpu": 128,
+ "package_tdp_w": 400.0,
+ "base_ghz": 3.3,
+ "max_turbo_ghz": 4.3,
+ "l3_cache_mb": 384.0,
+ "memory_channels_per_cpu": 12,
+ "pcie_lanes_per_cpu": 160,
+ "package_static_w": 28.0,
+ "uncore_io_baseline_w": 48.0,
+ "memory_controller_baseline_w": 34.0,
+ "uncore_dynamic_max_w": 14.0,
+ "memory_controller_dynamic_max_w": 14.0,
+ "core_curve_r": 1.62,
+ "vrm_efficiency": 0.92
+ },
+ "dram": {
+ "label": "MI355X host, 24x256GB DDR5-6400 RDIMM/MRDIMM",
+ "status": "DRAFT - pending human verification",
+ "n_sockets": 2,
+ "channels_per_socket": 12,
+ "n_dimm": 24,
+ "capacity_gb_per_dimm": 256.0,
+ "data_rate_mtps": 6400.0,
+ "dimm_type": "RDIMM/MRDIMM",
+ "background_w": 3.9549999999999996,
+ "refresh_w": 1.015,
+ "termination_w": 2.0300000000000002,
+ "io_dynamic_max_w": 4.95,
+ "core_dynamic_max_w": 6.05,
+ "activity_exponent": 1.0
+ },
+ "pollara400": {
+ "n_nic": 8,
+ "aggregate_gbps_per_nic": 400.0,
+ "pcie_generation": 5,
+ "pcie_lanes": 16,
+ "net_serdes_w": 10.5,
+ "pcie_serdes_w": 6.0,
+ "board_mgmt_w": 2.5,
+ "packet_engine_max_w": 12.5,
+ "packet_engine_floor_frac": 0.64,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "pcie_switch": {
+ "n_switch": 4,
+ "lanes_per_switch": 144,
+ "ports_per_switch": 72,
+ "pcie_gtps": 32.0,
+ "serdes_phy_floor_w": 30.5,
+ "control_leakage_clock_w": 6.0,
+ "fabric_datapath_floor_w": 4.0,
+ "fabric_datapath_variable_w": 8.5,
+ "stress_cap_per_switch_w": 70.0
+ },
+ "nvme": {
+ "n_front_u2": 8,
+ "front_u2_idle_w": 4.5,
+ "n_boot_m2": 2,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "residual": {
+ "bmc_ipmi_w": 12.0,
+ "onboard_10gbe_w": 8.0,
+ "motherboard_pch_aux_w": 16.0,
+ "cpld_tpm_superio_w": 8.0,
+ "clock_sensor_fru_w": 6.0,
+ "storage_backplane_idle_w": 16.0,
+ "front_panel_usb_led_w": 4.0,
+ "aux_margin_w": 12.0,
+ "normal_low_w": 60.0,
+ "normal_high_w": 120.0
+ },
+ "fans": {
+ "platform_label": "MI355X 10U air-cooled chassis",
+ "n_fan": 19,
+ "electrical_nameplate_w": 3600.0,
+ "min_pwm_frac": 0.25,
+ "normal_max_pwm_frac": 0.92,
+ "full_cooling_load_w": 15500.0,
+ "fan_curve_exponent": 1.04
+ },
+ "psu": {
+ "platform_label": "MI355X 10U chassis PSU bank",
+ "n_installed_psu": 6,
+ "n_load_sharing_psu": 6,
+ "n_redundant_capacity_psu": 4,
+ "psu_capacity_w": 6600.0,
+ "system_max_w": 26400.0,
+ "redundancy": "4+2",
+ "efficiency_curve": {
+ "0.05": 0.885,
+ "0.1": 0.92,
+ "0.2": 0.94,
+ "0.5": 0.96,
+ "1.0": 0.955
+ }
+ },
+ "pue": 1.2
+ }
+ },
+ "gb200": {
+ "functionName": "gb200_nvl72_rack_power",
+ "configFactory": "gb200_nvl72_rack_config",
+ "defaultConfig": {
+ "variant": "gb200",
+ "gpu_label": "GB200 Blackwell (NVL72)",
+ "n_compute_trays": 18,
+ "n_nvswitch_trays": 9,
+ "compute_tray": {
+ "n_gpu": 4,
+ "n_grace": 2,
+ "nic_generation": "connectx7",
+ "connectx7": {
+ "n_nic": 4,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "connectx8": {
+ "n_nic": 4,
+ "include_optic": true,
+ "network_serdes_static_w": 18.0,
+ "pcie_switch_static_w": 28.0,
+ "board_mgmt_static_w": 5.0,
+ "digital_static_w": 12.0,
+ "network_dynamic_max_w": 8.0,
+ "pcie_switch_dynamic_max_w": 5.0,
+ "digital_dynamic_max_w": 10.0,
+ "optic_idle_w": 15.0,
+ "optic_dynamic_max_w": 2.0,
+ "max_nic_slot_power_ref_w": 75.0,
+ "normal_low_w_per_nic_with_optic": 70.0,
+ "normal_high_w_per_nic_with_optic": 100.0
+ },
+ "dpu": {
+ "n_dpu": 2,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "nvme": {
+ "n_front_u2": 4,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 1,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "fans_w": 130.0,
+ "board_residual_w": 40.0
+ },
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "nvswitch_tray_residual_w": 50.0,
+ "tray_input_conversion_efficiency": 0.9725,
+ "n_management_switches": 2,
+ "management_switch_w": 100.0,
+ "power_shelf": {
+ "n_shelves": 8,
+ "n_psu_per_shelf": 6,
+ "psu_capacity_w": 5500.0,
+ "redundancy": "N+N",
+ "busbar_nominal_v": 50.0,
+ "efficiency_curve": {
+ "0.1": 0.9,
+ "0.2": 0.94,
+ "0.3": 0.965,
+ "1.0": 0.965
+ },
+ "peak_efficiency_ref": 0.975
+ },
+ "regulator_loss_frac_of_tdp": 0.15,
+ "regulator_allowance_includes_grace": false,
+ "pue": 1.2,
+ "cooling": "direct_liquid",
+ "gpu_tdp_example_w_per_gpu": 1200.0,
+ "grace_tdp_example_w_per_socket": 300.0,
+ "module_tdp_example_w_per_superchip": 2700.0
+ }
+ },
+ "gb300": {
+ "functionName": "gb200_nvl72_rack_power",
+ "configFactory": "gb300_nvl72_rack_config",
+ "defaultConfig": {
+ "variant": "gb300",
+ "gpu_label": "GB300 Blackwell Ultra (NVL72)",
+ "n_compute_trays": 18,
+ "n_nvswitch_trays": 9,
+ "compute_tray": {
+ "n_gpu": 4,
+ "n_grace": 2,
+ "nic_generation": "connectx8",
+ "connectx7": {
+ "n_nic": 4,
+ "net_serdes_w": 9.5,
+ "pcie_serdes_w": 5.5,
+ "board_w": 2.0,
+ "digital_max_w": 10.0,
+ "digital_floor_frac": 0.65,
+ "include_optic": true,
+ "optic_w": 8.0
+ },
+ "connectx8": {
+ "n_nic": 4,
+ "include_optic": true,
+ "network_serdes_static_w": 18.0,
+ "pcie_switch_static_w": 28.0,
+ "board_mgmt_static_w": 5.0,
+ "digital_static_w": 12.0,
+ "network_dynamic_max_w": 8.0,
+ "pcie_switch_dynamic_max_w": 5.0,
+ "digital_dynamic_max_w": 10.0,
+ "optic_idle_w": 15.0,
+ "optic_dynamic_max_w": 2.0,
+ "max_nic_slot_power_ref_w": 75.0,
+ "normal_low_w_per_nic_with_optic": 70.0,
+ "normal_high_w_per_nic_with_optic": 100.0
+ },
+ "dpu": {
+ "n_dpu": 2,
+ "idle_w_per_dpu": 65.0,
+ "public_max_power_cap_w": 150.0,
+ "max_modeled_u_dpu": 0.02
+ },
+ "nvme": {
+ "n_front_u2": 4,
+ "front_u2_idle_w": 5.0,
+ "n_boot_m2": 1,
+ "boot_m2_idle_w": 2.0,
+ "max_modeled_u_nvme": 0.02
+ },
+ "fans_w": 130.0,
+ "board_residual_w": 40.0
+ },
+ "nvswitch": {
+ "n_asic": 2,
+ "serdes_lanes": 144,
+ "serdes_lane_gbps": 200.0,
+ "serdes_pj_per_bit": 2.5,
+ "serdes_floor_frac": 0.95,
+ "digital_max_w": 190.0,
+ "digital_floor_frac": 0.4
+ },
+ "nvswitch_tray_residual_w": 50.0,
+ "tray_input_conversion_efficiency": 0.9725,
+ "n_management_switches": 2,
+ "management_switch_w": 100.0,
+ "power_shelf": {
+ "n_shelves": 8,
+ "n_psu_per_shelf": 6,
+ "psu_capacity_w": 5500.0,
+ "redundancy": "N+N",
+ "busbar_nominal_v": 50.0,
+ "efficiency_curve": {
+ "0.1": 0.9,
+ "0.2": 0.94,
+ "0.3": 0.965,
+ "1.0": 0.965
+ },
+ "peak_efficiency_ref": 0.975
+ },
+ "regulator_loss_frac_of_tdp": 0.15,
+ "regulator_allowance_includes_grace": false,
+ "pue": 1.2,
+ "cooling": "direct_liquid",
+ "gpu_tdp_example_w_per_gpu": 1400.0,
+ "grace_tdp_example_w_per_socket": 300.0,
+ "module_tdp_example_w_per_superchip": 0.0
+ }
+ }
+ },
"cases": [
{
"hardware": "h100",
diff --git a/packages/app/src/lib/system-power-model.test.ts b/packages/app/src/lib/system-power-model.test.ts
index d9d9ddbdc..90160cc6f 100644
--- a/packages/app/src/lib/system-power-model.test.ts
+++ b/packages/app/src/lib/system-power-model.test.ts
@@ -1,6 +1,11 @@
import { GPU_KEYS } from '@semianalysisai/inferencex-constants';
+import { createHash } from 'node:crypto';
+import { mkdtemp, mkdir, readFile, rm, writeFile } from 'node:fs/promises';
+import { tmpdir } from 'node:os';
+import { dirname, resolve } from 'node:path';
import { describe, expect, it } from 'vitest';
+import { buildSystemPowerProvenance } from '../../scripts/update-system-power-provenance';
import {
estimateChassisPower,
estimateRackPower,
@@ -8,6 +13,8 @@ import {
SUPPORTED_SYSTEM_POWER_HARDWARE,
SUPPORTED_SYSTEM_POWER_RACK_HARDWARE,
SYSTEM_POWER_MODEL_REVISION,
+ SYSTEM_POWER_MODEL_METADATA,
+ systemPowerSourceSha256,
} from './system-power-model';
import reference from './system-power-model.reference.json';
@@ -21,8 +28,7 @@ const rackInput = (row: (typeof reference.rackCases)[number]): RackMeasuredInput
: { basis: 'module', moduleWattsPerTray: row.moduleWattsPerTray };
describe('fixed 8k1k chassis model', () => {
- it('matches the pinned Python implementation across platforms, fan/PSU boundaries, and PUE', () => {
- expect(reference.modelRevision).toBe(SYSTEM_POWER_MODEL_REVISION);
+ it('matches the historical Python baseline across platforms, fan/PSU boundaries, and PUE', () => {
expect(new Set(reference.cases.map((row) => row.hardware))).toEqual(
new Set(SUPPORTED_SYSTEM_POWER_HARDWARE),
);
@@ -73,7 +79,7 @@ describe('fixed 8k1k chassis model', () => {
});
describe('NVL72 rack model with measured compute-module input', () => {
- it('matches the pinned Python rack model across variants, bases, shelf knots, and PUE', () => {
+ it('matches the historical Python baseline across variants, bases, shelf knots, and PUE', () => {
expect(new Set(reference.rackCases.map((row) => row.hardware))).toEqual(
new Set(SUPPORTED_SYSTEM_POWER_RACK_HARDWARE),
);
@@ -167,3 +173,73 @@ describe('NVL72 rack model with measured compute-module input', () => {
expect(module.facilityWatts).toBeCloseTo(split.facilityWatts, 0);
});
});
+
+describe('app-owned model provenance', () => {
+ it('identifies the actual app equations, parameters and admission/PUE policy', async () => {
+ const root = resolve(import.meta.dirname, '../../../..');
+ const metadata = SYSTEM_POWER_MODEL_METADATA;
+ expect(metadata).toMatchObject(await buildSystemPowerProvenance(root));
+ expect(metadata.source).toBe('https://github.com/SemiAnalysisAI/InferenceX-app');
+ expect(metadata.status).toBe('DRAFT / pending human verification');
+ expect(SYSTEM_POWER_MODEL_REVISION).toMatch(/^app-sha256:[0-9a-f]{64}$/u);
+ expect(Object.keys(metadata.sourceSha256)).toEqual([
+ 'packages/app/src/lib/modeled-system-power.ts',
+ 'packages/app/src/lib/system-power-model.profiles.json',
+ 'packages/app/src/lib/system-power-model.ts',
+ ]);
+ for (const [path, expectedHash] of Object.entries(metadata.sourceSha256)) {
+ const bytes = await readFile(resolve(root, path));
+ expect(expectedHash, path).toBe(createHash('sha256').update(bytes).digest('hex'));
+ }
+ for (const profile of Object.values({ ...metadata.profiles, ...metadata.rackProfiles })) {
+ expect(profile.modelPath).toBe('packages/app/src/lib/system-power-model.ts');
+ expect(profile).not.toHaveProperty('defaultConfig');
+ expect(profile).not.toHaveProperty('functionName');
+ expect(profile).not.toHaveProperty('configFactory');
+ expect(systemPowerSourceSha256(profile.modelPath)).toBe(
+ metadata.sourceSha256['packages/app/src/lib/system-power-model.ts'],
+ );
+ }
+ expect(systemPowerSourceSha256('toString')).toBeNull();
+ expect(systemPowerSourceSha256(reference.modelPaths.gb200)).toBeNull();
+ });
+
+ it('changes the version when profile parameters or admission/PUE policy change', async () => {
+ const root = resolve(import.meta.dirname, '../../../..');
+ const copy = await mkdtemp(resolve(tmpdir(), 'system-power-provenance-'));
+ try {
+ for (const path of Object.keys(SYSTEM_POWER_MODEL_METADATA.sourceSha256)) {
+ await mkdir(dirname(resolve(copy, path)), { recursive: true });
+ await writeFile(resolve(copy, path), await readFile(resolve(root, path)));
+ }
+ const baseline = await buildSystemPowerProvenance(copy);
+ expect(baseline.modelRevision).toBe(SYSTEM_POWER_MODEL_REVISION);
+ const profilePath = resolve(copy, 'packages/app/src/lib/system-power-model.profiles.json');
+ const profiles = JSON.parse(await readFile(profilePath, 'utf8'));
+ profiles.rackProfiles.gb200.computeTrayStaticDcWatts.compute_tray_fans_unverified += 1;
+ await writeFile(profilePath, JSON.stringify(profiles));
+ const parametersChanged = await buildSystemPowerProvenance(copy);
+ expect(parametersChanged.modelRevision).not.toBe(baseline.modelRevision);
+ const adapterPath = resolve(copy, 'packages/app/src/lib/modeled-system-power.ts');
+ const adapter = await readFile(adapterPath, 'utf8');
+ await writeFile(adapterPath, adapter.replace('DLC_SYSTEM_PUE = 1.1', 'DLC_SYSTEM_PUE = 1.2'));
+ const policyChanged = await buildSystemPowerProvenance(copy);
+ expect(policyChanged.modelRevision).not.toBe(parametersChanged.modelRevision);
+ } finally {
+ await rm(copy, { recursive: true, force: true });
+ }
+ });
+
+ it('keeps all independent historical expected values and their original lineage', () => {
+ expect(reference.modelRevision).toBe('6fcc086b77576d4cecb9d0c79637d6daf980308c');
+ expect(reference.source).toBe('https://github.com/SemiAnalysisAI/inferencex_power_model');
+ expect(reference.fixtureRole).toContain('Frozen historical Python migration baseline');
+ expect(reference.pythonConfigurations.gb200.defaultConfig.pue).toBe(1.2);
+ expect(reference.cases.length + reference.rackCases.length).toBe(496);
+ expect(
+ createHash('sha256')
+ .update(JSON.stringify([reference.cases, reference.rackCases]))
+ .digest('hex'),
+ ).toBe('98cb77c182c510b0e73fefaea1242f20b732e8623cc2cfcc8b6f367d9412d268');
+ });
+});
diff --git a/packages/app/src/lib/system-power-model.ts b/packages/app/src/lib/system-power-model.ts
index 15d7f3f43..6f1accfe4 100644
--- a/packages/app/src/lib/system-power-model.ts
+++ b/packages/app/src/lib/system-power-model.ts
@@ -1,6 +1,9 @@
import profileData from './system-power-model.profiles.json';
+import provenance from './system-power-model.provenance.json';
-export const SYSTEM_POWER_MODEL_REVISION = profileData.modelRevision;
+export const SYSTEM_POWER_MODEL_REVISION = provenance.modelRevision;
+export const SYSTEM_POWER_MODEL_SOURCE = provenance.source;
+export const SYSTEM_POWER_MODEL_METADATA = { ...provenance, ...profileData };
export const SYSTEM_POWER_ASSUMPTIONS = profileData.assumptions;
export const SYSTEM_POWER_PROFILES = profileData.profiles;
export type SystemPowerHardware = keyof typeof SYSTEM_POWER_PROFILES;
@@ -16,9 +19,9 @@ export const SUPPORTED_SYSTEM_POWER_RACK_HARDWARE = Object.keys(
export type RackMeasuredBasis = 'module' | 'gpu-plus-grace';
-/** SHA-256 of the pinned source file behind a profile's `modelPath`, for export provenance. */
+/** The full revision also covers profile parameters and the admission/PUE adapter. */
export function systemPowerSourceSha256(modelPath: string): string | null {
- const hashes: Readonly> = profileData.sourceSha256;
+ const hashes: Readonly> = provenance.sourceSha256;
return Object.hasOwn(hashes, modelPath) ? hashes[modelPath] : null;
}
@@ -87,7 +90,7 @@ function pythonRound(value: number, digits = 1): number {
return Number(value.toFixed(digits));
}
-/** Port of the source's `_interp_efficiency`: clamp outside the knots, linear between them. */
+/** Clamping avoids extrapolating beyond the profile's efficiency knots. */
function interpolateEfficiency(loadFraction: number, curve: number[][]): number {
const [first, last] = [curve[0], curve.at(-1)!];
if (loadFraction <= first[0]) return first[1];
@@ -103,10 +106,8 @@ function interpolateEfficiency(loadFraction: number, curve: number[][]): number
}
/**
- * Fixed README inference sweep at the pinned revision, for one complete 8-GPU chassis.
- * Profiles are generated by executing the original component models. Only their
- * load-dependent fan curve and PSU interpolation execute here; other components
- * stay fixed at the recorded assumptions. This function does no allocation scaling.
+ * One complete 8-GPU chassis. Only fan and PSU loads vary at runtime; the remaining
+ * component parameters stay fixed at the profile's recorded inference assumptions.
*/
export function estimateChassisPower(
hardware: string,
diff --git a/packages/app/src/lib/views-api/docs/extensions.ts b/packages/app/src/lib/views-api/docs/extensions.ts
index ccba97eae..80ee85a20 100644
--- a/packages/app/src/lib/views-api/docs/extensions.ts
+++ b/packages/app/src/lib/views-api/docs/extensions.ts
@@ -122,8 +122,8 @@ const PARAMETER_NOTES: Record = {
'模型许可或收入分成百分比,范围 0 至 100,默认值随模型变化。',
],
powerBasis: [
- 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. All in Measured uses measured GPU power plus modeled unmeasured components and PUE; it is not measured wall power. Eligible measured source rows are required; powerLabel identifies paired estimates and full-chassis extrapolation. NVL72 requires validated GPU and Grace/module telemetry with complete socket coverage; CPU rail alone is insufficient. powerSource records topology, measured basis, sensor, PUE and the model revision, path and profile hash. compare retains provisioned estimates when measured estimates are unavailable; skipped reasons distinguish missing CPU power and incompatible sensor bases. Missing coverage is not zero.',
- 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。整体实测功耗采用 GPU 实测值,加上未实测组件的功耗估算和 PUE,并非墙上电表读数。该估算需要符合条件的实测数据行;powerLabel 标明对比方式和整机外推。NVL72 需要通过验证的 GPU 与 Grace/模块遥测,并完整覆盖所有 socket;仅 CPU rail 读数不满足要求。powerSource 记录拓扑、实测口径、传感器、PUE 及模型版本、路径和 profile 哈希。compare 在实测估算不可用时保留预配估算;skipped 原因区分 CPU 功耗缺失与传感器口径不兼容。缺失数据不按零处理。',
+ 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. All in Measured uses measured GPU power plus modeled unmeasured components and PUE; it is not measured wall power. Eligible measured source rows are required; powerLabel identifies paired estimates and full-chassis extrapolation. NVL72 requires validated GPU and Grace/module telemetry with complete socket coverage; CPU rail alone is insufficient. powerSource records topology, measured basis, sensor, PUE and the app-owned model content revision, TypeScript source path and source hash. compare retains provisioned estimates when measured estimates are unavailable; skipped reasons distinguish missing CPU power and incompatible sensor bases. Missing coverage is not zero.',
+ 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。整体实测功耗采用 GPU 实测值,加上未实测组件的功耗估算和 PUE,并非墙上电表读数。该估算需要符合条件的实测数据行;powerLabel 标明对比方式和整机外推。NVL72 需要通过验证的 GPU 与 Grace/模块遥测,并完整覆盖所有 socket;仅 CPU rail 读数不满足要求。powerSource 记录拓扑、实测口径、传感器、PUE 及由应用维护的模型内容版本、TypeScript 源码路径和源码哈希。compare 在实测估算不可用时保留预配估算;skipped 原因区分 CPU 功耗缺失与传感器口径不兼容。缺失数据不按零处理。',
],
power: [
'Comma-separated certified and/or legacy power tiers. Omit for all tiers.',
diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md
index 22ffca2ae..43e4e5f96 100644
--- a/packages/skills/skills/inferencex-api/references/dashboard-views.md
+++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md
@@ -100,7 +100,10 @@ NVL72 estimates require valid GPU power plus validated Grace-socket or compute-m
power with complete socket coverage. CPU-rail-only readings do not establish the
Grace/LPDDR boundary. A module reading already includes GPU power; do not add GPU
watts again. Read `powerSource` for topology, measured basis, sensor, PUE, model
-revision/path and profile hash. `compare` preserves provisioned rows when a measured
+content revision, app TypeScript source path and source hash. Equations and parameters
+are maintained in InferenceX-app; the revision is a content digest, not a private-repository
+Git commit. Model-only updates recalculate retained valid measurements after deployment;
+they do not require telemetry backfill. `compare` preserves provisioned rows when a measured
estimate is unavailable; `skipped.reason` distinguishes `no-cpu-power` from
`incompatible-power-basis`. Modeled-only estimates never substitute provisioned
watts, and neither mode selects a different serving frontier to fill missing power.
From 35d15b90bb4a5af86e1d65f902b417ecd3de809f Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 13:03:33 -0700
Subject: [PATCH 13/22] test: expect the app-owned system power model in the
NVL72 CSV
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
989bb69f made InferenceX-app the source of truth for the NVL72 power model,
so the CSV's System power profile names system-power-model.ts with its
app-sha256 revision; the spec still looked for the retired Python file.
#1190 targets a feature branch, so the E2E workflow never ran this spec.
中文:989bb69f 起 NVL72 功耗模型以 InferenceX-app 为准,CSV 的 System power
profile 列改为标注 system-power-model.ts 及其 app-sha256 修订号;用例仍在
查找已停用的 Python 文件。#1190 基于功能分支,E2E 工作流不会运行该用例。
---
packages/app/cypress/e2e/profit-estimator.cy.ts | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts
index fbe67a183..81181b76c 100644
--- a/packages/app/cypress/e2e/profit-estimator.cy.ts
+++ b/packages/app/cypress/e2e/profit-estimator.cy.ts
@@ -405,7 +405,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
const measured = rows.find((row) => row.includes('measured module'));
expect(measured).to.contain('GPU + HBM + Grace + LPDDR5X; module sensor');
expect(measured).to.contain(',module,');
- expect(measured).to.contain('gb200_nvl72_rack_power_model.py @ ');
+ expect(measured).to.contain('system-power-model.ts @ app-sha256:');
expect(measured).to.contain(' sha256:');
});
});
From 55d43315b00d871a5d24009a76d85c7dd815c036 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 14:45:08 -0700
Subject: [PATCH 14/22] test: trim redundant PowerX modeling cases
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Remove 45 of 147 scoped PowerX unit cases (30.61%), retaining numerical,
sensor, topology, provenance, and financial regression coverage. Matched
V8 runs preserve every previously covered statement, branch, and function;
all 6,000 remaining app unit tests pass. Production code and coverage
configuration are unchanged.
中文:精简冗余的 PowerX 建模测试。删除 147 项相关单测中的 45 项
(30.61%),保留数值基准、传感器、拓扑、来源版本及财务回归覆盖。
前后使用相同代码及 V8 配置,四项覆盖率均不变,剩余 6,000 项应用
单元测试全部通过;未修改产品代码或覆盖率配置。
---
.../src/app/api/v1/views/extensions.test.ts | 4 +-
.../calculator/profit-power.test.ts | 32 +------
.../inference/utils/tooltip-utils.test.ts | 94 ++++--------------
.../lib/modeled-system-power-export.test.ts | 19 ++--
.../app/src/lib/modeled-system-power.test.ts | 95 ++++---------------
.../app/src/lib/system-power-model.test.ts | 50 ----------
6 files changed, 51 insertions(+), 243 deletions(-)
diff --git a/packages/app/src/app/api/v1/views/extensions.test.ts b/packages/app/src/app/api/v1/views/extensions.test.ts
index 9dbd26f30..069795f57 100644
--- a/packages/app/src/app/api/v1/views/extensions.test.ts
+++ b/packages/app/src/app/api/v1/views/extensions.test.ts
@@ -349,7 +349,7 @@ describe('new dashboard projections', () => {
expect([firstFetchCount, mocks.unofficial.mock.calls.length]).toEqual([1, 2]);
},
);
- it.each(['modeled', 'compare'])(
+ it.each(['modeled'])(
'uses exact comparison snapshots with %s power while keeping the primary date cutoff',
async (powerBasis) => {
const rows = [
@@ -478,7 +478,7 @@ describe('new dashboard projections', () => {
}
},
);
- it.each(['modeled', 'compare'])(
+ it.each(['modeled'])(
'labels full-chassis extrapolation for official and overlay %s estimates',
async (powerBasis) => {
const partial = agenticRow({
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index 8f109c88b..a408619f2 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -136,7 +136,7 @@ const withPoints = (base: InterpolatedResult, points: GPUDataPoint[]): Interpola
});
describe('profit power basis preview', () => {
- it.each([8, 4, 2])(
+ it.each([8])(
'keeps raw power attached through official and run-keyed %i-GPU frontiers',
(gpus) => {
const row = {
@@ -173,30 +173,6 @@ describe('profit power basis preview', () => {
},
);
- it('distinguishes partial chassis from NVL72 rows missing CPU telemetry', () => {
- for (const hardware of ['gb200', 'gb300']) {
- expect(
- modeledPowerAtTarget(
- { ...result, nearestPoints: [{ ...point, sourceRow: { ...source, hardware } }] },
- 45,
- ),
- ).toEqual({ reason: 'no-cpu-power' });
- }
- const partial = {
- ...source,
- prefill_tp: 4,
- decode_tp: 4,
- metrics: { ...source.metrics, avg_power_w: 500, avg_total_gpu_power_w: 2000 },
- };
- expect(modelSystemPower(partial, undefined, true)).toMatchObject({
- status: 'supported',
- chassisBasis: 'extrapolated',
- });
- expect(
- modeledPowerAtTarget({ ...result, nearestPoints: [{ ...point, sourceRow: partial }] }, 45),
- ).toMatchObject({ extrapolated: true });
- });
-
it('accepts fully measured NVL72 trays and records the measured basis behind the estimate', () => {
// 1.1 × the tray's amortised facility watts per GPU from the pinned GB200 rack profile.
const rack = estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 4300.75 }, 1.1)!;
@@ -367,7 +343,7 @@ describe('profit power basis preview', () => {
expect(modelSystemPower(source)).toMatchObject({ status: 'unsupported', reason: 'workload' });
});
- it.each([2, 4])(
+ it.each([2])(
'prices a validated %i-GPU allocation as a labeled full-chassis extrapolation',
(gpus) => {
const partial = {
@@ -546,10 +522,6 @@ describe('profit power basis preview', () => {
},
'unsupported-power-topology',
],
- [
- { metrics: { ...source.metrics, avg_power_w: 500, avg_total_gpu_power_w: 2000 } },
- 'no-measured-power',
- ],
[
{ metrics: { power_valid: 1, avg_power_w: 796.131, avg_total_gpu_power_w: 6369.045 } },
'no-measured-power',
diff --git a/packages/app/src/components/inference/utils/tooltip-utils.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
index 370d462b8..e1182ab79 100644
--- a/packages/app/src/components/inference/utils/tooltip-utils.test.ts
+++ b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
@@ -151,26 +151,23 @@ describe('modeled system-power tooltip', () => {
expect(html).not.toContain('Unmeasured chassis GPUs');
});
- it.each(['en', 'zh'] as const)(
- 'links %s model provenance to the deployed app source',
- (locale) => {
- const buildRef = 'b'.repeat(40);
- vi.stubEnv('NEXT_PUBLIC_APP_SOURCE_REF', buildRef);
- try {
- const html = generateTooltipContent(config({ locale }));
- const app = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${buildRef}`;
- expect(html).toContain(`${app}/${systemPower.modelPath}`);
- expect(html).toContain(`${app}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`);
- expect(html).toContain(locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions');
- expect(html).toContain(`title="${systemPower.modelRevision}"`);
- expect(html).toContain('h100 · aaaaaaaaaaaa');
- expect(html).not.toContain('inferencex_power_model');
- expect(html).not.toContain(`/blob/${systemPower.modelRevision}/`);
- } finally {
- vi.unstubAllEnvs();
- }
- },
- );
+ it.each(['zh'] as const)('links %s model provenance to the deployed app source', (locale) => {
+ const buildRef = 'b'.repeat(40);
+ vi.stubEnv('NEXT_PUBLIC_APP_SOURCE_REF', buildRef);
+ try {
+ const html = generateTooltipContent(config({ locale }));
+ const app = `https://github.com/SemiAnalysisAI/InferenceX-app/blob/${buildRef}`;
+ expect(html).toContain(`${app}/${systemPower.modelPath}`);
+ expect(html).toContain(`${app}/docs/powerx-system-power${locale === 'zh' ? '.zh' : ''}.md`);
+ expect(html).toContain(locale === 'zh' ? '功耗模型与假设' : 'Power model assumptions');
+ expect(html).toContain(`title="${systemPower.modelRevision}"`);
+ expect(html).toContain('h100 · aaaaaaaaaaaa');
+ expect(html).not.toContain('inferencex_power_model');
+ expect(html).not.toContain(`/blob/${systemPower.modelRevision}/`);
+ } finally {
+ vi.unstubAllEnvs();
+ }
+ });
it('labels an extrapolated partial chassis and reports the measured GPUs’ share', () => {
const data = pt({
@@ -299,17 +296,6 @@ describe('modeled system-power tooltip', () => {
);
});
- it('localizes the measurement boundary and occupancy assumptions', () => {
- const html = generateTooltipContent(config({ locale: 'zh' }));
- expect(html).toContain('GPU 实测功耗');
- expect(html).toContain('整个部署的机箱交流功耗估算');
- expect(html).toContain('数据中心功耗估算');
- expect(html).toContain('2 个完整八卡机箱 · 16 张 GPU');
- expect(html).toContain('CPU/DRAM 利用率:20%');
- expect(html).toContain('计入 GPU 机箱内的 CPU');
- expect(html).toContain('不计入独立的纯 CPU 前端或路由主机。');
- });
-
it('breaks normalization and host scope into two compact lines in pinned tooltips', () => {
for (const locale of ['en', 'zh'] as const) {
const html = generateTooltipContent(config({ locale }));
@@ -1103,15 +1089,6 @@ describe('generateGPUGraphTooltipContent', () => {
describe('measured-power withheld tooltip line', () => {
const reasons = ['sampling_gap_exceeded', 'expected_gpu_count_mismatch'];
- it('renders the withheld line with humanized codes (en)', () => {
- const html = generateTooltipContent(
- tooltipConfig({ data: pt({ power_valid: 0, power_invalid_reasons: reasons }) }),
- );
- expect(html).toContain('Measured power withheld');
- expect(html).toContain('sampling gap exceeded');
- expect(html).toContain('expected gpu count mismatch');
- });
-
it('renders the withheld line in Chinese on /zh surfaces', () => {
const html = generateTooltipContent(
tooltipConfig({
@@ -1144,11 +1121,7 @@ describe('measured-power withheld tooltip line', () => {
expect(html).not.toContain('Measured power withheld');
});
- it.each([
- ['absent reasons', pt({ power_valid: 0 })],
- ['empty reasons', pt({ power_valid: 0, power_invalid_reasons: [] })],
- ['valid row', pt({ power_valid: 1 })],
- ])('omits the line for %s', (_name, data) => {
+ it.each([['absent reasons', pt({ power_valid: 0 })]])('omits the line for %s', (_name, data) => {
const html = generateTooltipContent(tooltipConfig({ data }));
expect(html).not.toContain('Measured power withheld');
});
@@ -1240,15 +1213,6 @@ describe('worker power drilldown', () => {
expect(generateGPUGraphTooltipContent(config)).not.toContain('tooltip-worker-power');
});
- it('renders nothing when workers is absent or empty', () => {
- expect(generateTooltipContent(tooltipConfig({ isPinned: true }))).not.toContain(
- 'tooltip-worker-power',
- );
- expect(
- generateTooltipContent(tooltipConfig({ data: pt({ workers: [] }), isPinned: true })),
- ).not.toContain('tooltip-worker-power');
- });
-
it('caps the table at 8 rows with a "+N more workers" line', () => {
const many = Array.from({ length: 10 }, (_, i) => ({
role: 'decode',
@@ -1303,18 +1267,6 @@ describe('worker power drilldown', () => {
});
describe('power tier tooltip line', () => {
- it('states the tier for a legacy point on a measured axis', () => {
- const html = generateTooltipContent(
- tooltipConfig({
- selectedYAxisMetric: 'y_measuredJPerOutputToken',
- data: pt({ power_tier: 'legacy' }),
- }),
- );
- expect(html).toContain(
- 'Power Measurement: Historical (not validated under the current method)',
- );
- });
-
it('states the certified tier on a measured axis', () => {
const html = generateTooltipContent(
tooltipConfig({
@@ -1325,16 +1277,6 @@ describe('power tier tooltip line', () => {
expect(html).toContain('Power Measurement: Validated (current PowerX method)');
});
- it('omits the tier line on non-measured axes', () => {
- const html = generateTooltipContent(
- tooltipConfig({
- selectedYAxisMetric: 'y_tpPerGpu',
- data: pt({ power_tier: 'legacy' }),
- }),
- );
- expect(html).not.toContain('Power Measurement');
- });
-
it('omits the tier line when the point carries no tier', () => {
const html = generateTooltipContent(
tooltipConfig({ selectedYAxisMetric: 'y_measuredAvgPower', data: pt() }),
diff --git a/packages/app/src/lib/modeled-system-power-export.test.ts b/packages/app/src/lib/modeled-system-power-export.test.ts
index f308c32fa..e0fc5af57 100644
--- a/packages/app/src/lib/modeled-system-power-export.test.ts
+++ b/packages/app/src/lib/modeled-system-power-export.test.ts
@@ -162,17 +162,14 @@ describe('offline modeled PowerX comparisons', () => {
});
});
- it.each([0, -1, Infinity, NaN, 1.5])(
- 'withholds energy for invalid output denominator %s',
- (tokens) => {
- const source = input();
- source.rows[0].audit!.benchmark_window.total_output_tokens = tokens;
- expect(buildComparison(source).rows[0].estimated_energy).toEqual({
- status: 'unavailable',
- reason: 'audit-does-not-match-measured-input',
- });
- },
- );
+ it.each([0, 1.5])('withholds energy for invalid output denominator %s', (tokens) => {
+ const source = input();
+ source.rows[0].audit!.benchmark_window.total_output_tokens = tokens;
+ expect(buildComparison(source).rows[0].estimated_energy).toEqual({
+ status: 'unavailable',
+ reason: 'audit-does-not-match-measured-input',
+ });
+ });
it('withholds energy for missing, mismatched, and invalid audit receipts', () => {
for (const mutate of [
diff --git a/packages/app/src/lib/modeled-system-power.test.ts b/packages/app/src/lib/modeled-system-power.test.ts
index ab5445df6..e5397864a 100644
--- a/packages/app/src/lib/modeled-system-power.test.ts
+++ b/packages/app/src/lib/modeled-system-power.test.ts
@@ -1,7 +1,7 @@
import { describe, expect, it } from 'vitest';
import type { BenchmarkRow } from '@/lib/api';
-import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform';
+import { transformBenchmarkRows } from '@/lib/benchmark-transform';
import { modelSystemPower } from '@/lib/modeled-system-power';
import { estimateChassisPower, estimateRackPower } from '@/lib/system-power-model';
@@ -91,22 +91,6 @@ function nvl72Row(
}
describe('modeled system power admission and accounting', () => {
- it('defaults air-cooled chassis to PUE 1.3 and preserves explicit facility overrides', () => {
- // Pinned Python b200_chassis_power, fixed README utilization inputs.
- expect(modelSystemPower(row())).toMatchObject({
- pue: 1.3,
- chassisAcWatts: 4837.2,
- facilityWatts: 6288.4,
- measuredGpuWattsPerGpu: 349.859,
- });
- expect(modelSystemPower(row(), 1.1)).toMatchObject({
- pue: 1.1,
- chassisAcWatts: 4837.2,
- facilityWatts: 5320.9,
- measuredGpuWattsPerGpu: 349.859,
- });
- });
-
it('uses the validated physical count without summing aggregate aliases or multiplying by EP', () => {
const source = row({ num_prefill_gpu: 64, num_decode_gpu: 64, prefill_ep: 8, decode_ep: 8 });
const result = modelSystemPower(source);
@@ -129,22 +113,19 @@ describe('modeled system power admission and accounting', () => {
expect(source.metrics.joules_per_output_token).toBe(12.937902);
});
- it.each(['rtx6000pro', 'tpuv7', 'b200-nvl'])('does not substitute for %s', (hardware) => {
+ it.each(['b200-nvl'])('does not substitute for %s', (hardware) => {
expect(modelSystemPower(row({ hardware }))).toMatchObject({
status: 'unsupported',
reason: 'hardware',
});
});
- it.each(['gb200', 'gb300'])(
- 'never models the Grace side of %s from GPU-only telemetry',
- (hardware) => {
- expect(modelSystemPower(row({ hardware }))).toMatchObject({
- status: 'unsupported',
- reason: 'cpu-telemetry',
- });
- },
- );
+ it.each(['gb300'])('never models the Grace side of %s from GPU-only telemetry', (hardware) => {
+ expect(modelSystemPower(row({ hardware }))).toMatchObject({
+ status: 'unsupported',
+ reason: 'cpu-telemetry',
+ });
+ });
it('ignores CPU-side keys on x86 chassis rows', () => {
const source = row();
@@ -152,7 +133,7 @@ describe('modeled system power admission and accounting', () => {
expect(modelSystemPower(source)).toEqual(modelSystemPower(row()));
});
- it.each([{ benchmark_type: 'agentic_traces' }, { isl: 1024 }, { osl: 8192 }, { isl: null }])(
+ it.each([{ benchmark_type: 'agentic_traces' }, { isl: 1024 }, { osl: 8192 }])(
'keeps non-8k1k workloads unavailable: %j',
(overrides) => {
expect(modelSystemPower(row(overrides))).toMatchObject({
@@ -164,18 +145,10 @@ describe('modeled system power admission and accounting', () => {
it.each([
{ power_valid: 0 },
- { power_valid: undefined },
{ power_valid: '1' },
- { power_valid: true },
- { power_metric_schema_version: 1 },
{ power_metric_schema_version: 3 },
- { avg_power_w: undefined },
{ avg_power_w: 0 },
- { avg_power_w: -1 },
- { avg_power_w: Infinity },
- { avg_power_w: NaN },
{ avg_power_w: '349.859' },
- { avg_total_gpu_power_w: undefined },
{ avg_total_gpu_power_w: -1 },
])('rejects invalid measured inputs: %j', (overrides) => {
const source = row();
@@ -439,7 +412,7 @@ describe('modeled system power admission and accounting', () => {
});
});
- it.each([false, true])('requires schema-v2 for workers across hosts (disagg=%s)', (disagg) => {
+ it.each([false])('requires schema-v2 for workers across hosts (disagg=%s)', (disagg) => {
const source = row({
disagg,
is_multinode: true,
@@ -591,22 +564,6 @@ describe('modeled system power admission and accounting', () => {
Object.assign(frontend, { role: 'other', num_gpus: 0 });
expect(modelSystemPower(source)).toMatchObject({ status: 'unsupported' });
});
-
- it('shares official/overlay transforms without changing measured metrics or inventing modeled zeros', () => {
- const source = row();
- const entry = rowToAggDataEntry(source);
- expect(entry.avg_power_w).toBe(source.metrics.avg_power_w);
- expect(entry.joules_per_output_token).toBe(source.metrics.joules_per_output_token);
- const { chartData } = transformBenchmarkRows([source]);
- for (const points of chartData) {
- expect(points[0].modeledChassisPowerPerGpu?.y).toBeGreaterThan(source.metrics.avg_power_w);
- }
- const unsupported = transformBenchmarkRows([row({ hardware: 'gb200' })]);
- for (const points of unsupported.chartData) {
- expect(points[0].modeledChassisPowerPerGpu).toBeUndefined();
- expect(points[0].measuredAvgPower?.y).toBe(source.metrics.avg_power_w);
- }
- });
});
describe('NVL72 trays with measured compute-module power', () => {
@@ -917,27 +874,17 @@ describe('NVL72 trays with measured compute-module power', () => {
expect(modelSystemPower(twoTrays)).toMatchObject({ reason: 'topology' });
});
- it.each([
- { cpu_power_valid: undefined },
- { cpu_power_valid: 0 },
- { cpu_power_valid: '1' },
- { avg_total_cpu_power_w: undefined },
- { avg_total_cpu_power_w: 0 },
- { avg_total_cpu_power_w: -1 },
- { avg_cpu_socket_power_w: undefined },
- { avg_cpu_socket_power_w: 0 },
- { avg_total_module_power_w: 0 },
- { avg_total_module_power_w: -1 },
- { avg_total_module_power_w: NaN },
- { avg_total_module_power_w: '4300' },
- ])('keeps NVL72 rows without valid CPU-side telemetry unavailable: %j', (overrides) => {
- const source = nvl72Row();
- Object.assign(source.metrics, overrides);
- expect(modelSystemPower(source)).toMatchObject({
- status: 'unsupported',
- reason: 'cpu-telemetry',
- });
- });
+ it.each([{ cpu_power_valid: 0 }, { cpu_power_valid: '1' }, { avg_total_cpu_power_w: undefined }])(
+ 'keeps NVL72 rows without valid CPU-side telemetry unavailable: %j',
+ (overrides) => {
+ const source = nvl72Row();
+ Object.assign(source.metrics, overrides);
+ expect(modelSystemPower(source)).toMatchObject({
+ status: 'unsupported',
+ reason: 'cpu-telemetry',
+ });
+ },
+ );
it('requires schema-v2 GPU telemetry and stays within the shelf and facility domain', () => {
const legacy = nvl72Row({ ...MODULE, power_metric_schema_version: undefined });
diff --git a/packages/app/src/lib/system-power-model.test.ts b/packages/app/src/lib/system-power-model.test.ts
index 90160cc6f..62c9521fa 100644
--- a/packages/app/src/lib/system-power-model.test.ts
+++ b/packages/app/src/lib/system-power-model.test.ts
@@ -55,18 +55,6 @@ describe('fixed 8k1k chassis model', () => {
}
});
- it('keeps measured input, modeled chassis AC, and post-AC facility power separate', () => {
- const chassis = estimateChassisPower('H100', 4000, 1)!;
- const facility = estimateChassisPower('h100', 4000, 1.2)!;
- expect(chassis.measuredGpuWatts).toBe(4000);
- expect(chassis.chassisAcWatts).toBe(6228.2);
- expect(chassis.facilityWatts).toBe(chassis.chassisAcWatts);
- expect(facility.chassisAcWatts).toBe(chassis.chassisAcWatts);
- expect(facility.facilityWatts).toBe(7473.8);
- // Fixed chassis overhead and nonlinear fan/PSU behavior forbid proportional GPU scaling.
- expect(estimateChassisPower('h100', 8000)!.chassisAcWatts).not.toBe(chassis.chassisAcWatts * 2);
- });
-
it('names every chassis and rack profile by its hardware registry key', () => {
for (const hardware of [
...SUPPORTED_SYSTEM_POWER_HARDWARE,
@@ -134,44 +122,6 @@ describe('NVL72 rack model with measured compute-module input', () => {
}
expect(estimateChassisPower('gb200', 4000)).toBeNull();
});
-
- it('applies PUE once to the rounded rack AC and amortises the rack over all 72 GPUs', () => {
- const input: RackMeasuredInput = { basis: 'module', moduleWattsPerTray: 5400 };
- const rack = estimateRackPower('GB200', input, 1)!;
- const facility = estimateRackPower('gb200', input, 1.1)!;
- expect(rack.hardware).toBe('gb200');
- expect(rack.facilityWatts).toBe(rack.rackAcWatts);
- expect(facility.rackAcWatts).toBe(rack.rackAcWatts);
- expect(facility.facilityWatts).toBeCloseTo(rack.rackAcWatts * 1.1, 0);
- expect(facility.perGpuAcWatts).toBeCloseTo(rack.rackAcWatts / 72, 0);
- expect(facility.perGpuFacilityWatts).toBeCloseTo(facility.facilityWatts / 72, 0);
- expect(facility.gpuCount).toBe(72);
- expect(facility.computeTrayCount).toBe(18);
- // Static trays, switch trays, and the shelf curve forbid proportional scaling.
- expect(
- estimateRackPower('gb200', { basis: 'module', moduleWattsPerTray: 2700 }, 1)!.rackAcWatts * 2,
- ).not.toBe(rack.rackAcWatts);
- });
-
- it('adds the sourced regulator allowance only on the GPU-board share of the split basis', () => {
- const split = estimateRackPower('gb300', {
- basis: 'gpu-plus-grace',
- gpuBoardWattsPerTray: 4800,
- graceSocketWattsPerTray: 600,
- })!;
- const module = estimateRackPower('gb300', {
- basis: 'module',
- moduleWattsPerTray: split.measuredWattsPerTray + split.regulatorAllowanceWattsPerTray,
- })!;
- expect(split.basis).toBe('gpu-plus-grace');
- expect(split.measuredWattsPerTray).toBe(5400);
- expect(split.regulatorAllowanceWattsPerTray).toBeGreaterThan(0);
- expect(module.basis).toBe('module');
- expect(module.regulatorAllowanceWattsPerTray).toBe(0);
- // The electrical equivalent module reading reproduces the split-basis rack.
- expect(module.rackAcWatts).toBeCloseTo(split.rackAcWatts, 0);
- expect(module.facilityWatts).toBeCloseTo(split.facilityWatts, 0);
- });
});
describe('app-owned model provenance', () => {
From b41c0e41127b6ac7eb66be620f2ba4ac8ffb74b4 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 15:28:23 -0700
Subject: [PATCH 15/22] fix: align NVL72 chart explanations with measured
inputs
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Correct measured-power eligibility, energy scaling, and cooling-specific PUE in chart copy and model documentation. Show the app model digest and cover English/Chinese desktop and mobile states.
中文:修正 NVL72 图表说明,使其与实测输入一致。更新实测功耗准入、能耗换算及不同冷却方式的 PUE 说明,显示模型摘要,并覆盖中英文桌面和手机状态。
---
docs/data-transforms.md | 4 +-
docs/powerx-permanent-view.md | 16 ++---
.../cypress/component/scatter-graph.cy.tsx | 69 +++++++++++++++++++
.../components/inference/ui/ChartDisplay.tsx | 16 +++--
packages/app/src/lib/power-basis.ts | 4 +-
5 files changed, 90 insertions(+), 19 deletions(-)
diff --git a/docs/data-transforms.md b/docs/data-transforms.md
index ebf7e541e..22325d3ca 100644
--- a/docs/data-transforms.md
+++ b/docs/data-transforms.md
@@ -79,8 +79,8 @@ Returns `{ chartData: InferenceData[][], hardwareConfig: HardwareConfig }`.
| B4 utility modeled (measured → chassis → PUE) | `modeledSystemPower.deploymentFacilityWatts ÷ modeledSystemPower.gpuCount` | `B1 J/out × (B4 W ÷ B1 W)` | `utilityModeledWatts`, `utilityModeledJPerOutputToken` |
- **N_alloc and total throughput** (`powerBasisNormalization`). Aggregate rows report output per allocated GPU, so `N_alloc` cancels and the builder uses `W ÷ output_tput_per_gpu` without trusting display counts (legacy ingest can encode TP × EP twice). Fixed-sequence disaggregated rows (`disagg && benchmark_type === 'single_turn'`) report output per decode GPU while the deployment also powers the prefill pool, so `total output tok/s = output_tput_per_gpu × num_decode_gpu` and `N_alloc = num_prefill_gpu + num_decode_gpu`; the article's 4P+4D counts eight GPUs, and B3 energy is `(P + D) / D` times `jOutput` on such rows. Other disaggregated benchmark types emit no provisioned energy because it is not verifiable in-app whether AgentX throughput already divides by all GPUs.
-- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. `deploymentFacilityWatts` is chassis AC × PUE with PUE applied exactly once inside `estimateChassisPower`, divided by the physical measured GPU count (not `modeledGpuCount`, which over-counts partially allocated chassis: a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8, and the unit test pins that divisor). This is a different quantity from the existing `modeledChassisPowerPerGpu` (chassis AC ÷ modeled GPU count, no PUE). Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator).
-- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware such as GB200/GB300 NVL72 (`reason: 'hardware'`), and telemetry/topology failures withhold it. The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here).
+- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. B4 W/GPU is `deploymentFacilityWatts ÷ gpuCount`. PUE is applied exactly once to rounded AC power inside `estimateChassisPower` or `estimateRackPower`. `modelSystemPower` combines the resulting chassis or tray shares and attributes partially allocated units to their measured GPUs; NVL72 tray shares come from one rack evaluated at the measured trays’ mean input. Divide the deployment total by the physical measured GPU count, not `modeledGpuCount` (a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8). The existing `modeledChassisPowerPerGpu` instead uses `chassisAcWatts ÷ modeledGpuCount`, without PUE. Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator).
+- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware, and telemetry/topology failures withhold it. GB200/GB300 NVL72 are supported when validated GPU telemetry is accompanied by complete, matching Grace or module telemetry (`reason: 'cpu-telemetry'` when that CPU-side evidence is missing or invalid). The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here).
- **Ordering invariant** (unit-tested on real B200 telemetry): B3 ≥ B4 ≥ B1 and B2 ≥ B1 for both W/GPU and J/out on the same point.
- **Reconstructed prefill energy** (`reconstructedPrefillJPerOutputToken`, `utils/role-energy.ts`). For validated (`power_valid === 1`, schema 2) disaggregated rows, `prefill_joules_per_input_token × (joules_per_output_token ÷ joules_per_input_token)` carries the prefill pool's energy onto the output-token axis: the ratio is the served input:output token count because schema-2 aggregate energy has one numerator. With `decode_joules_per_output_token` it sums back to the deployment's J/out. It feeds only the `i_pcompare=roles` comparison on the energy axis (PowerX Figure 7) and is never a y-axis of its own; aggregate rows and rows missing any of the four inputs omit it.
- Historical Trends substitutes `output_tput_per_gpu := tput_per_gpu` for legacy rows lacking output throughput; provisioned energies in trends inherit that fallback.
diff --git a/docs/powerx-permanent-view.md b/docs/powerx-permanent-view.md
index 4c730b37e..14c3516e4 100644
--- a/docs/powerx-permanent-view.md
+++ b/docs/powerx-permanent-view.md
@@ -23,17 +23,17 @@ group stays out of the selector otherwise. The boundary metrics are members of t
## Boundaries
-| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source |
-| --------------------- | ---------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------- |
-| `gpu-measured` | GPU measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged |
-| `gpu-provisioned` | GPU provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` |
-| `utility-provisioned` | Utility provisioned (all-in) | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) |
-| `utility-modeled` | Utility modeled (PUE) | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` chassis AC × PUE 1.3 (applied once), ÷ measured GPUs |
+| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source |
+| --------------------- | --------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------------------------- |
+| `gpu-measured` | GPU Level Measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged |
+| `gpu-provisioned` | GPU Level Provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` |
+| `utility-provisioned` | All in Provisioned | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) |
+| `utility-modeled` | All in Measured | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` deployment AC ÷ measured GPUs × PUE (1.3 air, 1.1 NVL72; applied once) |
Formulas, the all-GPU normalization (`N_alloc` = prefill + decode GPUs for disaggregated
rows) and the null rules are specified in
-[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis model
-itself in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its
+[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis and NVL72 rack models
+in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its
per-decode-GPU normalization; its labels and the boundary metric's `all GPUs` label keep the
two distinguishable in the selector, the availability list and CSV headers.
diff --git a/packages/app/cypress/component/scatter-graph.cy.tsx b/packages/app/cypress/component/scatter-graph.cy.tsx
index 7ab5903c7..2598d8632 100644
--- a/packages/app/cypress/component/scatter-graph.cy.tsx
+++ b/packages/app/cypress/component/scatter-graph.cy.tsx
@@ -8,6 +8,7 @@ import {
import ScatterGraph from '@/components/inference/ui/ScatterGraph';
import { useParetoHighlightToggle } from '@/components/inference/hooks/useParetoHighlightToggle';
import ChartDisplay from '@/components/inference/ui/ChartDisplay';
+import chartDefinitions from '@/components/inference/metric-registry';
import { mountWithProviders } from '../support/test-utils';
import { expandLegendAdvanced } from '../support/legend-advanced';
import {
@@ -466,6 +467,7 @@ describe('ScatterGraph', () => {
});
it('explains why All in Measured has no points', () => {
+ cy.viewport(1280, 720);
mountWithProviders(
{
'be.visible',
);
cy.contains('No measurements to plot for this selection.').should('not.exist');
+ cy.contains('NVL72 also needs complete Grace or module telemetry').should('be.visible');
+ cy.contains('not NVL72 systems').should('not.exist');
+ cy.screenshot('nvl72-empty-en-desktop', { overwrite: true });
});
it('localizes the All in Measured explanation', () => {
+ cy.viewport(390, 720);
mountWithProviders(
@@ -521,6 +527,9 @@ describe('ScatterGraph', () => {
cy.contains('当前选择没有可用的整体实测功耗数值。').should('be.visible');
cy.contains('当前选择没有可绘制的测量数据。').should('not.exist');
+ cy.contains('NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测').should('be.visible');
+ cy.contains('不含 NVL72 系统').should('not.exist');
+ cy.screenshot('nvl72-empty-zh-mobile', { overwrite: true });
});
for (const selectedYAxisMetric of ['y_tpPerGpu', 'y_measuredPrefillJPerInputToken'] as const) {
@@ -1786,6 +1795,66 @@ describe('ScatterGraph', () => {
});
});
+describe('ChartDisplay modeled power disclosures', () => {
+ for (const locale of ['en', 'zh'] as const) {
+ it(`explains NVL72 telemetry and cooling PUE without hiding its modeled point (${locale})`, () => {
+ const point = createMockInferenceData({
+ hwKey: 'gb200',
+ hw: 'NVIDIA GB200',
+ model: Model.Qwen3_5,
+ y: 1000,
+ utilityModeledWatts: { y: 1000, roof: false },
+ });
+ mountWithProviders(
+
+
+
+
+ ,
+ {
+ inference: {
+ selectedModel: Model.Qwen3_5,
+ selectedYAxisMetric: 'y_utilityModeledWatts',
+ activeHwTypes: new Set(['gb200']),
+ hwTypesWithData: new Set(['gb200']),
+ hardwareConfig: { gb200: { name: 'gb200', label: 'GB200', suffix: '', gpu: 'GB200' } },
+ graphs: [
+ {
+ model: Model.Qwen3_5,
+ sequence: Sequence.EightK_OneK,
+ chartDefinition: chartDefinitions[0],
+ data: [point],
+ },
+ ],
+ },
+ globalFilters: { selectedModel: Model.Qwen3_5 },
+ unofficial: {},
+ },
+ );
+ for (const width of [1280, 390]) {
+ cy.viewport(width, 900);
+ cy.get('[data-testid="power-basis-assumptions"]')
+ .should('be.visible')
+ .and('contain.text', 'PUE 1.3')
+ .and('contain.text', 'PUE 1.1')
+ .and('contain.text', 'Grace')
+ .and('not.contain.text', 'app-sha')
+ .and(
+ 'not.contain.text',
+ locale === 'en'
+ ? 'NVL72 systems (GB200, GB300) and points without values are omitted'
+ : 'NVL72 系统(GB200、GB300)及缺少数值的数据点不绘制',
+ )
+ .should(($note) => {
+ expect($note[0].scrollWidth).to.be.at.most($note[0].clientWidth);
+ });
+ cy.get('.dot-group .visible-shape').should('exist');
+ cy.screenshot(`nvl72-boundary-${locale}-${width}`, { overwrite: true });
+ }
+ });
+ }
+});
+
describe('ChartDisplay responsive status notes', () => {
for (const locale of ['en', 'zh'] as const) {
for (const width of [390, 1440]) {
diff --git a/packages/app/src/components/inference/ui/ChartDisplay.tsx b/packages/app/src/components/inference/ui/ChartDisplay.tsx
index eb8e0a6e3..5a99dbd9e 100644
--- a/packages/app/src/components/inference/ui/ChartDisplay.tsx
+++ b/packages/app/src/components/inference/ui/ChartDisplay.tsx
@@ -14,7 +14,7 @@ import chartDefinitions, {
} from '@/components/inference/metric-registry';
import { metricRowLabel } from '@/components/inference/axis-metric-explanations';
import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config';
-import { AIR_COOLED_SYSTEM_PUE } from '@/lib/modeled-system-power';
+import { AIR_COOLED_SYSTEM_PUE, DLC_SYSTEM_PUE } from '@/lib/modeled-system-power';
import { SYSTEM_POWER_MODEL_REVISION } from '@/lib/system-power-model';
import { ALL_IN_MEASURED_EMPTY, ALL_IN_MEASURED_NOTE } from '@/lib/power-basis';
import {
@@ -121,6 +121,8 @@ import WorkflowInfoDisplay from './WorkflowInfoDisplay';
type InferenceViewMode = 'chart' | 'table';
+const modelRevisionLabel = SYSTEM_POWER_MODEL_REVISION.replace(/^app-sha256:/u, '').slice(0, 12);
+
const STRINGS = {
en: {
inferencePerformance: 'Inference Performance',
@@ -138,14 +140,14 @@ const STRINGS = {
e2eNormIntvtyDisclaimer:
'E2E Normalized Interactivity requires persisted per-request traces, so unofficial-run overlays are unavailable for this experimental view.',
systemPowerAssumptions:
- '8k1k estimate from validated GPU telemetry · CPU/DRAM utilization 20% · Eight-GPU chassis models; a partially allocated chassis is extrapolated to a full chassis at the measured per-GPU power. Chassis AC includes platform overheads; PUE is applied separately for facility power. Click a point for measured GPU power, topology, and power model provenance. Unsupported inputs are omitted.',
+ '8K / 1K estimates from validated telemetry. Eight-GPU chassis assume 20% CPU/DRAM utilization; partial allocations are extrapolated to a full chassis at the measured per-GPU power. NVL72 uses measured Grace or module power plus modeled rack overhead. Facility power applies PUE once: 1.3 for air-cooled chassis, 1.1 for NVL72. Click a point for measurement and model provenance. Unsupported inputs are omitted.',
completedSequenceLengths: (count: string) =>
`Completed requests across all resident points (n=${count})`,
viewMode: 'View mode',
noChartData:
'No benchmark data matches the current model, scenario, and filter selection. Adjust the filters above to see results.',
noSystemPowerData:
- 'No system-power estimates are available for this selection. Choose 8K / 1K with validated GPU telemetry, supported hardware, and known eight-GPU chassis placement. Measured GPU power remains available separately where telemetry exists.',
+ 'No system-power estimates are available for this selection. Choose 8K / 1K with validated GPU telemetry and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Measured GPU power remains available separately where telemetry exists.',
// Boundary disclosures for the derived power axes (lib/power-basis.ts).
// Formulas in words; constants named so a screenshot records its method.
powerBasisAssumptions: {
@@ -153,7 +155,7 @@ const STRINGS = {
'GPU Level Provisioned (TDP) · Watts are the rated TDP per GPU from the hardware registry, so the power curve is flat per hardware. Joules per output token = TDP × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together. Hardware without a published TDP is omitted.',
'utility-provisioned':
'All in Provisioned · Watts are the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model), so the power curve is flat per hardware. Joules per output token = all-in W × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together, unlike the ungated All-in Provisioned J per Output Token, which divides per decode GPU.',
- 'utility-modeled': `All in Measured · Measured GPU power carried through the modeled chassis (CPU, DRAM, platform, PSU losses) to the utility meter: modeled chassis AC × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled, applied once), divided by the measured GPUs; joules per output token scale measured joules by the same ratio. Chassis power model revision ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}. Available for 8K / 1K with validated telemetry on supported hardware only; NVL72 systems (GB200, GB300) and points without values are omitted.`,
+ 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K on supported hardware; incomplete inputs are omitted.`,
},
vsTtft: (word: string) => `vs. ${word} Time To First Token`,
vsE2eLatency: (pctl?: string) =>
@@ -175,18 +177,18 @@ const STRINGS = {
e2eNormIntvtyDisclaimer:
'端到端归一化交互性需要持久化的逐请求 trace 数据,因此该实验性视图不支持非官方运行覆盖。',
systemPowerAssumptions:
- '基于已验证 GPU 遥测的 8k1k 估算 · CPU/DRAM 利用率 20% · 采用八卡机箱模型;仅使用部分 GPU 的机箱按实测每卡功耗外推至满机箱。机箱交流功耗包含平台开销;数据中心功耗另行应用 PUE。点击数据点可查看 GPU 实测功耗、拓扑和功耗模型来源。不支持的输入不绘制。',
+ '基于已验证遥测的 8K / 1K 估算。八卡机箱假设 CPU/DRAM 利用率为 20%;仅使用部分 GPU 时,按实测每卡功耗外推至满机箱。NVL72 使用实测 Grace 或 module 功耗,加上模型估算的机架开销。数据中心功耗只应用一次 PUE:风冷机箱为 1.3,NVL72 为 1.1。点击数据点可查看测量与模型来源。不支持的输入不绘制。',
completedSequenceLengths: (count: string) => `当前所有数据点的已完成请求(n=${count})`,
viewMode: '视图模式',
noChartData: '当前模型、场景与筛选条件下没有匹配的基准测试数据。请调整上方筛选条件查看结果。',
noSystemPowerData:
- '当前选择没有可用的系统功耗估算。请选择 8K / 1K 场景;估算仅覆盖 GPU 遥测已验证、硬件受支持、八卡机箱位置已知的运行。存在遥测数据时,仍可单独查看 GPU 实测功耗。',
+ '当前选择没有可用的系统功耗估算。请选择 8K / 1K 场景,并确保 GPU 遥测已验证、机箱或机架功耗模型受支持。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。存在遥测数据时,仍可单独查看 GPU 实测功耗。',
powerBasisAssumptions: {
'gpu-provisioned':
'GPU 额定功耗(TDP)· 功率取硬件注册表中每 GPU 的额定 TDP,因此每种硬件的功率曲线为水平线。每输出 token 能耗 = TDP × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入。未公布 TDP 的硬件不绘制。',
'utility-provisioned':
'整体预配功耗 · 功率取硬件注册表中每 GPU 的全电源配置(all-in)市电功率(来源:SemiAnalysis Datacenter Industry Model),因此每种硬件的功率曲线为水平线。每输出 token 能耗 = all-in 功率 × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入,这与未加门控的“每输出 token 全电源配置能耗”按 decode GPU 计算不同。',
- 'utility-modeled': `整体实测功耗 · 将 GPU 实测功耗经机箱功耗模型(CPU、DRAM、平台开销、PSU 损耗)推算至市电侧:机箱交流功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷,仅应用一次),再除以实测 GPU 数;每输出 token 能耗按同一比例放大实测能耗。机箱功耗模型版本 ${SYSTEM_POWER_MODEL_REVISION.slice(0, 7)}。仅适用于 8K / 1K、遥测已验证且硬件受支持的运行;NVL72 系统(GB200、GB300)及缺少数值的数据点不绘制。`,
+ 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。仅适用于受支持硬件的 8K / 1K 场景,输入不完整的数据点不绘制。`,
},
vsTtft: (word: string) => `vs. ${word === 'Median' ? '中位' : word} 首 token 延迟(TTFT)`,
vsE2eLatency: (pctl?: string) => (pctl ? `vs. ${pctl} 端到端延迟` : 'vs. 端到端延迟'),
diff --git a/packages/app/src/lib/power-basis.ts b/packages/app/src/lib/power-basis.ts
index 5b5f6e61b..5a30ab8c6 100644
--- a/packages/app/src/lib/power-basis.ts
+++ b/packages/app/src/lib/power-basis.ts
@@ -40,8 +40,8 @@ export const ALL_IN_MEASURED_NOTE = {
};
export const ALL_IN_MEASURED_EMPTY = {
- en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K, validated GPU telemetry, and hardware covered by the chassis power model (not NVL72 systems). Choose another boundary to keep the points.',
- zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 场景、已验证的 GPU 遥测,且硬件在机箱功耗模型覆盖范围内(不含 NVL72 系统)。可切换到其他功耗边界以保留数据点。',
+ en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K, validated GPU telemetry, and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Choose another boundary to keep the points.',
+ zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 场景、已验证的 GPU 遥测,以及受支持的机箱或机架功耗模型。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。可切换到其他功耗边界以保留数据点。',
};
/** InferenceData keys per derived basis and quantity. B1 lives on the measured* fields. */
From 3278b48bde386d43e11b98709e757da98b955a66 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 22:24:50 -0700
Subject: [PATCH 16/22] fix: keep dense profit comparison charts readable
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Reserve readable bar spacing, retain the SVG across scrolling thresholds, and export the full plot without moving the live chart.
中文:为密集利润对比图保留足够的柱形间距,在窄屏中启用图内横向滚动,并保持完整 PNG 导出与固定说明区域。切换滚动状态时保留 SVG,避免临界宽度下图形消失。
---
docs/dashboard-readonly-views.md | 4 +
.../component/profit-estimator-chart.cy.tsx | 38 ++++-
.../app/cypress/e2e/profit-estimator.cy.ts | 151 +++++++++++++++++-
.../calculator/ProfitEstimatorChart.tsx | 31 +++-
.../src/components/ui/d3-chart-wrapper.tsx | 147 ++++++++++-------
packages/app/src/hooks/useChartExport.test.ts | 31 ++++
packages/app/src/hooks/useChartExport.ts | 8 +-
.../app/src/lib/d3-chart/D3Chart/D3Chart.tsx | 2 +
.../app/src/lib/d3-chart/D3Chart/types.ts | 2 +
9 files changed, 339 insertions(+), 75 deletions(-)
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index c65916275..e5fe639a8 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -68,6 +68,10 @@ Provisioned, and All in Measured. The last combines measured GPU power with mode
unmeasured components and PUE; it is not a wall-meter measurement. These labels and
collapsed power-assumption/availability notes do not change metric IDs, API selectors,
or calculations. Profit comparison `powerLabel` display text follows the same names.
+Dense profit charts reserve readable space per bar and scroll within the plot on narrow
+screens; captions and controls stay fixed. This is presentation-only: API selectors,
+calculations, source identities and CSV rows are unchanged. PNG export includes the full
+plot regardless of its current scroll position, so no API or skills contract change is needed.
The GPU statistics table includes startup and warmup for all chips in the selected
series, regardless of chip visibility. It is separate from serving-window power,
J/token and selected-time-window calculations. Run telemetry is DB-first with an
diff --git a/packages/app/cypress/component/profit-estimator-chart.cy.tsx b/packages/app/cypress/component/profit-estimator-chart.cy.tsx
index be46f8df6..16da8ef3e 100644
--- a/packages/app/cypress/component/profit-estimator-chart.cy.tsx
+++ b/packages/app/cypress/component/profit-estimator-chart.cy.tsx
@@ -58,7 +58,7 @@ function colorForRow(row: ProfitEstimatorRow): string {
function mountChart(widthPx: number, rows: ProfitEstimatorRow[] = ROWS) {
cy.mount(
-
+
{
});
}
-/** Left-to-right boxes of every revenue figure and margin line above the bars. */
+/** Combined revenue and margin bounds for each bar, in left-to-right order. */
function labelBoxes(): Cypress.Chainable {
return cy
- .get('[data-testid="profit-estimator-chart"] .revenue-label tspan')
+ .get('[data-testid="profit-estimator-chart"] .revenue-label')
.then(($tspans) =>
[...$tspans]
.filter((el) => (el.textContent ?? '') !== '')
@@ -102,8 +102,31 @@ function overlaps(a: DOMRect, b: DOMRect): boolean {
}
describe('ProfitEstimatorChart revenue labels', () => {
+ it('preserves the rendered SVG at the exact scrolling threshold', () => {
+ cy.viewport(1280, 900);
+ mountChart(390, ROWS.slice(0, 3));
+ cy.get('[data-testid="d3-chart-svg"]').then(($svg) => {
+ const svg = $svg[0];
+ const width = svg.getBoundingClientRect().width;
+ cy.get('[data-testid="profit-chart-container"]').invoke('css', 'width', `${width + 32}px`);
+ cy.get('[data-chart-scroll][tabindex]').should('not.exist');
+ cy.get('[data-testid="d3-chart-svg"]').should(($current) => {
+ expect($current[0], 'same SVG at the threshold').to.equal(svg);
+ expect($current[0].getBoundingClientRect().width).to.equal(width);
+ expect($current.find('.revenue-label')).to.have.length(3);
+ });
+ cy.get('[data-chart-scroll]').should('not.have.attr', 'tabindex');
+ cy.get('[data-testid="profit-chart-container"]').invoke('css', 'width', `${width + 31}px`);
+ cy.get('[data-chart-scroll]').should('have.attr', 'tabindex', '0');
+ cy.get('[data-testid="d3-chart-svg"]').should(($current) => {
+ expect($current[0], 'same SVG below the threshold').to.equal(svg);
+ expect($current.find('.revenue-label')).to.have.length(3);
+ });
+ });
+ });
+
it('never lets neighbouring revenue figures overlap on a phone', () => {
- // iPhone 15/16 CSS viewport; the card padding leaves the chart ~361px.
+ // iPhone 15/16 viewport: the 361px card viewport scrolls across the wider plot.
cy.viewport(393, 852);
mountChart(393);
cy.screenshot('profit-estimator-chart-phone', { overwrite: true });
@@ -125,9 +148,12 @@ describe('ProfitEstimatorChart revenue labels', () => {
mountChart(393, LOSING_ROWS);
cy.get('[data-testid="profit-estimator-chart"] .loss-label').should('have.length', 2);
cy.screenshot('profit-estimator-chart-phone-loss', { overwrite: true });
- // The wide figure loses the word and its decimal; the narrow one keeps its decimal.
+ // The scrollable plot has room to retain the loss label as well as the signed amount.
cy.get('[data-testid="profit-estimator-chart"] .loss-label').then(($labels) => {
- expect([...$labels].map((el) => el.textContent)).to.deep.equal(['-$561M', '-$4.9B']);
+ expect([...$labels].map((el) => el.textContent)).to.deep.equal([
+ 'Loss -$561M',
+ 'Loss -$4.9B',
+ ]);
});
cy.get('[data-testid="profit-estimator-chart"] rect.bar')
.first()
diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts
index 0e704dc00..8a5407e18 100644
--- a/packages/app/cypress/e2e/profit-estimator.cy.ts
+++ b/packages/app/cypress/e2e/profit-estimator.cy.ts
@@ -381,12 +381,14 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
chart().find('text.revenue-label').should('have.length', 9);
chart().scrollIntoView();
cy.screenshot('profit-nvl72-compare-mobile', { capture: 'viewport', overwrite: true });
- chartSvg().scrollIntoView().should('be.visible');
- chartSvg().then(($svg) => {
- const bounds = $svg[0].getBoundingClientRect();
- expect(bounds.left).to.be.at.least(0);
- expect(bounds.right).to.be.at.most(393);
- });
+ chart().find('[data-chart-scroll]').scrollIntoView().scrollTo('left').should('be.visible');
+ chart()
+ .find('[data-chart-scroll]')
+ .then(($scroll) => {
+ const bounds = $scroll[0].getBoundingClientRect();
+ expect(bounds.left).to.be.at.least(0);
+ expect(bounds.right).to.be.at.most(393);
+ });
cy.screenshot('profit-nvl72-chart-mobile', { capture: 'viewport', overwrite: true });
cy.get('[data-testid="export-button"]').first().click();
cy.get('[data-testid="export-csv-button"]').click();
@@ -407,6 +409,143 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
});
});
+ for (const locale of ['en', 'zh'] as const) {
+ it(`keeps all nine power comparison bars readable and reachable on mobile (${locale})`, () => {
+ stubOpenRouter();
+ cy.viewport(390, 900);
+ cy.intercept('GET', '/api/v1/benchmarks*', {
+ body: [
+ ...profitBenchmarkRows().map((row) => ({
+ ...row,
+ metrics: {
+ ...row.metrics,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ },
+ })),
+ ...profitNvl72Rows(),
+ ],
+ });
+ cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=compare`, {
+ onBeforeLoad: unlockPowerGate,
+ });
+ chart()
+ .find('text.revenue-label')
+ .should('have.length', 9)
+ .should(($labels) => {
+ const boxes = [...$labels].map((label) => label.getBoundingClientRect());
+ for (let i = 1; i < boxes.length; i++) {
+ expect(
+ boxes[i].left - boxes[i - 1].right,
+ 'space between revenue and margin labels',
+ ).to.be.at.least(4);
+ }
+ });
+ chart()
+ .find('image.bar-vendor-mark')
+ .should(($marks) => {
+ const boxes = [...$marks].map((mark) => mark.getBoundingClientRect());
+ for (let i = 1; i < boxes.length; i++) {
+ expect(boxes[i].left - boxes[i - 1].right, 'space between vendor marks').to.be.at.least(
+ 4,
+ );
+ }
+ });
+ const scroller = () => chart().find('[data-chart-scroll]');
+ scroller().should('have.attr', 'tabindex', '0').and('have.attr', 'role', 'region');
+ scroller()
+ .scrollIntoView({ offset: { top: -70, left: 0 } })
+ .focus()
+ .should('have.focus');
+ scroller().then(($scroll) => {
+ const el = $scroll[0];
+ const event = new el.ownerDocument.defaultView!.KeyboardEvent('keydown', {
+ key: 'ArrowRight',
+ bubbles: true,
+ cancelable: true,
+ });
+ el.dispatchEvent(event);
+ expect(
+ event.defaultPrevented,
+ `handled key on ${el.clientWidth}/${el.scrollWidth}`,
+ ).to.equal(true);
+ expect(el.scrollLeft, 'synchronous scroll').to.be.greaterThan(0);
+ });
+ scroller().should(($scroll) => expect($scroll[0].scrollLeft).to.be.greaterThan(0));
+ scroller().scrollTo('right');
+ chart().find('rect.bar-profit').last().should('be.visible');
+ chart().find('rect.bar-profit').last().click({ scrollBehavior: false });
+ cy.get('[data-chart-tooltip="profit-estimator-chart"]')
+ .should('be.visible')
+ .and('contain', 'MI355X');
+ // Clear the pinned tooltip for screenshots; the SVG center is outside the scroll viewport.
+ chartSvg().trigger('click', { force: true, scrollBehavior: false });
+ cy.get('[data-chart-tooltip="profit-estimator-chart"]').should('not.be.visible');
+ cy.document().should((doc) => {
+ expect(doc.documentElement.scrollWidth).to.be.at.most(390);
+ });
+ cy.get('[data-testid="profit-caption"]').should(($caption) => {
+ const bounds = $caption[0].getBoundingClientRect();
+ expect(bounds.left).to.be.at.least(0);
+ expect(bounds.right).to.be.at.most(390);
+ });
+ scroller()
+ .scrollIntoView({ offset: { top: -70, left: 0 } })
+ .scrollTo('left');
+ chart().find('rect.bar-tco').first().should('be.visible');
+ cy.screenshot(`profit-dense-${locale}-mobile-left`, { capture: 'viewport', overwrite: true });
+ scroller().scrollTo('right');
+ cy.screenshot(`profit-dense-${locale}-mobile-right`, {
+ capture: 'viewport',
+ overwrite: true,
+ });
+ let exportedPng = '';
+ let exportedLabels: string[] = [];
+ cy.window().then((win) => {
+ win.HTMLAnchorElement.prototype.click = function () {
+ if (this.download.endsWith('.png')) {
+ exportedPng = this.href;
+ exportedLabels = [
+ ...win.document.querySelectorAll('#profit-estimator-chart-export text.revenue-label'),
+ ].map((label) => label.textContent ?? '');
+ }
+ };
+ });
+ cy.get('[data-testid="export-button"]').first().click();
+ cy.get('[data-testid="export-png-button"]').click();
+ cy.window()
+ .should(() => expect(exportedPng).to.match(/^data:image\/png;base64,/u))
+ .then((win) => {
+ expect(exportedLabels).to.have.length(9);
+ cy.writeFile(
+ `cypress/downloads/profit-dense-${locale}.png`,
+ exportedPng.split(',')[1],
+ 'base64',
+ );
+ return new Cypress.Promise((resolve, reject) => {
+ const png = new win.Image();
+ png.addEventListener('load', () => {
+ expect(png.naturalWidth).to.be.greaterThan(1500);
+ resolve();
+ });
+ png.addEventListener('error', () => reject(new Error('Profit PNG did not decode')));
+ png.src = exportedPng;
+ });
+ });
+ scroller().should(($scroll) => expect($scroll[0].scrollLeft).to.be.greaterThan(0));
+ cy.viewport(1280, 900);
+ scroller().should('not.have.attr', 'tabindex');
+ scroller().should(($scroll) => {
+ expect($scroll[0].scrollWidth).to.equal($scroll[0].clientWidth);
+ });
+ chart().find('text.revenue-label').should('have.length', 9);
+ chartSvg().scrollIntoView({ offset: { top: -70, left: 0 } });
+ cy.screenshot(`profit-dense-${locale}-desktop`, { capture: 'viewport', overwrite: true });
+ });
+ }
+
it('keeps the benchmark settings and restores the original chart after unavailable power', () => {
stubOpenRouter();
cy.visit('/profit-estimator-per-gigawatt', { onBeforeLoad: unlockPowerGate });
diff --git a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
index 5b9ff0a7d..5d0d0d06b 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorChart.tsx
@@ -97,6 +97,8 @@ const GLYPH_WIDTH_EM = 0.55;
const LABEL_SIDE_PAD_PX = 4;
/** The labels above a bar may borrow this much of the gap to each neighbour, in px. */
const X_GAP_ALLOWANCE = 12;
+/** Keep comparison labels and vendor marks readable when bars outgrow the viewport. */
+const MIN_BAR_STEP_PX = 96;
/**
* Horizontal room the labels above a bar may use, in px. A label may overhang
@@ -274,6 +276,7 @@ const STRINGS = {
'measured GPU board + Grace socket (Grace socket sensor), regulator loss modeled',
},
noData: 'No SKU can be priced for the current selection.',
+ scrollHint: 'Scroll horizontally to view the full chart.',
},
zh: {
yAxisModeled: '每吉瓦设施总功耗对应的年收入(美元)',
@@ -307,6 +310,7 @@ const STRINGS = {
'实测 GPU 板卡 + Grace socket 功耗(Grace socket 传感器),稳压损耗由模型估算',
},
noData: '当前选择下没有可定价的 SKU。',
+ scrollHint: '横向滚动查看完整图表。',
},
} as const;
@@ -1032,12 +1036,22 @@ export default function ProfitEstimatorChart({
);
const baseMargin = compact ? CHART_MARGIN_COMPACT : CHART_MARGIN;
+ const minimumMargin = slantedMargins(
+ [...labelMap.values()],
+ MIN_BAR_STEP_PX,
+ CHART_TYPE.axisLabelSub,
+ baseMargin,
+ );
+ const chartWidth = Math.max(
+ dimensions.width,
+ rows.length * MIN_BAR_STEP_PX + minimumMargin.left + minimumMargin.right,
+ );
// Upright two-line labels when each SKU has room for them; slanted otherwise.
const labelLayout = useMemo(() => {
- const plotWidth = dimensions.width - baseMargin.left - baseMargin.right;
+ const plotWidth = chartWidth - baseMargin.left - baseMargin.right;
const slot = rows.length > 0 ? plotWidth / rows.length : 0;
return xLabelLayout([...labelMap.values()], slot, CHART_TYPE.axisLabelSub);
- }, [dimensions.width, baseMargin, rows.length, labelMap]);
+ }, [chartWidth, baseMargin, rows.length, labelMap]);
const margin = useMemo(() => {
if (labelLayout === 'stacked') {
const dated = [...labelMap.values()].some((label) => splitHistoryLabel(label)[1] !== '');
@@ -1046,15 +1060,15 @@ export default function ProfitEstimatorChart({
bottom: X_LABEL_STACKED_BOTTOM + (dated ? X_LABEL_HISTORY_LINE_PX : 0),
};
}
- const plotWidth = dimensions.width - baseMargin.left - baseMargin.right;
+ const plotWidth = chartWidth - baseMargin.left - baseMargin.right;
const slot = rows.length > 0 ? plotWidth / rows.length : 0;
return slantedMargins([...labelMap.values()], slot, CHART_TYPE.axisLabelSub, baseMargin);
- }, [baseMargin, labelLayout, dimensions.width, rows.length, labelMap]);
+ }, [baseMargin, labelLayout, chartWidth, rows.length, labelMap]);
const plotHeight = chartHeight - margin.top - margin.bottom;
// The vendor mark grows with the bar, so the headroom above the tallest stack
// has to be sized from the same band width the renderer will see.
const bandwidth = useMemo(() => {
- const plotWidth = dimensions.width - margin.left - margin.right;
+ const plotWidth = chartWidth - margin.left - margin.right;
if (plotWidth <= 0 || rows.length === 0) return 0;
return d3
.scaleBand()
@@ -1062,7 +1076,7 @@ export default function ProfitEstimatorChart({
.range([0, plotWidth])
.padding(BAND_PADDING)
.bandwidth();
- }, [dimensions.width, margin, rows]);
+ }, [chartWidth, margin, rows]);
const yDomain = useMemo(
() => profitYDomain(rows, plotHeight, stackHeadroomPx(barMarkHeight(bandwidth))),
[rows, plotHeight, bandwidth],
@@ -1189,6 +1203,11 @@ export default function ProfitEstimatorChart({
instructions=""
legendElement={legendElement}
caption={caption}
+ scrollablePlot={{
+ minWidth: chartWidth,
+ label: t.scrollHint,
+ enabled: chartWidth > dimensions.width,
+ }}
/>
);
diff --git a/packages/app/src/components/ui/d3-chart-wrapper.tsx b/packages/app/src/components/ui/d3-chart-wrapper.tsx
index 1d91d5103..5880cb96e 100644
--- a/packages/app/src/components/ui/d3-chart-wrapper.tsx
+++ b/packages/app/src/components/ui/d3-chart-wrapper.tsx
@@ -65,6 +65,7 @@ export interface D3ChartWrapperProps {
instructions?: string;
testId?: string;
grabCursor?: boolean;
+ scrollablePlot?: { minWidth: number; label: string; enabled: boolean };
}
export function D3ChartWrapper({
@@ -83,83 +84,117 @@ export function D3ChartWrapper({
instructions,
testId,
grabCursor = true,
+ scrollablePlot,
}: D3ChartWrapperProps) {
const locale = useLocale();
const resolvedInstructions = instructions ?? DEFAULT_CHART_INSTRUCTIONS[locale];
- return (
-
- {caption &&
{caption} }
-
-
-
- {/* Stable hook for tests. `[data-testid="scatter-graph"] svg` also
+ const plot = (
+
+
+
+ {/* Stable hook for tests. `[data-testid="scatter-graph"] svg` also
matches every Lucide icon inside the card — dozens of them —
so picking "the first svg" silently grabs an icon whenever the
selected metric renders one above the chart. */}
-
{
- (e.currentTarget as SVGSVGElement).style.cursor = 'grabbing';
- }
- : undefined
- }
- onMouseUp={
- grabCursor
- ? (e) => {
- (e.currentTarget as SVGSVGElement).style.cursor = 'grab';
- }
- : undefined
+ {
+ (e.currentTarget as SVGSVGElement).style.cursor = 'grabbing';
+ }
+ : undefined
+ }
+ onMouseUp={
+ grabCursor
+ ? (e) => {
+ (e.currentTarget as SVGSVGElement).style.cursor = 'grab';
+ }
+ : undefined
+ }
+ onClick={() => {
+ if (isPinned()) {
+ dismissTooltip();
+ hideTooltipElements(tooltipRef, svgRef);
}
- onClick={() => {
- if (isPinned()) {
- dismissTooltip();
- hideTooltipElements(tooltipRef, svgRef);
- }
- }}
- />
- {/* Tooltip is portalled to with position:fixed so it can
+ }}
+ />
+ {/* Tooltip is portalled to with position:fixed so it can
rise above sibling chart cards' stacking contexts. The d3 layer
writes viewport-coords into style.left/top — see
computeTooltipPosition. */}
-
- {noDataOverlay}
-
- {resolvedInstructions && (
-
- {resolvedInstructions}
-
- )}
-
+
+ {noDataOverlay}
- {legendElement && (
- /* Sizes to the legend content: when the sidebar legend panel is open
+ {resolvedInstructions && (
+
+ {resolvedInstructions}
+
+ )}
+
+
+ {legendElement && (
+ /* Sizes to the legend content: when the sidebar legend panel is open
(.sidebar-legend present) the column grows to fit the widest
legend label (capped) so full names display without truncation,
while still sitting next to the plot without overlapping it; when
closed the legend renders only a small reopen button and the
chart reclaims the width. Height belongs to the legend itself:
short lists should not reserve an empty chart-height column. */
+
+ {legendElement}
+
+ )}
+
+ );
+
+ return (
+
+ {caption &&
{caption} }
+ {scrollablePlot ? (
+ <>
+ {scrollablePlot.enabled && (
+
{scrollablePlot.label}
+ )}
{
+ if (!scrollablePlot.enabled) return;
+ if (event.target !== event.currentTarget) return;
+ if (event.key !== 'ArrowLeft' && event.key !== 'ArrowRight') return;
+ event.preventDefault();
+ event.currentTarget.scrollBy({
+ left: (event.key === 'ArrowRight' ? 1 : -1) * event.currentTarget.clientWidth * 0.8,
+ });
+ }}
>
- {legendElement}
+ {plot}
- )}
-
+ >
+ ) : (
+ plot
+ )}
);
}
diff --git a/packages/app/src/hooks/useChartExport.test.ts b/packages/app/src/hooks/useChartExport.test.ts
index 0c28af0cf..2fb740f01 100644
--- a/packages/app/src/hooks/useChartExport.test.ts
+++ b/packages/app/src/hooks/useChartExport.test.ts
@@ -131,6 +131,37 @@ describe('useChartExport failure messages', () => {
},
);
+ it.each([false, true])(
+ 'exports the whole plot with scrolling=%s without moving the live chart',
+ async (scrolling) => {
+ const plot =
+ '
';
+ chart.innerHTML = `
Power comparison ${scrolling ? `
${plot}
` : plot}`;
+ const liveScroller = chart.querySelector
('[data-chart-scroll]');
+ if (liveScroller) liveScroller.scrollLeft = 250;
+ const original = chart.innerHTML;
+ let snapshot: HTMLElement;
+ exportMocks.toPng.mockImplementationOnce((element: HTMLElement) => {
+ snapshot = element.cloneNode(true) as HTMLElement;
+ throw new Error('stop after capture');
+ });
+ vi.spyOn(window, 'alert').mockImplementation(() => {});
+ vi.spyOn(console, 'error').mockImplementation(() => {});
+ await act(() => current.exportToImage());
+ expect(exportMocks.toPng).toHaveBeenCalledOnce();
+ expect(snapshot!.textContent).toContain('Power comparison');
+ expect(snapshot!.textContent).toContain('First bar');
+ expect(snapshot!.textContent).toContain('Last bar');
+ const exportedScroller = snapshot!.querySelector('[data-chart-scroll]');
+ if (scrolling) {
+ expect(exportedScroller?.style.overflow).toBe('visible');
+ expect(exportedScroller?.scrollLeft).toBe(0);
+ } else expect(exportedScroller).toBeNull();
+ expect(chart.innerHTML).toBe(original);
+ if (liveScroller) expect(liveScroller.scrollLeft).toBe(250);
+ },
+ );
+
it('includes the legend again after line labels are toggled off', async () => {
chart.innerHTML =
'';
diff --git a/packages/app/src/hooks/useChartExport.ts b/packages/app/src/hooks/useChartExport.ts
index 514d7e82a..a105c20a6 100644
--- a/packages/app/src/hooks/useChartExport.ts
+++ b/packages/app/src/hooks/useChartExport.ts
@@ -356,7 +356,13 @@ export function useChartExport({
// Layout: force side-by-side flex row for export
applyStyles(exportElement, { width: 'fit-content', overflow: 'visible', padding: '16px' });
- const flexContainer = clone.querySelector(':scope > .flex') as HTMLElement | null;
+ for (const scroller of clone.querySelectorAll('[data-chart-scroll]')) {
+ applyStyles(scroller, { width: 'fit-content', overflow: 'visible' });
+ scroller.scrollLeft = 0;
+ }
+ const flexContainer = clone.querySelector(
+ ':scope > .flex, [data-chart-scroll] > .flex',
+ ) as HTMLElement | null;
applyStyles(flexContainer, {
flexDirection: 'row',
width: 'fit-content',
diff --git a/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx b/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx
index 6553b1611..21bf95998 100644
--- a/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx
+++ b/packages/app/src/lib/d3-chart/D3Chart/D3Chart.tsx
@@ -39,6 +39,7 @@ function D3ChartInner(
legendElement,
noDataOverlay,
caption,
+ scrollablePlot,
onRender,
onDisplayUpdate,
}: D3ChartProps,
@@ -159,6 +160,7 @@ function D3ChartInner(
legendElement={legendElement}
noDataOverlay={noDataOverlay}
caption={caption}
+ scrollablePlot={scrollablePlot}
/>
);
}
diff --git a/packages/app/src/lib/d3-chart/D3Chart/types.ts b/packages/app/src/lib/d3-chart/D3Chart/types.ts
index e60543d0d..e8de3b7b9 100644
--- a/packages/app/src/lib/d3-chart/D3Chart/types.ts
+++ b/packages/app/src/lib/d3-chart/D3Chart/types.ts
@@ -294,6 +294,8 @@ export interface D3ChartProps {
legendElement?: React.ReactNode;
noDataOverlay?: React.ReactNode;
caption?: React.ReactNode;
+ /** Keep the SVG mounted so D3 retains its layout when scrolling toggles. */
+ scrollablePlot?: { minWidth: number; label: string; enabled: boolean };
/** Called after all layers render. Useful for one-off DOM manipulations. */
onRender?: (ctx: RenderContext) => void;
From 84b9d53d932aec11e718dd87dfa3c0466b5abe6d Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Wed, 30 Sep 2026 22:38:57 -0700
Subject: [PATCH 17/22] fix: distinguish unpriced profit estimates from missing
power
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Group unavailable notices by the exact retained provisioned result, preserving configuration and history identity without changing estimates or API data.
中文:按实际保留的预配置功耗估算区分未定价 SKU 与缺少实测加建模结果的提示,避免将没有基准估算的配置误标为仅缺少实测功耗。保留现有 GPU 与 Grace 功耗回退行为。
---
docs/dashboard-readonly-views.md | 3 +
.../app/cypress/e2e/profit-estimator.cy.ts | 55 +++++++++++++++++++
.../calculator/ProfitEstimatorDisplay.tsx | 37 ++++++++-----
3 files changed, 80 insertions(+), 15 deletions(-)
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index e5fe639a8..6cea316de 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -72,6 +72,9 @@ Dense profit charts reserve readable space per bar and scroll within the plot on
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
plot regardless of its current scroll position, so no API or skills contract change is needed.
+Unavailable-estimate notices distinguish unpriced SKUs from missing measured-plus-modeled
+estimates using the retained provisioned result identity. This explanatory grouping preserves
+the API's existing rows, skip reasons, selectors and calculations.
The GPU statistics table includes startup and warmup for all chips in the selected
series, regardless of chip visibility. It is separate from serving-window power,
J/token and selected-time-window calculations. Run telemetry is DB-first with an
diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts
index 8a5407e18..4e7aef39a 100644
--- a/packages/app/cypress/e2e/profit-estimator.cy.ts
+++ b/packages/app/cypress/e2e/profit-estimator.cy.ts
@@ -110,6 +110,61 @@ function assertDisclosureOpen(testId: string, open: boolean) {
// Clear the preceding chart before each case changes the viewport.
describe('Profit estimator power option', { testIsolation: true }, () => {
for (const locale of ['en', 'zh'] as const) {
+ it(`distinguishes unpriced SKUs from missing measured-power estimates (${locale})`, () => {
+ stubOpenRouter();
+ const width = locale === 'en' ? 1280 : 390;
+ cy.viewport(width, 900);
+ cy.intercept('GET', '/api/v1/benchmarks*', {
+ body: profitBenchmarkRows().map((row) => ({
+ ...row,
+ metrics: {
+ ...row.metrics,
+ power_valid: row.hardware === 'b300' ? 0 : 1,
+ power_metric_schema_version: 2,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ },
+ })),
+ });
+ cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=compare`, {
+ onBeforeLoad: unlockPowerGate,
+ });
+ chart().find('text.revenue-label').should('have.length', 6);
+ chart()
+ .find('.x-axis')
+ .should('not.contain', 'H200')
+ .and('contain', 'B300')
+ .and('contain', 'GB300');
+ const unpriced = locale === 'en' ? 'Not priced:' : '未定价:';
+ const measured =
+ locale === 'en' ? 'Measured + modeled unavailable:' : '实测加建模估算不可用:';
+ cy.get('[data-testid="profit-power-unavailable"] > summary').click();
+ cy.get('[data-testid="profit-power-unavailable"] > p')
+ .should('be.visible')
+ .should(($notice) => {
+ const [baseline, measurement] = $notice.text().split(measured);
+ expect(baseline).to.contain(unpriced).and.to.contain('H200');
+ expect(baseline).not.to.contain('B300');
+ expect(measurement).to.contain('B300').and.to.contain('GB300');
+ expect(measurement).not.to.contain('H200');
+ expect(measurement).to.contain(
+ locale === 'en' ? 'no usable measured power' : '同一组基准测试数据点缺少有效功耗',
+ );
+ expect(measurement).to.contain(
+ locale === 'en'
+ ? 'missing complete Grace or module power'
+ : '缺少完整的 Grace 或 module 功耗',
+ );
+ const bounds = $notice[0].getBoundingClientRect();
+ expect(bounds.left).to.be.at.least(0);
+ expect(bounds.right).to.be.at.most(width);
+ });
+ cy.get('[data-testid="profit-power-unavailable"]').scrollIntoView({
+ offset: { top: -70, left: 0 },
+ });
+ cy.screenshot(`profit-unavailable-basis-${locale}`, { capture: 'viewport', overwrite: true });
+ });
+
it(`prices DeepSeek Flash partial chassis with a one-line power note and CSV labels (${locale})`, () => {
stubOpenRouter();
cy.viewport(locale === 'en' ? 1280 : 393, 900);
diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
index 166d16d0a..ea8122428 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
@@ -1371,21 +1371,28 @@ function ProfitEstimatorInner({
historyCurrentRunIds,
]);
- const powerUnavailable = useMemo(
- () =>
- (powerBasis === 'compare' ? t.modeledUnavailable : t.skipped)(
- fullEstimate.skipped
- .map((row) => {
- const label = rowLabel(
- { ...row, dateLabel: row.date ? historyEntryLabel(row.date) : undefined },
- hardwareConfig,
- );
- return `${label}: ${t.skipReason[row.reason]}`;
- })
- .join('; '),
- ),
- [fullEstimate.skipped, hardwareConfig, historyEntryLabel, powerBasis, t],
- );
+ const powerUnavailable = useMemo(() => {
+ const unpriced: string[] = [];
+ const measuredUnavailable: string[] = [];
+ for (const row of fullEstimate.skipped) {
+ const label = rowLabel(
+ { ...row, dateLabel: row.date ? historyEntryLabel(row.date) : undefined },
+ hardwareConfig,
+ );
+ const entries =
+ powerBasis === 'compare' &&
+ fullEstimate.rows.some((priced) => priced.resultKey === `${row.resultKey}__provisioned`)
+ ? measuredUnavailable
+ : unpriced;
+ entries.push(`${label}: ${t.skipReason[row.reason]}`);
+ }
+ return [
+ unpriced.length > 0 ? t.skipped(unpriced.join('; ')) : '',
+ measuredUnavailable.length > 0 ? t.modeledUnavailable(measuredUnavailable.join('; ')) : '',
+ ]
+ .filter(Boolean)
+ .join(' ');
+ }, [fullEstimate, hardwareConfig, historyEntryLabel, powerBasis, t]);
const powerBasisNotes = useMemo(() => {
const notes = new Map();
From c80f7f3e0c185206df33de77cf392bed21ee5a6f Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Thu, 1 Oct 2026 12:51:44 -0700
Subject: [PATCH 18/22] fix: retain AgentX estimates in all-in power views
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Enable the existing AgentX model option in shared chart transforms while
keeping the standalone chassis AC metric scoped to 8K/1K. Preserve valid
multi-node rows across charts, tables, history, overlays and the public view.
简体中文:整体实测功耗复用现有 AgentX 模型选项,保留符合条件的多节点
曲线与表格记录;独立机箱交流功耗指标仍仅适用于 8K/1K。精简中英文指南,
说明操作步骤、遥测要求和模型假设,移除 PR 与交付历史。
---
docs/dashboard-readonly-views.md | 6 +
docs/index.md | 2 +-
docs/powerx-system-power.md | 845 +++++-------------
docs/powerx-system-power.zh.md | 717 ++++-----------
packages/app/cypress/e2e/powerx-compare.cy.ts | 138 +++
.../components/inference/ui/ChartDisplay.tsx | 21 +-
.../inference/utils/tooltip-utils.test.ts | 20 +
.../inference/utils/tooltipUtils.ts | 10 +-
packages/app/src/lib/api-route-catalog.ts | 2 +-
packages/app/src/lib/benchmark-transform.ts | 2 +-
packages/app/src/lib/chart-utils.ts | 8 +-
packages/app/src/lib/power-basis.test.ts | 113 ++-
packages/app/src/lib/power-basis.ts | 14 +-
.../app/src/lib/views-api/docs/inference.ts | 4 +-
.../skills/inferencex-api/integrity.json | 2 +-
.../references/dashboard-views.md | 6 +
16 files changed, 744 insertions(+), 1166 deletions(-)
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index 6cea316de..3f98dd8eb 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -68,6 +68,12 @@ Provisioned, and All in Measured. The last combines measured GPU power with mode
unmeasured components and PUE; it is not a wall-meter measurement. These labels and
collapsed power-assumption/availability notes do not change metric IDs, API selectors,
or calculations. Profit comparison `powerLabel` display text follows the same names.
+All in Measured watts and energy accept validated 8K/1K and AgentX rows through the
+shared chart/API transform, including historical and unofficial rows. AgentX reuses
+the chassis or rack model without independent workload calibration. Telemetry and
+topology gates still apply; NVL72 needs complete Grace or module power. The standalone
+Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope.
+
Dense profit charts reserve readable space per bar and scroll within the plot on narrow
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
diff --git a/docs/index.md b/docs/index.md
index a3f5427bd..a6ce4bb46 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -8,7 +8,7 @@ Design rationale and non-obvious conventions. See [CLAUDE.md](../CLAUDE.md) for
- [Pareto Boundary API](./pareto-api.md): Query frontier and hinterland observations, preserve provenance, and distinguish API scope from chart highlights.
- [API Skill Examples](./inferencex-api-examples.md) — Install the public skill, query benchmarks, export measured PowerX data, and explain empty results
-- [PowerX System Power](./powerx-system-power.md) / [简体中文](./powerx-system-power.zh.md) — Smart provisioning walkthrough, Kimi K3 missing-row diagnosis, NVL72/chassis models, assumptions, and reproducible exports
+- [PowerX System Power](./powerx-system-power.md) / [简体中文](./powerx-system-power.zh.md) — Measured-curve steps, chassis/NVL72 requirements, a worked planning example, missing-data diagnosis, and reproducible exports
- [PowerX Permanent View](./powerx-permanent-view.md) — Power boundaries as gated Measured Energy metrics, `i_metric`/`i_rulers` share links, missing-value states
- [PowerX Persistence and Recovery](./powerx-persistence-recovery.md) — Telemetry receipts, migration prerequisites, and targeted repair
- [API Skill Releases](./inferencex-skills-release.md) — Prepare an immutable package, verify clean installations and agent exports, and publish through the package-specific workflow
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 66b0907e0..0808b0ce8 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -2,632 +2,241 @@
[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md)
-PowerX starts with benchmark measurements and estimates how much facility power
-is needed to run the measured workload. Smart provisioning uses that estimate to
-calculate capacity and economics for a fixed facility power budget.
-[PR #1190](https://github.com/SemiAnalysisAI/InferenceX-app/pull/1190) extends this
-path to GB200/GB300 NVL72 racks. It also preserves valid provisioned comparisons
-when the corresponding measured estimate is unavailable.
-
-The ordinary modeled-power chart supports the non-agentic 8192-input/1024-output
-workload. The Profit Estimator explicitly opts into AgentX estimates, including
-Kimi K3; that does not establish AgentX calibration. The app, derived views API,
-and offline exporter reuse the model. The producer and raw benchmark API retain
-the original measurements.
-
-Start with [the Kimi K3 example](#why-kimi-k3-can-have-power-data-but-no-smart-provisioning-row),
-then [the worked calculation](#a-worked-nvl72-calculation). The
-[architecture diagrams](#nvl72-architecture-walkthrough) and
-[code map](#where-each-part-lives) connect the explanation to implementation.
-The later sections retain the full admission, topology and export contracts.
-
-The model belongs to InferenceX-app. `system-power-model.ts` owns the equations,
-including nonlinear fan curves, PSU efficiency interpolation, intermediate
-rounding, and PUE ordering. `system-power-model.profiles.json` owns the editable
-component parameters, hardware mapping, platform configuration, and metadata
-describing the fixed inference scenario. These checked-in files are the source of truth for the frontend,
-shared views API, and offline exporter. `system-power-model.provenance.json`
-records their app model revision and source hashes.
-
-The checked-in reference cases retain the historical Python baseline for
-regression comparison; Python and its former private repository are not required
-to develop, build, or deploy the app model. The model remains **DRAFT / pending
-human verification**. Matching a numerical baseline establishes implementation
-consistency, not empirical chassis calibration.
-
-## What is measured, modeled, and provisioned?
-
-These are different inputs to a calculation, not interchangeable labels:
-
-| Quantity | Meaning | Used for |
-| --------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------- |
-| GPU Level Measured | Validated GPU-board watts during the benchmark window | GPU power charts and the measured GPU input to system modeling |
-| GPU Level Provisioned (TDP) | The configured hardware TDP reference | GPU-level reference comparisons; not a measurement of this run |
-| All in Provisioned | The hardware registry's fixed facility kW/GPU allowance | Baseline capacity and profit planning |
-| All in Measured | Measured inputs plus modeled unmeasured components and PUE | Workload-dependent system-power estimates and smart provisioning |
-
-“All in Measured” still includes modeled components. The chart must retain that
-qualification. Smart provisioning changes the estimated GPU capacity per GW;
-it does not set GPU power limits or prove that average power plus a reserve is a
-safe electrical peak limit.
-
-The seven eight-GPU chassis profiles measure GPU boards and model CPU, DRAM and
-other chassis overhead. **NVL72 has a different boundary:** it requires measured
-compute-module power, or measured GPU-board power plus complete Grace-socket
-power. Grace and its LPDDR5X are not filled in with a CPU utilization estimate.
-Rack networking, switches, tray residuals and conversion losses are modeled.
-
-## Why Kimi K3 can have power data but no smart-provisioning row
-
-A GPU-power chart answers “what GPU power was measured?” The profit calculation
-answers “what whole-system power belongs to the performance point that serves
-this target?” A visible GPU measurement is only one of those required inputs.
-
-The September 29, 2026 reproduction used 93 saved public Kimi K3 benchmark rows,
-AgentX P90 at **45 tok/s/user**, and automatic FP4 selection. It reproduced the
-reported two priced configurations: B200 Dynamo-vLLM and MI355X ATOM. This is a
-historical reproduction of the reported screenshot, not a statement that today's
-live database has the same coverage. At #1190 head
-[cc86afd6](https://github.com/SemiAnalysisAI/InferenceX-app/commit/cc86afd6b179cceeb550cbf579ff60df46a47ac3),
-the model and admission rules explain those saved inputs as follows:
-
-| Configuration | Why GPU data does not produce a measured profit estimate | Work needed |
-| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| GB200 NVL72 | The 13 saved rows had valid GPU power but no accepted CPU/module metrics or CPU audit. #1190 supplies the rack model, but cannot supply those missing measurements. | Collect complete Grace/module and GPU telemetry for the same benchmark windows, validate it, and ingest the matched results. |
-| GB300 NVL72 | The 11 saved rows had no CPU/module evidence; only five had valid GPU power. The selected frontier also used GPU-invalid row 442101. | Recover valid GPU evidence where retained artifacts permit it, and obtain the missing Grace/module measurements. Both requirements must pass. |
-| MI355X vLLM | The original performance bracket used rows 443223 and 443220; 443220 had invalid power. Other valid GPU-power points belonged to a different bracket. | Diagnose the original point's retained telemetry. Reprocess only if it contains sufficient valid evidence; otherwise collect and publish replacement benchmark results with matched performance and power. |
-| B300 | The selected rows 439941/439935 had no measured power. | Supply qualified measurements for the serving curve used by the estimator. |
-| H200 | The available serving curve did not reach the requested 45 tok/s/user. | Use a target inside that curve's supported range, or obtain a qualified curve covering the requested target. |
-
-The screenshots also selected different engines: the power chart hid ATOM and
-showed MI355X vLLM, while the priced AMD result was ATOM. Comparisons must match
-model, workload, date/run, engine, precision, percentile and target before their
-row counts are compared.
-
-The AMD example is concrete: the performance interpolation at 45 uses
-**14.832 and 47.596 tok/s/user**, with invalid power at the second point. A valid
-power sample at **60.386 tok/s/user** cannot replace it without also changing
-which performance points support the estimate. Filtering out all bad-power rows
-and rebuilding the frontier would define a different comparison methodology.
-#1190 deliberately retains the original serving frontier.
-
-The saved invalid rows do not identify the collection failure's cause. Their
-public audit field was null, so this evidence does not justify blaming an
-exporter, deleting all AMD history, or promising that re-ingestion will repair
-it. New CPU readings from another run also cannot be attached to historical
+Use this guide to inspect measured curves, estimate system power, and compare
+capacity under a fixed facility power budget. Start with [the dashboard steps](#inspect-a-measured-curve),
+then use [the hardware requirements](#hardware-and-telemetry-requirements) or
+[troubleshooting](#when-a-curve-or-estimate-is-missing) when a result is unavailable.
+
+The system model is **DRAFT / pending human verification**. Its regression
+fixtures establish numerical consistency, not empirical calibration. AgentX
+estimates, including Kimi K3, are planning previews; they do not establish AgentX
+calibration or safe peak-load provisioning.
+
+## Choose the power boundary
+
+| Dashboard boundary | What the value represents |
+| --------------------------- | ---------------------------------------------------------------------------------------------------- |
+| GPU Level Measured | Validated GPU-board power during the benchmark window. |
+| GPU Level Provisioned (TDP) | Hardware TDP; a reference, not a reading from this run. |
+| All in Provisioned | The hardware registry's fixed facility kW/GPU allowance. |
+| All in Measured | Measured inputs plus modeled system components and facility overhead. It is not measured wall power. |
+
+The examples below use W per GPU and J per output token. GPU measurements remain
+available independently of whether the system model can accept the row. Measured
+P75/P90 values are time-weighted percentiles of synchronized fleet GPU power,
+divided by GPU count; neither the average nor individual-device percentiles
+substitute for them.
+
+## Inspect a measured curve
+
+**Prerequisites:** a benchmark selection with retained power data. Unlock the
+experimental power controls with ↑↑↓↓ if they are hidden.
+
+1. Open `/inference`, select the model (for example, Kimi K3), workload, date/run,
+ engine, precision, and hardware. Keep those selections fixed when comparing
+ boundaries; different engines or historical runs are different curves.
+2. Choose a measured-power metric and **GPU Level Measured**. Use **Table** to
+ inspect numeric rows and select a chart point to inspect its measurement
+ provenance. Start with average W/GPU: energy also requires a valid token
+ denominator, and P75/P90 require retained percentile measurements.
+3. Switch to **All in Measured** to inspect facility estimates. This boundary
+ supports 8K/1K single-turn results and AgentX previews. The separate
+ **Modeled Chassis AC** metric remains limited to 8K/1K single-turn results.
+4. For a capacity comparison, open `/profit-estimator-per-gigawatt`, choose the
+ model and a supported interactivity target, then select **Compare both** in
+ Benchmark Config. Match each result's workload, engine and precision to the
+ inference selection. Expand **Unavailable estimates** for missing results;
+ hover or select a bar for its power basis and read the formula notes below.
+
+**Expected result:** the inference table includes rows with a valid selected
+metric, including supported B200/H200 multi-node deployments. Selecting All in
+Measured does not include every valid GPU measurement: it also requires the
+inputs below. The Profit Estimator adds target-range and financial requirements
+and uses only official frontier points. Inference charts and tables also
+support unofficial-run overlays.
+
+## Hardware and telemetry requirements
+
+| Hardware / deployment | System estimate requires |
+| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| H100, H200, B200, B300, MI300X, MI325X, MI355X | Valid GPU-board telemetry; an eight-GPU chassis profile models CPU, DRAM, networking, storage, board, fans and PSU losses. |
+| B200/H200 and other supported chassis across multiple hosts | Consistent GPU counts and watts. Per-worker data must identify one chassis per distinct host. Aggregate non-disaggregated rows without workers may use complete eight-GPU hosts at the deployment mean (`uniform-hosts`). |
+| Prefill/decode disaggregation | Per-worker host placement and role power consistent with deployment totals. A role average alone cannot establish each host's load. Separate CPU-only frontend/router hosts are outside the estimate. |
+| GB200 / GB300 NVL72 | Valid GPU telemetry **and** complete, independently valid compute-module or Grace-socket telemetry from the same window: four GPUs and two sockets per compute tray. |
+
+The normal contract is numeric `power_valid=1`,
+`power_metric_schema_version=2`, positive `avg_power_w` (W/GPU) and
+`avg_total_gpu_power_w` (deployment W), and a consistent physical GPU count.
+The model retains a legacy validated single-node exception with no schema
+marker as `validated-unversioned-single-node`; this does not admit unversioned
+multi-node, disaggregated or NVL72 rows. Profit planning always requires schema 2.
+
+**NVL72 sensor boundary:** `cpu_power_valid=1` and `power_audit.cpu` must record
+matching expected/observed socket counts, two per tray. Either:
+
+- `avg_total_module_power_w` with `sensor_kind: module` covers GPU, HBM, Grace
+ and LPDDR5X together. Do not add GPU or Grace watts again.
+- `avg_total_cpu_power_w` and `avg_cpu_socket_power_w` with
+ `sensor_kind: grace_socket` must agree with the socket count. The model adds
+ GPU-board power and a regulator allowance of GPU W × 0.15 / 0.85.
+
+CPU-rail-only, missing or unknown sensor provenance is insufficient. A present
+but invalid module measurement remains unavailable; it does not silently fall
+back to Grace readings. Model support does not imply that every producer version
+collects these fields.
+
+**Partial allocations:** a one-to-seven-GPU chassis is modeled as eight GPUs at
+the measured per-GPU load and labeled `extrapolated`. Deployment totals retain
+only the measured GPUs' share. Profit planning accepts single-node 1/2/4-GPU
+chassis allocations as whole-replica extrapolations, assuming co-location leaves
+performance and power unchanged. Other partial layouts, including partial NVL72
+trays, are not accepted for profit planning. A module sensor already covers its
+whole tray, so that reading is never scaled to fill unmeasured GPUs.
+
+## One worked NVL72 example
+
+**Input:** the [GB300 reference fixture](../packages/app/src/lib/system-power-model.reference.json)
+with **3,000.75 W of module power per complete tray** and PUE 1.1.
+This is a numerical fixture, not a measured Kimi K3 result. An actual benchmark
+must separately satisfy both GPU and CPU/module validation.
+
+1. Scale the measured mean module input to 18 compute trays. Add modeled tray
+ components, nine NVSwitch trays, conversion losses and management switches.
+2. Evaluate the power-shelf efficiency once at the combined rack load. Every
+ measured tray receives the same 1/18 rack share; this assumes the remaining
+ trays run at that mean load, rather than measuring actual rack occupancy.
+3. Apply PUE once to rack AC power, then divide by 72 GPUs.
+4. Apply the separate 10% planning reserve for the Profit Estimator.
+
+| Stage | Watts |
+| ----------------------------------------------- | -------: |
+| Module input × 18 trays | 54,013.5 |
+| Modeled compute-tray components | 11,466.0 |
+| Modeled NVSwitch trays | 4,107.6 |
+| Tray conversion losses | 1,967.8 |
+| Rack DC, including 200 W of management switches | 71,754.9 |
+| Rack AC, after shelf losses | 74,904.6 |
+| Facility power after PUE 1.1 | 82,395.1 |
+
+```text
+Planning kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
+GPU capacity per GW = 1,000,000 / 1.258814 ≈ 794,399
+GPU-hours per GW-year = GPU capacity × 8,760
+```
+
+An eight-GPU benchmark on two complete trays receives
+82,395.1 × 8 / 72 ≈ 9,155.0 W of facility power, giving the same per-GPU result.
+It has not measured all 72 GPUs. Intermediate values are rounded; summing the
+displayed components can differ by 0.1 W. Planning uses the retained facility
+total, not the fixture's rounded per-GPU display value.
+
+## Assumptions behind the estimate
+
+- **PUE:** the dashboard applies 1.3 to air-cooled chassis and 1.1 to NVL72,
+ after AC conversion losses. The historical profile default of 1.2 is not the
+ dashboard default. Cooling describes the model, not verified site cooling.
+ Changing PUE does not convert an air-cooled chassis model into a DLC model.
+- **Chassis overhead:** coefficients reflect fixed assumptions of 20% CPU/DRAM
+ utilization, 5% PCIe utilization and idle NVMe. These are not live utilization
+ readings. Each host's nonlinear fan/PSU model is evaluated at its own load;
+ `uniform-hosts` explicitly substitutes the deployment mean for every host.
+- **NVL72 overhead:** Grace and LPDDR5X are measured. Rack networking, switches,
+ fans, board residuals, conversion losses and power shelves are modeled.
+ [Profiles](../packages/app/src/lib/system-power-model.profiles.json) retain
+ component values, source status and ranges for unverified parameters. Rack DC
+ above the 264 kW installed shelf capacity is outside the model domain.
+- **Planning reserve:** facility kW/GPU × 1.10 is a separate capacity buffer.
+ Average power plus this reserve is not a validated electrical peak limit.
+- **Matched comparison:** profit modes keep the same original performance
+ frontier, target, token prices, utilization, license share and per-GPU-hour
+ costs. Between frontier points, planning power is interpolated only between
+ those original points, with compatible model revision, PUE, topology and sensor
+ basis. No extrapolation or replacement by another power-valid point occurs.
+
+Lower planning power increases GPU capacity per GW. Revenue, compute cost and
+license fees scale with that capacity under the fixed per-GPU assumptions; profit
+margin and per-chip-hour economics do not improve. Electricity expense is not
+recomputed separately.
+
+## When a curve or estimate is missing
+
+Compare the same model, workload, date/run, engine, precision and metric first.
+Chart/table rows and target-based profit estimates answer different questions.
+
+| Symptom / reason | Check and next action |
+| -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
+| GPU curve exists, All in Measured is absent | Check supported workload/hardware and the row's system-model status. Valid GPU power alone does not establish a system estimate. |
+| B200/H200 multi-node row is absent (`topology`, `role-power`, `gpu-count`) | Inspect physical GPU count, host placement and total/role watts. Use the original producer topology; do not infer chassis placement from a display label or sum TP/EP aliases. |
+| NVL72 reports `cpu-telemetry` / `no-cpu-power` | Inspect the same-window CPU audit, sensor kind and complete socket coverage. GPU validity remains independent. |
+| `telemetry` / `no-measured-power` | Check the original validation audit and raw samples. Reprocess only when the retained evidence supports the original window; otherwise collect replacement performance and power together. |
+| `outside-measured-range` or power-invalid target bracket | Choose a target supported by the selected serving curve. Both original bounding points need valid power; another valid point elsewhere on the curve cannot fill the gap. |
+| `incompatible-power-basis` | Do not interpolate between module and GPU-plus-Grace readings, or different model/PUE bases. |
+| No cost, token mix or provisioned power | Inspect the financial inputs. This can prevent both profit estimates even when power is valid. |
+| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. |
+
+In **Compare both**, a valid provisioned result remains when its measured estimate
+is unavailable. Configurations that cannot be priced at all are listed separately
+from missing measured estimates. Measured-only mode never substitutes provisioned
+watts. See [persistence and recovery](./powerx-persistence-recovery.md) for retained
+telemetry and targeted repair; new power readings cannot be attached to old
throughput results.
-### What #1190 fixes, and the remaining fix
-
-The PR fixes the unsupported NVL72 topology, validates the two accepted sensor
-boundaries, and keeps a valid **All in Provisioned** result in **Compare both**
-when the measured estimate is absent. It already emits a specific skip reason;
-**All in Measured** alone continues to omit unqualified numeric results.
-
-The existing collapsed **Unavailable estimates** disclosure already lists each
-configuration and its reason. A possible presentation follow-up is to make that
-information easier to find beside the results and link each entry to its
-measurement or target. This is a proposed improvement, not behavior added by
-this document. It should not turn missing measurements into zeroes or
-provisioned values under a measured label.
-
-To produce the missing numeric rows, first inspect the original run artifacts.
-If the required samples and provenance exist, validate and reprocess those same
-windows. If they were never collected, collect a replacement performance-and-power
-result together. Complete Grace/module coverage is required for NVL72; fixing
-only the rack model or only GPU validity is insufficient.
-
-## A worked NVL72 calculation
-
-This is an **illustrative reference fixture, not a measured Kimi K3 result**.
-Use the stored GB300 module-basis case with **3,000.75 W per complete tray**,
-four GPUs and two Grace sockets per tray, and PUE 1.1. A real admitted benchmark
-must separately pass the GPU and CPU/module audit checks described below.
-
-1. **Choose the sensor boundary.** Here the module reading already covers the
- GPU/Grace compute module, so the model does not add GPU or Grace watts again.
- On the alternative GPU-plus-Grace path, the current profile adds the measured
- GPU and Grace totals, then a regulator allowance of GPU watts × 0.15 / 0.85.
- That allowance is a model assumption, not an independently measured rail.
-2. **Evaluate one rack.** Replicate the measured mean across 18 compute trays,
- add the profile's tray components and nine NVSwitch trays, then apply tray
- conversion losses and two management switches. Evaluate the power-shelf
- efficiency at this combined load, rather than evaluating a separate rack for
- each measured tray.
-3. **Apply facility overhead once.** Multiply rounded rack AC power by PUE.
- Allocate the resulting rack share over 72 GPUs. This assumes a rack whose
- remaining trays run at the measured trays' mean load; it does not measure
- the actual occupancy or power of an entire rack.
-4. **Apply planning headroom separately.** Multiply facility kW/GPU by 1.10.
- PUE 1.1 and the 10% reserve are different factors with different purposes.
-
-The checked-in reference fixture, originally captured from the Python baseline,
-contains these intermediate values:
-
-| Stage | Watts |
-| ---------------------------------------------- | -------: |
-| Measured module input × 18 trays | 54,013.5 |
-| Modeled compute-tray components | 11,466.0 |
-| Modeled NVSwitch trays | 4,107.6 |
-| Tray conversion losses | 1,967.8 |
-| Rack DC, including 200 W management switches | 71,754.9 |
-| Rack AC, after load-dependent shelf efficiency | 74,904.6 |
-| Facility power after PUE 1.1 | 82,395.1 |
-
-Displayed intermediate values are rounded. The calculation retains the model's
-summation and rounding order, so adding displayed values can differ by 0.1 W.
-The profile's baseline defaults include PUE 1.2; the dashboard wrapper supplies
-**1.1 for NVL72** and **1.3 for the supported air-cooled chassis**.
-
- Planning kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
- GPU capacity per GW = 1,000,000 / 1.258814 ≈ 794,399
- GPU-hours per GW-year = GPU capacity × 8,760
-
-For an eight-GPU benchmark on two full trays, the assigned facility share would
-be 82,395.1 × 8 / 72 ≈ 9,155.0 W. Dividing that share by the eight measured GPUs
-gives the same per-GPU planning value. The estimator does not claim that this
-small benchmark measured all 72 GPUs.
-
-At a non-exact interactivity target, the existing performance frontier determines
-the two bounding benchmark points. Both must have valid planning power and the
-same model revision, PUE, topology and sensor basis. The code linearly
-interpolates planning kW/GPU between those original points. It does not replace
-the existing throughput interpolation or extrapolate beyond its measured range.
-
-The economics then reuse the same throughput, token-price schedule, utilization,
-license share and per-GPU-hour cost in both power modes:
-
- Revenue = revenue per active GPU-hour × GPU-hours × utilization
- Cost = cost per GPU-hour × GPU-hours
- License share = revenue × license fraction
- Profit = revenue − cost − license share
-
-Changing planning power changes capacity and therefore total revenue, cost and
-profit per GW. Under these fixed per-GPU assumptions it does not improve profit
-margin, throughput per GPU, or profit per chip-hour. This is a planning projection,
-not a prediction that workload behavior will remain unchanged at facility scale.
-
-## Where each part lives
-
-The links below are repository-relative so they follow the reviewed checkout.
-The walkthrough was checked against app head cc86afd6; the example comes from
-the committed reference JSON. No new hardware measurements were taken for it.
-
-| Responsibility | Implementation |
-| ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| Preserve CPU/GPU metrics and audit provenance during ingest | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts), [power-publication.ts](../packages/db/src/etl/power-publication.ts) |
-| Edit component parameters; refresh app model revision and source hashes | [profiles](../packages/app/src/lib/system-power-model.profiles.json), [provenance](../packages/app/src/lib/system-power-model.provenance.json), [update-system-power-provenance.ts](../packages/app/scripts/update-system-power-provenance.ts) |
-| Apply workload, validity, sensor and topology admission; allocate deployment shares | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
-| Calculate nonlinear chassis/rack power, losses and PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
-| Match power to the original frontier and retain provisioned comparisons | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
-| Convert planning kW/GPU into capacity, revenue, cost and profit | [profit-estimator.ts](../packages/app/src/components/calculator/profit-estimator.ts) |
-| Display values, provenance, unavailable reasons and CSV fields | [ProfitEstimatorDisplay.tsx](../packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx), [ProfitEstimatorChart.tsx](../packages/app/src/components/calculator/ProfitEstimatorChart.tsx) |
-| Reuse the same economics through the read-only views API | [calculator-extensions.ts](../packages/app/src/lib/views-api/calculator-extensions.ts) |
-| Export modeled benchmark rows offline | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
-| Check calculation parity and admission behavior | [reference fixtures](../packages/app/src/lib/system-power-model.reference.json), [model tests](../packages/app/src/lib/system-power-model.test.ts), [admission tests](../packages/app/src/lib/modeled-system-power.test.ts), [planning tests](../packages/app/src/components/calculator/profit-power.test.ts) |
-
-The equations and editable parameters are reviewed and released together in
-InferenceX-app. Publishing an app change publishes the model it uses; there is
-no separate Python-source publication step or private-repository access gate.
-The reference fixtures record the earlier Python implementation as historical
-lineage only. App model versions and file hashes identify subsequent changes;
-neither moving ownership nor passing regression checks establishes calibration.
-See the update procedure below for historical results and frozen exports.
-
-## NVL72 architecture walkthrough
-
-### Measurements, ingestion and system power
-
-```mermaid
-flowchart TB
- subgraph Producer["Producer — InferenceX #3296"]
- GPU["Measured GPU-board power"]
- CPU["Measured Grace / compute-module power"]
- ART["Benchmark results and power_audit.cpu Matching measurement window"]
- GPU --> ART
- CPU --> ART
- end
- ART --> ETL["benchmark-mapper Retain validity, metrics and sensor provenance"]
- ETL --> DB[("Benchmark rows and metrics")]
- DB --> API["Benchmark API"]
- API --> GATE{"modelSystemPower Valid GPU verdict and CPU/module audit? GPU, socket and worker topology consistent?"}
- GATE -->|"No"| NONE["Unavailable with reason"]
- GATE -->|"Yes"| TRAY["Compute trays 4 GPUs + 2 Grace sockets per full tray"]
- TRAY --> BASIS{"Measured basis"}
- BASIS -->|"Module sensor"| MODULE["Measured module watts No separate Grace watts required No duplicate GPU / Grace addition"]
- BASIS -->|"Grace socket sensor"| SUM["Measured GPU-board + Grace watts Profile regulator allowance on GPU share"]
- MODULE --> RACK
- SUM --> RACK
- PROFILE["App-owned GB200 / GB300 rack profiles Static loads and conversion assumptions"] -.-> RACK
- RACK["Mean tray load scaled to an 18-tray rack Modeled rack residual + power-shelf losses"]
- RACK --> AC["Rack AC power"]
- AC --> FAC["Apply PUE once NVL72 default: 1.1"]
- FAC --> DEP["Allocate rack share to measured deployment Retain basis, profile and provenance"]
-```
+## Provenance and reproducible exports
-A valid module sensor already covers GPU, HBM, Grace and LPDDR5X; that path
-requires the complete module/socket audit but no redundant Grace-watts fields.
-The GPU-plus-Grace path instead requires valid socket measurements. Neither
-path substitutes a CPU estimate for missing CPU/module telemetry. A present but
-invalid module measurement is unavailable rather than silently falling back.
-
-### Planning, comparison and output
-
-```mermaid
-flowchart TB
- PERF["Official benchmark performance Selected percentile and target"] --> FRONT["Original serving frontier Exact point or original bounding points"]
- POWER["Validated deployment facility power From diagram 1"] --> ACCEPT
- FRONT --> ACCEPT{"Full measured NVL72 trays? Target inside measured range? Same basis and sensor between points?"}
- ACCEPT -->|"No"| SKIP["Measured estimate unavailable Keep the specific reason"]
- ACCEPT -->|"Yes"| POINT["Per-point planning power Facility kW / GPU x 1.10"]
- POINT --> SMART["Matched planning kW/GPU Interpolate at the original throughput knots Never swap points to fill missing power"]
- SPEC["Provisioned kW/GPU"] --> CAP
- SMART --> CAP["Same facility budget Compute deployable GPU capacity"]
- FRONT --> INPUT["Same throughput, token prices, utilization and unit costs"]
- INPUT --> ECON
- CAP --> ECON["Revenue, costs and profit"]
- ECON --> UI["All in Provisioned / All in Measured / Compare both Chart, tooltips, details and CSV"]
- SKIP --> KEEP["Compare retains valid provisioned bars Measured-only mode does not substitute"]
- KEEP --> UI
- META["Sensor basis, PUE, 10% reserve App model revision and source hash"] -.-> UI
-```
+Inspect the point's measurement source and the estimate's model revision, PUE,
+sensor basis and topology. Profit tooltips identify the power basis; formula
+notes and CSV columns `Power basis`, `Power sensor` and `System power profile`
+provide the assumptions and source details.
+The model revision is an `app-sha256:` digest, not a benchmark run ID or Git commit.
-Solid arrows show data flow; dashed arrows supply assumptions or provenance.
-Planning uses official frontier points, not `?unofficialrun=` overlays. The
-NVL72 model supplies a system-power calculation; it does not create missing
-measurements or choose new performance points. The detailed gates below also
-cover eight-GPU chassis and their existing replica extrapolation.
-
-## Updating the model for historical results
-
-Modeled power is derived from retained measurements when the browser or a shared
-views API transforms a benchmark row. Changing the model does not rewrite the
-original GPU measurements or require a per-run database backfill.
-
-1. Edit equations in `packages/app/src/lib/system-power-model.ts` and active
- parameters in `system-power-model.profiles.json`: `fixedComponentsDcWatts`,
- `fan`, `psu`, and the coefficients in `rackProfiles`. The per-profile
- `assumptions` describe the scenario used to derive those coefficients;
- changing `u_cpu`, `u_ram` or similar metadata alone does not recalculate watts.
- A new utilization scenario needs justified coefficients or new equations,
- matching assumption metadata, and regression acceptance. The workload,
- validity, topology and PUE policy live in `modeled-system-power.ts`; update
- that adapter when the intended behavior changes there. All three files are
- reviewed in InferenceX-app, without a separate model-repository change.
-2. Refresh the committed provenance manifest with the command below. Its
- `modelRevision` is `app-sha256:<64 hex>`, derived from the actual SHA-256 hashes
- of those three files. `modelPath` identifies
- `packages/app/src/lib/system-power-model.ts`. Commit the refreshed manifest
- with the model changes; `--check` reports drift without writing files. The
- digest is a model identity, not a Git commit. Source links use the deployment's
- `VERCEL_GIT_COMMIT_SHA` or `GITHUB_SHA`, falling back to `master`.
-
- ```sh
- bun packages/app/scripts/update-system-power-provenance.ts
- bun packages/app/scripts/update-system-power-provenance.ts --check
- ```
-
-3. Review the numerical changes and run the relevant model, admission, planning,
- views API and export regression checks. Keep the 496 historical reference
- cases as a frozen baseline. An intentional model change needs independently
- justified expected values and explicit regression acceptance; the provenance
- command never regenerates expected numbers to match the current code.
- Regression acceptance does not establish empirical calibration.
-4. Deploy the app with the accepted model and profiles. Historical rows with
- sufficient, matched raw telemetry are recalculated when they pass through the
- updated model; no model-only database backfill is required. Existing browser
- sessions need the updated bundle. Derived API responses need the normal
- authenticated cache invalidation or cache expiry; deployment alone does not
- establish that every cached response uses the new revision.
-5. Regenerate frozen CSV/JSON exports separately. If the revised model needs
- inputs that were never recorded, those rows stay unavailable until the input
- gap is resolved. A new benchmark's power must not be attached to an older
- benchmark's throughput.
-
-## Boundary and assumptions
-
-The input is measured mean GPU power during a validated serving window. The
-modeled chassis AC output adds the app profile's CPU, DRAM, networking, storage,
-board, fans, and PSU conversion losses. Facility power is a separate estimate:
-PUE is applied after chassis AC, preserving the model's rounding order.
-
-The app profiles retain the fixed inference assumptions `u_cpu=0.20`,
-`u_ram=0.20`, `u_pcie=0.05`, and `u_nvme=0.0`. The profile baseline includes PUE
-`1.2`; PowerX uses
-`1.3` for its supported air-cooled chassis profiles. Utility power = critical IT
-power × PUE (`1.3` air, `1.1` DLC).
-The factor applies after chassis AC; measured GPU power and chassis AC do not change.
-Cooling describes the modeled chassis, not verified benchmark-site cooling.
-The chassis profiles do not model DLC; `--pue` remains an explicit facility-factor
-override and does not convert an air-cooled chassis model into a DLC model. The
-NVL72 rack profiles below are direct-liquid-cooled and default to `1.1`.
-Platform-specific network assumptions,
-fan control, component counts, and chassis defaults are preserved in the
-editable app profile; every JSON export includes that profile and every CSV row
-includes its applicable assumptions, model revision and source hash. Utilization
-assumptions such as `u_cpu`, `u_ram` and `u_ib` record the scenario from which the
-active coefficients were derived; they are not live utilization controls or
-measured CPU/DRAM utilization. Editing these labels alone does not change the
-fixed watts or curves.
-
-These are the active keys in
-[system-power-model.profiles.json](../packages/app/src/lib/system-power-model.profiles.json).
-Their equations are implemented by the named functions in
-[system-power-model.ts](../packages/app/src/lib/system-power-model.ts).
-
-| Hardware identity | Editable app profile | TypeScript estimator |
-| ----------------- | -------------------- | ---------------------- |
-| `h100` | `profiles.h100` | `estimateChassisPower` |
-| `h200` | `profiles.h200` | `estimateChassisPower` |
-| `b200` | `profiles.b200` | `estimateChassisPower` |
-| `b300` | `profiles.b300` | `estimateChassisPower` |
-| `mi300x` | `profiles.mi300x` | `estimateChassisPower` |
-| `mi325x` | `profiles.mi325x` | `estimateChassisPower` |
-| `mi355x` | `profiles.mi355x` | `estimateChassisPower` |
-
-All listed profiles describe a complete eight-GPU chassis. Their topology is not
-substituted onto GB200 or GB300, which use the NVL72 rack profiles instead:
-
-| Hardware identity | Editable app profile | TypeScript estimator |
-| ----------------- | -------------------- | -------------------- |
-| `gb200` | `rackProfiles.gb200` | `estimateRackPower` |
-| `gb300` | `rackProfiles.gb300` | `estimateRackPower` |
-
-The [reference fixture](../packages/app/src/lib/system-power-model.reference.json)
-retains historical source/revision metadata, component hashes, and
-`pythonConfigurations` for baseline lineage only. Those records are not active
-app parameters or an ongoing Python dependency.
-
-The rack profiles (`rackProfiles`) take **measured** compute-module watts per tray
-as their input: the module sensor total (`avg_total_module_power_w`) when the
-producer publishes it, otherwise GPU-board watts plus the Grace-socket total
-(`avg_total_cpu_power_w`) with the app profile's regulator-loss allowance on the GPU
-share. The Grace CPU and LPDDR5X are never modelled; rows without
-`cpu_power_valid=1` and complete module or Grace provenance stay unavailable (`cpu-telemetry`).
-Each measured worker host is one compute tray (four GPUs, two Grace sockets); an
-aggregate multinode row without a per-worker array is `gpuCount / 4` trays at the
-deployment mean, cross-checked against the Grace-socket count and the CPU leg's
-`power_audit.cpu.observed_sockets`. The
-measured trays are folded into one rack of 18 trays matching their mean
-compute-module input, the power-shelf efficiency curve is evaluated once at that
-rack's DC load, using one mean per-tray input. Every tray takes the same 1/18
-share, so NVSwitch trays,
-power shelves, and management switches are amortised over 72 GPUs. Chassis, by
-contrast, own their fans and PSUs and are each evaluated at their own load. The
-result carries `topologyBasis: 'nvl72-trays'`, `measuredBasis`, and `sensorKind`.
-A partially allocated tray extrapolates only the GPU-board share (a module reading
-already covers the whole tray) and is labeled `extrapolated`. The app-owned
-model remains DRAFT / pending human verification. The NVL72 section below lists
-the measured input, the modeled residual and the Profit Estimator gate rules.
-
-A partially allocated chassis (one to seven measured GPUs on one host) is
-modeled at measured per-GPU power × 8. This `n_gpu × W/GPU` input is retained
-from the historical baseline and assumes the unmeasured GPUs run the same
-workload. The estimate is labeled `chassisBasis:
-'extrapolated'`: per-GPU values divide by the modeled chassis GPU count
-(`modeledGpuCount`), while `deploymentAcWatts` / `deploymentFacilityWatts` keep
-only the measured GPUs' share of each chassis. This is not a proportional share
-of a chassis evaluated at partial load; fixed components, the fan curve, and PSU
-efficiency are all evaluated at full-chassis load. Missing or invalid telemetry,
-inconsistent counts, missing host placement, more than one chassis per host, and
-model-domain overflow remain unavailable.
-
-For a single-node deployment, the producer's physical width is `TP * PP * PCP`.
-EP partitions that width. Some existing API configuration aliases contain
-`TP * EP`; the model cross-checks the physical width against measured total and
-per-GPU watts instead of trusting or summing those aliases. Multi-node and
-disaggregated inputs require one chassis (one to eight GPUs) per measured
-worker, distinct worker hosts, and consistent total/role watts. A role average
-alone cannot establish physical placement or evaluate each host's nonlinear
-model. CPU-only frontend workers are excluded from GPU-chassis counting. Separate CPU-only
-frontend/router hosts are outside this estimate; CPU power within GPU chassis
-still uses fixed coefficients derived for the app profile's 20% utilization
-scenario.
-
-The default measured contract is numeric `power_valid=1` and metric schema 2.
-The original validated single-node producer predates the schema marker but
-already uses the schema-2 definitions for both watts fields. This path retains the absent
-schema and reports `validated-unversioned-single-node`; it does not upgrade the
-source or admit unversioned disaggregated power. The article receipt additionally
-pins the producer checkout and retains each original audit artifact.
-
-## NVL72 rack estimate (GB200, GB300)
-
-**Measured input.** Every compute tray is fed a measured compute-module figure; the
-Grace CPU and LPDDR5X are never modeled. The producer's CPU power leg (srt-slurm,
-ACPI hwmon) publishes, over the same formal window as GPU energy,
-`avg_cpu_socket_power_w`, `avg_total_cpu_power_w`, `total_cpu_energy_j`, and, when
-the `Module Power Socket` sensor exists on every socket, `avg_total_module_power_w`
-and `total_module_energy_j`, with the independent verdict `cpu_power_valid` and the
-`power_audit.cpu` block (sensor kind, collector, socket coverage, reason codes).
-Admission requires `power_valid=1`, schema 2, `cpu_power_valid=1`, and
-`power_audit.cpu` with matching expected/observed socket counts: two per tray.
-A module reading requires `sensor_kind: module` and does not need redundant Grace
-metrics. The GPU-plus-Grace path requires `sensor_kind: grace_socket`, positive
-Grace watts and total/mean watts consistent with the audited socket count.
-CPU-rail-only and missing or unknown sensor provenance stay unavailable. Basis selection: `module` when `avg_total_module_power_w` is present (the
-reading already contains the GPU boards, so it is never scaled), otherwise
-`gpu-plus-grace` (GPU-board watts × 4 plus the Grace-socket total per tray, with the
-profile's regulator-loss allowance `regulatorLossFracOfTdp / (1 − frac)` on the GPU
-share only). A present-but-invalid module key makes the row unavailable
-(`cpu-telemetry`); it never falls back to the Grace socket silently.
-
-**Modeled residual.** Everything outside the compute modules comes from the app
-profile (`rackProfiles`), evaluated once for a rack of 18 trays at the measured
-trays' mean input (the shelf curve sees the whole rack's DC load, never one tray's)
-and amortised over 72 GPUs; the parameters marked UNVERIFIED carry a documented
-range in `unverifiedParameters` and no published rail:
-
-| Component (per rack unless noted) | GB200 | GB300 | Source status |
-| -------------------------------------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------ | ---------------------------------------- |
-| NVSwitch tray silicon (9 trays) | 406.4 W / tray at `u_nvlink` 0.5 | same | `blackwell_nvswitch` model |
-| NVSwitch tray residual | 50 W / tray | same | UNVERIFIED (20–80 W) |
-| Compute-tray NICs with optics | ConnectX-7, 4 × 31.5 W = 126 W / tray | ConnectX-8 integrated PCIe, 4 NICs totaling 315 W (78.8 W each, rounded) | `generic/connectx7`, `generic/connectx8` |
-| Compute-tray BlueField-3 DPUs | 2 × 65 W idle = 130 W / tray | same | `generic/dpu`, idle only |
-| Compute-tray NVMe | 22 W / tray idle | same | `generic/nvme`, idle only |
-| Compute-tray fans | 130 W / tray | same | UNVERIFIED (40–220 W) |
-| Compute-tray board residual | 40 W / tray | same | UNVERIFIED (20–60 W) |
-| Management switches | 2 × 100 W | same | profile constant |
-| Tray 50 V → 12 V conversion | efficiency 0.9725 on tray loads | same | UNVERIFIED (0.96–0.985) |
-| Regulator allowance (`gpu-plus-grace`) | GPU-board W × 0.15 / 0.85; GPU share only | same | Grace tuning guide |
-| Power shelf | 264 kW installed, 132 kW redundant; efficiency 0.90 → 0.94 → 0.965 at 10/20/30% load | same | profile curve |
-| Facility PUE | 1.1 (direct liquid cooling) | same | PowerX policy, applied once to rack AC |
-
-Rack DC above the installed shelf capacity (264 kW) overflows the efficiency curve
-and the row is unavailable (`model-domain`). The model rounds rack AC to 0.1 W
-before PUE. Checked-in `rackCases` retain the historical Python baseline for
-both variants, both bases, every shelf knot and PUE 1.0–1.2. They are regression
-references, not calibration evidence or a dependency on the former repository.
-
-**Gate rules (Profit Estimator).** Planning kW/GPU = deployment facility watts ÷
-measured GPUs ÷ 1000 × 1.1. It accepts fully measured eight-GPU chassis
-(`chassisBasis: 'full'` on the `single-node`, `worker-hosts`, or `uniform-hosts`
-basis; see the Profit Estimator power basis section), or an `nvl72-trays` estimate
-whose trays are all fully measured (four GPUs and two sockets each: one tray per
-measured worker host, or, for an aggregate multinode row without a per-worker
-array, `gpuCount / 4` trays at the deployment mean, cross-checked against the
-Grace-socket count and `power_audit.cpu.observed_sockets`). Partial trays are
-extrapolated in the chart but rejected here. Single-node 1/2/4-GPU chassis
-allocations use the replica extrapolation below. Between two
-frontier knots both must share the same measured basis and sensor kind; a module
-knot beside a Grace-socket knot stays unavailable rather than blending sensors. The
-bar tooltip, the collapsed Power assumptions disclosure and the CSV columns `Power
-basis`, `Power sensor`, `System power profile` name the basis (measured module, or
-measured GPU board + Grace socket with regulator loss modeled), the sensor kind, and
-the app profile (`modelPath @ modelRevision sha256:`) per row.
-The path and revision identify the app-owned model; the source hash identifies
-the TypeScript equation file, and the revision also covers the editable profile
-and admission/PUE adapter.
-The `?unofficialrun=` overlay rule does not apply to the Profit Estimator basis
-control: the estimator prices official frontier points only.
-
-## Profit Estimator power basis
-
-The per-GW Profit Estimator offers provisioned power, measured + modeled power,
-and a paired comparison in Benchmark Config. Provisioned remains the default.
-The control uses the existing insider feature gate and is hidden while locked.
-Unlock with ↑↑↓↓ (`inferencex-feature-gate=1` in local storage). While locked,
-`c_power` cannot activate an alternative calculation or fetch full power rows;
-relocking restores provisioned estimates immediately.
-The alternative reuses the same hardware, P90 target, throughput frontier,
-token mix, prices, utilization, and per-GPU-hour costs. It changes only the
-facility kW/GPU used to calculate capacity per GW. Consequently, revenue,
-compute expense, license fee, and profit scale together; profit margin does not
-change. Electricity expense is not recomputed separately.
-
-This opt-in AgentX estimate requires validated schema-v2 telemetry and a supported
-system profile. It accepts full eight-GPU chassis on the single-node or measured
-worker-host basis. Aggregate multinode rows without per-worker telemetry may use
-the explicit `uniform-hosts` assumption: each complete chassis is evaluated at
-the measured deployment-mean GPU watts. Disaggregated rows require per-host role
-power. NVL72 requires complete four-GPU trays with the CPU provenance above.
-
-Validated single-node 1/2/4-GPU allocations retain full-chassis extrapolation:
-fill an eight-GPU server with whole replicas at the measured per-GPU power and
-throughput, then divide modeled facility power by eight. This assumes replica
-co-location does not change performance or power; it is not a measurement of a
-partly idle server. The chart, tooltip, and CSV label every extrapolated estimate,
-including interpolation with one partial knot. Other partial layouts and missing
-or invalid measurements remain unavailable with distinct reasons.
-The ordinary 8K/1K transformation keeps its existing admission policy.
-
-Compare retains each valid provisioned estimate even when measured + modeled
-power is unavailable. Its notice identifies the missing measured estimate;
-modeled-only mode never substitutes provisioned watts.
-
-At an exact frontier point, use that point's modeled power. Between points,
-estimate power linearly using the same two knots as the existing throughput
-interpolation; never select a different point to fill a power gap, and never blend
-two knots measured on different bases. The estimate uses PowerX's PUE policy (1.3 for
-air-cooled chassis, 1.1 for DLC NVL72 racks) and an additional 10% planning margin.
-These assumptions, including the fixed CPU/DRAM utilization above, are not validated
-peak-load provisioning or AgentX system calibration. The UI keeps a short measured-versus-modeled note visible; detailed assumptions
-and NVL72 measured inputs are in a collapsed disclosure. The CSV retains the full
-assumptions and, per row, the measured basis, sensor kind and profile.
-`c_power=modeled` and `c_power=compare` preserve the selection in share URLs.
-Unavailable historical estimates use the hardware registry when a chip is absent
-from today's results and include the source date/run label.
-
-## Offline comparison export
-
-The exporter reads a local cohort envelope and writes a **new** output directory:
+The [offline exporter](../packages/app/scripts/export-modeled-system-power.ts)
+accepts a local `ComparisonInput` envelope with original `BenchmarkRow` entries,
+metadata, and optional original artifacts/audits. Use the script's type for the
+complete shape; preserve run, attempt, producer revision, capture time and hashes.
+The exporter follows the ordinary 8K/1K model policy, not the AgentX preview opt-in.
+Run from the repository root with installed dependencies and a new output path:
```sh
bun packages/app/scripts/export-modeled-system-power.ts \
- --input /path/to/original-qwen-article.input.json \
- --output /path/to/new-original-qwen-comparison
-
-bun packages/app/scripts/export-modeled-system-power.ts \
- --input /path/to/qwen35-current.input.json \
- --output /path/to/new-current-qwen-comparison --pue 1.3
+ --input /path/to/cohort.input.json \
+ --output /path/to/new-comparison
```
-Without `--pue`, each row uses the dashboard's default for its hardware (`1.3` for
-the air-cooled chassis profiles, `1.1` for the DLC NVL72 rack profiles) through the
-same `modelSystemPower` path the chart uses, so article figures match chart hovers;
-`metadata.pue_override` records an explicit `--pue`, which then applies to every
-row, and `metadata.pue_defaults` records the per-hardware defaults. NVL72 rows also
-carry `measured_basis`, `sensor_kind`, the Grace-socket and module measured inputs
-under their own `cpu_power_valid`, the rack profile's `model_path` and assumptions,
-and rack-specific `calculation_boundary` / `extrapolation_note` text; x86 rows are
-unchanged.
-
-The maintained input shape is `ComparisonInput` in the script:
-
-```ts
-{
- cohort: string,
- metadata: { /* source URLs, capture times, hashes and cohort selection */ },
- rows: [{
- id: string, // stable observation identity
- cell?: string, // optional group of original replicates
- benchmark: BenchmarkRow, // original API row or existing ETL output
- rawInput?: unknown, // original artifact before ETL normalization
- source?: object, // run, attempt, producer revision and artifact receipt
- audit?: object // matching original power-validation sidecar
- }]
-}
+Outputs are `comparison.json`, `comparison.csv` and optional `cells.csv`. JSON
+retains original rows, audits, validity, model outputs and provenance; unavailable
+CSV values stay blank. Optional `--pue 1.3` overrides the facility factor for
+**every** row and is recorded in metadata. Without it, per-hardware defaults apply.
+See [API examples](./inferencex-api-examples.md) for measured-data extraction.
+
+Modeled energy requires a matching audit with exact telemetry duration, physical
+GPU count and successful request/token denominators. It is modeled deployment
+power × duration, not integrated measured wall energy. No kernel-level
+prefill/decode energy is inferred. Each replicate is modeled before aggregation;
+a cell mean remains unavailable if any scoped replicate is unavailable.
+
+## Maintain the model
+
+| Responsibility | Source |
+| --------------------------------------------------- | ------------------------------------------------------------------------------- |
+| Equations, nonlinear curves and rounding | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
+| Component parameters, assumptions and source status | [profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| Workload, telemetry, topology and PUE admission | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
+| Matched frontier and planning reserve | [profit-power.ts](../packages/app/src/components/calculator/profit-power.ts) |
+| Frozen numerical baseline | [reference fixtures](../packages/app/src/lib/system-power-model.reference.json) |
+| Model identity and source hashes | [provenance](../packages/app/src/lib/system-power-model.provenance.json) |
+
+Edit equations or active coefficients together with justified expected values and
+assumption metadata. Changing labels such as `u_cpu` alone does not change watts.
+Refresh and check the manifest:
+
+```sh
+bun packages/app/scripts/update-system-power-provenance.ts
+bun packages/app/scripts/update-system-power-provenance.ts --check
```
-Use the existing `normalizeArtifactRows` / `mapBenchmarkRow` for raw producer
-aggregates. Keep the original aggregate as `rawInput`, retain original schema
-markers, and check that its measured metrics survive normalization unchanged.
-Use complete raw API responses for current snapshots, then select the exact
-`single_turn`, `isl=8192`, `osl=1024` workload locally. Retain every scoped row,
-including unsupported hardware and missing/invalid power; never mix a current
-snapshot into the frozen article campaign.
-
-Outputs are `comparison.json`, `comparison.csv`, and optional `cells.csv`.
-The JSON includes original input rows, measured validity, modeled outputs,
-assumptions, audit windows, and source provenance. CSV includes separate measured
-and modeled columns; unavailable numbers are blank. Invalid/unverified raw
-values remain in the raw-input record and are not labeled valid measurements.
-Metadata records input and implementation SHA-256 hashes, model and application
-revisions, worktree state, generation time, and full profile provenance. Generate
-the final release export from the intended application commit; file hashes also
-identify any local changes during development.
-
-Modeled energy is available only when a matching valid audit sidecar supplies an
-exact telemetry duration, physical GPU count, and successful request/token
-denominators. It is modeled deployment power (the measured GPUs' share of each
-chassis) multiplied by that duration, not a time integral of measured wall power. Actual output-token counts are used;
-nominal `1024` tokens per query never replace recorded counts. No kernel-level
-prefill/decode energy is inferred. API snapshots without these sidecars receive
-power estimates only.
-
-Each replicate is modeled before aggregation. A cell mean averages its modeled
-replicate outputs; it does not evaluate the model at mean watts. If any replicate
-is unavailable, the corresponding mean remains unavailable rather than silently
-dropping that replicate.
-
-## Measured P75 and P90 GPU power
-
-`y_measuredP75Power` and `y_measuredP90Power` show the time-weighted P75 and P90 of
-synchronized fleet GPU-board power over the validated load window, divided by GPU
-count. They share the regular measured-power chart path for official points and
-unofficial overlays. Missing or unvalidated percentile data remains unavailable;
-average power is never used as a substitute. These metrics are separate from modeled
-chassis AC power and individual-device percentiles.
-
-P75 and P90 backfills use the same 34 original validated traces and exact windows
-recorded in `docs/data/power-p90-backfill.json`.
+Run affected model, admission, planning, views API and export checks. Keep the
+496 historical reference cases frozen; baseline parity does not establish
+calibration. The app owns the model and parameters; no private Python repository
+is required. Historical measurements are modeled on read, so a model-only change
+needs an updated app bundle and API cache refresh/expiry, not a raw-data backfill.
+Regenerate frozen exports separately; missing source measurements remain missing.
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index 10b46959b..64b96e14f 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -2,533 +2,208 @@
[English](./powerx-system-power.md) | [简体中文](./powerx-system-power.zh.md)
-PowerX 以基准测试的实测数据为起点,估算运行该负载所需的设施功率。智能容量规划
-(smart provisioning)再根据这一估算,计算固定设施功率预算下的 GPU 容量及经济指标。
-[PR #1190](https://github.com/SemiAnalysisAI/InferenceX-app/pull/1190) 将这条路径扩展到
-GB200/GB300 NVL72 机架。当相应的实测加建模估算不可用时,对比模式仍保留有效的预配结果。
-
-常规建模功耗图支持非 AgentX 的 8192 输入 / 1024 输出负载。利润估算器显式启用
-AgentX 估算,包括 Kimi K3;这不代表模型已经过 AgentX 校准。应用、派生 views API
-和离线导出器共用同一模型。生产端和原始 benchmark API 保留原始测量值。
-
-建议先看 [Kimi K3 的例子](#为什么-kimi-k3-有-gpu-功耗却没有智能容量规划结果),再看
-[完整计算示例](#nvl72-计算示例)。[架构图](#nvl72-架构导览) 和
-[代码位置](#各部分代码在哪里) 将说明与实现对应起来。后面的章节完整记录了数据接纳、
-拓扑和导出约定。
-
-模型由 InferenceX-app 维护。`system-power-model.ts` 定义公式,包括非线性风扇曲线、
-PSU 效率插值、中间值舍入和 PUE 应用顺序。`system-power-model.profiles.json` 保存可
-直接编辑的组件参数、硬件映射、平台配置,以及说明固定推理场景的元数据。前端、
-共享 views API 和离线导出器均以这些仓库内文件中的模型定义为准。`system-power-model.provenance.json`
-记录应用模型版本和源码哈希。
-
-已提交的参考用例保留历史 Python 基线,用于回归对照;开发、构建和部署应用模型
-都不需要 Python 或原私有仓库。模型仍为 **DRAFT / pending human verification
-(草稿,待人工核验)**。与数值基线一致只能证明实现的一致性,不能证明已经过实机机箱校准。
-
-## 实测、建模和预配分别指什么?
-
-它们是计算中的不同输入,不能互换名称:
-
-| 指标 | 含义 | 用途 |
-| --------------------------- | ---------------------------------------- | ---------------------------------------- |
-| GPU Level Measured | 基准测试窗口内通过验证的 GPU 板卡功率 | GPU 功耗图,以及系统模型的实测 GPU 输入 |
-| GPU Level Provisioned (TDP) | 硬件配置中的 TDP 参考值 | GPU 层面的参考对比;不是本次运行的测量值 |
-| All in Provisioned | 硬件注册表中固定的设施 kW/GPU 配额 | 容量和利润规划的基线 |
-| All in Measured | 实测输入,加上未实测组件的模型估算及 PUE | 随负载变化的系统功耗估算和智能容量规划 |
-
-“All in Measured” 仍包含建模组件,图表必须保留这一限定。智能容量规划改变的是每 GW
-能容纳的 GPU 数量估算;它不会设置 GPU 功率上限,也不能证明“平均功耗加余量”就是
-安全的供电峰值上限。
-
-七种八卡机箱 profile 使用 GPU 板卡实测值,并对 CPU、DRAM 和其他机箱开销建模。
-**NVL72 的测量边界不同:**它要求实测计算模块功耗,或实测 GPU 板卡功耗加上完整的
-Grace socket 功耗。不能用 CPU 利用率假设补齐 Grace 及其 LPDDR5X 功耗。机架网络、
-交换机、tray 其余组件和转换损耗由模型估算。
-
-## 为什么 Kimi K3 有 GPU 功耗,却没有智能容量规划结果?
-
-GPU 功耗图回答的是“测到了多少 GPU 功耗”;利润计算回答的是“满足这个性能目标的
-数据点,对应多少整机功耗”。图上有 GPU 实测值,只能说明其中一项输入存在。
-
-2026 年 9 月 29 日的复现使用了保存的 93 条公开 Kimi K3 基准测试记录,设置为 AgentX
-P90、**45 tok/s/user**、自动选择 FP4。它复现了截图中两个有价格结果的配置:B200 Dynamo-vLLM
-和 MI355X ATOM。这是对当时截图的历史复现,不代表今天的在线数据库仍有相同的数据覆盖。
-按 #1190 的
-[cc86afd6](https://github.com/SemiAnalysisAI/InferenceX-app/commit/cc86afd6b179cceeb550cbf579ff60df46a47ac3)
-版本,其模型和接纳规则可以这样解释这批输入:
-
-| 配置 | 为什么有 GPU 数据,却算不出实测加建模的利润结果 | 需要做什么 |
-| ----------- | ------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- |
-| GB200 NVL72 | 保存的 13 条记录都有有效 GPU 功耗,但没有符合要求的 CPU/module 指标或 CPU 审计。#1190 提供了机架模型,无法补出缺失测量。 | 在相同基准测试窗口内采集完整 Grace/module 和 GPU 遥测,通过验证后将匹配的结果入库。 |
-| GB300 NVL72 | 保存的 11 条记录都没有 CPU/module 证据,其中只有五条 GPU 功耗有效。所选性能前沿还使用了 GPU 功耗无效的记录 442101。 | 若保留的产物足以证明 GPU 功耗有效,则恢复该证据;同时补齐 Grace/module 测量。两项要求都必须满足。 |
-| MI355X vLLM | 原性能插值区间使用记录 443223 和 443220;443220 的功耗无效。其他功耗有效的数据点属于另一个区间。 | 检查原数据点保留的遥测。只有存在充分有效证据时才重新处理;否则重新采集并发布性能与功耗匹配的基准测试结果。 |
-| B300 | 所选记录 439941/439935 没有实测功耗。 | 为估算器使用的服务性能曲线补充通过验证的测量。 |
-| H200 | 现有服务性能曲线达不到所要求的 45 tok/s/user。 | 将目标设在曲线支持的范围内,或取得覆盖该目标且通过验证的新曲线。 |
-
-这些截图选择的引擎也不同:功耗图隐藏了 ATOM,显示 MI355X vLLM;有价格结果的 AMD
-配置则是 ATOM。比较记录数量前,应先对齐模型、负载、日期/运行、引擎、精度、分位数
-和目标值。
-
-AMD 的例子更具体:45 tok/s/user 的性能插值使用 **14.832 和 47.596 tok/s/user** 两个点,
-第二个点的功耗无效。不能直接借用 **60.386 tok/s/user** 处的有效功耗,否则支撑估算的
-性能点也随之改变。如果先删掉所有功耗无效的记录,再重建前沿,就采用了另一套比较方法。
-#1190 明确保留原有服务性能前沿。
-
-保存的无效记录不能说明采集失败的原因:它们的公开审计字段为 null。因此,这份证据
-不足以归咎于某个 exporter、删除全部 AMD 历史数据,或承诺重新入库就能修复。
-也不能把另一轮运行新采集的 CPU 功耗附到历史吞吐量上。
-
-### #1190 修复了什么,还缺什么?
-
-该 PR 支持了原先不支持的 NVL72 拓扑,验证两种可接纳的传感器测量边界,并在实测
-估算缺失时,让 **Compare both** 继续保留有效的 **All in Provisioned** 结果。
-代码已经输出具体的跳过原因;单独选择 **All in Measured** 时,仍不显示不符合要求的数值。
-
-现有默认折叠的 **Unavailable estimates** 详情已逐项列出配置及不可用原因。后续可以
-改善展示,让读者更容易在结果旁找到这些信息,并跳转到对应测量或目标。这只是建议的
-展示改进,不是本文新增的行为。不能把缺失测量改成零,也不能在实测标签下填入预配值。
-
-要补齐缺失的数值结果,应先检查原运行的产物。若所需样本和来源记录齐全,就验证并
-重新处理原测量窗口;若从未采集,则需要一起重测性能和功耗,形成替代结果。
-NVL72 必须具备完整 Grace/module 覆盖;仅修机架模型或仅修 GPU 有效性都不够。
-
-## NVL72 计算示例
-
-以下是**用于说明计算过程的参考测试数据,不是 Kimi K3 实测结果**。采用已保存的 GB300
-module 口径用例:**每个完整 tray 为 3,000.75 W**,每 tray 四张 GPU、两个 Grace socket,
-PUE 为 1.1。真实基准测试还必须分别通过后文列出的 GPU 和 CPU/module 审计。
-
-1. **确定传感器测量边界。**本例的 module 读数已包含 GPU/Grace 计算模块,模型不会再
- 加一遍 GPU 或 Grace 功耗。另一条 GPU 加 Grace 路径则先相加两者的实测总功耗,再按
- GPU 功耗 × 0.15 / 0.85 加入稳压损耗余量。这个余量是模型假设,不是独立测得的供电轨功耗。
-2. **计算一个整机架。**把实测均值扩展到 18 个计算 tray,加入 profile 中的 tray 组件和
- 九个 NVSwitch tray,再计入 tray 转换损耗和两台管理交换机。在合计负载处计算电源架
- 效率,而不是为每个实测 tray 单独计算一整套机架。
-3. **只应用一次设施开销。**将舍入后的机架交流功率乘以 PUE,再按 72 张 GPU 分摊。
- 这里假设机架其余 tray 的负载与实测 tray 的平均负载相同,并未测量整个机架的实际
- 占用情况或总功耗。
-4. **单独计入规划余量。**将设施 kW/GPU 乘以 1.10。PUE 1.1 与 10% 规划余量是不同
- 的系数,作用也不同。
-
-已提交的参考测试数据最初取自 Python 基线,包含以下中间值:
-
-| 阶段 | 功率(W) |
-| ------------------------------------------ | --------: |
-| 实测 module 输入 × 18 个 tray | 54,013.5 |
-| 建模的计算 tray 组件 | 11,466.0 |
-| 建模的 NVSwitch tray | 4,107.6 |
-| tray 转换损耗 | 1,967.8 |
-| 机架直流功率,含 200 W 管理交换机 | 71,754.9 |
-| 计入随负载变化的电源架效率后的机架交流功率 | 74,904.6 |
-| 乘以 PUE 1.1 后的设施功率 | 82,395.1 |
-
-表中中间值已舍入。计算保留模型的求和与舍入顺序,因此直接相加表中数值可能差
-0.1 W。profile 的基线默认 PUE 为 1.2;仪表板封装层对 **NVL72 使用 1.1**,
-对**当前支持的风冷机箱使用 1.3**。
-
- 规划 kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
- 每 GW 的 GPU 容量 = 1,000,000 / 1.258814 ≈ 794,399
- 每 GW 每年的 GPU 小时数 = GPU 容量 × 8,760
-
-若基准测试使用两个完整 tray、共八张 GPU,则分摊的设施功率为
-82,395.1 × 8 / 72 ≈ 9,155.0 W。再除以八张实测 GPU,得到的每 GPU 规划功率相同。
-估算器并不声称这次小规模测试测量了全部 72 张 GPU。
-
-当交互性能目标不恰好落在实测点上时,由现有性能前沿决定两侧的基准测试点。两点都
-必须具备有效规划功率,且模型版本、PUE、拓扑和传感器口径相同。代码在原有两点之间
-线性插值规划 kW/GPU,不替换原吞吐量插值,也不向实测范围之外外推。
-
-随后,两种功耗模式共用相同的吞吐量、token 定价、利用率、模型许可分成和每 GPU 小时成本:
-
- 收入 = 每个活跃 GPU 小时的收入 × GPU 小时数 × 利用率
- 成本 = 每 GPU 小时成本 × GPU 小时数
- 模型许可分成 = 收入 × 许可分成比例
- 利润 = 收入 − 成本 − 模型许可分成
-
-规划功率改变容量,因而改变每 GW 的总收入、总成本和总利润。在每 GPU 假设固定的
-情况下,它不会提高利润率、每 GPU 吞吐量或每芯片小时利润。这是容量规划推算,
-并不预测负载扩展到整个设施后仍保持相同行为。
-
-## 各部分代码在哪里
-
-以下链接均为仓库内相对链接,随当前检出的版本变化。本导览依据应用版本 cc86afd6
-核对,示例来自已提交的参考 JSON;没有为本文新增硬件测量。
-
-| 职责 | 实现 |
-| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
-| 入库时保留 CPU/GPU 指标和审计来源 | [benchmark-mapper.ts](../packages/db/src/etl/benchmark-mapper.ts)、[power-publication.ts](../packages/db/src/etl/power-publication.ts) |
-| 编辑组件参数;更新应用模型版本和源码哈希 | [profiles](../packages/app/src/lib/system-power-model.profiles.json)、[来源清单](../packages/app/src/lib/system-power-model.provenance.json)、[update-system-power-provenance.ts](../packages/app/scripts/update-system-power-provenance.ts) |
-| 检查负载、有效性、传感器和拓扑条件;分摊部署功耗 | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
-| 计算非线性机箱/机架功耗、损耗和 PUE | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
-| 将功耗匹配到原前沿,并保留预配对比结果 | [modeledPowerAtTarget / estimateProfitByPower](../packages/app/src/components/calculator/profit-power.ts) |
-| 将规划 kW/GPU 换算为容量、收入、成本和利润 | [profit-estimator.ts](../packages/app/src/components/calculator/profit-estimator.ts) |
-| 显示数值、来源、不可用原因和 CSV 字段 | [ProfitEstimatorDisplay.tsx](../packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx)、[ProfitEstimatorChart.tsx](../packages/app/src/components/calculator/ProfitEstimatorChart.tsx) |
-| 通过只读 views API 复用同一经济指标计算 | [calculator-extensions.ts](../packages/app/src/lib/views-api/calculator-extensions.ts) |
-| 离线导出建模后的基准测试记录 | [export-modeled-system-power.ts](../packages/app/scripts/export-modeled-system-power.ts) |
-| 验证计算一致性和接纳行为 | [参考测试数据](../packages/app/src/lib/system-power-model.reference.json)、[模型测试](../packages/app/src/lib/system-power-model.test.ts)、[接纳规则测试](../packages/app/src/lib/modeled-system-power.test.ts)、[规划测试](../packages/app/src/components/calculator/profit-power.test.ts) |
-
-公式和可编辑参数在 InferenceX-app 中一同审阅、一同发布。发布应用变更就会发布该
-版本采用的模型,不需要单独发布 Python 源码,也不以私有仓库权限作为前提。
-参考测试数据只将早期 Python 实现记作历史来源。后续变化由应用模型版本和文件哈希
-标识;维护归属的改变或回归检查通过,都不代表完成了实测校准。历史结果和冻结导出
-的更新方式见后文。
-
-## NVL72 架构导览
-
-### 测量、入库与系统功耗
-
-```mermaid
-flowchart TB
- subgraph Producer["生产端 — InferenceX #3296"]
- GPU["实测 GPU 板卡功耗"]
- CPU["实测 Grace / 计算模块功耗"]
- ART["基准测试结果与 power_audit.cpu 测量窗口一致"]
- GPU --> ART
- CPU --> ART
- end
- ART --> ETL["benchmark-mapper 保留有效性、指标与传感器来源"]
- ETL --> DB[("基准测试记录与指标")]
- DB --> API["Benchmark API"]
- API --> GATE{"modelSystemPower GPU 验证结论与 CPU/module 审计有效? GPU、socket、worker 拓扑一致?"}
- GATE -->|"否"| NONE["不可用,并给出原因"]
- GATE -->|"是"| TRAY["计算 tray 每个完整 tray 为 4 张 GPU + 2 个 Grace socket"]
- TRAY --> BASIS{"实测口径"}
- BASIS -->|"Module 传感器"| MODULE["实测 module 功耗 无需单独提供 Grace 功耗 不重复加入 GPU / Grace 功耗"]
- BASIS -->|"Grace socket 传感器"| SUM["实测 GPU 板卡 + Grace 功耗 按 profile 对 GPU 份额计入稳压损耗余量"]
- MODULE --> RACK
- SUM --> RACK
- PROFILE["应用维护的 GB200 / GB300 机架 profile 静态负载与转换假设"] -.-> RACK
- RACK["按 tray 平均负载扩展至 18 tray 机架 加入机架其余组件和电源架损耗模型"]
- RACK --> AC["机架交流功率"]
- AC --> FAC["只应用一次 PUE NVL72 默认值:1.1"]
- FAC --> DEP["向实测部署分摊机架功耗 保留口径、profile 与来源"]
+本指南介绍如何查看实测曲线、估算系统功耗,以及比较固定设施功率预算下的 GPU 容量。
+先按[仪表板操作步骤](#查看实测曲线)选择数据;结果缺失时,再查阅
+[硬件要求](#硬件与遥测要求)和[排查说明](#曲线或估算结果缺失时如何排查)。
+
+系统模型仍为 **DRAFT / pending human verification(草稿,待人工核验)**。
+回归用例只能证明数值实现一致,不能证明已完成实机校准。包括 Kimi K3 在内的
+AgentX 估算属于容量规划预览,不代表已完成 AgentX 校准,也不能作为安全峰值供电规划的依据。
+
+## 选择功耗边界
+
+| 仪表板边界 | 数值含义 |
+| --------------------------- | ---------------------------------------------------------------- |
+| GPU Level Measured | 基准测试窗口内通过验证的 GPU 板卡实测功率。 |
+| GPU Level Provisioned (TDP) | 硬件 TDP 参考值,不是本次运行的测量值。 |
+| All in Provisioned | 硬件注册表中固定的设施 kW/GPU 配额。 |
+| All in Measured | 实测输入加上系统组件和设施开销的模型估算,不是墙上功率计的读数。 |
+
+下文示例采用 W/GPU 和 J/输出 token。系统模型不支持某条记录,并不影响
+该记录有效 GPU 测量值的展示。实测 P75/P90 是全部 GPU 同步汇总功率的时间加权
+分位数,再除以 GPU 数量;平均值和单卡分位数都不能替代。
+
+## 查看实测曲线
+
+**前提:**所选基准测试保留了功耗数据。若实验性功耗控件未显示,用 ↑↑↓↓ 解锁。
+
+1. 打开 `/inference`,选择模型(如 Kimi K3)、工作负载、日期/运行、引擎、精度和
+ 硬件。比较不同功耗边界时保持这些选项一致;不同引擎或历史运行属于不同曲线。
+2. 选择实测功耗指标和 **GPU Level Measured**,通过 **Table** 查看数值记录,点击
+ 图表上的数据点查看测量来源。先看平均 W/GPU:能耗还需要有效的 token 分母,
+ P75/P90 则需要保留相应分位数测量值。
+3. 切换到 **All in Measured** 查看设施功率估算。该边界支持 8K/1K 单轮结果和
+ AgentX 预览;独立的 **Modeled Chassis AC** 指标仍仅支持 8K/1K 单轮结果。
+4. 要比较容量,打开 `/profit-estimator-per-gigawatt`,选择模型和曲线支持的
+ 交互性目标,再在 Benchmark Config 中选择 **Compare both**。逐项核对结果的
+ 工作负载、引擎和精度是否与推理页面所选一致。展开 **Unavailable estimates**
+ 查看缺失原因;悬停或选中柱形查看功耗依据,并阅读图表下方的公式说明。
+
+**预期结果:**推理表格列出所选指标有效的记录,包括满足要求的 B200/H200 多节点
+部署。All in Measured 不会包含所有 GPU 功耗有效的记录,还需满足下述输入要求。
+利润估算器另有目标范围和财务输入要求,且只使用官方性能前沿上的数据点;推理
+图表和表格同时支持非官方运行叠加。
+
+## 硬件与遥测要求
+
+| 硬件 / 部署方式 | 系统估算所需输入 |
+| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| H100、H200、B200、B300、MI300X、MI325X、MI355X | 有效 GPU 板卡遥测;八卡机箱模型估算 CPU、DRAM、网络、存储、主板、风扇和 PSU 损耗。 |
+| B200/H200 等受支持机箱的多主机部署 | GPU 数量与功率须一致。逐 worker 数据须明确每台独立主机对应一个机箱。没有 worker 数据的非分离式多节点记录,可按完整八卡主机使用部署平均负载(`uniform-hosts`)。 |
+| Prefill/decode 分离式部署 | 逐 worker 的主机分布和角色功率须与部署总量一致。仅有角色平均值无法确定各主机负载。独立的纯 CPU frontend/router 主机不在估算范围内。 |
+| GB200 / GB300 NVL72 | 有效 GPU 遥测,**以及**同一窗口内完整且独立通过验证的计算模块或 Grace socket 遥测:每个计算 tray 为四张 GPU、两个 socket。 |
+
+常规输入要求数值型 `power_valid=1`、`power_metric_schema_version=2`,
+`avg_power_w`(W/GPU)和 `avg_total_gpu_power_w`(部署总 W)均为正值,且物理
+GPU 数量一致。模型对已验证但没有 schema 标记的历史单节点记录保留例外,并标为
+`validated-unversioned-single-node`;该例外不适用于未标版本的多节点、分离式或
+NVL72 记录。利润规划始终要求 schema 2。
+
+**NVL72 传感器边界:**要求 `cpu_power_valid=1`,且 `power_audit.cpu` 记录的预期
+和实际 socket 数量一致,每个 tray 两个 socket。支持两种输入:
+
+- `avg_total_module_power_w` 配合 `sensor_kind: module`,已覆盖 GPU、HBM、Grace
+ 和 LPDDR5X,不能再次加上 GPU 或 Grace 功率。
+- `avg_total_cpu_power_w` 和 `avg_cpu_socket_power_w` 配合
+ `sensor_kind: grace_socket`,总功率须与单 socket 平均功率及 socket 数量相符。
+ 模型另加 GPU 板卡功率,
+ 以及 GPU W × 0.15 / 0.85 的稳压损耗估算。
+
+仅有 CPU rail 读数、缺少来源信息或传感器类型未知,都不满足要求。module 字段
+存在但无效时,结果保持不可用,不会自动改用 Grace 读数。模型支持某字段,并不
+代表每个生产端版本都已采集该字段。
+
+**部分 GPU 分配:**单机箱仅测量一至七张 GPU 时,按实测单卡负载外推到八卡机箱,
+标为 `extrapolated`;部署总量只计入已测 GPU 的份额。利润规划接受单节点 1/2/4 卡
+机箱配置的整副本外推,前提是假设副本共置不改变性能和功耗。其他部分分配方式,
+包括未测满的 NVL72 tray,不用于利润规划。module 传感器已覆盖整个 tray,因此
+不会放大其读数来补足未测 GPU。
+
+## 一个 NVL72 计算示例
+
+**输入:**[GB300 参考用例](../packages/app/src/lib/system-power-model.reference.json)
+中,每个完整 tray 的 module 功率为 **3,000.75 W**,PUE 为 1.1。这是数值回归
+用例,不是 Kimi K3 实测结果。实际基准测试还须分别通过 GPU 和 CPU/module 验证。
+
+1. 将实测平均 module 输入扩展到 18 个计算 tray,加上模型中的 tray 组件、九个
+ NVSwitch tray、转换损耗及管理交换机。
+2. 在整个机架的合计负载上计算一次电源架效率。每个已测 tray 分得 1/18 的机架
+ 功率;这是假设其余 tray 也处于相同平均负载,并非测量了实际机架占用情况。
+3. 对机架 AC 功率应用一次 PUE,再除以 72 张 GPU。
+4. 利润估算器另加 10% 的规划余量。
+
+| 阶段 | 功率(W) |
+| ------------------------------ | --------: |
+| Module 输入 × 18 个 tray | 54,013.5 |
+| 建模的计算 tray 组件 | 11,466.0 |
+| 建模的 NVSwitch tray | 4,107.6 |
+| Tray 转换损耗 | 1,967.8 |
+| 机架 DC,包含 200 W 管理交换机 | 71,754.9 |
+| 计入电源架损耗后的机架 AC | 74,904.6 |
+| 应用 PUE 1.1 后的设施功率 | 82,395.1 |
+
+```text
+规划 kW/GPU = 82,395.1 / 72 / 1,000 × 1.10 ≈ 1.258814
+每 GW 的 GPU 容量 = 1,000,000 / 1.258814 ≈ 794,399
+每 GW 每年的 GPU 小时 = GPU 容量 × 8,760
```
-有效的 module 传感器已覆盖 GPU、HBM、Grace 和 LPDDR5X;这条路径仍要求完整的
-module/socket 审计,但无需重复提供 Grace 功耗字段。GPU 加 Grace 路径则要求有效的
-socket 测量。两条路径都不会用 CPU 估算值补齐缺失的 CPU/module 遥测。
-module 测量存在但无效时,结果保持不可用,不会悄然回退到另一条路径。
-
-### 规划、对比与输出
-
-```mermaid
-flowchart TB
- PERF["正式基准测试性能 所选分位数和目标"] --> FRONT["原服务性能前沿 精确点或原有区间两端点"]
- POWER["通过验证的部署设施功率 来自图 1"] --> ACCEPT
- FRONT --> ACCEPT{"NVL72 tray 均完整实测? 目标在实测范围内? 两点的口径与传感器相同?"}
- ACCEPT -->|"否"| SKIP["实测加建模估算不可用 保留具体原因"]
- ACCEPT -->|"是"| POINT["逐点计算规划功率 设施 kW / GPU x 1.10"]
- POINT --> SMART["匹配目标的规划 kW/GPU 在原吞吐量数据点之间插值 不换点填补缺失功耗"]
- SPEC["预配 kW/GPU"] --> CAP
- SMART --> CAP["相同设施功率预算 计算可部署的 GPU 容量"]
- FRONT --> INPUT["相同吞吐量、token 价格、 利用率和单位成本"]
- INPUT --> ECON
- CAP --> ECON["收入、成本和利润"]
- ECON --> UI["All in Provisioned / All in Measured / Compare both 图表、提示框、详情和 CSV"]
- SKIP --> KEEP["对比模式保留有效预配柱子 仅实测模式不以预配值替代"]
- KEEP --> UI
- META["传感器口径、PUE、10% 余量 应用模型版本和源码哈希"] -.-> UI
-```
-
-实线表示数据流,虚线提供假设或来源信息。规划采用正式性能前沿点,不采用
-`?unofficialrun=` 叠加层。NVL72 模型提供系统功耗计算,不会创造缺失测量,也不会重新
-选择性能点。下文的详细接纳规则还涵盖八卡机箱及其现有的实例外推方式。
-
-## 更新模型后,历史结果如何更新?
-
-浏览器或共享 views API 转换基准测试记录时,会根据保留的测量值推导建模功耗。
-修改模型不会改写原始 GPU 测量值,也不需要逐 run 回填数据库。
-
-1. 在 `packages/app/src/lib/system-power-model.ts` 中修改公式,在
- `system-power-model.profiles.json` 中修改参与计算的参数:`fixedComponentsDcWatts`、
- `fan`、`psu` 及 `rackProfiles` 中的系数。每个 profile 的 `assumptions` 记录推导这些
- 系数时采用的场景;只改 `u_cpu`、`u_ram` 等元数据,不会重新计算功率。要支持新的
- 利用率场景,需要有依据的新系数或新公式、与之对应的假设记录,并通过回归验收。
- 负载、有效性、拓扑和 PUE 策略位于 `modeled-system-power.ts`;若这些行为需要改变,
- 应同步修改该适配层。三个文件都在 InferenceX-app 中审阅,无需另改一个模型仓库。
-2. 用以下命令更新已提交的来源清单。`modelRevision` 的格式为 `app-sha256:<64 hex>`,
- 根据上述三个文件实际内容的 SHA-256 哈希生成。`modelPath` 指向
- `packages/app/src/lib/system-power-model.ts`。更新后的清单应与模型修改一起提交;
- `--check` 只检查清单是否与当前文件一致,不写文件。这个哈希摘要是模型标识,
- 不是 Git 提交。源码链接使用部署的 `VERCEL_GIT_COMMIT_SHA` 或 `GITHUB_SHA`,
- 两者均缺失时使用 `master`。
-
- ```sh
- bun packages/app/scripts/update-system-power-provenance.ts
- bun packages/app/scripts/update-system-power-provenance.ts --check
- ```
-
-3. 审阅数值变化,运行相关模型、接纳规则、规划、views API 和导出回归检查。
- 保留 496 个历史参考用例作为冻结基线。有意改变模型时,需要提供独立论证的预期值,
- 并明确验收回归结果;来源清单命令不会根据当前代码重新生成预期数字。
- 回归验收通过不代表完成了实测校准。
-4. 部署已验收的模型和 profiles。历史记录只要具有充分、匹配的原始遥测,就会在经过
- 新模型时重新计算;单纯修改模型无需回填数据库。已有浏览器会话需要加载新 bundle;
- 派生 API 响应需要走正常的认证缓存失效流程,或等待缓存过期。仅完成部署,不能证明
- 所有缓存响应都已采用新版本。
-5. 冻结的 CSV/JSON 导出需单独重新生成。若新模型需要从未记录的输入,相应记录应
- 保持不可用,直至输入缺口解决。不能把新基准测试的功耗附到旧基准测试的吞吐量上。
-
-## 计算边界与假设
-
-输入是已验证服务窗口内的平均 GPU 实测功率。机箱交流功耗模型在此基础上,加入
-应用 profile 中的 CPU、DRAM、网络、存储、主板、风扇和 PSU 转换损耗。设施功率另行
-估算:先算机箱交流功率,再应用 PUE,并保留模型的舍入顺序。
-
-应用 profiles 保留固定推理假设:`u_cpu=0.20`、`u_ram=0.20`、`u_pcie=0.05` 和
-`u_nvme=0.0`。profile 基线包含默认 PUE `1.2`;PowerX 对当前支持的风冷
-机箱 profile 使用 `1.3`。市电侧功率 = IT 负载功率 × PUE(风冷 `1.3`,直接液冷 DLC
-`1.1`)。该系数作用于机箱交流功率之后,不改变 GPU 实测功率或机箱交流功率。
-这里的冷却方式指模型中的机箱,并非已经核实的基准测试站点冷却配置。机箱 profile
-不支持 DLC;`--pue` 只是显式覆盖设施功率系数,不会把风冷机箱模型转成液冷模型。
-下文的 NVL72 机架 profile 为直接液冷,默认使用 `1.1`。
-
-应用内可编辑的 profile 保留各平台的网络假设、风扇控制、组件数量和机箱默认值。每份 JSON
-导出包含完整 profile,每行 CSV 包含适用假设、模型版本和源码哈希。`u_cpu`、
-`u_ram`、`u_ib` 等利用率假设记录的是推导当前参与计算的系数时采用的场景,不是实时利用率
-控件,也不是 CPU/DRAM 利用率实测值。只修改这些标签,不会改变固定功率或曲线。
-
-下表列出
-[system-power-model.profiles.json](../packages/app/src/lib/system-power-model.profiles.json)
-中参与计算的配置项,对应公式由
-[system-power-model.ts](../packages/app/src/lib/system-power-model.ts) 中的函数实现。
-
-| 硬件标识 | 应用内可编辑的 profile | TypeScript 计算函数 |
-| -------- | ---------------------- | ---------------------- |
-| `h100` | `profiles.h100` | `estimateChassisPower` |
-| `h200` | `profiles.h200` | `estimateChassisPower` |
-| `b200` | `profiles.b200` | `estimateChassisPower` |
-| `b300` | `profiles.b300` | `estimateChassisPower` |
-| `mi300x` | `profiles.mi300x` | `estimateChassisPower` |
-| `mi325x` | `profiles.mi325x` | `estimateChassisPower` |
-| `mi355x` | `profiles.mi355x` | `estimateChassisPower` |
-
-上述 profile 均描述完整八卡机箱。GB200 和 GB300 使用独立的 NVL72 机架 profile,
-不套用这些机箱拓扑:
-
-| 硬件标识 | 应用内可编辑的 profile | TypeScript 计算函数 |
-| -------- | ---------------------- | ------------------- |
-| `gb200` | `rackProfiles.gb200` | `estimateRackPower` |
-| `gb300` | `rackProfiles.gb300` | `estimateRackPower` |
-
-[参考测试数据](../packages/app/src/lib/system-power-model.reference.json) 保留历史源码及
-版本信息、组件哈希和 `pythonConfigurations`,仅用于追溯基线来源。这些记录不是应用
-当前使用的参数,也不意味着应用仍依赖 Python。
-
-机架 profile(`rackProfiles`)的输入是每个 tray 的**实测**计算模块功耗:生产端发布
-`avg_total_module_power_w` 时使用 module 传感器总值;否则使用 GPU 板卡功耗加
-Grace socket 总功耗(`avg_total_cpu_power_w`),并按应用 profile 对 GPU 份额计入稳压损耗
-余量。Grace CPU 和 LPDDR5X 从不由模型补算。缺少 `cpu_power_valid=1` 或完整 module /
-Grace 来源记录的行保持不可用(`cpu-telemetry`)。
-
-每个实测 worker 主机对应一个计算 tray(四张 GPU、两个 Grace socket)。没有逐 worker
-数组的聚合多节点记录,按 `gpuCount / 4` 推算 tray 数,各 tray 采用部署平均值,并与
-Grace socket 数及 CPU 采集记录中的 `power_audit.cpu.observed_sockets` 交叉校验。
-按实测 tray 的平均计算模块输入,构造一个包含 18 个同等负载 tray 的机架;电源架效率
-曲线只在整机架直流负载处求值一次,使用单一的 tray 平均输入。每个 tray 分摊 1/18,
-因此 NVSwitch tray、电源架和管理
-交换机按 72 张 GPU 分摊。机箱则各自拥有风扇和 PSU,按各自的负载单独求值。
-结果包含 `topologyBasis: 'nvl72-trays'`、`measuredBasis` 和 `sensorKind`。
-
-部分分配的 tray 只外推 GPU 板卡份额,因为 module 读数本身已覆盖整个 tray;结果标记
-为 `extrapolated`。应用维护的模型仍为 DRAFT / pending human verification。
-下文列出 NVL72 的实测输入、其余组件模型,以及利润估算器的接纳规则。
-
-部分分配的机箱,即单台主机上实测一至七张 GPU,按“实测每 GPU 功率 × 8”建模。
-这一 `n_gpu × W/GPU` 输入沿用历史基线,并假设未实测的 GPU 运行相同负载。估算标记为 `chassisBasis: 'extrapolated'`:每 GPU 数值按建模机箱 GPU 数
-(`modeledGpuCount`)分摊;`deploymentAcWatts` / `deploymentFacilityWatts` 只保留
-各机箱中实测 GPU 的份额。这不是先在部分负载处计算机箱,再按比例分摊;固定组件、
-风扇曲线和 PSU 效率都在满机箱负载处求值。缺失或无效遥测、数量不一致、缺少主机
-放置记录、每主机多于一个机箱,以及超出模型适用范围的情况,仍保持不可用。
-
-单节点部署的物理宽度按生产端的 `TP * PP * PCP` 计算,EP 在该宽度内划分。
-部分已有 API 配置别名含有 `TP * EP`;模型不直接相信或相加这些别名,而是用实测
-总功率和每 GPU 功率交叉核对物理宽度。多节点和分离式输入要求每个实测 worker
-对应一个机箱(一至八张 GPU)、worker 位于不同主机,并且总功率与角色功率一致。
-只有角色平均值,无法证明物理放置方式,也无法计算各主机的非线性模型。
-纯 CPU frontend worker 不计入 GPU 机箱数量。独立的纯 CPU frontend/router 主机
-不在估算范围内;GPU 机箱内的 CPU 功耗仍采用固定系数,这些系数按应用 profile 中 20% 利用率场景推导得出。
-
-默认实测约定要求数值型 `power_valid=1` 和指标 schema 2。原有通过验证的单节点生产端
-早于 schema 标记,但两个 watts 字段的定义已与 schema 2 相同。这条路径保留 schema 缺失的
-原状,并报告 `validated-unversioned-single-node`;不会升级源数据版本,也不会接纳
-无版本的分离式功耗。文章的验证回执还会固定生产端 checkout,并保留各原始审计产物。
-
-## NVL72 机架估算(GB200、GB300)
-
-**实测输入。**每个计算 tray 都使用实测计算模块功耗,Grace CPU 和 LPDDR5X 不由模型
-估算。生产端的 CPU 功耗采集(srt-slurm,ACPI hwmon)在与 GPU 能耗相同的正式窗口内
-输出 `avg_cpu_socket_power_w`、`avg_total_cpu_power_w` 和 `total_cpu_energy_j`,并附带
-独立验证结论 `cpu_power_valid` 及 `power_audit.cpu`(传感器类型、采集器、socket 覆盖情况、
-原因码)。若每个 socket 都有 `Module Power Socket` 传感器,还会输出
-`avg_total_module_power_w` 和 `total_module_energy_j`。
-
-接纳条件为 `power_valid=1`、schema 2、`cpu_power_valid=1`,且 `power_audit.cpu` 中
-预期和实测 socket 数一致,每 tray 两个。module 读数要求 `sensor_kind: module`,
-无需重复提供 Grace 指标。GPU 加 Grace 路径要求 `sensor_kind: grace_socket`、Grace
-功率为正,且总功率和平均功率与审计 socket 数一致。仅有 CPU rail 测量,或传感器
-来源缺失、未知的记录,保持不可用。
-
-口径选择:存在 `avg_total_module_power_w` 时采用 `module`,因为读数已包含 GPU
-板卡,所以不会再缩放;否则采用 `gpu-plus-grace`,即每 tray 的每 GPU 板卡功率 × 4
-加 Grace socket 总功率,并且只对 GPU 份额应用 profile 的稳压损耗余量
-`regulatorLossFracOfTdp / (1 − frac)`。module 字段存在但无效时,该行不可用
-(`cpu-telemetry`),不会悄然回退到 Grace socket。
-
-**其余组件的模型估算。**计算模块以外的部分均来自应用 profile(`rackProfiles`)。
-按实测 tray 的平均输入构造 18 tray 机架,整体求值一次,再按 72 张 GPU 分摊。
-电源架曲线使用整机架的直流负载,不能只使用单个 tray 的负载。标为 UNVERIFIED 的
-参数在 `unverifiedParameters` 中记录了范围,但没有已发布的供电轨测量:
-
-| 组件(未注明时按每机架计) | GB200 | GB300 | 来源状态 |
-| -------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------- | ---------------------------------------- |
-| NVSwitch tray 芯片(9 个 tray) | `u_nvlink` 为 0.5 时,每 tray 406.4 W | 相同 | `blackwell_nvswitch` 模型 |
-| NVSwitch tray 其余组件 | 每 tray 50 W | 相同 | UNVERIFIED(20–80 W) |
-| 计算 tray 网卡及光模块 | ConnectX-7,每 tray 4 × 31.5 W = 126 W | 集成 PCIe 的 ConnectX-8,4 个 NIC 合计 315 W(每个 78.8 W,已舍入) | `generic/connectx7`、`generic/connectx8` |
-| 计算 tray BlueField-3 DPU | 每 tray 空闲功耗 2 × 65 W = 130 W | 相同 | `generic/dpu`,仅空闲功耗 |
-| 计算 tray NVMe | 每 tray 空闲功耗 22 W | 相同 | `generic/nvme`,仅空闲功耗 |
-| 计算 tray 风扇 | 每 tray 130 W | 相同 | UNVERIFIED(40–220 W) |
-| 计算 tray 主板其余组件 | 每 tray 40 W | 相同 | UNVERIFIED(20–60 W) |
-| 管理交换机 | 2 × 100 W | 相同 | profile 常量 |
-| tray 内 50 V → 12 V 转换 | tray 负载处效率 0.9725 | 相同 | UNVERIFIED(0.96–0.985) |
-| 稳压损耗余量(`gpu-plus-grace`) | GPU 板卡功率 × 0.15 / 0.85;仅计 GPU 板卡功率,不含 Grace | 相同 | Grace 调优指南 |
-| 电源架 | 装机容量 264 kW,冗余容量 132 kW;在 10/20/30% 负载处效率为 0.90 → 0.94 → 0.965 | 相同 | profile 曲线 |
-| 设施 PUE | 1.1(直接液冷) | 相同 | PowerX 策略,仅对机架交流功率应用一次 |
-
-机架直流功率超过电源架装机容量(264 kW)时,超出效率曲线适用范围,该行不可用
-(`model-domain`)。模型先将机架交流功率舍入到 0.1 W,再应用 PUE。已提交的
-`rackCases` 保留历史 Python 基线,涵盖两种变体、两种口径、全部电源架曲线节点和
-PUE 1.0–1.2。它们用于回归对照,不是校准证据,也不构成对原仓库的依赖。
-
-**利润估算器的接纳规则。**规划 kW/GPU = 部署设施功率 ÷ 实测 GPU 数 ÷ 1000 × 1.1。
-接受完整实测的八卡机箱(`single-node`、`worker-hosts` 或 `uniform-hosts` 拓扑下的
-`chassisBasis: 'full'`,见下一节),或者所有 tray 都完整实测的 `nvl72-trays` 估算。
-后者要求每 tray 四张 GPU、两个 socket:每个实测 worker 主机对应一个 tray;没有逐
-worker 数组的聚合多节点记录则按 `gpuCount / 4` 推算 tray 数,各 tray 采用部署平均值,
-并与 Grace socket 数和 `power_audit.cpu.observed_sockets` 交叉校验。部分 tray 可在
-功耗图中外推,但不进入利润规划。单节点 1/2/4 卡机箱采用下文的实例外推。
-
-两个前沿点必须采用相同实测口径和传感器类型;若一端是 module、另一端是 Grace
-socket,结果不可用,不混合两种传感器。柱形提示框、默认折叠的 Power assumptions
-详情,以及 CSV 中的 `Power basis`、`Power sensor`、`System power profile` 列,会逐行
-列明口径、传感器类型和应用 profile。口径为实测 module,或实测 GPU 板卡 + Grace
-socket 并由模型估算稳压损耗;profile 表示为
-`modelPath @ modelRevision sha256:`。路径和版本标识应用维护的模型;
-源码哈希标识 TypeScript 公式文件,模型版本还涵盖可编辑 profile 和接纳/PUE 适配层。
-`?unofficialrun=` 叠加层规则不适用于利润估算器的功耗口径控件;估算器只对正式前沿点计价。
-
-## 利润估算器的功耗口径
-
-每 GW 利润估算器在 Benchmark Config 中提供预配功耗、实测加建模功耗,以及两者的
-成对对比,默认仍采用预配功耗。控件由现有内部功能开关控制,锁定时隐藏。
-按 ↑↑↓↓ 解锁(本地存储 `inferencex-feature-gate=1`)。锁定时,`c_power` 不会启用
-其他计算方式或请求完整功耗记录;重新锁定后立即恢复预配估算。
-
-另一种功耗口径复用相同硬件、P90 目标、吞吐量前沿、token 比例、价格、利用率和每 GPU
-小时成本,只改变换算每 GW 容量时采用的设施 kW/GPU。因此,收入、计算成本、模型
-许可费和利润按相同比例变化,利润率不变。不会另行重新计算电费。
-
-这项显式启用的 AgentX 估算要求通过验证的 schema-v2 遥测和适用的系统 profile。
-完整八卡机箱可采用单节点或逐 worker 主机的实测功耗。没有逐 worker 遥测的聚合
-多节点记录,可以显式采用 `uniform-hosts` 假设:每个完整机箱都按部署平均 GPU 功率
-计算。分离式记录要求逐主机、逐角色功耗。NVL72 要求完整四卡 tray,并具有上文所述
-CPU 来源证据。
-
-通过验证的单节点 1/2/4 卡配置保留整机箱外推:按实测的每 GPU 功耗和吞吐量,用完整实例
-填满一台八卡服务器,再将建模设施功耗除以八。该假设认为实例同机部署不影响性能
-或功耗,不代表测量了部分 GPU 闲置的服务器。图表、提示框和 CSV 对所有外推估算
-作出标注,包括仅一端为部分分配点的插值。其他部分分配方式以及缺失、无效测量仍不可用,
-并给出不同原因。常规 8K/1K 转换路径保留原有接纳策略。
-
-即使实测加建模功耗不可用,对比模式仍保留每个有效预配估算,并在提示中指出缺失的
-实测加建模估算;只看建模功耗的模式不会用预配功耗代替。
-
-目标恰好落在前沿点上时,使用该点的建模功耗;目标在两点之间时,采用原吞吐量插值
-的同一对点,线性估算功耗。不换用其他点填补缺失,也不混合测量口径不同的两点。
-估算采用 PowerX 的 PUE 策略(风冷机箱 1.3,直接液冷 NVL72 机架 1.1),另加 10%
-规划余量。这些假设,包括上述固定 CPU/DRAM 利用率,不构成峰值供电容量验证,也未
-经过 AgentX 系统校准。
-
-界面保留简短的实测与建模说明,详细假设和 NVL72 实测输入放在默认折叠的详情中。
-CSV 保留完整假设,并逐行记录实测口径、传感器类型和 profile。分享链接通过
-`c_power=modeled` 或 `c_power=compare` 保留选择。历史估算不可用且当天结果不含该
-芯片时,从硬件注册表获取芯片信息,并显示来源日期/运行标签。
-
-## 离线对比导出
-
-导出器读取本地数据组封装文件,并写入一个**新的**输出目录:
+若基准测试使用两个完整 tray、共八张 GPU,其设施功率份额为
+82,395.1 × 8 / 72 ≈ 9,155.0 W,归一到每张 GPU 后结果相同。这不代表测量了全部
+72 张 GPU。中间值经过舍入,展示值相加可能相差 0.1 W。规划计算使用保留的设施
+总功率,不使用参考用例中已舍入的单卡展示值。
+
+## 估算采用的假设
+
+- **PUE:**仪表板对风冷机箱使用 1.3,对 NVL72 使用 1.1,均在 AC 转换损耗之后
+ 应用。历史 profile 的默认值 1.2 不是仪表板默认值。冷却类型描述的是模型,不是
+ 经核实的现场实际冷却方式;改变 PUE 不会把风冷机箱模型变成 DLC 模型。
+- **机箱开销:**系数采用固定假设:CPU/DRAM 利用率 20%、PCIe 利用率 5%、NVMe
+ 空闲。这些不是实时利用率读数。各主机的非线性风扇/PSU 模型按本机负载计算;
+ `uniform-hosts` 则明确假设每台主机都采用部署平均负载。
+- **NVL72 开销:**Grace 和 LPDDR5X 使用实测值;机架网络、交换机、风扇、主板
+ 其余开销、转换损耗和电源架由模型估算。[Profiles](../packages/app/src/lib/system-power-model.profiles.json)
+ 保留组件参数、来源状态和未核实参数的范围。机架 DC 超过已安装电源架容量
+ 264 kW 时,超出模型定义域。
+- **规划余量:**设施 kW/GPU × 1.10 是独立的容量缓冲。平均功率加此余量不等于
+ 经过验证的供电峰值上限。
+- **匹配比较:**各利润模式保持原性能前沿、目标、token 价格、利用率、授权分成
+ 和每 GPU 小时成本一致。目标位于两个前沿数据点之间时,只在这两个原始点之间
+ 插值规划功率,且模型版本、PUE、拓扑和传感器边界须兼容;不做范围外推,也不
+ 换用其他功耗有效的数据点。
+
+较低的规划功率会提高每 GW 可容纳的 GPU 数量。在固定单卡假设下,收入、计算成本
+和授权费用都随容量变化;利润率和每芯片小时的经济指标不会改善,也不会另行重算电费。
+
+## 曲线或估算结果缺失时如何排查
+
+先确认比较的是同一模型、工作负载、日期/运行、引擎、精度和指标。图表/表格记录与
+按性能目标计算的利润估算,回答的问题不同。
+
+| 现象 / 原因 | 检查项与下一步 |
+| ----------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
+| 有 GPU 曲线,但没有 All in Measured | 检查工作负载、硬件是否受支持,以及该记录的系统模型状态。GPU 功耗有效不等于系统估算可用。 |
+| B200/H200 多节点记录缺失(`topology`、`role-power`、`gpu-count`) | 检查物理 GPU 数量、主机分布和总功率/角色功率。使用原生产端拓扑,不从展示名称推断机箱分布,也不累加 TP/EP 别名。 |
+| NVL72 返回 `cpu-telemetry` / `no-cpu-power` | 检查同一窗口的 CPU 审计、传感器类型和完整 socket 覆盖;GPU 有效性独立判断。 |
+| `telemetry` / `no-measured-power` | 检查原验证审计及原始样本。只有保留证据足以支持原窗口时才重处理,否则须重新采集匹配的性能和功耗。 |
+| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前所选曲线支持的目标。两个原始端点的功耗都须有效,曲线上其他位置的有效点不能补齐这个缺口。 |
+| `incompatible-power-basis` | 不在 module 读数与 GPU 加 Grace 读数之间插值,也不在不同模型版本或 PUE 取值之间插值。 |
+| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 |
+| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 |
+
+**Compare both** 在实测估算不可用时仍保留有效预配结果。完全无法定价的配置,与
+仅缺少实测估算的配置分开列出;仅实测模式不会用预配功率代替。保留遥测和定向修复
+方式见[持久化与恢复](./powerx-persistence-recovery.md);新采集的功率不能附到旧吞吐量上。
+
+## 来源与可复现导出
+
+核对数据点的测量来源,以及估算的模型版本、PUE、传感器边界和拓扑。利润图表提示框
+标明功耗依据;公式说明及 CSV 中的 `Power basis`、`Power sensor`、
+`System power profile` 列提供假设和来源细节。模型版本是 `app-sha256:` 摘要,不是基准
+测试运行 ID,也不是 Git commit。
+
+[离线导出器](../packages/app/scripts/export-modeled-system-power.ts)读取本地
+`ComparisonInput`,其中包含原始 `BenchmarkRow`、元数据,以及可选的原始产物/
+审计。完整结构以脚本中的类型为准;保留运行、attempt、生产端版本、获取时间和哈希。
+导出器采用常规 8K/1K 模型策略,不启用 AgentX 预览。安装项目依赖后,从仓库根目录
+运行,并指定一个新输出路径:
```sh
bun packages/app/scripts/export-modeled-system-power.ts \
- --input /path/to/original-qwen-article.input.json \
- --output /path/to/new-original-qwen-comparison
-
-bun packages/app/scripts/export-modeled-system-power.ts \
- --input /path/to/qwen35-current.input.json \
- --output /path/to/new-current-qwen-comparison --pue 1.3
+ --input /path/to/cohort.input.json \
+ --output /path/to/new-comparison
```
-未指定 `--pue` 时,每行通过与图表相同的 `modelSystemPower` 路径,采用该硬件在
-仪表板中的默认值:风冷机箱 profile 为 `1.3`,直接液冷 NVL72 机架 profile 为 `1.1`。
-因此文章数值与图表悬停值一致。显式传入 `--pue` 时,它适用于所有行,并记录在
-`metadata.pue_override` 中;`metadata.pue_defaults` 记录各硬件的默认值。
-
-NVL72 行还包含 `measured_basis`、`sensor_kind`、按独立 `cpu_power_valid` 保留的
-Grace socket 与 module 实测输入、机架 profile 的 `model_path` 和假设,以及机架专用的
-`calculation_boundary` / `extrapolation_note`。x86 行保持不变。
-
-脚本中维护的输入类型为 `ComparisonInput`:
-
-```ts
-{
- cohort: string,
- metadata: { /* source URLs, capture times, hashes and cohort selection */ },
- rows: [{
- id: string, // stable observation identity
- cell?: string, // optional group of original replicates
- benchmark: BenchmarkRow, // original API row or existing ETL output
- rawInput?: unknown, // original artifact before ETL normalization
- source?: object, // run, attempt, producer revision and artifact receipt
- audit?: object // matching original power-validation sidecar
- }]
-}
+输出包括 `comparison.json`、`comparison.csv` 和可选的 `cells.csv`。JSON 保留原始
+记录、审计、有效性、模型输出和来源;CSV 中不可用数值留空。可选参数 `--pue 1.3`
+会覆盖**所有**记录的设施系数,并写入元数据;不传时采用各硬件默认值。实测数据提取
+方法见 [API 示例](./inferencex-api-examples.md)。
+
+建模能耗要求匹配的审计提供精确遥测时长、物理 GPU 数量,以及成功请求/token 分母。
+计算为建模部署功率 × 时长,不是实测墙上功率的积分,也不会推断 kernel 级
+prefill/decode 能耗。各次重复测试先独立建模再聚合;只要范围内任一次不可用,该单元
+的均值就保持不可用。
+
+## 维护模型
+
+| 职责 | 源码 |
+| ----------------------------------- | ---------------------------------------------------------------------------- |
+| 公式、非线性曲线和舍入 | [system-power-model.ts](../packages/app/src/lib/system-power-model.ts) |
+| 组件参数、假设和来源状态 | [profiles](../packages/app/src/lib/system-power-model.profiles.json) |
+| 工作负载、遥测、拓扑和 PUE 接纳规则 | [modelSystemPower](../packages/app/src/lib/modeled-system-power.ts) |
+| 匹配前沿和规划余量 | [profit-power.ts](../packages/app/src/components/calculator/profit-power.ts) |
+| 冻结的数值基线 | [参考用例](../packages/app/src/lib/system-power-model.reference.json) |
+| 模型标识和源码哈希 | [provenance](../packages/app/src/lib/system-power-model.provenance.json) |
+
+修改公式或生效系数时,同时提供有依据的期望值和假设元数据。仅修改 `u_cpu` 等标签
+不会改变功率。刷新并检查 manifest:
+
+```sh
+bun packages/app/scripts/update-system-power-provenance.ts
+bun packages/app/scripts/update-system-power-provenance.ts --check
```
-代码保留原类型和注释:`metadata` 记录来源 URL、采集时间、哈希和数据组选择条件;
-`id` 是稳定的观测标识,`cell` 可将原始重复测量分组;`benchmark` 是原 API 行或现有
-ETL 输出;`rawInput` 保留 ETL 规范化前的产物;`source` 记录运行、attempt、生产端版本
-和产物回执;`audit` 是与该观测匹配的原始功耗验证 sidecar。
-
-原始生产端聚合数据使用现有 `normalizeArtifactRows` / `mapBenchmarkRow` 处理。
-将原聚合数据保留为 `rawInput`,保留原 schema 标记,并确认规范化没有改变实测指标。
-当前快照应使用完整原始 API 响应,再在本地筛选精确的 `single_turn`、`isl=8192`、
-`osl=1024` 负载。保留范围内所有行,包括不支持的硬件以及缺失/无效功耗;不得把
-当前快照混入文章冻结的数据集。
-
-输出为 `comparison.json`、`comparison.csv`,以及可选的 `cells.csv`。JSON 包含原始
-输入行、实测有效性、建模输出、假设、审计窗口和来源记录。CSV 将实测与建模字段
-分开,不可用数值留空。无效或未经验证的原始值仍保留在原始输入记录中,不标为有效
-测量。元数据记录输入与实现的 SHA-256 哈希、模型与应用版本、worktree 状态、生成时间
-和完整 profile 来源。最终发布导出应从指定应用提交生成;文件哈希也能标识开发期间
-的本地修改。
-
-只有匹配且有效的审计 sidecar 提供了精确遥测时长、物理 GPU 数,以及成功请求/token
-分母时,才能估算能耗。计算方式是建模部署功率,即实测 GPU 在各机箱中的份额,乘以
-该时长;它不是对实测墙插功率做时间积分。输出 token 使用实际计数,不能用每次查询
-名义上的 `1024` token 代替。不推导 kernel 层面的 prefill/decode 能耗。
-没有这些 sidecar 的 API 快照只导出功耗估算。
-
-每次重复测量先独立建模,再求聚合值。cell 均值取各重复测量的建模输出平均值,不能
-先平均功耗再运行模型。任何一次重复测量不可用时,对应均值也保持不可用,不能默默
-丢掉那次测量后再平均。
-
-## GPU 实测 P75 和 P90 功耗
-
-`y_measuredP75Power` 和 `y_measuredP90Power` 分别表示:在已验证负载窗口内,对同步
-采集的整组 GPU 板卡功耗计算时间加权 P75、P90,再除以 GPU 数量。正式数据点与
-非正式运行叠加层共用常规实测功耗图路径。分位数缺失或未通过验证时保持不可用,
-绝不用平均功耗替代。这些指标与建模机箱交流功耗、单设备功耗分位数不同。
-
-P75 和 P90 回填使用 `docs/data/power-p90-backfill.json` 中记录的同一批 34 份原始
-有效遥测,以及完全相同的测量窗口。
+运行受影响的模型、接纳规则、规划、views API 和导出检查。保留 496 个历史参考用例,
+不要重写其基线;数值一致不代表完成校准。模型和参数由应用维护,不依赖私有 Python
+仓库。历史测量数据在读取时建模,因此仅修改模型需要更新应用包并刷新 API 缓存或
+等待其过期,不需要回填原始数据。冻结的导出文件需另行生成,缺失的源测量仍然缺失。
diff --git a/packages/app/cypress/e2e/powerx-compare.cy.ts b/packages/app/cypress/e2e/powerx-compare.cy.ts
index 159b01683..24290d316 100644
--- a/packages/app/cypress/e2e/powerx-compare.cy.ts
+++ b/packages/app/cypress/e2e/powerx-compare.cy.ts
@@ -216,3 +216,141 @@ describe('PowerX article panels', () => {
});
});
});
+
+// Aggregate role counts describe shared devices: B200 TP8 × PP2, H200 TP16 × 2 workers.
+const agenticRows = [
+ ['b200', 'dynamo-vllm', 8, 2, 1, 16, 760],
+ ['h200', 'vllm', 16, 1, 2, 32, 168],
+ ['gb200', 'trt', 4, 1, 1, 4, 450],
+ ['b300', 'vllm', 8, 1, 1, 8, 0],
+].map(([hardware, framework, tp, pp, replicas, chips, watts], index) => ({
+ ...rows(null, 'b200')[0],
+ id: 990100 + index,
+ model: 'kimik3',
+ hardware,
+ framework,
+ benchmark_type: 'agentic_traces',
+ disagg: false,
+ isl: null,
+ osl: null,
+ prefill_tp: tp,
+ decode_tp: tp,
+ prefill_num_workers: replicas,
+ decode_num_workers: replicas,
+ num_prefill_gpu: chips,
+ num_decode_gpu: chips,
+ metrics: {
+ power_valid: watts ? 1 : 0,
+ power_metric_schema_version: 2,
+ avg_power_w: watts,
+ avg_total_gpu_power_w: Number(watts) * Number(chips),
+ prefill_pp: pp,
+ decode_pp: pp,
+ p90_itl: 0.02 + index * 0.01,
+ median_itl: 0.01 + index * 0.01,
+ median_intvty: 100 / (index + 1),
+ tput_per_gpu: 200 + index * 100,
+ output_tput_per_gpu: 100 + index * 50,
+ joules_per_output_token: 2 + index,
+ },
+}));
+
+describe('AgentX All in Measured chart and table', () => {
+ for (const [locale, width] of [
+ ['en', 1280],
+ ['zh', 390],
+ ] as const) {
+ it(`retains B200/H200 multinode rows and export at ${locale} ${width}px`, () => {
+ cy.viewport(width, 900);
+ cy.intercept('GET', '/api/v1/availability', { body: agenticRows }).as('agenticAvailability');
+ cy.intercept('GET', '/api/v1/benchmarks*', { body: agenticRows }).as('agenticBenchmarks');
+ cy.intercept('GET', '/api/v1/workflow-info*', {
+ body: { runs: [], changelogs: [], configs: [] },
+ });
+ cy.intercept('GET', '/api/v1/trace-availability*', { body: {} });
+ cy.intercept('GET', '/api/v1/log-availability*', { body: {} });
+ cy.intercept('GET', '/api/v1/resident-sequence-lengths*', { body: {} });
+ cy.intercept('GET', '/api/unofficial-run*', {
+ body: {
+ runInfos: [
+ {
+ id: OVERLAY_RUN_ID,
+ name: 'agentic-power',
+ branch: 'agentic-power',
+ sha: 'abc000',
+ createdAt: `${DATE}T00:00:00Z`,
+ url: OVERLAY_RUN_URL,
+ conclusion: 'success',
+ status: 'completed',
+ isNonMainBranch: true,
+ },
+ ],
+ benchmarks: [{ ...agenticRows[1], id: 0, run_url: OVERLAY_RUN_URL }],
+ evaluations: [],
+ },
+ }).as('agenticOverlay');
+ let csvBlob: Blob | undefined;
+ cy.visit(
+ `${locale === 'zh' ? '/zh' : ''}/inference?g_model=Kimi-K3&i_seq=agentic-traces&i_prec=fp4&i_pctl=p90&i_metric=y_utilityModeledWatts&i_optimal=0&i_best=0&unofficialrun=${OVERLAY_RUN_ID}`,
+ {
+ onBeforeLoad(win) {
+ win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now()));
+ win.localStorage.setItem('inferencex-feature-gate', '1');
+ win.URL.createObjectURL = (object) => {
+ if (object instanceof win.Blob) csvBlob = object;
+ return 'blob:agentic-power';
+ };
+ win.HTMLAnchorElement.prototype.click = () => {};
+ },
+ },
+ );
+ cy.wait(['@agenticAvailability', '@agenticBenchmarks', '@agenticOverlay']);
+ cy.get('[data-testid="chart-figure"]').first().find('.dot-group').should('have.length', 2);
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .find('.unofficial-overlay-pt')
+ .should('have.length', 1);
+ cy.get('[data-testid="power-agentic-model-note"]')
+ .first()
+ .should(
+ 'contain.text',
+ locale === 'en' ? 'not been independently calibrated' : '尚未针对 AgentX',
+ );
+ // A fixed page header otherwise repeats over the stitched element capture.
+ cy.get('header').invoke('css', 'visibility', 'hidden');
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .scrollIntoView()
+ .screenshot(`agentic-all-in-${locale}-chart`, { overwrite: true });
+ cy.get('[data-testid="inference-table-view-btn"]').first().click();
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .find('tbody tr')
+ .should('have.length', 3)
+ .then(($rows) => {
+ expect($rows.text()).to.contain('B200').and.contain('H200');
+ expect($rows.text()).not.to.contain('GB200').and.not.to.contain('B300');
+ });
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .screenshot(`agentic-all-in-${locale}-table`, { overwrite: true });
+ cy.get('header').invoke('css', 'visibility', '');
+ cy.get('[data-testid="export-button"]').first().click();
+ cy.get('[data-testid="export-csv-button"]').click();
+ cy.then(() => csvBlob!.text()).then((csv) => {
+ const [header, ...data] = csv.split('\n').filter((line) => !line.startsWith('#'));
+ const columns = header.split(',');
+ const values = data.map((line) => line.split(','));
+ expect(values.map((row) => row[columns.indexOf('Hardware')]).sort()).to.deep.equal([
+ 'b200',
+ 'h200',
+ 'h200',
+ ]);
+ expect(
+ values.map((row) => Number(row[columns.indexOf('Physical Chips')])).sort((a, b) => a - b),
+ ).to.deep.equal([16, 32, 32]);
+ });
+ cy.document().then((doc) => expect(doc.documentElement.scrollWidth).to.be.at.most(width));
+ });
+ }
+});
diff --git a/packages/app/src/components/inference/ui/ChartDisplay.tsx b/packages/app/src/components/inference/ui/ChartDisplay.tsx
index 923dbc9be..372decb65 100644
--- a/packages/app/src/components/inference/ui/ChartDisplay.tsx
+++ b/packages/app/src/components/inference/ui/ChartDisplay.tsx
@@ -16,7 +16,11 @@ import { metricRowLabel } from '@/components/inference/axis-metric-explanations'
import { getMeasuredMetricConfig } from '@/components/inference/measured-metric-config';
import { AIR_COOLED_SYSTEM_PUE, DLC_SYSTEM_PUE } from '@/lib/modeled-system-power';
import { SYSTEM_POWER_MODEL_REVISION } from '@/lib/system-power-model';
-import { ALL_IN_MEASURED_EMPTY, ALL_IN_MEASURED_NOTE } from '@/lib/power-basis';
+import {
+ ALL_IN_MEASURED_AGENTIC_NOTE,
+ ALL_IN_MEASURED_EMPTY,
+ ALL_IN_MEASURED_NOTE,
+} from '@/lib/power-basis';
import {
applyTokenRevenuePricing,
cachedInputPricePerMillion,
@@ -154,7 +158,7 @@ const STRINGS = {
'GPU Level Provisioned (TDP) · Watts are the rated TDP per GPU from the hardware registry, so the power curve is flat per hardware. Joules per output token = TDP × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together. Hardware without a published TDP is omitted.',
'utility-provisioned':
'All in Provisioned · Watts are the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model), so the power curve is flat per hardware. Joules per output token = all-in W × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together, unlike the ungated All-in Provisioned J per Output Token, which divides per decode GPU.',
- 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K on supported hardware; incomplete inputs are omitted.`,
+ 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K and AgentX on supported hardware; incomplete inputs are omitted.`,
},
vsTtft: (word: string) => `vs. ${word} Time To First Token`,
vsE2eLatency: (pctl?: string) =>
@@ -187,7 +191,7 @@ const STRINGS = {
'GPU 额定功耗(TDP)· 功率取硬件注册表中每 GPU 的额定 TDP,因此每种硬件的功率曲线为水平线。每输出 token 能耗 = TDP × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入。未公布 TDP 的硬件不绘制。',
'utility-provisioned':
'整体预配功耗 · 功率取硬件注册表中每 GPU 的全电源配置(all-in)市电功率(来源:SemiAnalysis Datacenter Industry Model),因此每种硬件的功率曲线为水平线。每输出 token 能耗 = all-in 功率 × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入,这与未加门控的“每输出 token 全电源配置能耗”按 decode GPU 计算不同。',
- 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。仅适用于受支持硬件的 8K / 1K 场景,输入不完整的数据点不绘制。`,
+ 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。适用于受支持硬件的 8K / 1K 和 AgentX 场景,输入不完整的数据点不绘制。`,
},
vsTtft: (word: string) => `vs. ${word === 'Median' ? '中位' : word} 首 token 延迟(TTFT)`,
vsE2eLatency: (pctl?: string) => (pctl ? `vs. ${pctl} 端到端延迟` : 'vs. 端到端延迟'),
@@ -1220,6 +1224,17 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
{ALL_IN_MEASURED_NOTE[locale]}
)}
+ {isAgenticSequence &&
+ selectedPowerBasis &&
+ (selectedPowerBasis === 'utility-modeled' ||
+ powerCompare === 'boundaries') && (
+
+ {ALL_IN_MEASURED_AGENTIC_NOTE[locale]}
+
+ )}
{isUnofficialRun &&
selectedXAxisMode === 'e2e-normalized-interactivity' && (
diff --git a/packages/app/src/components/inference/utils/tooltip-utils.test.ts b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
index e1182ab79..6a3709271 100644
--- a/packages/app/src/components/inference/utils/tooltip-utils.test.ts
+++ b/packages/app/src/components/inference/utils/tooltip-utils.test.ts
@@ -134,6 +134,26 @@ describe('modeled system-power tooltip', () => {
...overrides,
});
+ it.each(['en', 'zh'] as const)(
+ 'discloses the AgentX estimate in the %s All in Measured tooltip',
+ (locale) => {
+ const html = generateTooltipContent(
+ config({
+ locale,
+ selectedYAxisMetric: 'y_utilityModeledWatts',
+ data: pt({ modeledSystemPower: systemPower, benchmark_type: 'agentic_traces' }),
+ }),
+ );
+ expect(html).toContain('AgentX');
+ expect(html).toContain(
+ locale === 'en'
+ ? 'not been independently calibrated'
+ : '尚未针对 AgentX 工作负载进行独立校准',
+ );
+ expect(html).not.toContain('8k1k');
+ },
+ );
+
it('separates measured input, normalized chassis AC, and whole-deployment facility power', () => {
const html = generateTooltipContent(config());
expect(html).toContain('500 W/GPU');
diff --git a/packages/app/src/components/inference/utils/tooltipUtils.ts b/packages/app/src/components/inference/utils/tooltipUtils.ts
index 9ebfbb110..3590e5dda 100644
--- a/packages/app/src/components/inference/utils/tooltipUtils.ts
+++ b/packages/app/src/components/inference/utils/tooltipUtils.ts
@@ -8,6 +8,7 @@ import type { Locale } from '@/lib/i18n';
import { isKvOffloadEnabled } from '@/lib/kv-offload';
import { chartStateHref } from '@/lib/url-state';
import { chipCounts } from '@/lib/chip-counts';
+import { ALL_IN_MEASURED_AGENTIC_NOTE } from '@/lib/power-basis';
import type {
SystemPowerSensorKind,
SystemPowerUnsupportedReason,
@@ -15,6 +16,7 @@ import type {
import type { HardwareConfig, InferenceData, OverlayData } from '@/components/inference/types';
import {
+ isAllInMeasuredConfigKey,
isMeasuredEnergyConfigKey,
isModeledSystemPowerConfigKey,
} from '@/components/inference/metric-registry';
@@ -260,6 +262,7 @@ const escapeHtml = (s: string): string =>
const SYSTEM_POWER_STRINGS = {
en: {
heading: 'Draft System-Power Model · 8k1k',
+ agenticHeading: 'Draft System-Power Model · AgentX',
measuredGpu: 'Measured GPU power',
normalizedAc: 'Modeled chassis AC per GPU',
deploymentAc: 'Modeled deployment chassis AC',
@@ -311,6 +314,7 @@ const SYSTEM_POWER_STRINGS = {
},
zh: {
heading: '系统功耗模型(草案)· 8k1k',
+ agenticHeading: '系统功耗模型(草案)· AgentX',
measuredGpu: 'GPU 实测功耗',
normalizedAc: '每 GPU 分摊的机箱交流功耗估算',
deploymentAc: '整个部署的机箱交流功耗估算',
@@ -367,7 +371,8 @@ const modeledSystemPowerHTML = (
if (
!estimate ||
(!isMeasuredEnergyConfigKey(selectedYAxisMetric) &&
- !isModeledSystemPowerConfigKey(selectedYAxisMetric))
+ !isModeledSystemPowerConfigKey(selectedYAxisMetric) &&
+ !isAllInMeasuredConfigKey(selectedYAxisMetric))
) {
return '';
}
@@ -399,7 +404,8 @@ const modeledSystemPowerHTML = (
]
: [t.assumptions, t.platformAssumptions, t.normalization, t.boundary];
return `
-
${t.heading}
+
${d.benchmark_type === 'agentic_traces' ? t.agenticHeading : t.heading}
+ ${d.benchmark_type === 'agentic_traces' ? `
${ALL_IN_MEASURED_AGENTIC_NOTE[locale]}
` : ''}
${tooltipLine(t.measuredGpu, `${fmt(estimate.measuredGpuWattsPerGpu)} W/GPU`)}
${tooltipLine(t.normalizedAc, `${fmt(estimate.chassisAcWattsPerGpu)} W/GPU`)}
${
diff --git a/packages/app/src/lib/api-route-catalog.ts b/packages/app/src/lib/api-route-catalog.ts
index b989cadd6..9b849efd3 100644
--- a/packages/app/src/lib/api-route-catalog.ts
+++ b/packages/app/src/lib/api-route-catalog.ts
@@ -885,7 +885,7 @@ export const apiContractSourceDigests = [
},
{
source: 'src/lib/benchmark-transform.ts',
- sourceSha256: '53fb9102b1ddd0c597b3b2a414894564d2deda3da5ea2cf0059f997b2db852c8',
+ sourceSha256: 'e92e215ead748eafa2a6b496201adcd0f9387110d0843cb8fe7ebfdbcd8a59ef',
reviewArea: {
en: 'Raw benchmark means and derived reciprocal mean-TPOT interactivity used by Dashboard and read-only views.',
zh: '仪表板和只读视图共用的原始 benchmark 均值与 mean TPOT 倒数形式的 interactivity。',
diff --git a/packages/app/src/lib/benchmark-transform.ts b/packages/app/src/lib/benchmark-transform.ts
index 3c7263a56..788f5a382 100644
--- a/packages/app/src/lib/benchmark-transform.ts
+++ b/packages/app/src/lib/benchmark-transform.ts
@@ -236,7 +236,7 @@ export function rowToAggDataEntry(row: BenchmarkRow): AggDataEntry {
? row.power_invalid_reasons
: undefined,
power_metric_schema_version: m.power_metric_schema_version,
- modeledSystemPower: modelSystemPower(row),
+ modeledSystemPower: modelSystemPower(row, undefined, true),
power_tier: resolvePowerTier({
powerValid: m.power_valid,
wholeDeploymentSemantics: hasWholeDeploymentEnergySemantics,
diff --git a/packages/app/src/lib/chart-utils.ts b/packages/app/src/lib/chart-utils.ts
index a1ae03ac2..e9aae745f 100644
--- a/packages/app/src/lib/chart-utils.ts
+++ b/packages/app/src/lib/chart-utils.ts
@@ -442,7 +442,13 @@ export function buildDerivedChartFields(
if (wants(key)) fields[key] = value;
}
- if (wants('modeledChassisPowerPerGpu') && entry.modeledSystemPower?.status === 'supported') {
+ if (
+ wants('modeledChassisPowerPerGpu') &&
+ entry.benchmark_type === 'single_turn' &&
+ entry.isl === 8192 &&
+ entry.osl === 1024 &&
+ entry.modeledSystemPower?.status === 'supported'
+ ) {
fields.modeledChassisPowerPerGpu = chartMetric(entry.modeledSystemPower.chassisAcWattsPerGpu);
}
diff --git a/packages/app/src/lib/power-basis.test.ts b/packages/app/src/lib/power-basis.test.ts
index 2ba7c6e00..aee0a6d8f 100644
--- a/packages/app/src/lib/power-basis.test.ts
+++ b/packages/app/src/lib/power-basis.test.ts
@@ -4,6 +4,9 @@ import type { BenchmarkRow } from '@/lib/api';
import { rowToAggDataEntry, transformBenchmarkRows } from '@/lib/benchmark-transform';
import { buildDerivedChartFields, getHardwareKey } from '@/lib/chart-utils';
import { POWER_BASIS_FIELDS } from '@/lib/power-basis';
+import { Sequence } from '@/lib/data-mappings';
+import { buildInferenceSeries } from '@/lib/views-api/series';
+import { rowToLightweightPoint } from '@/components/inference/hooks/interpolated-trend-core';
// Qwen3.5 B200 c1, run 34175132645: actual rounded telemetry, eight GPUs
// (same fixture as modeled-system-power.test.ts) plus an output rate.
@@ -60,17 +63,105 @@ function derive(source: BenchmarkRow) {
}
describe('power boundaries through the derived-field builder', () => {
- it('keeps measured GPU boundaries when NVL72 CPU telemetry is unavailable', () => {
- const { entry, fields } = derive(row({ hardware: 'gb200' }));
- expect(entry.modeledSystemPower).toMatchObject({
- status: 'unsupported',
- reason: 'cpu-telemetry',
- });
- expect(fields.measuredAvgPower).toBeDefined();
- expect(fields.gpuProvisionedWatts).toBeDefined();
- expect(fields.utilityProvisionedWatts).toBeDefined();
- expect(fields.utilityModeledWatts).toBeUndefined();
- expect(fields.utilityModeledJPerOutputToken).toBeUndefined();
+ it.each([
+ ['b200', 8, 2, 1, 16, 759.752, 12156.029],
+ ['h200', 16, 1, 2, 32, 167.357, 5355.413],
+ ] as const)(
+ 'retains AgentX %s multinode estimates across chart and public view',
+ (hardware, tp, pp, replicas, gpus, watts, totalWatts) => {
+ // Topologies from Kimi K3 rows 441678 / 441866; aggregate role counts share devices.
+ const source = row({
+ hardware,
+ model: 'kimik3',
+ precision: 'fp4',
+ framework: hardware === 'b200' ? 'dynamo-vllm' : 'vllm',
+ benchmark_type: 'agentic_traces',
+ isl: null,
+ osl: null,
+ is_multinode: true,
+ prefill_tp: tp,
+ decode_tp: tp,
+ prefill_ep: hardware === 'h200' ? 32 : 1,
+ decode_ep: hardware === 'h200' ? 32 : 1,
+ prefill_dp_attention: hardware === 'h200',
+ decode_dp_attention: hardware === 'h200',
+ prefill_num_workers: replicas,
+ decode_num_workers: replicas,
+ num_prefill_gpu: gpus,
+ num_decode_gpu: gpus,
+ metrics: {
+ ...row().metrics,
+ avg_power_w: watts,
+ avg_total_gpu_power_w: totalWatts,
+ prefill_pp: pp,
+ decode_pp: pp,
+ p90_itl: 0.05,
+ },
+ });
+ const { entry, fields } = derive(source);
+ expect(entry.modeledSystemPower).toMatchObject({ status: 'supported', gpuCount: gpus });
+ expect(fields.utilityModeledWatts?.y).toBeGreaterThan(watts);
+ expect(fields.utilityModeledJPerOutputToken?.y).toBeGreaterThan(
+ source.metrics.joules_per_output_token,
+ );
+ expect(fields.modeledChassisPowerPerGpu).toBeUndefined();
+ const historical = rowToLightweightPoint({ ...source, date: '2026-08-01' }, [
+ 'utilityModeledWatts',
+ 'utilityModeledJPerOutputToken',
+ ]);
+ expect(historical?.utilityModeledWatts).toEqual(fields.utilityModeledWatts);
+ expect(historical?.utilityModeledJPerOutputToken).toEqual(
+ fields.utilityModeledJPerOutputToken,
+ );
+ const { chartData } = transformBenchmarkRows([source], 'p90', 'external');
+ expect(chartData[0][0].utilityModeledWatts).toEqual(fields.utilityModeledWatts);
+ const result = buildInferenceSeries([source], {
+ sequence: Sequence.AgenticTraces,
+ percentile: 'p90',
+ precisions: ['fp4'],
+ metricConfigKey: 'y_utilityModeledWatts',
+ xmode: 'interactivity',
+ xmetric: 'p90_ttft',
+ gpus: [],
+ quickFilters: { vendors: [], frameworks: [], deployment: [], spec: [], power: [] },
+ optimal: false,
+ best: false,
+ });
+ expect(result.count).toBe(1);
+ expect(result.series[0].points[0].y).toBe(fields.utilityModeledWatts?.y);
+ },
+ );
+
+ it.each(['single_turn', 'agentic_traces'] as const)(
+ 'keeps %s GPU boundaries when NVL72 CPU telemetry is unavailable',
+ (benchmark_type) => {
+ const { entry, fields } = derive(row({ hardware: 'gb200', benchmark_type }));
+ expect(entry.modeledSystemPower).toMatchObject({
+ status: 'unsupported',
+ reason: 'cpu-telemetry',
+ });
+ expect(fields.measuredAvgPower).toBeDefined();
+ expect(fields.gpuProvisionedWatts).toBeDefined();
+ expect(fields.utilityProvisionedWatts).toBeDefined();
+ expect(fields.utilityModeledWatts).toBeUndefined();
+ expect(fields.utilityModeledJPerOutputToken).toBeUndefined();
+ },
+ );
+
+ it.each(['power_valid', 'avg_power_w', 'avg_total_gpu_power_w'])(
+ 'omits AgentX All in Measured when GPU telemetry lacks %s',
+ (missingMetric) => {
+ const metrics = { ...row().metrics };
+ delete metrics[missingMetric];
+ const { entry, fields } = derive(row({ benchmark_type: 'agentic_traces', metrics }));
+ expect(entry.modeledSystemPower?.status).toBe('unsupported');
+ expect(fields.utilityModeledWatts).toBeUndefined();
+ expect(fields.utilityModeledJPerOutputToken).toBeUndefined();
+ },
+ );
+
+ it('retains the standalone 8K/1K chassis AC metric', () => {
+ expect(derive(row()).fields.modeledChassisPowerPerGpu?.y).toBeGreaterThan(0);
});
it('serves the same fields to ?unofficialrun= overlays through transformBenchmarkRows', () => {
diff --git a/packages/app/src/lib/power-basis.ts b/packages/app/src/lib/power-basis.ts
index 5a30ab8c6..649ae9743 100644
--- a/packages/app/src/lib/power-basis.ts
+++ b/packages/app/src/lib/power-basis.ts
@@ -39,9 +39,14 @@ export const ALL_IN_MEASURED_NOTE = {
zh: 'GPU 功耗来自实测;未实测的组件功耗由模型估算,并计入数据中心 PUE。',
};
+export const ALL_IN_MEASURED_AGENTIC_NOTE = {
+ en: 'AgentX estimates reuse the chassis or rack model; they have not been independently calibrated for AgentX workloads.',
+ zh: 'AgentX 估算复用机箱或机架功耗模型,尚未针对 AgentX 工作负载进行独立校准。',
+};
+
export const ALL_IN_MEASURED_EMPTY = {
- en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K, validated GPU telemetry, and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Choose another boundary to keep the points.',
- zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 场景、已验证的 GPU 遥测,以及受支持的机箱或机架功耗模型。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。可切换到其他功耗边界以保留数据点。',
+ en: 'No values are available for All in Measured in this selection. This boundary needs 8K / 1K or AgentX, validated GPU telemetry, and a supported chassis or rack power model. NVL72 also needs complete Grace or module telemetry from the same measurement window. Choose another boundary to keep the points.',
+ zh: '当前选择没有可用的整体实测功耗数值。该边界需要 8K / 1K 或 AgentX 场景、已验证的 GPU 遥测,以及受支持的机箱或机架功耗模型。NVL72 还需要同一测量窗口内完整的 Grace 或 module 遥测。可切换到其他功耗边界以保留数据点。',
};
/** InferenceData keys per derived basis and quantity. B1 lives on the measured* fields. */
@@ -180,8 +185,9 @@ export function powerBasisNormalization(
* telemetry admission: `modelSystemPower` requires `power_valid === 1` plus
* schema v2, or the validated unversioned single-node producer it records as
* `telemetryBasis: 'validated-unversioned-single-node'`. That is the same
- * population the app plots as B1 (`measuredAvgPower`) and as
- * `modeledChassisPowerPerGpu`, so B4 renders exactly where they do. The public
+ * measured population the app plots as B1 (`measuredAvgPower`). All in Measured
+ * also admits AgentX estimates; the separate `modeledChassisPowerPerGpu` metric
+ * retains its 8K/1K workload restriction. The public
* API's stricter `strictV2` row filter is not re-applied here; it is not
* applied to the chart's B1 either.
*/
diff --git a/packages/app/src/lib/views-api/docs/inference.ts b/packages/app/src/lib/views-api/docs/inference.ts
index 5530d2b22..8c9811e07 100644
--- a/packages/app/src/lib/views-api/docs/inference.ts
+++ b/packages/app/src/lib/views-api/docs/inference.ts
@@ -685,8 +685,8 @@ export const operations: ApiOperation[] = [
path: '/api/v1/views/inference',
summary: text('Get the main inference chart view', '获取主推理图表视图'),
description: text(
- 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection.',
- '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。',
+ 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection. All in Measured watts and energy support 8K/1K and AgentX with validated telemetry and supported topology; AgentX reuses the model without independent workload calibration. The standalone Modeled Chassis AC metric remains limited to 8K/1K.',
+ '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。整体实测功耗及能耗支持遥测已验证、拓扑受支持的 8K/1K 和 AgentX 数据;AgentX 复用同一模型,尚未针对该工作负载单独校准。单独列出的每 GPU 分摊的机箱交流功耗估算指标仍仅支持 8K/1K。',
),
audience: 'public',
stability: 'beta',
diff --git a/packages/skills/skills/inferencex-api/integrity.json b/packages/skills/skills/inferencex-api/integrity.json
index 164cffc62..6e4dc88cd 100644
--- a/packages/skills/skills/inferencex-api/integrity.json
+++ b/packages/skills/skills/inferencex-api/integrity.json
@@ -8,7 +8,7 @@
"references/cli-contract.md": "fa44faa38d889b4fbdee5ba42752b47758ee3706fab150513bd4e6db87a86cb0",
"references/cli.md": "96b228f34cb3600f4548f4dc84df8506531747286763c56ca834c77bee05e1eb",
"references/collectivex.md": "eb794f9c28d4a27bec4db80c42c4685ff3b1d204fd9b12a6788511fb258a789f",
- "references/dashboard-views.md": "42c21df31ee399116944a0fab6f29e6946423a2bc822efa43d2e8a7be5c91154",
+ "references/dashboard-views.md": "afa2422706e8aa0d8833ead3c4c3fccc0bee11c8ab45f73b2e54622f69d1100c",
"references/offline-exports.md": "95aa565dcd4c9e592159baa560a9a58217bc61e395de1c6daf0576795c7f4ec9",
"references/pareto.md": "1b4d2d163f982e3f2d97addce310789eae9c1e5badd82dd7501708c9e0385027",
"references/powerx.md": "cf8adcc5dfe395c5fb659908819980a872475f862f22b723815af43d363cedad",
diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md
index 7d40ed176..00e492558 100644
--- a/packages/skills/skills/inferencex-api/references/dashboard-views.md
+++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md
@@ -96,6 +96,12 @@ wall power. Metric IDs and API selector values are unchanged. Profit `powerBasis
still accepts `provisioned`, `modeled`, or `compare`; `powerLabel` is display text.
Expanding assumptions or unavailable-estimate details does not change returned data.
+All in Measured watts and energy accept validated 8K/1K and AgentX rows through the
+shared chart/API transform, including historical and unofficial rows. AgentX reuses
+the chassis or rack model without independent workload calibration. Telemetry and
+topology gates still apply; NVL72 needs complete Grace or module power. The standalone
+Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope.
+
NVL72 estimates require valid GPU power plus validated Grace-socket or compute-module
power with complete socket coverage. CPU-rail-only readings do not establish the
Grace/LPDDR boundary. A module reading already includes GPU power; do not add GPU
From 78585abe63f942c1e08e0aa18beef6f2b45d1031 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Thu, 1 Oct 2026 13:47:05 -0700
Subject: [PATCH 19/22] fix: retain measured rows without all-in estimates
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:整体功耗估算不可用时保留实测记录。
---
docs/dashboard-readonly-views.md | 15 ++-
docs/powerx-system-power.md | 13 +-
docs/powerx-system-power.zh.md | 5 +-
packages/app/cypress/e2e/powerx-compare.cy.ts | 110 ++++++++++++----
.../app/api/v1/views/inference/route.test.ts | 62 +++++++++
.../src/app/api/v1/views/inference/route.ts | 10 +-
.../components/inference/InferenceContext.tsx | 22 +++-
.../inference/hooks/useChartData.ts | 11 ++
.../app/src/components/inference/types.ts | 3 +
.../components/inference/ui/ChartDisplay.tsx | 77 ++++++++---
.../inference/ui/InferenceTable.tsx | 58 +++++++--
.../inference/ui/inference-table-sort.ts | 8 +-
.../src/components/inference/utils.test.ts | 32 +++++
.../app/src/components/inference/utils.ts | 18 ++-
.../utils/inference-table-data.test.ts | 107 ++++++++++++++++
.../inference/utils/inference-table-data.ts | 66 ++++++++++
packages/app/src/lib/api-route-catalog.ts | 12 +-
packages/app/src/lib/csv-export-helpers.ts | 24 +++-
.../app/src/lib/views-api/docs/inference.ts | 87 ++++++++++++-
packages/app/src/lib/views-api/series.test.ts | 121 ++++++++++++++++++
packages/app/src/lib/views-api/series.ts | 72 +++++++++++
.../skills/inferencex-api/integrity.json | 2 +-
.../references/dashboard-views.md | 10 +-
23 files changed, 863 insertions(+), 82 deletions(-)
create mode 100644 packages/app/src/components/inference/utils/inference-table-data.test.ts
create mode 100644 packages/app/src/components/inference/utils/inference-table-data.ts
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index 3f98dd8eb..cababa570 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -74,6 +74,13 @@ the chassis or rack model without independent workload calibration. Telemetry an
topology gates still apply; NVL72 needs complete Grace or module power. The standalone
Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope.
+For All in Measured, `tableRows` retains every GPU-valid observation in the selected
+scope and best-series selection, including axis-clipped and non-frontier points.
+Missing system estimates use `y: null`, `status: "unavailable"`, and
+`unavailableReason`; `measuredGpuWatts` remains available. CSV exports these rows
+with blank missing values. Numeric `series` and `count` are unchanged. Each date
+comparison and unofficial overlay has its own `tableRows`; latest does not pool history.
+
Dense profit charts reserve readable space per bar and scroll within the plot on narrow
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
@@ -214,8 +221,7 @@ point identity. Fewer than three distinct output rates return `fit: null` with
power, and R² is null when power did not vary.
These analytical results are JSON-only: `format=csv` with any analysis enabled returns
-400, rather than silently exporting only the primary chart. Ordinary CSV retains
-its existing plotted-point contract.
+400, rather than silently exporting only the primary chart. CSV exports plotted points except for All in Measured, which exports `tableRows`.
| Surface | Dashboard control / share parameter | Read-only API coverage |
| ---------------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------- |
@@ -246,6 +252,11 @@ those properties.
@semianalysisai/inferencex-skills 包。所有仪表板路由(含隐藏和功能开关控制的
视图)均在覆盖表中登记;上表列出各只读接口接受的全部查询参数名。
接口复用现有计算函数,公开运行与非官方叠加数据保留各自来源。
+整体实测指标的 `tableRows` 保留当前筛选范围和最优曲线选择内所有 GPU 遥测有效的观测点,
+不按前沿或坐标轴显示范围裁剪。估算不可用时 `y` 为 null,`status` 为 `unavailable`,
+`unavailableReason` 给出原因;`measuredGpuWatts` 保留实测 GPU 功耗。CSV 导出同一组行,缺失值留空。
+`series` 和 `count` 保持不变,仍只包含可绘制的数值点;各日期对比和非官方叠加分别返回自己的 `tableRows`,
+Latest 不会合并历史数据。
私有上传、密钥、提示词、反馈及管理操作不作为公开读取接口。
OperatorX 的入口受功能开关控制,页面使用专属的 `/api/v1/operatorx/*`
接口;目前没有发布 `/api/v1/views/operatorx` 契约。
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 0808b0ce8..456fb8d84 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -48,12 +48,13 @@ experimental power controls with ↑↑↓↓ if they are hidden.
inference selection. Expand **Unavailable estimates** for missing results;
hover or select a bar for its power basis and read the formula notes below.
-**Expected result:** the inference table includes rows with a valid selected
-metric, including supported B200/H200 multi-node deployments. Selecting All in
-Measured does not include every valid GPU measurement: it also requires the
-inputs below. The Profit Estimator adds target-range and financial requirements
-and uses only official frontier points. Inference charts and tables also
-support unofficial-run overlays.
+**Expected result:** the All in Measured table keeps every GPU-valid record in
+the selected scope, including B200/H200 multi-node deployments. It shows measured
+GPU power even when an all-in estimate is unavailable; the estimate displays
+`—` with a reason and stays blank in CSV. The graph plots numeric estimates only.
+The Profit Estimator adds target-range and financial requirements and uses only
+official frontier points. Inference charts and tables also support unofficial-run
+overlays.
## Hardware and telemetry requirements
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index 64b96e14f..7a370d502 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -39,8 +39,9 @@ AgentX 估算属于容量规划预览,不代表已完成 AgentX 校准,也
工作负载、引擎和精度是否与推理页面所选一致。展开 **Unavailable estimates**
查看缺失原因;悬停或选中柱形查看功耗依据,并阅读图表下方的公式说明。
-**预期结果:**推理表格列出所选指标有效的记录,包括满足要求的 B200/H200 多节点
-部署。All in Measured 不会包含所有 GPU 功耗有效的记录,还需满足下述输入要求。
+**预期结果:**All in Measured 表格保留当前筛选范围内所有 GPU 功耗有效的记录,
+包括 B200/H200 多节点部署。即使无法计算整体功耗估算,仍会显示实测 GPU 功率;
+缺失的估算显示为 `—` 并注明原因,CSV 中对应数值留空。图表只绘制有数值的估算。
利润估算器另有目标范围和财务输入要求,且只使用官方性能前沿上的数据点;推理
图表和表格同时支持非官方运行叠加。
diff --git a/packages/app/cypress/e2e/powerx-compare.cy.ts b/packages/app/cypress/e2e/powerx-compare.cy.ts
index 24290d316..648aaa2b6 100644
--- a/packages/app/cypress/e2e/powerx-compare.cy.ts
+++ b/packages/app/cypress/e2e/powerx-compare.cy.ts
@@ -270,25 +270,28 @@ describe('AgentX All in Measured chart and table', () => {
cy.intercept('GET', '/api/v1/trace-availability*', { body: {} });
cy.intercept('GET', '/api/v1/log-availability*', { body: {} });
cy.intercept('GET', '/api/v1/resident-sequence-lengths*', { body: {} });
- cy.intercept('GET', '/api/unofficial-run*', {
- body: {
- runInfos: [
- {
- id: OVERLAY_RUN_ID,
- name: 'agentic-power',
- branch: 'agentic-power',
- sha: 'abc000',
- createdAt: `${DATE}T00:00:00Z`,
- url: OVERLAY_RUN_URL,
- conclusion: 'success',
- status: 'completed',
- isNonMainBranch: true,
- },
- ],
- benchmarks: [{ ...agenticRows[1], id: 0, run_url: OVERLAY_RUN_URL }],
- evaluations: [],
- },
- }).as('agenticOverlay');
+ const overlayBody = {
+ runInfos: [
+ {
+ id: OVERLAY_RUN_ID,
+ name: 'agentic-power',
+ branch: 'agentic-power',
+ sha: 'abc000',
+ createdAt: `${DATE}T00:00:00Z`,
+ url: OVERLAY_RUN_URL,
+ conclusion: 'success',
+ status: 'completed',
+ isNonMainBranch: true,
+ },
+ ],
+ benchmarks: [agenticRows[1], agenticRows[2]].map((row) => ({
+ ...row,
+ id: 0,
+ run_url: OVERLAY_RUN_URL,
+ })),
+ evaluations: [],
+ };
+ cy.intercept('GET', '/api/unofficial-run*', { body: overlayBody }).as('agenticOverlay');
let csvBlob: Blob | undefined;
cy.visit(
`${locale === 'zh' ? '/zh' : ''}/inference?g_model=Kimi-K3&i_seq=agentic-traces&i_prec=fp4&i_pctl=p90&i_metric=y_utilityModeledWatts&i_optimal=0&i_best=0&unofficialrun=${OVERLAY_RUN_ID}`,
@@ -326,14 +329,43 @@ describe('AgentX All in Measured chart and table', () => {
cy.get('[data-testid="chart-figure"]')
.first()
.find('tbody tr')
- .should('have.length', 3)
+ .should('have.length', 5)
.then(($rows) => {
expect($rows.text()).to.contain('B200').and.contain('H200');
- expect($rows.text()).not.to.contain('GB200').and.not.to.contain('B300');
+ expect($rows.text()).to.contain('GB200').and.not.to.contain('B300');
+ const missing = [...$rows].filter((row) => row.textContent!.includes('GB200'));
+ expect(missing).to.have.length(2);
+ for (const row of missing) {
+ expect(row.textContent).to.contain('450').and.contain('—');
+ expect(row.textContent).to.contain(
+ locale === 'en'
+ ? 'Grace or module telemetry missing or invalid'
+ : 'Grace 或 module 遥测缺失或无效',
+ );
+ }
});
cy.get('[data-testid="chart-figure"]')
.first()
.screenshot(`agentic-all-in-${locale}-table`, { overwrite: true });
+ if (locale === 'zh') {
+ cy.get('[data-testid="inference-results-table"]')
+ .first()
+ .contains('th', '整体估算状态')
+ .then(($status) => {
+ const scroll = $status[0].closest('table')!.parentElement!;
+ const pinnedWidth =
+ $status[0].parentElement!.firstElementChild!.getBoundingClientRect().width;
+ cy.wrap(scroll).scrollTo($status[0].offsetLeft - pinnedWidth, 0);
+ cy.wrap($status).should(($cell) => {
+ expect($cell[0].getBoundingClientRect().right).to.be.at.most(
+ scroll.getBoundingClientRect().right + 1,
+ );
+ });
+ });
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .screenshot('agentic-all-in-zh-table-status', { overwrite: true });
+ }
cy.get('header').invoke('css', 'visibility', '');
cy.get('[data-testid="export-button"]').first().click();
cy.get('[data-testid="export-csv-button"]').click();
@@ -343,14 +375,48 @@ describe('AgentX All in Measured chart and table', () => {
const values = data.map((line) => line.split(','));
expect(values.map((row) => row[columns.indexOf('Hardware')]).sort()).to.deep.equal([
'b200',
+ 'gb200',
+ 'gb200',
'h200',
'h200',
]);
expect(
values.map((row) => Number(row[columns.indexOf('Physical Chips')])).sort((a, b) => a - b),
- ).to.deep.equal([16, 32, 32]);
+ ).to.deep.equal([4, 4, 16, 32, 32]);
+ const missing = values.filter((row) => row[columns.indexOf('Hardware')] === 'gb200');
+ for (const row of missing) {
+ expect(row[10]).to.equal('');
+ expect(row[columns.indexOf('Measured GPU Power (W/chip)')]).to.equal('450');
+ expect(row[columns.indexOf('All-in Estimate Status')]).to.equal(
+ 'Grace or module telemetry missing or invalid',
+ );
+ }
});
cy.document().then((doc) => expect(doc.documentElement.scrollWidth).to.be.at.most(width));
+ // A selection containing only GPU measurements must still have a usable table.
+ cy.intercept('GET', '/api/v1/benchmarks*', { body: [agenticRows[2]] }).as('missingOnly');
+ cy.intercept('GET', '/api/unofficial-run*', {
+ body: {
+ ...overlayBody,
+ benchmarks: [{ ...agenticRows[2], id: 0, run_url: OVERLAY_RUN_URL }],
+ },
+ }).as('missingOverlay');
+ cy.reload();
+ cy.wait(['@missingOnly', '@missingOverlay']);
+ cy.get('[data-testid="inference-table-view-btn"]').first().click();
+ cy.get('[data-testid="chart-figure"]')
+ .first()
+ .find('tbody tr')
+ .should('have.length', 2)
+ .and('contain.text', 'GB200');
+ cy.location('href').then((href) => {
+ const url = new URL(href);
+ url.searchParams.set('i_best', '1');
+ url.searchParams.set('i_xmode', 'concurrency');
+ cy.visit(url.toString());
+ });
+ cy.get('[data-testid="inference-table-view-btn"]').first().click();
+ cy.get('[data-testid="chart-figure"]').first().find('tbody tr').should('have.length', 2);
});
}
});
diff --git a/packages/app/src/app/api/v1/views/inference/route.test.ts b/packages/app/src/app/api/v1/views/inference/route.test.ts
index 662b9717a..a7d3a8d29 100644
--- a/packages/app/src/app/api/v1/views/inference/route.test.ts
+++ b/packages/app/src/app/api/v1/views/inference/route.test.ts
@@ -102,6 +102,68 @@ beforeEach(() => {
});
describe('GET /api/v1/views/inference', () => {
+ it('keeps unavailable all-in table rows in official, compared and unofficial scopes and CSV', async () => {
+ const gpuOnly = makeRow({
+ hardware: 'gb200',
+ metrics: {
+ ...makeRow().metrics,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ },
+ });
+ const historical = { ...gpuOnly, id: 998, date: '2026-02-28' };
+ const unofficial = {
+ ...gpuOnly,
+ id: 999,
+ run_url: 'https://github.com/org/repo/actions/runs/999',
+ };
+ mockGetLatestBenchmarks.mockImplementation((_db, _models, date) =>
+ Promise.resolve(date === historical.date ? [historical] : [gpuOnly]),
+ );
+ mockUnofficialRun.mockImplementation(() =>
+ Response.json({ benchmarks: [unofficial], evaluations: [] }),
+ );
+ const query =
+ '/api/v1/views/inference?model=DeepSeek-R1-0528&metric=utilityModeledWatts&best=false&optimal=true&dates=2026-02-28&unofficialrun=999';
+ const response = await GET(request(query));
+ const body = await response.json();
+ expect(response.status).toBe(200);
+ for (const [scope, id] of [
+ [body, gpuOnly.id],
+ [body.comparisons[0], historical.id],
+ [body.overlays[0], unofficial.id],
+ ]) {
+ expect(scope.series).toEqual([]);
+ expect(scope.count).toBe(0);
+ expect(scope.tableRows).toMatchObject([
+ {
+ id,
+ y: null,
+ measuredGpuWatts: 500,
+ status: 'unavailable',
+ unavailableReason: 'cpu-telemetry',
+ },
+ ]);
+ expect(scope).not.toHaveProperty('observedPoints');
+ }
+ expect(body.comparisons[0].tableRows[0].date).toBe(historical.date);
+ expect(body.overlays[0].tableRows[0].runId).toBe(999);
+
+ const csvResponse = await GET(request(`${query}&format=csv`));
+ const csv = await csvResponse.text();
+ const [header, ...lines] = csv.trim().split('\r\n');
+ const columns = header.split(',');
+ expect(lines).toHaveLength(3);
+ for (const line of lines) {
+ const values = line.split(',');
+ expect(values[columns.indexOf('y')]).toBe('');
+ expect(values[columns.indexOf('measuredGpuWatts')]).toBe('500');
+ expect(values[columns.indexOf('unavailableReason')]).toBe('cpu-telemetry');
+ }
+ });
+
it('compares stitched observations before frontier pruning and preserves each producer endpoint', async () => {
const rows = ['h200', 'mi300x'].flatMap((hardware, index) =>
[20, 60].map((x, position) =>
diff --git a/packages/app/src/app/api/v1/views/inference/route.ts b/packages/app/src/app/api/v1/views/inference/route.ts
index 8ae6f0d7d..7fc666470 100644
--- a/packages/app/src/app/api/v1/views/inference/route.ts
+++ b/packages/app/src/app/api/v1/views/inference/route.ts
@@ -197,7 +197,14 @@ function buildView(
return { resolvedPrecisions, result };
}
-function csvRows(data: { result: Pick }) {
+function csvRows(data: { result: Pick }) {
+ if (data.result.tableRows)
+ return data.result.tableRows.map(({ metrics, ...row }) => ({
+ ...row,
+ ...Object.fromEntries(
+ Object.entries(metrics).map(([key, value]) => [`metric_${key}`, value]),
+ ),
+ }));
return data.result.series.flatMap((entry) =>
entry.points.map((point) => ({
hwKey: entry.hwKey,
@@ -583,6 +590,7 @@ export function GET(request: NextRequest) {
frontier: data.result.frontier,
hardware: data.result.hardware,
series: data.result.series,
+ ...(data.result.tableRows === undefined ? {} : { tableRows: data.result.tableRows }),
count: data.result.count,
comparisons,
overlays,
diff --git a/packages/app/src/components/inference/InferenceContext.tsx b/packages/app/src/components/inference/InferenceContext.tsx
index d29bc5759..38759dbea 100644
--- a/packages/app/src/components/inference/InferenceContext.tsx
+++ b/packages/app/src/components/inference/InferenceContext.tsx
@@ -1245,6 +1245,14 @@ export function InferenceProvider({
const wantedType = selectedXAxisMode === 'interactivity' ? 'interactivity' : 'e2e';
const graph = graphs.find((candidate) => candidate.chartDefinition.chartType === wantedType);
if (!graph) return hwTypesWithData;
+ // All in Measured can have GPU-valid table rows without a numeric estimate to rank.
+ const unrankedHwTypes = graph.tableData
+ ? new Set(
+ graph.tableData
+ .filter((point) => effectivePrecisions.includes(point.precision))
+ .map(extractHwKey),
+ )
+ : hwTypesWithData;
const direction =
graph.chartDefinition[
`${selectedYAxisMetric}_roofline` as keyof typeof graph.chartDefinition
@@ -1255,11 +1263,19 @@ export function InferenceProvider({
direction !== 'lower_left' &&
direction !== 'lower_right'
) {
- return hwTypesWithData;
+ return unrankedHwTypes;
}
const best = bestSeriesPerSku(graph.data, direction);
- return best.size > 0 ? best : hwTypesWithData;
- }, [graphs, hwTypesWithData, selectedXAxisMode, selectedYAxisMetric]);
+ if (best.size > 0) return best;
+ return unrankedHwTypes;
+ }, [
+ graphs,
+ hwTypesWithData,
+ selectedXAxisMode,
+ selectedYAxisMetric,
+ effectivePrecisions,
+ extractHwKey,
+ ]);
const setBestPerSkuAndApply = useCallback(
(enabled: boolean, options?: { applySelection?: boolean }) => {
diff --git a/packages/app/src/components/inference/hooks/useChartData.ts b/packages/app/src/components/inference/hooks/useChartData.ts
index 3401a6002..281c50b75 100644
--- a/packages/app/src/components/inference/hooks/useChartData.ts
+++ b/packages/app/src/components/inference/hooks/useChartData.ts
@@ -16,8 +16,10 @@ import { rowToSequence } from '@semianalysisai/inferencex-constants';
import { useQueries, useQuery } from '@tanstack/react-query';
import chartDefinitions, {
+ isAllInMeasuredConfigKey,
tokenMetricTypeForConfigKey,
} from '@/components/inference/metric-registry';
+import { allInMeasuredTableData } from '@/components/inference/utils/inference-table-data';
import {
applyTokenRevenuePricing,
usesTokenSalePricing,
@@ -597,6 +599,15 @@ export function useChartData(
chartDefinition,
data: processedData,
clippedData,
+ ...(isAllInMeasuredConfigKey(selectedYAxisMetric)
+ ? {
+ tableData: expandPowerCompareSeries(
+ allInMeasuredTableData(filteredData, metricKey, xAxisField),
+ selectedYAxisMetric,
+ powerCompare,
+ ),
+ }
+ : {}),
};
},
);
diff --git a/packages/app/src/components/inference/types.ts b/packages/app/src/components/inference/types.ts
index 33a491b16..296c5d140 100644
--- a/packages/app/src/components/inference/types.ts
+++ b/packages/app/src/components/inference/types.ts
@@ -494,6 +494,8 @@ export interface RenderableGraph {
sequence: string;
chartDefinition: ChartDefinition;
data: InferenceData[];
+ /** All GPU-valid rows for the All in Measured table, including unavailable estimates. */
+ tableData?: InferenceData[];
clippedData?: ClippedInferenceData[];
}
/**
@@ -511,6 +513,7 @@ export interface RenderableGraph {
export interface OverlayData {
/** The data points to overlay */
data: InferenceData[];
+ tableData?: InferenceData[];
/** Overlay points hidden by the same display limits as official data. */
clippedData?: ClippedInferenceData[];
/** Hardware configuration for the overlay data (may have different hardware types) */
diff --git a/packages/app/src/components/inference/ui/ChartDisplay.tsx b/packages/app/src/components/inference/ui/ChartDisplay.tsx
index 372decb65..544c188dc 100644
--- a/packages/app/src/components/inference/ui/ChartDisplay.tsx
+++ b/packages/app/src/components/inference/ui/ChartDisplay.tsx
@@ -158,7 +158,7 @@ const STRINGS = {
'GPU Level Provisioned (TDP) · Watts are the rated TDP per GPU from the hardware registry, so the power curve is flat per hardware. Joules per output token = TDP × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together. Hardware without a published TDP is omitted.',
'utility-provisioned':
'All in Provisioned · Watts are the all-in provisioned utility power per GPU from the hardware registry (SemiAnalysis Datacenter Industry Model), so the power curve is flat per hardware. Joules per output token = all-in W × allocated GPUs ÷ whole-deployment output tok/s; disaggregated configurations count prefill and decode GPUs together, unlike the ungated All-in Provisioned J per Output Token, which divides per decode GPU.',
- 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K and AgentX on supported hardware; incomplete inputs are omitted.`,
+ 'utility-modeled': `All in Measured · Validated GPU telemetry with unmeasured components modeled. NVL72 additionally requires complete measured Grace or module power; rack overhead is modeled. Facility watts per GPU = modeled IT watts per GPU × PUE ${AIR_COOLED_SYSTEM_PUE} (air-cooled) or PUE ${DLC_SYSTEM_PUE} (NVL72), applied once. Measured GPU energy per output token scales by facility W/GPU divided by measured GPU W/GPU. Model revision ${modelRevisionLabel}. Available for 8K / 1K and AgentX on supported hardware; unavailable estimates stay in the table with their measured GPU power and reason.`,
},
vsTtft: (word: string) => `vs. ${word} Time To First Token`,
vsE2eLatency: (pctl?: string) =>
@@ -191,7 +191,7 @@ const STRINGS = {
'GPU 额定功耗(TDP)· 功率取硬件注册表中每 GPU 的额定 TDP,因此每种硬件的功率曲线为水平线。每输出 token 能耗 = TDP × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入。未公布 TDP 的硬件不绘制。',
'utility-provisioned':
'整体预配功耗 · 功率取硬件注册表中每 GPU 的全电源配置(all-in)市电功率(来源:SemiAnalysis Datacenter Industry Model),因此每种硬件的功率曲线为水平线。每输出 token 能耗 = all-in 功率 × 分配的 GPU 数 ÷ 整个部署的输出 tok/s;分离式配置将 prefill 与 decode GPU 一并计入,这与未加门控的“每输出 token 全电源配置能耗”按 decode GPU 计算不同。',
- 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。适用于受支持硬件的 8K / 1K 和 AgentX 场景,输入不完整的数据点不绘制。`,
+ 'utility-modeled': `整体实测功耗 · GPU 遥测已验证,未实测组件由模型估算。NVL72 还需完整的 Grace 或 module 实测功耗,机架开销由模型估算。每 GPU 分摊的数据中心功耗 = 每 GPU 分摊的 IT 功耗估算 × PUE ${AIR_COOLED_SYSTEM_PUE}(风冷)或 PUE ${DLC_SYSTEM_PUE}(NVL72);PUE 只应用一次。每输出 token 的实测 GPU 能耗按“每卡数据中心功耗 ÷ 每卡实测 GPU 功耗”的比例换算。模型版本 ${modelRevisionLabel}。适用于受支持硬件的 8K / 1K 和 AgentX 场景。估算不可用的数据点仍保留在表格中,并显示实测 GPU 功耗和不可用原因。`,
},
vsTtft: (word: string) => `vs. ${word === 'Median' ? '中位' : word} 首 token 延迟(TTFT)`,
vsE2eLatency: (pctl?: string) => (pctl ? `vs. ${pctl} 端到端延迟` : 'vs. 端到端延迟'),
@@ -503,7 +503,11 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
let overlayPoints = processed.data;
let clippedOverlayPoints = processed.clippedData;
+ let tableOverlayPoints = processed.tableData;
if (compareGpuPair?.length === 2) {
+ tableOverlayPoints = tableOverlayPoints?.filter((p) =>
+ hardwareKeyMatchesAnyBase(String(p.hwKey), compareGpuPair),
+ );
overlayPoints = overlayPoints.filter((p) =>
hardwareKeyMatchesAnyBase(String(p.hwKey), compareGpuPair),
);
@@ -512,10 +516,15 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
);
}
- if (overlayPoints.length === 0 && clippedOverlayPoints.length === 0) return null;
+ if (
+ overlayPoints.length === 0 &&
+ clippedOverlayPoints.length === 0 &&
+ !tableOverlayPoints?.length
+ )
+ return null;
const keySet = new Set([
- ...overlayPoints.map((p) => String(p.hwKey)),
+ ...(tableOverlayPoints ?? overlayPoints).map((p) => String(p.hwKey)),
...clippedOverlayPoints.map(({ point }) => String(point.hwKey)),
]);
const hardwareConfigFiltered = Object.fromEntries(
@@ -525,6 +534,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
return {
data: overlayPoints,
clippedData: clippedOverlayPoints,
+ tableData: tableOverlayPoints,
hardwareConfig: hardwareConfigFiltered,
label: unofficialRunInfo.branch,
runUrl: unofficialRunInfo.url,
@@ -559,7 +569,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
const eligibleKeys = new Set();
for (const overlay of [overlayDataByChartType.e2e, overlayDataByChartType.interactivity]) {
const points = [
- ...(overlay?.data ?? []),
+ ...(overlay?.tableData ?? overlay?.data ?? []),
...(overlay?.clippedData ?? []).map((entry) => entry.point),
];
for (const point of points) {
@@ -577,7 +587,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
const officialScope = useMemo(() => {
const eligibleKeys = new Set();
for (const graph of graphs) {
- const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)];
+ const points = graph.tableData ?? [
+ ...graph.data,
+ ...(graph.clippedData ?? []).map((entry) => entry.point),
+ ];
for (const point of points) {
if (
selectedPrecisions.includes(point.precision) &&
@@ -736,6 +749,8 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
const effectiveGraphs = useMemo(() => {
if (graphs.length > 0) return graphs;
const hasOverlay =
+ (overlayDataByChartType.e2e?.tableData?.length ?? 0) > 0 ||
+ (overlayDataByChartType.interactivity?.tableData?.length ?? 0) > 0 ||
(overlayDataByChartType.e2e?.data.length ?? 0) > 0 ||
(overlayDataByChartType.e2e?.clippedData?.length ?? 0) > 0 ||
(overlayDataByChartType.interactivity?.data.length ?? 0) > 0 ||
@@ -747,6 +762,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
chartDefinition,
data: [] as InferenceData[],
clippedData: [],
+ tableData: undefined,
}));
}, [graphs, overlayDataByChartType, selectedModel, selectedSequence]);
@@ -761,7 +777,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
if (!isAgenticSequence) return [] as number[];
const ids = new Set();
for (const graph of visibleGraphs) {
- const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)];
+ const points = graph.tableData ?? [
+ ...graph.data,
+ ...(graph.clippedData ?? []).map((entry) => entry.point),
+ ];
for (const point of points) {
if (
selectedPrecisions.includes(point.precision) &&
@@ -791,7 +810,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
if (!useDerivedXAxis) return [] as number[];
const ids = new Set();
for (const graph of visibleGraphs) {
- const points = [...graph.data, ...(graph.clippedData ?? []).map((entry) => entry.point)];
+ const points = graph.tableData ?? [
+ ...graph.data,
+ ...(graph.clippedData ?? []).map((entry) => entry.point),
+ ];
for (const point of points) {
if (point.benchmark_type === 'agentic_traces' && isPersistedBenchmarkId(point.id)) {
ids.add(point.id);
@@ -815,7 +837,12 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
// Legacy AgentX axes can still render transient/non-persisted rows, which
// have no ids to request.
if (!derivedSpec && derivedTargetIds.length === 0) return visibleGraphs;
- return visibleGraphs.map((graph) => ({ ...graph, data: [], clippedData: [] }));
+ return visibleGraphs.map((graph) => ({
+ ...graph,
+ data: [],
+ clippedData: [],
+ tableData: graph.tableData ? [] : undefined,
+ }));
}
return visibleGraphs.map((graph) => {
const rooflineKey = `${selectedYAxisMetric}_roofline` as keyof typeof graph.chartDefinition;
@@ -844,7 +871,10 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
})
.filter((entry): entry is NonNullable => entry !== null);
- if (!derivedSpec) return { ...graph, data, clippedData };
+ const tableData = graph.tableData?.map(
+ (point) => preparePoint(point) ?? { ...point, x: NaN },
+ );
+ if (!derivedSpec) return { ...graph, data, clippedData, tableData };
const chartDefinition = {
...graph.chartDefinition,
@@ -854,7 +884,7 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
y_latency_limit: undefined,
...(derivedCorner ? { [rooflineKey]: derivedCorner } : {}),
};
- return { ...graph, chartDefinition, data, clippedData };
+ return { ...graph, chartDefinition, data, clippedData, tableData };
});
}, [
isAgenticSequence,
@@ -1013,12 +1043,29 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
graph.chartDefinition.chartType,
overlayDataByChartType,
);
+ const tableMode = getViewMode(graphIndex) === 'table';
+ const exportData = tableMode
+ ? (graph.tableData ?? [
+ ...graph.data,
+ ...(graph.clippedData ?? []).map((entry) => entry.point),
+ ])
+ : graph.data;
+ const exportOverlay =
+ tableMode && overlay
+ ? {
+ ...overlay,
+ data: overlay.tableData ?? [
+ ...overlay.data,
+ ...(overlay.clippedData ?? []).map((entry) => entry.point),
+ ],
+ }
+ : overlay;
const {
officialRows: visibleData,
overlayRows: visibleOverlayRowsForExport,
} = isGpuComparison
- ? visibleDateComparisonRows(graph.data, overlay)
- : visibleComparisonRows(graph.data, overlay);
+ ? visibleDateComparisonRows(exportData, exportOverlay)
+ : visibleComparisonRows(exportData, exportOverlay);
const { headers, rows } = inferenceChartToCsv(
visibleData,
graph.model,
@@ -1255,14 +1302,14 @@ export default function ChartDisplay({ embedded = false }: { embedded?: boolean
// must not silently remove measured rows from the table.
// Restore both official and unofficial clipped points before
// applying the shared precision, quick-filter, and legend gates.
- const tableOfficialData = [
+ const tableOfficialData = graph.tableData ?? [
...graph.data,
...(graph.clippedData ?? []).map((entry) => entry.point),
];
const tableOverlay = overlay
? {
...overlay,
- data: [
+ data: overlay.tableData ?? [
...overlay.data,
...(overlay.clippedData ?? []).map((entry) => entry.point),
],
diff --git a/packages/app/src/components/inference/ui/InferenceTable.tsx b/packages/app/src/components/inference/ui/InferenceTable.tsx
index 0d4e8ab74..3eb31a667 100644
--- a/packages/app/src/components/inference/ui/InferenceTable.tsx
+++ b/packages/app/src/components/inference/ui/InferenceTable.tsx
@@ -5,8 +5,15 @@ import { useMemo } from 'react';
import type { ChartDefinition, InferenceData } from '@/components/inference/types';
import { type DataTableColumn, DataTable } from '@/components/ui/data-table';
import { chipCounts } from '@/lib/chip-counts';
-import { getNestedYValue, metricLabel, xAxisLabel } from '@/lib/chart-utils';
-import { isModeledSystemPowerConfigKey } from '@/components/inference/metric-registry';
+import { metricLabel, xAxisLabel } from '@/lib/chart-utils';
+import {
+ isAllInMeasuredConfigKey,
+ isModeledSystemPowerConfigKey,
+} from '@/components/inference/metric-registry';
+import {
+ allInMeasuredStatusLabel,
+ inferenceTableYValue,
+} from '@/components/inference/utils/inference-table-data';
import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare';
import { sortRowsByYMetric } from '@/components/inference/ui/inference-table-sort';
import { type Precision, getPrecisionLabel } from '@/lib/data-mappings';
@@ -22,7 +29,8 @@ interface InferenceTableProps {
}
/** Format a number for table display — picks sensible precision and groups thousands. */
-export function formatInferenceTableNumber(value: number, decimals?: number): string {
+export function formatInferenceTableNumber(value: number | null, decimals?: number): string {
+ if (value === null || !Number.isFinite(value)) return '—';
const fixedDecimals =
decimals ??
(Math.abs(value) >= 100 ? 0 : Math.abs(value) >= 1 ? 1 : Math.abs(value) >= 0.01 ? 3 : 4);
@@ -48,6 +56,8 @@ export function inferenceTableHeaderLabels(
yMetric: metricLabel(chartDefinition, selectedYAxisMetric, locale),
xMetric: xAxisLabel(chartDefinition, locale),
throughput: locale === 'zh' ? '单芯片吞吐量 (tok/s)' : 'Throughput/Chip (tok/s)',
+ measuredGpuPower: locale === 'zh' ? 'GPU 实测功耗 (W/芯片)' : 'Measured GPU Power (W/chip)',
+ estimateStatus: locale === 'zh' ? '整体估算状态' : 'All-in Estimate Status',
};
}
@@ -59,6 +69,7 @@ export default function InferenceTable({
const locale = useLocale();
const yPath = chartDefinition[selectedYAxisMetric as keyof ChartDefinition] as string | undefined;
const showModeledPower = isModeledSystemPowerConfigKey(selectedYAxisMetric);
+ const showAllInMeasured = isAllInMeasuredConfigKey(selectedYAxisMetric);
const headers = useMemo(
() => inferenceTableHeaderLabels(chartDefinition, selectedYAxisMetric, locale),
[chartDefinition, selectedYAxisMetric, locale],
@@ -148,14 +159,33 @@ export default function InferenceTable({
header: headers.yMetric,
align: 'right',
// Comparison clones keep the source metrics; y holds the plotted role/boundary.
- cell: (row) =>
- formatInferenceTableNumber(
- row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath),
- ),
- sortValue: (row) => (row.powerVariant || !yPath ? row.y : getNestedYValue(row, yPath)),
- className: 'tabular-nums',
+ cell: (row) => formatInferenceTableNumber(inferenceTableYValue(row, yPath)),
+ sortValue: (row) => inferenceTableYValue(row, yPath) ?? '',
+ className: showAllInMeasured ? 'tabular-nums min-w-36' : 'tabular-nums',
importance: 'key',
},
+ ...(showAllInMeasured
+ ? [
+ {
+ header: headers.measuredGpuPower,
+ align: 'right' as const,
+ cell: (row: InferenceData) =>
+ formatInferenceTableNumber(row.measuredAvgPower?.y ?? null),
+ sortValue: (row: InferenceData) => row.measuredAvgPower?.y ?? '',
+ className: 'tabular-nums min-w-32',
+ importance: 'key' as const,
+ },
+ {
+ header: headers.estimateStatus,
+ cell: (row: InferenceData) =>
+ allInMeasuredStatusLabel(row, selectedYAxisMetric.slice(2), locale),
+ sortValue: (row: InferenceData) =>
+ allInMeasuredStatusLabel(row, selectedYAxisMetric.slice(2), locale),
+ className: 'w-32 min-w-32 sm:w-auto sm:min-w-48',
+ importance: 'key' as const,
+ },
+ ]
+ : []),
{
header: headers.xMetric,
align: 'right',
@@ -173,7 +203,15 @@ export default function InferenceTable({
importance: 'key',
},
],
- [yPath, headers, showModeledPower, powerCompare, selectedYAxisMetric, locale],
+ [
+ yPath,
+ headers,
+ showModeledPower,
+ showAllInMeasured,
+ powerCompare,
+ selectedYAxisMetric,
+ locale,
+ ],
);
return (
diff --git a/packages/app/src/components/inference/ui/inference-table-sort.ts b/packages/app/src/components/inference/ui/inference-table-sort.ts
index 112ec5062..6a384a768 100644
--- a/packages/app/src/components/inference/ui/inference-table-sort.ts
+++ b/packages/app/src/components/inference/ui/inference-table-sort.ts
@@ -1,5 +1,5 @@
import type { ChartDefinition, InferenceData } from '@/components/inference/types';
-import { getNestedYValue } from '@/lib/chart-utils';
+import { inferenceTableYValue } from '@/components/inference/utils/inference-table-data';
/**
* Default row order for the inference table.
@@ -25,8 +25,10 @@ export function sortRowsByYMetric(
const yAscending = rooflineDir?.startsWith('lower');
return [...data].toSorted((a, b) => {
- const ay = a.powerVariant ? a.y : getNestedYValue(a, yPath);
- const by = b.powerVariant ? b.y : getNestedYValue(b, yPath);
+ const ay = inferenceTableYValue(a, yPath);
+ const by = inferenceTableYValue(b, yPath);
+ if (ay === null) return by === null ? 0 : 1;
+ if (by === null) return -1;
return yAscending ? ay - by : by - ay;
});
}
diff --git a/packages/app/src/components/inference/utils.test.ts b/packages/app/src/components/inference/utils.test.ts
index 2b07210fb..4f9168c56 100644
--- a/packages/app/src/components/inference/utils.test.ts
+++ b/packages/app/src/components/inference/utils.test.ts
@@ -25,6 +25,38 @@ describe('selectUnofficialOverlayForMode', () => {
);
});
});
+
+describe('All in Measured table coverage', () => {
+ it('retains GPU-valid overlay rows with missing all-in power without plotting a fallback', () => {
+ const result = processOverlayChartDataWithClipping(
+ [
+ pt({
+ id: 1,
+ measuredAvgPower: { y: 450, roof: false },
+ utilityModeledWatts: { y: 700, roof: false },
+ }),
+ pt({
+ id: 2,
+ measuredAvgPower: { y: 500, roof: false },
+ modeledSystemPower: {
+ status: 'unsupported',
+ reason: 'cpu-telemetry',
+ modelRevision: 'test',
+ },
+ }),
+ pt({ id: 3 }),
+ ],
+ 'interactivity',
+ 'y_utilityModeledWatts',
+ null,
+ );
+ expect(result.data.map((point) => [point.id, point.y])).toEqual([[1, 700]]);
+ expect(result.tableData?.map((point) => [point.id, point.y])).toEqual([
+ [1, 700],
+ [2, NaN],
+ ]);
+ });
+});
// ---------------------------------------------------------------------------
// fixture factories
// ---------------------------------------------------------------------------
diff --git a/packages/app/src/components/inference/utils.ts b/packages/app/src/components/inference/utils.ts
index f459b5639..4dadf08e3 100644
--- a/packages/app/src/components/inference/utils.ts
+++ b/packages/app/src/components/inference/utils.ts
@@ -5,7 +5,8 @@ import { getGpuSpecs, type TcoBasis } from '@/lib/constants';
* For Pareto front calculations, see @/lib/chart-utils
*/
-import chartDefinitions from '@/components/inference/metric-registry';
+import chartDefinitions, { isAllInMeasuredConfigKey } from '@/components/inference/metric-registry';
+import { allInMeasuredTableData } from '@/components/inference/utils/inference-table-data';
import {
resolveXAxisField,
type FixedSequenceStatistic,
@@ -40,6 +41,7 @@ export function selectUnofficialOverlayForMode(
export interface ProcessedChartData {
data: InferenceData[];
clippedData: ClippedInferenceData[];
+ tableData?: InferenceData[];
}
/**
@@ -260,7 +262,7 @@ export function processOverlayChartDataWithClipping(
}
}
- return partitionChartDataByLimits(
+ const partition = partitionChartDataByLimits(
processedData,
{ ...chartDef, x_scale_field: xAxisField },
selectedYAxisMetric,
@@ -269,4 +271,16 @@ export function processOverlayChartDataWithClipping(
isAgentic,
},
);
+ return {
+ ...partition,
+ ...(isAllInMeasuredConfigKey(selectedYAxisMetric)
+ ? {
+ tableData: expandPowerCompareSeries(
+ allInMeasuredTableData(sourceData, metricKey, xAxisField),
+ selectedYAxisMetric,
+ options?.powerCompare ?? 'none',
+ ),
+ }
+ : {}),
+ };
}
diff --git a/packages/app/src/components/inference/utils/inference-table-data.test.ts b/packages/app/src/components/inference/utils/inference-table-data.test.ts
new file mode 100644
index 000000000..bcf85b42e
--- /dev/null
+++ b/packages/app/src/components/inference/utils/inference-table-data.test.ts
@@ -0,0 +1,107 @@
+import { describe, expect, it } from 'vitest';
+import { createElement } from 'react';
+import { renderToStaticMarkup } from 'react-dom/server';
+import type { InferenceData } from '@/components/inference/types';
+import { chartDefinitions } from '@/components/inference/metric-registry';
+import InferenceTable from '@/components/inference/ui/InferenceTable';
+import { sortRowsByYMetric } from '@/components/inference/ui/inference-table-sort';
+import { inferenceChartToCsv } from '@/lib/csv-export-helpers';
+import {
+ allInMeasuredTableData,
+ allInMeasuredUnavailableReason,
+ inferenceTableYValue,
+} from './inference-table-data';
+
+const point = (overrides: Partial = {}): InferenceData =>
+ ({
+ id: 1,
+ x: 20,
+ y: 999,
+ date: '2026-09-01',
+ tp: 4,
+ conc: 8,
+ hwKey: 'gb200_trt',
+ hw: 'gb200',
+ precision: 'fp4',
+ measuredAvgPower: { y: 450, roof: false },
+ modeledSystemPower: { status: 'unsupported', reason: 'cpu-telemetry', modelRevision: 'test' },
+ ...overrides,
+ }) as InferenceData;
+
+describe('All in Measured table values', () => {
+ it('retains only positive GPU measurements and never substitutes throughput for a missing estimate', () => {
+ const rows = allInMeasuredTableData(
+ [
+ point(),
+ point({ measuredAvgPower: undefined }),
+ point({ measuredAvgPower: { y: 0, roof: false } }),
+ ],
+ 'utilityModeledWatts',
+ 'median_intvty',
+ );
+ expect(rows).toHaveLength(1);
+ expect(rows[0].y).toBeNaN();
+ expect(inferenceTableYValue(rows[0], 'utilityModeledWatts.y')).toBeNull();
+ expect(allInMeasuredUnavailableReason(rows[0], 'utilityModeledWatts')).toBe('cpu-telemetry');
+ expect(
+ allInMeasuredUnavailableReason(
+ { ...rows[0], y: 1400, powerVariant: { kind: 'basis', id: 'gpu-provisioned' } },
+ 'utilityModeledWatts',
+ ),
+ ).toBe('cpu-telemetry');
+ });
+
+ it('sorts unavailable estimates last and renders their measured watts and reason', () => {
+ const rows = allInMeasuredTableData(
+ [point(), point({ id: 2, utilityModeledWatts: { y: 700, roof: false } })],
+ 'utilityModeledWatts',
+ 'median_intvty',
+ );
+ expect(
+ sortRowsByYMetric(rows, chartDefinitions[0], 'y_utilityModeledWatts').map((row) => row.id),
+ ).toEqual([2, 1]);
+ const html = renderToStaticMarkup(
+ createElement(InferenceTable, {
+ data: rows,
+ chartDefinition: chartDefinitions[0],
+ selectedYAxisMetric: 'y_utilityModeledWatts',
+ }),
+ );
+ expect(html).toContain('Grace or module telemetry missing or invalid');
+ expect(html).toContain('450');
+ expect(html).toContain('—');
+ expect(html).not.toContain('NaN');
+ expect(html).not.toContain('999');
+ });
+
+ it('exports unavailable estimates as blanks alongside GPU watts and reason', () => {
+ const rows = allInMeasuredTableData(
+ [point({ x: NaN })],
+ 'utilityModeledWatts',
+ 'median_intvty',
+ );
+ const csv = inferenceChartToCsv(rows, 'Kimi-K3', 'agentic-traces', [], {
+ yHeader: 'All in Measured',
+ yPath: 'utilityModeledWatts.y',
+ xHeader: 'Interactivity',
+ });
+ expect(csv.rows[0][csv.headers.indexOf('All in Measured')]).toBe('');
+ expect(csv.rows[0][csv.headers.indexOf('Measured GPU Power (W/chip)')]).toBe(450);
+ expect(csv.rows[0][csv.headers.indexOf('All-in Estimate Status')]).toBe(
+ 'Grace or module telemetry missing or invalid',
+ );
+ expect(csv.rows[0][csv.headers.indexOf('Interactivity')]).toBe('');
+ const cloneCsv = inferenceChartToCsv(
+ [{ ...rows[0], powerVariant: { kind: 'basis', id: 'utility-modeled' } }],
+ 'Kimi-K3',
+ 'agentic-traces',
+ [],
+ {
+ yHeader: 'All in Measured',
+ yPath: 'utilityModeledWatts.y',
+ xHeader: 'Interactivity',
+ },
+ );
+ expect(cloneCsv.rows[0][cloneCsv.headers.indexOf('All in Measured')]).toBe('');
+ });
+});
diff --git a/packages/app/src/components/inference/utils/inference-table-data.ts b/packages/app/src/components/inference/utils/inference-table-data.ts
new file mode 100644
index 000000000..f1228483d
--- /dev/null
+++ b/packages/app/src/components/inference/utils/inference-table-data.ts
@@ -0,0 +1,66 @@
+import type { AggDataEntry, InferenceData, YAxisMetricKey } from '@/components/inference/types';
+import { remapInferencePoint } from '@/lib/chart-utils';
+import type { Locale } from '@/lib/i18n';
+import type { SystemPowerUnsupportedReason } from '@/lib/modeled-system-power';
+
+/** Table candidates keep GPU measurements even when the all-in model cannot run. */
+export function allInMeasuredTableData(
+ data: readonly InferenceData[],
+ metricKey: YAxisMetricKey,
+ xAxisField: keyof AggDataEntry,
+): InferenceData[] {
+ return data
+ .filter((point) => Number.isFinite(point.measuredAvgPower?.y) && point.measuredAvgPower!.y > 0)
+ .map((point) => ({
+ ...remapInferencePoint(point, metricKey, xAxisField),
+ // InferenceData requires a number. This sentinel is table-only; exports use null/blank.
+ y: point[metricKey]?.y ?? NaN,
+ }));
+}
+
+export function inferenceTableYValue(point: InferenceData, yPath?: string): number | null {
+ const key = yPath?.split('.')[0] as YAxisMetricKey | undefined;
+ const value = point.powerVariant || !key ? point.y : point[key]?.y;
+ return typeof value === 'number' && Number.isFinite(value) ? value : null;
+}
+
+type UnavailableReason = SystemPowerUnsupportedReason | 'energy' | 'model-unavailable';
+
+export function allInMeasuredUnavailableReason(
+ point: InferenceData,
+ metricKey: string,
+): UnavailableReason | null {
+ if (Number.isFinite(point[metricKey as YAxisMetricKey]?.y)) return null;
+ if (point.modeledSystemPower?.status === 'unsupported') return point.modeledSystemPower.reason;
+ if (
+ point.modeledSystemPower?.status === 'supported' &&
+ metricKey === 'utilityModeledJPerOutputToken'
+ )
+ return 'energy';
+ return 'model-unavailable';
+}
+
+const REASONS: Record> = {
+ workload: { en: 'Unsupported workload', zh: '不支持此工作负载' },
+ hardware: { en: 'Unsupported hardware', zh: '不支持此硬件' },
+ telemetry: { en: 'GPU telemetry missing or invalid', zh: 'GPU 遥测缺失或无效' },
+ 'cpu-telemetry': {
+ en: 'Grace or module telemetry missing or invalid',
+ zh: 'Grace 或 module 遥测缺失或无效',
+ },
+ 'gpu-count': { en: 'GPU count missing or inconsistent', zh: 'GPU 数量缺失或不一致' },
+ topology: { en: 'Unsupported topology', zh: '不支持此拓扑' },
+ 'role-power': { en: 'Role power incomplete', zh: 'Prefill 或 decode 功耗不完整' },
+ 'model-domain': { en: 'Outside model range', zh: '超出模型适用范围' },
+ energy: { en: 'Measured energy unavailable', zh: '缺少实测能耗' },
+ 'model-unavailable': { en: 'All-in estimate unavailable', zh: '整体功耗估算不可用' },
+};
+
+export function allInMeasuredStatusLabel(
+ point: InferenceData,
+ metricKey: string,
+ locale: Locale,
+): string {
+ const reason = allInMeasuredUnavailableReason(point, metricKey);
+ return reason === null ? (locale === 'zh' ? '可用' : 'Available') : REASONS[reason][locale];
+}
diff --git a/packages/app/src/lib/api-route-catalog.ts b/packages/app/src/lib/api-route-catalog.ts
index 9b849efd3..826b2f081 100644
--- a/packages/app/src/lib/api-route-catalog.ts
+++ b/packages/app/src/lib/api-route-catalog.ts
@@ -139,7 +139,7 @@ export const apiRouteCatalog = [
method: 'GET',
classification: 'published-read',
operationId: 'get-inference-view',
- sourceSha256: 'bafac08dbe6e6e4b75dd15987e79f84b514f3e93351149ad44b296f6dd2a9662',
+ sourceSha256: '428fab797ad0d304ee7ee193fd2e115219e986e87294d4f561953ab64e0fd5ef',
},
{
source: 'src/app/api/v1/views/options/route.ts',
@@ -843,6 +843,14 @@ export interface ApiContractSourceDigest {
* touching a route module. Digest changes require an explicit documentation review.
*/
export const apiContractSourceDigests = [
+ {
+ source: 'src/components/inference/utils/inference-table-data.ts',
+ sourceSha256: '816eb0dbc416f6c322fd3955af13fc03f8c6988704d82f36a9bfd8ec9abc9669',
+ reviewArea: {
+ en: 'All in Measured table eligibility, nullable values and unavailable reasons shared by UI, CSV and public views.',
+ zh: '界面、CSV 和公开视图共用的整体实测表格行筛选、可空数值与不可用原因。',
+ },
+ },
{
source: 'src/components/inference/utils/resolveXAxisField.ts',
sourceSha256: '4783579c7b3c1a21b91968cb03e9c35a57a85251f4992cb5667a855ba7c77497',
@@ -1033,7 +1041,7 @@ export const apiContractSourceDigests = [
{
source: 'src/lib/views-api/series.ts',
- sourceSha256: '9974fe5166ac847f4b9284eda8a901ee0f9c8b432567ef184890453c715d2f4f',
+ sourceSha256: '9bf521e64e111965b597cb0f55d2f2b4b765b09f8168e8399e9c4316921cc643',
reviewArea: {
en: 'Dashboard read-only selector and calculation parity.',
zh: '仪表板只读接口的选择项与计算一致性。',
diff --git a/packages/app/src/lib/csv-export-helpers.ts b/packages/app/src/lib/csv-export-helpers.ts
index a58ba948f..b70bfa4ec 100644
--- a/packages/app/src/lib/csv-export-helpers.ts
+++ b/packages/app/src/lib/csv-export-helpers.ts
@@ -7,7 +7,11 @@
* plotted x/y axes.
*/
-import { METRIC_REGISTRY } from '@/components/inference/metric-registry';
+import { METRIC_REGISTRY, isAllInMeasuredConfigKey } from '@/components/inference/metric-registry';
+import {
+ allInMeasuredStatusLabel,
+ inferenceTableYValue,
+} from '@/components/inference/utils/inference-table-data';
import type { InferenceData, TrendDataPoint } from '@/components/inference/types';
import { inferPowerCompare, powerSeriesLabel } from '@/components/inference/utils/power-compare';
import { chipCounts } from '@/lib/chip-counts';
@@ -31,9 +35,9 @@ function nestedMetric(point: InferenceData, path: string): number | '' {
const value = point[key as keyof InferenceData];
if (nestedKey && typeof value === 'object' && value !== null && nestedKey in value) {
const nestedValue = (value as Record)[nestedKey];
- return typeof nestedValue === 'number' ? nestedValue : '';
+ return typeof nestedValue === 'number' && Number.isFinite(nestedValue) ? nestedValue : '';
}
- return typeof value === 'number' ? value : '';
+ return typeof value === 'number' && Number.isFinite(value) ? value : '';
}
/** Preserve a real zero while leaving source metrics that were not measured blank. */
@@ -64,6 +68,7 @@ export function inferenceChartToCsv(
const powerCompare = inferPowerCompare(allPoints);
const showPowerSeries = powerCompare !== 'none';
const plottedMetric = displayedMetrics ? `y_${displayedMetrics.yPath.split('.')[0]}` : '';
+ const showAllInMeasured = isAllInMeasuredConfigKey(plottedMetric);
const headers = [
'Model',
'ISL',
@@ -119,6 +124,7 @@ export function inferenceChartToCsv(
'DP',
...(showModeledPower ? ['Configured Chip Count'] : []),
...(showPowerSeries ? ['Power Series'] : []),
+ ...(showAllInMeasured ? ['Measured GPU Power (W/chip)', 'All-in Estimate Status'] : []),
];
const displayedColumns = displayedMetrics
@@ -126,9 +132,14 @@ export function inferenceChartToCsv(
{
header: displayedMetrics.yHeader,
value: (point: InferenceData) =>
- point.powerVariant ? point.y : nestedMetric(point, displayedMetrics.yPath),
+ point.powerVariant
+ ? (inferenceTableYValue(point) ?? '')
+ : nestedMetric(point, displayedMetrics.yPath),
+ },
+ {
+ header: displayedMetrics.xHeader,
+ value: (point: InferenceData) => (Number.isFinite(point.x) ? point.x : ''),
},
- { header: displayedMetrics.xHeader, value: (point: InferenceData) => point.x },
].filter(
(column, index, columns) =>
!headers.includes(column.header) &&
@@ -187,6 +198,9 @@ export function inferenceChartToCsv(
d.dp ?? '',
...(showModeledPower ? [chips.configured] : []),
...(showPowerSeries ? [powerSeriesLabel(d, plottedMetric, powerCompare, 'en')] : []),
+ ...(showAllInMeasured
+ ? [d.measuredAvgPower?.y ?? '', allInMeasuredStatusLabel(d, plottedMetric.slice(2), 'en')]
+ : []),
];
row.splice(10, 0, ...displayedColumns.map((column) => column.value(d)));
return row;
diff --git a/packages/app/src/lib/views-api/docs/inference.ts b/packages/app/src/lib/views-api/docs/inference.ts
index 8c9811e07..bdbed2ac9 100644
--- a/packages/app/src/lib/views-api/docs/inference.ts
+++ b/packages/app/src/lib/views-api/docs/inference.ts
@@ -293,8 +293,8 @@ const parameters: readonly ApiParameter[] = [
required: false,
type: 'enum',
description: text(
- 'Response encoding. csv returns one flat row per plotted point; serviceCompare, roleShare and powerFit analytical results require JSON and return 400 with CSV.',
- '响应编码。csv 为每个图表点返回一行平面数据;serviceCompare、roleShare 和 powerFit 分析结果仅支持 JSON,与 CSV 同用时返回 400。',
+ 'Response encoding. csv returns plotted points, or tableRows for All in Measured with unavailable y cells blank. serviceCompare, roleShare and powerFit analytical results require JSON and return 400 with CSV.',
+ '响应编码。csv 返回图表数据点;整体实测指标则导出 tableRows,不可用的 y 留空。serviceCompare、roleShare 和 powerFit 分析结果仅支持 JSON,与 CSV 同用时返回 400。',
),
schema: { type: 'string', enum: ['json', 'csv'], default: 'json' },
example: 'csv',
@@ -345,6 +345,68 @@ const seriesSchema = objectSchema(
],
);
+const tableRowSchema = objectSchema(
+ {
+ ...pointSchema.properties,
+ id: { type: ['integer', 'null'] },
+ hwKey: stringSchema,
+ gpu: stringSchema,
+ framework: stringSchema,
+ specMethod: stringSchema,
+ label: stringSchema,
+ vendor: stringSchema,
+ deployment: stringSchema,
+ kvOffload: booleanSchema,
+ x: { type: ['number', 'null'], description: 'Selected service axis; null when unavailable.' },
+ y: { type: ['number', 'null'], description: 'Selected all-in metric; null when unavailable.' },
+ measuredGpuWatts: { ...numberSchema, description: 'Validated measured mean W/GPU.' },
+ status: { type: 'string', enum: ['available', 'unavailable'] },
+ unavailableReason: {
+ oneOf: [
+ {
+ type: 'string',
+ enum: [
+ 'workload',
+ 'hardware',
+ 'telemetry',
+ 'cpu-telemetry',
+ 'gpu-count',
+ 'topology',
+ 'role-power',
+ 'model-domain',
+ 'energy',
+ 'model-unavailable',
+ ],
+ },
+ { type: 'null' },
+ ],
+ },
+ },
+ [
+ 'id',
+ 'precision',
+ 'hwKey',
+ 'gpu',
+ 'framework',
+ 'specMethod',
+ 'label',
+ 'deployment',
+ 'kvOffload',
+ 'x',
+ 'y',
+ 'concurrency',
+ 'topologyKey',
+ 'tp',
+ 'date',
+ 'frontier',
+ 'bestPerSku',
+ 'metrics',
+ 'measuredGpuWatts',
+ 'status',
+ 'unavailableReason',
+ ],
+);
+
const sourceSchema = objectSchema({ key: stringSchema, label: stringSchema }, ['key', 'label']);
const identitySchema = objectSchema({
id: { type: ['integer', 'null'] },
@@ -459,6 +521,11 @@ const responseSchema = objectSchema(
]),
),
series: arraySchema(seriesSchema),
+ tableRows: {
+ ...arraySchema(tableRowSchema),
+ description:
+ 'All in Measured only: GPU-valid observations in the selected scope, including unavailable estimates. Honors hardware/best selection, but not optimal or chart axis clipping. Also returned per comparison and overlay. count remains the numeric series point count.',
+ },
count: integerSchema,
serviceSources: arraySchema(sourceSchema),
equalServiceComparison: { ...comparisonSchema, type: ['object', 'null'] },
@@ -546,8 +613,16 @@ const responseSchema = objectSchema(
}),
),
pricing: { type: ['object', 'null'], additionalProperties: true },
- comparisons: arraySchema({ type: 'object', additionalProperties: true }),
- overlays: arraySchema({ type: 'object', additionalProperties: true }),
+ comparisons: arraySchema({
+ type: 'object',
+ properties: { tableRows: arraySchema(tableRowSchema) },
+ additionalProperties: true,
+ }),
+ overlays: arraySchema({
+ type: 'object',
+ properties: { tableRows: arraySchema(tableRowSchema) },
+ additionalProperties: true,
+ }),
},
['view', 'apiVersion', 'params', 'metric', 'xAxis', 'frontier', 'series', 'count'],
);
@@ -685,8 +760,8 @@ export const operations: ApiOperation[] = [
path: '/api/v1/views/inference',
summary: text('Get the main inference chart view', '获取主推理图表视图'),
description: text(
- 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection. All in Measured watts and energy support 8K/1K and AgentX with validated telemetry and supported topology; AgentX reuses the model without independent workload calibration. The standalone Modeled Chassis AC metric remains limited to 8K/1K.',
- '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。整体实测功耗及能耗支持遥测已验证、拓扑受支持的 8K/1K 和 AgentX 数据;AgentX 复用同一模型,尚未针对该工作负载单独校准。单独列出的每 GPU 分摊的机箱交流功耗估算指标仍仅支持 8K/1K。',
+ 'Returns the chart-ready series the /inference scatter chart renders: per hardware config, x/y points at each measured concurrency for the selected metric, sequence, precisions and x-axis mode, with boundary and best-per-SKU flags computed by the same code the dashboard runs. Filters mirror the dashboard quick filters (gpus, vendors, framework families, deployment, spec). Use optimal=true for boundary points or best=true for the best series per GPU SKU. Measured-power boundaries follow the higher-power outer envelope: frontier.direction describes that boundary, while metric.direction remains the optimization direction used by best-per-SKU selection. All in Measured watts and energy support 8K/1K and AgentX with validated telemetry and supported topology; AgentX reuses the model without independent workload calibration. The standalone Modeled Chassis AC metric remains limited to 8K/1K. All in Measured adds tableRows with every GPU-valid observation in the selected scope and best-series selection, before optimal or axis clipping; unavailable estimates have y=null and a reason. series and count remain numeric chart points.',
+ '返回 /inference 散点图所用的序列:按硬件配置分组,在所选指标、序列、精度与 x 轴模式下给出各并发档位的 x/y 数据点,并复用仪表板代码计算边界与 best-per-SKU 标记。筛选参数与仪表板快捷筛选一致(gpus、vendors、框架系列、部署模式、投机解码)。设置 optimal=true 可只保留边界点,best=true 可只保留每个 GPU SKU 的最优曲线。实测功耗使用较高功耗侧的外包络:frontier.direction 描述这一边界,metric.direction 则保留 best-per-SKU 选择所用的优化方向。整体实测功耗及能耗支持遥测已验证、拓扑受支持的 8K/1K 和 AgentX 数据;AgentX 复用同一模型,尚未针对该工作负载单独校准。单独列出的每 GPU 分摊的机箱交流功耗估算指标仍仅支持 8K/1K。整体实测指标另返回 tableRows,保留当前筛选范围和最优曲线选择内所有 GPU 遥测有效的观测点,不按 optimal 或坐标轴显示范围裁剪;估算不可用时 y 为 null,并给出原因。series 和 count 仍仅包含可绘制的数值点。',
),
audience: 'public',
stability: 'beta',
diff --git a/packages/app/src/lib/views-api/series.test.ts b/packages/app/src/lib/views-api/series.test.ts
index 78a4c5c7a..e25c68d6b 100644
--- a/packages/app/src/lib/views-api/series.test.ts
+++ b/packages/app/src/lib/views-api/series.test.ts
@@ -116,6 +116,127 @@ function powerSweepRows(): BenchmarkRow[] {
}
describe('buildInferenceSeries', () => {
+ it('retains GPU-valid all-in table rows without adding unavailable estimates to chart series', () => {
+ const measured = metrics({
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ pp: 1,
+ pcp_size: 1,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ joules_per_output_token: 2,
+ });
+ const supported = makeRow({ metrics: measured });
+ const noCpu = makeRow({ hardware: 'gb200', metrics: measured });
+ const invalid = makeRow({ hardware: 'b200', metrics: { ...measured, power_valid: 0 } });
+ const otherPrecision = makeRow({ hardware: 'gb300', precision: 'fp4', metrics: measured });
+ const result = buildInferenceSeries([supported, noCpu, invalid, otherPrecision], {
+ ...BASE_OPTIONS,
+ metricConfigKey: 'y_utilityModeledWatts',
+ optimal: true,
+ });
+
+ expect(result.count).toBe(1);
+ expect(result.series.flatMap((series) => series.points.map((point) => point.id))).toEqual([
+ supported.id,
+ ]);
+ expect(result.tableRows).toHaveLength(2);
+ expect(result.tableRows?.find((point) => point.id === supported.id)).toMatchObject({
+ y: result.series[0].points[0].y,
+ measuredGpuWatts: 500,
+ status: 'available',
+ unavailableReason: null,
+ });
+ expect(result.tableRows?.find((point) => point.id === noCpu.id)).toMatchObject({
+ hwKey: 'gb200_trt',
+ y: null,
+ measuredGpuWatts: 500,
+ status: 'unavailable',
+ unavailableReason: 'cpu-telemetry',
+ frontier: false,
+ });
+ });
+
+ it('keeps all-in energy missing, honors selected hardware, and leaves ordinary views unchanged', () => {
+ const row = makeRow({
+ metrics: metrics({
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ pp: 1,
+ pcp_size: 1,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ }),
+ });
+ const result = buildInferenceSeries([row], {
+ ...BASE_OPTIONS,
+ metricConfigKey: 'y_utilityModeledJPerOutputToken',
+ });
+ expect(result.series).toEqual([]);
+ expect(result.tableRows).toMatchObject([
+ { id: row.id, y: null, status: 'unavailable', unavailableReason: 'energy' },
+ ]);
+ expect(
+ buildInferenceSeries([row], {
+ ...BASE_OPTIONS,
+ metricConfigKey: 'y_utilityModeledWatts',
+ gpus: ['gb200'],
+ }).tableRows,
+ ).toEqual([]);
+ expect(buildInferenceSeries([row], BASE_OPTIONS)).not.toHaveProperty('tableRows');
+ });
+
+ it('keeps non-frontier transient rows in the table without assigning them frontier flags', () => {
+ const rows = [
+ [1, 10, 400],
+ [2, 20, 500],
+ ].map(([conc, x, watts]) => {
+ const row = makeRow({
+ conc,
+ metrics: metrics({
+ median_intvty: x,
+ avg_power_w: watts,
+ avg_total_gpu_power_w: watts * 8,
+ power_valid: 1,
+ power_metric_schema_version: 2,
+ pp: 1,
+ pcp_size: 1,
+ }),
+ });
+ Reflect.deleteProperty(row, 'id');
+ return row;
+ });
+ const result = buildInferenceSeries(rows, {
+ ...BASE_OPTIONS,
+ metricConfigKey: 'y_utilityModeledWatts',
+ optimal: true,
+ });
+ expect(result.count).toBe(1);
+ expect(result.tableRows).toHaveLength(2);
+ expect(result.tableRows?.map((point) => [point.concurrency, point.frontier])).toEqual([
+ [1, false],
+ [2, true],
+ ]);
+ });
+
+ it('returns null for missing or invalid normalized axes without dropping table rows', () => {
+ const row = makeRow({
+ benchmark_type: 'agentic_traces',
+ metrics: metrics({ avg_power_w: 500, power_valid: 1, power_metric_schema_version: 2 }),
+ });
+ const result = buildInferenceSeries([row], {
+ ...BASE_OPTIONS,
+ sequence: Sequence.AgenticTraces,
+ metricConfigKey: 'y_utilityModeledWatts',
+ xmode: 'e2e-normalized-interactivity',
+ derivedMetrics: {
+ [row.id]: { id: row.id, p75_e2e_norm_intvty: null, p90_e2e_norm_intvty: 0 },
+ },
+ });
+ expect(result.series).toEqual([]);
+ expect(result.tableRows).toMatchObject([{ id: row.id, x: null, y: null }]);
+ });
+
it('assembles one series per hardware config with x-sorted points', () => {
const result = buildInferenceSeries(fixtureRows(), BASE_OPTIONS);
diff --git a/packages/app/src/lib/views-api/series.ts b/packages/app/src/lib/views-api/series.ts
index 54cc033e4..40fb3fa18 100644
--- a/packages/app/src/lib/views-api/series.ts
+++ b/packages/app/src/lib/views-api/series.ts
@@ -9,6 +9,7 @@ import {
type RooflineDirection,
} from '@/components/inference/hooks/chart-data-core';
import chartDefinitions, {
+ isAllInMeasuredConfigKey,
METRIC_REGISTRY,
tokenMetricTypeForConfigKey,
type MetricConfigKey,
@@ -25,6 +26,10 @@ import type {
} from '@/components/inference/types';
import { partitionChartDataByLimits } from '@/components/inference/utils';
import { bestSeriesPerSku } from '@/components/inference/utils/best-series-per-sku';
+import {
+ allInMeasuredTableData,
+ allInMeasuredUnavailableReason,
+} from '@/components/inference/utils/inference-table-data';
import {
isMeasuredPowerCurveMetric,
upperPowerEnvelope,
@@ -127,6 +132,22 @@ export interface InferenceSeriesEntry {
readonly points: readonly InferenceSeriesPoint[];
}
+export interface InferenceTableRow extends Omit {
+ readonly hwKey: string;
+ readonly gpu: string;
+ readonly framework: string;
+ readonly specMethod: string;
+ readonly label: string;
+ readonly vendor?: string;
+ readonly deployment: string;
+ readonly kvOffload: boolean;
+ readonly x: number | null;
+ readonly y: number | null;
+ readonly measuredGpuWatts: number;
+ readonly status: 'available' | 'unavailable';
+ readonly unavailableReason: ReturnType;
+}
+
export interface InferenceSeriesMetricMeta {
readonly key: MetricKey;
readonly configKey: MetricConfigKey;
@@ -139,6 +160,8 @@ export interface InferenceSeriesMetricMeta {
export interface InferenceSeriesResult {
readonly series: readonly InferenceSeriesEntry[];
+ /** All in Measured table population, including unavailable system estimates. */
+ readonly tableRows?: readonly InferenceTableRow[];
readonly hardware: readonly { key: string; label: string; vendor?: string }[];
readonly frontier: { direction: ParetoDirection | null; points: number };
readonly metric: InferenceSeriesMetricMeta;
@@ -408,9 +431,58 @@ export function buildInferenceSeries(
}
const count = series.reduce((total, entry) => total + entry.points.length, 0);
+ // Remapping retains nested metric objects, including for transient rows without IDs.
+ const frontierMetrics = new Set([...frontierPoints].map((point) => point[metricKey]));
+ const tableRows = isAllInMeasuredConfigKey(metricConfigKey)
+ ? allInMeasuredTableData(scoped, metricKey, resolved.xAxisField)
+ .filter((point) => !best || bestHwKeys.size === 0 || bestHwKeys.has(point.hwKey))
+ .map((point): InferenceTableRow => {
+ const derived = point.id === undefined ? undefined : options.derivedMetrics?.[point.id];
+ const x =
+ xmode === 'e2e-normalized-interactivity'
+ ? percentile === 'p75'
+ ? derived?.p75_e2e_norm_intvty
+ : derived?.p90_e2e_norm_intvty
+ : point.x;
+ const unavailableReason = allInMeasuredUnavailableReason(point, metricKey);
+ return {
+ id: point.id ?? null,
+ precision: point.precision,
+ hwKey: point.hwKey,
+ gpu: point.hwKey.split('_')[0],
+ framework: point.framework ?? '',
+ specMethod: point.spec_decoding ?? 'none',
+ label: hardwareLegendLabel(point.hwKey, point.model),
+ vendor: GPU_VENDORS[point.hwKey.split('_')[0]],
+ deployment: pointDeploymentMode(point),
+ kvOffload: isKvOffloadEnabled(point),
+ x:
+ typeof x === 'number' &&
+ Number.isFinite(x) &&
+ (xmode !== 'e2e-normalized-interactivity' || x > 0)
+ ? x
+ : null,
+ y: Number.isFinite(point.y) ? point.y : null,
+ concurrency: point.conc ?? 0,
+ topologyKey: pointTopologyKey(point),
+ tp: point.tp ?? 0,
+ date: point.date ?? '',
+ ...(runIdFromUrl(point.run_url) === undefined
+ ? {}
+ : { runId: runIdFromUrl(point.run_url) }),
+ frontier: unavailableReason === null && frontierMetrics.has(point[metricKey]),
+ bestPerSku: bestHwKeys.has(point.hwKey),
+ metrics: pointMetrics(point, metricKey),
+ measuredGpuWatts: point.measuredAvgPower!.y,
+ status: unavailableReason === null ? 'available' : 'unavailable',
+ unavailableReason,
+ };
+ })
+ : undefined;
return {
series,
+ ...(tableRows === undefined ? {} : { tableRows }),
hardware: series.map((entry) => ({
key: entry.hwKey,
label: entry.label,
diff --git a/packages/skills/skills/inferencex-api/integrity.json b/packages/skills/skills/inferencex-api/integrity.json
index 6e4dc88cd..a0640e448 100644
--- a/packages/skills/skills/inferencex-api/integrity.json
+++ b/packages/skills/skills/inferencex-api/integrity.json
@@ -8,7 +8,7 @@
"references/cli-contract.md": "fa44faa38d889b4fbdee5ba42752b47758ee3706fab150513bd4e6db87a86cb0",
"references/cli.md": "96b228f34cb3600f4548f4dc84df8506531747286763c56ca834c77bee05e1eb",
"references/collectivex.md": "eb794f9c28d4a27bec4db80c42c4685ff3b1d204fd9b12a6788511fb258a789f",
- "references/dashboard-views.md": "afa2422706e8aa0d8833ead3c4c3fccc0bee11c8ab45f73b2e54622f69d1100c",
+ "references/dashboard-views.md": "61d995291e9aa7280a5ced5af1a672ad5e0b001336db7c56a7655aa1a024ec64",
"references/offline-exports.md": "95aa565dcd4c9e592159baa560a9a58217bc61e395de1c6daf0576795c7f4ec9",
"references/pareto.md": "1b4d2d163f982e3f2d97addce310789eae9c1e5badd82dd7501708c9e0385027",
"references/powerx.md": "cf8adcc5dfe395c5fb659908819980a872475f862f22b723815af43d363cedad",
diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md
index 00e492558..b2765ff2c 100644
--- a/packages/skills/skills/inferencex-api/references/dashboard-views.md
+++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md
@@ -102,6 +102,13 @@ the chassis or rack model without independent workload calibration. Telemetry an
topology gates still apply; NVL72 needs complete Grace or module power. The standalone
Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope.
+For All in Measured, `tableRows` retains every GPU-valid observation in the selected
+scope and best-series selection, including axis-clipped and non-frontier points.
+Missing system estimates use `y: null`, `status: "unavailable"`, and
+`unavailableReason`; `measuredGpuWatts` remains available. CSV exports these rows
+with blank missing values. Numeric `series` and `count` are unchanged. Each date
+comparison and unofficial overlay has its own `tableRows`; latest does not pool history.
+
NVL72 estimates require valid GPU power plus validated Grace-socket or compute-module
power with complete socket coverage. CPU-rail-only readings do not establish the
Grace/LPDDR boundary. A module reading already includes GPU power; do not add GPU
@@ -169,8 +176,7 @@ n, x-range, registry `tdpWatts` and point identities. Fewer than three distinct
rates return `fit: null`, `reason: "too-few-points"`. Call `P₀` an extrapolated
intercept, not idle power, and do not read the line outside its x-range.
-These analytical results require JSON; enabling any with CSV returns 400. Existing CSV remains
-a plotted-point export.
+These analytical results require JSON; enabling any with CSV returns 400. CSV exports plotted points except for All in Measured, which exports `tableRows`.
同等服务对比与同并发诊断仅通过 API 提供,查询参数为 `serviceCompare`、`serviceBaseline`、
`serviceComparator` 和 `serviceTarget`;仪表板没有对应的服务对比控件、来源选择、目标值输入、
From cd0739426ec8ee54fa2d9322ac7917bfd1569be0 Mon Sep 17 00:00:00 2001
From: Wenyao Gao
Date: Thu, 1 Oct 2026 14:08:46 -0700
Subject: [PATCH 20/22] docs: clarify unavailable multi-node system estimates
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:明确多节点系统估算不可用的含义。
---
docs/powerx-system-power.md | 20 ++++++++++----------
docs/powerx-system-power.zh.md | 20 ++++++++++----------
2 files changed, 20 insertions(+), 20 deletions(-)
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 456fb8d84..9ad1d24e6 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -164,16 +164,16 @@ recomputed separately.
Compare the same model, workload, date/run, engine, precision and metric first.
Chart/table rows and target-based profit estimates answer different questions.
-| Symptom / reason | Check and next action |
-| -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
-| GPU curve exists, All in Measured is absent | Check supported workload/hardware and the row's system-model status. Valid GPU power alone does not establish a system estimate. |
-| B200/H200 multi-node row is absent (`topology`, `role-power`, `gpu-count`) | Inspect physical GPU count, host placement and total/role watts. Use the original producer topology; do not infer chassis placement from a display label or sum TP/EP aliases. |
-| NVL72 reports `cpu-telemetry` / `no-cpu-power` | Inspect the same-window CPU audit, sensor kind and complete socket coverage. GPU validity remains independent. |
-| `telemetry` / `no-measured-power` | Check the original validation audit and raw samples. Reprocess only when the retained evidence supports the original window; otherwise collect replacement performance and power together. |
-| `outside-measured-range` or power-invalid target bracket | Choose a target supported by the selected serving curve. Both original bounding points need valid power; another valid point elsewhere on the curve cannot fill the gap. |
-| `incompatible-power-basis` | Do not interpolate between module and GPU-plus-Grace readings, or different model/PUE bases. |
-| No cost, token mix or provisioned power | Inspect the financial inputs. This can prevent both profit estimates even when power is valid. |
-| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. |
+| Symptom / reason | Check and next action |
+| ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
+| GPU curve exists, All in Measured is absent | Check supported workload/hardware and the row's system-model status. Valid GPU power alone does not establish a system estimate. |
+| B200/H200 multi-node system estimate is unavailable (`topology`, `role-power`, `gpu-count`) | Inspect physical GPU count, host placement and total/role watts. Use the original producer topology; do not infer chassis placement from a display label or sum TP/EP aliases. |
+| NVL72 reports `cpu-telemetry` / `no-cpu-power` | Inspect the same-window CPU audit, sensor kind and complete socket coverage. GPU validity remains independent. |
+| `telemetry` / `no-measured-power` | Check the original validation audit and raw samples. Reprocess only when the retained evidence supports the original window; otherwise collect replacement performance and power together. |
+| `outside-measured-range` or power-invalid target bracket | Choose a target supported by the selected serving curve. Both original bounding points need valid power; another valid point elsewhere on the curve cannot fill the gap. |
+| `incompatible-power-basis` | Do not interpolate between module and GPU-plus-Grace readings, or different model/PUE bases. |
+| No cost, token mix or provisioned power | Inspect the financial inputs. This can prevent both profit estimates even when power is valid. |
+| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. |
In **Compare both**, a valid provisioned result remains when its measured estimate
is unavailable. Configurations that cannot be priced at all are listed separately
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index 7a370d502..2c739e4cd 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -141,16 +141,16 @@ NVL72 记录。利润规划始终要求 schema 2。
先确认比较的是同一模型、工作负载、日期/运行、引擎、精度和指标。图表/表格记录与
按性能目标计算的利润估算,回答的问题不同。
-| 现象 / 原因 | 检查项与下一步 |
-| ----------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
-| 有 GPU 曲线,但没有 All in Measured | 检查工作负载、硬件是否受支持,以及该记录的系统模型状态。GPU 功耗有效不等于系统估算可用。 |
-| B200/H200 多节点记录缺失(`topology`、`role-power`、`gpu-count`) | 检查物理 GPU 数量、主机分布和总功率/角色功率。使用原生产端拓扑,不从展示名称推断机箱分布,也不累加 TP/EP 别名。 |
-| NVL72 返回 `cpu-telemetry` / `no-cpu-power` | 检查同一窗口的 CPU 审计、传感器类型和完整 socket 覆盖;GPU 有效性独立判断。 |
-| `telemetry` / `no-measured-power` | 检查原验证审计及原始样本。只有保留证据足以支持原窗口时才重处理,否则须重新采集匹配的性能和功耗。 |
-| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前所选曲线支持的目标。两个原始端点的功耗都须有效,曲线上其他位置的有效点不能补齐这个缺口。 |
-| `incompatible-power-basis` | 不在 module 读数与 GPU 加 Grace 读数之间插值,也不在不同模型版本或 PUE 取值之间插值。 |
-| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 |
-| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 |
+| 现象 / 原因 | 检查项与下一步 |
+| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- |
+| 有 GPU 曲线,但没有 All in Measured | 检查工作负载、硬件是否受支持,以及该记录的系统模型状态。GPU 功耗有效不等于系统估算可用。 |
+| B200/H200 多节点系统估算不可用(`topology`、`role-power`、`gpu-count`) | 检查物理 GPU 数量、主机分布和总功率/角色功率。使用原生产端拓扑,不从展示名称推断机箱分布,也不累加 TP/EP 别名。 |
+| NVL72 返回 `cpu-telemetry` / `no-cpu-power` | 检查同一窗口的 CPU 审计、传感器类型和完整 socket 覆盖;GPU 有效性独立判断。 |
+| `telemetry` / `no-measured-power` | 检查原验证审计及原始样本。只有保留证据足以支持原窗口时才重处理,否则须重新采集匹配的性能和功耗。 |
+| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前所选曲线支持的目标。两个原始端点的功耗都须有效,曲线上其他位置的有效点不能补齐这个缺口。 |
+| `incompatible-power-basis` | 不在 module 读数与 GPU 加 Grace 读数之间插值,也不在不同模型版本或 PUE 取值之间插值。 |
+| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 |
+| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 |
**Compare both** 在实测估算不可用时仍保留有效预配结果。完全无法定价的配置,与
仅缺少实测估算的配置分开列出;仅实测模式不会用预配功率代替。保留遥测和定向修复
From 9a627452e1bb7bcd62329c7f021a60649b32d8ba Mon Sep 17 00:00:00 2001
From: Wenyao Gao <105094497+edwingao28@users.noreply.github.com>
Date: Thu, 1 Oct 2026 14:17:31 -0700
Subject: [PATCH 21/22] fix: remove unavailable estimate lists (#1249)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:移除无法估算配置列表。
---
docs/dashboard-readonly-views.md | 7 +-
docs/powerx-system-power.md | 11 ++-
docs/powerx-system-power.zh.md | 9 +-
docs/tco-calculator.md | 6 +-
.../app/cypress/e2e/profit-estimator.cy.ts | 83 +++----------------
.../calculator/ProfitEstimatorDisplay.tsx | 81 ------------------
6 files changed, 27 insertions(+), 170 deletions(-)
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index cababa570..6ae649aa3 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -66,7 +66,7 @@ state. GPU interactive downsampling does not alter returned raw data or statisti
Power boundary labels are GPU Level Measured, GPU Level Provisioned (TDP), All in
Provisioned, and All in Measured. The last combines measured GPU power with modeled
unmeasured components and PUE; it is not a wall-meter measurement. These labels and
-collapsed power-assumption/availability notes do not change metric IDs, API selectors,
+power-assumption notes do not change metric IDs, API selectors,
or calculations. Profit comparison `powerLabel` display text follows the same names.
All in Measured watts and energy accept validated 8K/1K and AgentX rows through the
shared chart/API transform, including historical and unofficial rows. AgentX reuses
@@ -85,9 +85,8 @@ Dense profit charts reserve readable space per bar and scroll within the plot on
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
plot regardless of its current scroll position, so no API or skills contract change is needed.
-Unavailable-estimate notices distinguish unpriced SKUs from missing measured-plus-modeled
-estimates using the retained provisioned result identity. This explanatory grouping preserves
-the API's existing rows, skip reasons, selectors and calculations.
+Profit charts omit the per-configuration unavailable-estimate list. This is presentation-only;
+the API retains skipped rows and reasons, and calculations and CSV exports are unchanged.
The GPU statistics table includes startup and warmup for all chips in the selected
series, regardless of chip visibility. It is separate from serving-window power,
J/token and selected-time-window calculations. Run telemetry is DB-first with an
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index 9ad1d24e6..c64bcf3f7 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -45,8 +45,8 @@ experimental power controls with ↑↑↓↓ if they are hidden.
4. For a capacity comparison, open `/profit-estimator-per-gigawatt`, choose the
model and a supported interactivity target, then select **Compare both** in
Benchmark Config. Match each result's workload, engine and precision to the
- inference selection. Expand **Unavailable estimates** for missing results;
- hover or select a bar for its power basis and read the formula notes below.
+ inference selection. Hover or select a bar for its power basis and read the
+ formula notes below.
**Expected result:** the All in Measured table keeps every GPU-valid record in
the selected scope, including B200/H200 multi-node deployments. It shows measured
@@ -176,10 +176,9 @@ Chart/table rows and target-based profit estimates answer different questions.
| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. |
In **Compare both**, a valid provisioned result remains when its measured estimate
-is unavailable. Configurations that cannot be priced at all are listed separately
-from missing measured estimates. Measured-only mode never substitutes provisioned
-watts. See [persistence and recovery](./powerx-persistence-recovery.md) for retained
-telemetry and targeted repair; new power readings cannot be attached to old
+is unavailable. Measured-only mode never substitutes provisioned watts. See
+[persistence and recovery](./powerx-persistence-recovery.md) for retained telemetry
+and targeted repair; new power readings cannot be attached to old
throughput results.
## Provenance and reproducible exports
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index 2c739e4cd..ec76d8c66 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -36,8 +36,8 @@ AgentX 估算属于容量规划预览,不代表已完成 AgentX 校准,也
AgentX 预览;独立的 **Modeled Chassis AC** 指标仍仅支持 8K/1K 单轮结果。
4. 要比较容量,打开 `/profit-estimator-per-gigawatt`,选择模型和曲线支持的
交互性目标,再在 Benchmark Config 中选择 **Compare both**。逐项核对结果的
- 工作负载、引擎和精度是否与推理页面所选一致。展开 **Unavailable estimates**
- 查看缺失原因;悬停或选中柱形查看功耗依据,并阅读图表下方的公式说明。
+ 工作负载、引擎和精度是否与推理页面所选一致。悬停或选中柱形查看功耗依据,
+ 并阅读图表下方的公式说明。
**预期结果:**All in Measured 表格保留当前筛选范围内所有 GPU 功耗有效的记录,
包括 B200/H200 多节点部署。即使无法计算整体功耗估算,仍会显示实测 GPU 功率;
@@ -152,9 +152,8 @@ NVL72 记录。利润规划始终要求 schema 2。
| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 |
| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 |
-**Compare both** 在实测估算不可用时仍保留有效预配结果。完全无法定价的配置,与
-仅缺少实测估算的配置分开列出;仅实测模式不会用预配功率代替。保留遥测和定向修复
-方式见[持久化与恢复](./powerx-persistence-recovery.md);新采集的功率不能附到旧吞吐量上。
+**Compare both** 在实测估算不可用时仍保留有效预配结果。仅实测模式不会用预配
+功率代替。保留遥测和定向修复方式见[持久化与恢复](./powerx-persistence-recovery.md);新采集的功率不能附到旧吞吐量上。
## 来源与可复现导出
diff --git a/docs/tco-calculator.md b/docs/tco-calculator.md
index 756ccf4d8..4b03e3393 100644
--- a/docs/tco-calculator.md
+++ b/docs/tco-calculator.md
@@ -1135,7 +1135,7 @@ and the two can be collapsed into one once both are on master.
input / $0.06 cached / $1.20 output per M tok, the permanent 50%-off rate on its
pay-as-you-go page); the OpenRouter aggregate also sits below it. At 83
tok/s/user the B200, B300, GB200, and MI355X agentic curves are priced, and the
- H100, H200, MI300X, and MI325X curves top out below it and list as not priced.
+ H100, H200, MI300X, and MI325X curves top out below it and are omitted at that target.
GLM 5.2/5.3 opens on a 10% model license fee and MiniMax M3 on 20%; Kimi K3
opens on the 30% `DEFAULT_LAB_CUT_PCT`. DeepSeek V4 Pro opens on 24 tok/s/user,
the speed DeepSeek's own API serves at, DeepSeek's peak-hour list price for
@@ -1144,14 +1144,14 @@ and the two can be collapsed into one once both are on master.
and a 0% model license fee, since the weights ship under the MIT license. At 24
tok/s/user the B200, B300, and MI355X agentic curves are priced; the
GB200, GB300, and H200 curves bottom out above it (their lowest measured
- points sit at roughly 40, 30, and 27 tok/s/user) and list as not priced until
+ points sit at roughly 40, 30, and 27 tok/s/user) and are omitted until
a lower-interactivity run lands. DeepSeek V4.1 Flash opens on 125 tok/s/user,
the speed DeepSeek's own API serves the Flash tier at, DeepSeek's peak-hour
list price for `deepseek-flash` ($0.30 input / $0.006 cached / $1.20 output
per M tok; off-peak is half that), and a 0% model license fee, since the
weights ship under the MIT license. It entered the fleet on AgentX only (InferenceX#2961), so the page is
wired ahead of the first published rows; SKUs whose agentic curves stop short
- of 125 tok/s/user list as not priced rather than extrapolated. A
+ of 125 tok/s/user are omitted rather than extrapolated. A
model with a list price gets a third Token Price option, ` list
price`, next to OpenRouter and Custom; the caption names the source in force and
links the lab's pricing page when the list price is used. Switching to Custom
diff --git a/packages/app/cypress/e2e/profit-estimator.cy.ts b/packages/app/cypress/e2e/profit-estimator.cy.ts
index 4e7aef39a..dfb6648ed 100644
--- a/packages/app/cypress/e2e/profit-estimator.cy.ts
+++ b/packages/app/cypress/e2e/profit-estimator.cy.ts
@@ -95,22 +95,10 @@ const chart = () => cy.get('[data-testid="profit-estimator-chart"]');
const chartSvg = () => chart().find('svg').filter(':has(.chart-root)').first();
const bars = () => chart().find('rect.bar');
-function assertDisclosureOpen(testId: string, open: boolean) {
- cy.get(`[data-testid="${testId}"]`).should(($details) => {
- expect($details[0].open, `${testId} native disclosure state`).to.equal(open);
- const content = $details[0].querySelector('p');
- expect(content, `${testId} content`).not.to.equal(null);
- // Cypress visibility omits native closed-details rendering in some browsers.
- if (content && typeof content.checkVisibility === 'function') {
- expect(content.checkVisibility(), `${testId} browser visibility`).to.equal(open);
- }
- });
-}
-
// Clear the preceding chart before each case changes the viewport.
describe('Profit estimator power option', { testIsolation: true }, () => {
for (const locale of ['en', 'zh'] as const) {
- it(`distinguishes unpriced SKUs from missing measured-power estimates (${locale})`, () => {
+ it(`keeps available estimates without a per-configuration warning list (${locale})`, () => {
stubOpenRouter();
const width = locale === 'en' ? 1280 : 390;
cy.viewport(width, 900);
@@ -135,34 +123,12 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
.should('not.contain', 'H200')
.and('contain', 'B300')
.and('contain', 'GB300');
- const unpriced = locale === 'en' ? 'Not priced:' : '未定价:';
- const measured =
- locale === 'en' ? 'Measured + modeled unavailable:' : '实测加建模估算不可用:';
- cy.get('[data-testid="profit-power-unavailable"] > summary').click();
- cy.get('[data-testid="profit-power-unavailable"] > p')
- .should('be.visible')
- .should(($notice) => {
- const [baseline, measurement] = $notice.text().split(measured);
- expect(baseline).to.contain(unpriced).and.to.contain('H200');
- expect(baseline).not.to.contain('B300');
- expect(measurement).to.contain('B300').and.to.contain('GB300');
- expect(measurement).not.to.contain('H200');
- expect(measurement).to.contain(
- locale === 'en' ? 'no usable measured power' : '同一组基准测试数据点缺少有效功耗',
- );
- expect(measurement).to.contain(
- locale === 'en'
- ? 'missing complete Grace or module power'
- : '缺少完整的 Grace 或 module 功耗',
- );
- const bounds = $notice[0].getBoundingClientRect();
- expect(bounds.left).to.be.at.least(0);
- expect(bounds.right).to.be.at.most(width);
- });
- cy.get('[data-testid="profit-power-unavailable"]').scrollIntoView({
- offset: { top: -70, left: 0 },
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
+ chart().scrollIntoView();
+ cy.screenshot(`profit-no-unavailable-list-${locale}`, {
+ capture: 'viewport',
+ overwrite: true,
});
- cy.screenshot(`profit-unavailable-basis-${locale}`, { capture: 'viewport', overwrite: true });
});
it(`prices DeepSeek Flash partial chassis with a one-line power note and CSV labels (${locale})`, () => {
@@ -202,10 +168,6 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
},
);
const label = locale === 'en' ? 'Full-chassis extrapolation' : '整机外推';
- const cpuReason =
- locale === 'en'
- ? 'missing complete Grace or module power'
- : '缺少完整的 Grace 或 module 功耗';
cy.get('#profit-target').should('have.value', '125');
chart().find('text.revenue-label').should('have.length', 3);
chart()
@@ -213,19 +175,11 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
.and('contain', 'B300')
.and('contain', 'MI355X')
.and('contain', label);
- cy.get('[data-testid="profit-power-unavailable"]')
- .should('contain', 'GB300')
- .and('contain', cpuReason);
- assertDisclosureOpen('profit-power-unavailable', false);
- cy.get('[data-testid="profit-power-unavailable"] > summary').click();
- assertDisclosureOpen('profit-power-unavailable', true);
- cy.get('[data-testid="profit-power-unavailable"] > p').should('be.visible');
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
cy.get('[data-testid="profit-power-note"]')
.should('contain', locale === 'en' ? 'All in Measured' : '整体实测功耗')
.and('not.contain', locale === 'en' ? 'unmeasured components' : '未实测的组件');
cy.get('[data-testid="profit-power-assumptions"]').should('not.exist');
- cy.get('[data-testid="profit-power-unavailable"] > summary').click();
- assertDisclosureOpen('profit-power-unavailable', false);
cy.get('[data-testid="profit-power-note"]').then(($note) => {
const box = $note[0].getBoundingClientRect();
expect(box.left).to.be.at.least(0);
@@ -293,7 +247,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
});
for (const currentValid of [true, false]) {
- it(`dates historical power skips without current hardware metadata (${currentValid ? 'with current bars' : 'provisioned bars only'})`, () => {
+ it(`keeps historical provisioned bars when measured power is unavailable (${currentValid ? 'with current bars' : 'provisioned bars only'})`, () => {
stubOpenRouter();
cy.intercept('GET', '/api/v1/benchmarks*', (req) => {
const historical = req.query['date'] === PROFIT_HISTORY_DATE;
@@ -318,14 +272,6 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
`/profit-estimator-per-gigawatt?c_power=compare&i_gpus=b200_sglang,b300_vllm&i_dstart=${PROFIT_HISTORY_DATE}&i_dend=${PROFIT_HISTORY_DATE}`,
{ onBeforeLoad: unlockPowerGate },
);
- cy.get('[data-testid="profit-power-unavailable"]').should(
- 'contain',
- `B300 (vLLM) (FP4) • ${PROFIT_HISTORY_DATE}`,
- );
- cy.get('[data-testid="profit-power-unavailable"]').should(
- 'contain',
- `B200 (SGLang) (FP4) • ${PROFIT_HISTORY_DATE}`,
- );
chart()
.find('text.revenue-label')
.should('have.length', currentValid ? 4 : 3);
@@ -361,7 +307,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
chart().find('text.revenue-label').should('have.length', 7);
chart().should('contain', 'B200').and('contain', 'B300').and('contain', 'MI355X');
chart().should('contain', 'All in Measured').and('contain', 'All in Provisioned');
- cy.get('[data-testid="profit-power-unavailable"]').should('contain', 'GB300');
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
});
it('prices a GB200 NVL72 tray on its measured compute module and names the basis', () => {
@@ -421,10 +367,7 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
.and('contain', 'DLC PUE 1.1');
cy.get('body').type('{esc}');
cy.get('[data-testid="option-help-content-profit-power"]').should('not.exist');
- // Only GB300's measured estimate is unavailable; its provisioned estimate remains visible.
- cy.get('[data-testid="profit-power-unavailable"]')
- .should('contain', 'GB300')
- .and('not.contain', 'GB200');
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
chart().scrollIntoView();
cy.screenshot('profit-nvl72-compare-desktop', { capture: 'viewport', overwrite: true });
cy.viewport(393, 900);
@@ -611,10 +554,8 @@ describe('Profit estimator power option', { testIsolation: true }, () => {
cy.get('#profit-power').click();
cy.get('[role="option"]').contains('All in Measured').click();
// These existing fixtures intentionally have throughput but no validated power.
- cy.get('[data-testid="profit-power-unavailable"]').should(
- 'contain',
- 'no usable measured power',
- );
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
+ cy.contains('No SKU can be priced for the current selection.').should('be.visible');
cy.get('#profit-target').should('have.value', '45');
cy.get('[data-testid="profit-model-selector"]').should('contain', 'Kimi K3');
cy.get('[data-testid="profit-price-source-selector"]').should('contain', 'Moonshot');
diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
index ea8122428..c59d3ae10 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
@@ -91,7 +91,6 @@ import {
profitModelDefaults,
type ProfitBasis,
type ProfitEstimatorRow,
- type ProfitEstimatorSkipReason,
} from './profit-estimator';
import { powerBasisLabel, profitEstimatorChartStrings, rowLabel } from './ProfitEstimatorChart';
import { estimateProfitByPower, powerSourceKey, type ProfitPowerBasis } from './profit-power';
@@ -220,7 +219,6 @@ const STRINGS = {
powerPreview: `${ALL_IN_MEASURED_NOTE.en} AgentX system power is not yet qualified.`,
powerDetails:
'GPU power is interpolated between the same throughput points. Includes PUE 1.3 for air-cooled chassis or 1.1 for NVL72, and 10% headroom. Aggregate multinode hosts use the measured deployment mean. Full-chassis extrapolation fills an eight-GPU server with replicas of the measured 1/2/4-GPU workload at the same per-GPU power and throughput; it does not measure a partly idle server.',
- unavailableEstimates: (count: number) => `Unavailable estimates (${count})`,
powerNvl72Note: (hardware: string, basis: string, pue: number) =>
`${hardware}: ${basis}. Modeled: NVSwitch trays, NICs/DPUs, NVMe, power shelves, DLC PUE ${pue}.`,
csvPowerHeaders: ['Power basis', 'Power sensor', 'System power profile'],
@@ -300,21 +298,6 @@ const STRINGS = {
'Revenue ($/GPU/hr, 100% util)',
],
},
- skipped: (entries: string) => `Not priced: ${entries}.`,
- modeledUnavailable: (entries: string) => `Measured + modeled unavailable: ${entries}.`,
- skipReason: {
- 'outside-measured-range': 'no measured point at the target interactivity',
- 'no-power': 'no all-in power figure',
- 'no-measured-power': 'no usable measured power for these benchmark points',
- 'no-cpu-power': 'missing complete Grace or module power for these benchmark points',
- 'incompatible-power-basis': 'bounding points use different power measurement bases',
- 'unsupported-power-hardware': 'no system power model for this hardware',
- 'unsupported-power-topology':
- 'GPU counts, physical hosts or role power do not support this system model',
- 'outside-power-model': 'these benchmark points are outside the supported power model',
- 'no-cost': 'no TCO for this tier',
- 'no-token-mix': 'no input/output token mix recorded',
- } satisfies Record,
compareHistory: 'Compare history',
gpuConfig: 'Chip Config',
gpuConfigTooltip: `Select up to ${PROFIT_HISTORY_MAX_GPUS} chip configurations to compare how their estimated revenue and profit have moved over time. Each config is priced again on every compared date (the ends of the date range, plus any date or run added from the Config Changelog below) using the run measured then, so software updates show up as a change in the bar.`,
@@ -350,7 +333,6 @@ const STRINGS = {
powerPreview: `${ALL_IN_MEASURED_NOTE.zh} AgentX 系统功耗模型尚未完成验证。`,
powerDetails:
'GPU 功耗在相同的吞吐量数据点间插值,风冷机箱 PUE 为 1.3,NVL72 为 1.1,另加 10% 功耗余量。聚合多节点按部署平均功耗估算各台服务器。整机外推假设八卡服务器部署多个相同的实测单卡、双卡或四卡实例,每卡功耗和吞吐量保持不变;它不代表部分 GPU 闲置时的整机实测功耗。',
- unavailableEstimates: (count: number) => `无法估算(${count} 项)`,
powerNvl72Note: (hardware: string, basis: string, pue: number) =>
`${hardware}:${basis}。建模部分:NVSwitch tray、网卡/DPU、NVMe、电源架,液冷 PUE ${pue}。`,
csvPowerHeaders: ['功耗口径', '功耗传感器', '系统功耗 profile'],
@@ -430,20 +412,6 @@ const STRINGS = {
'收入($/GPU/hr,100% 利用率)',
],
},
- skipped: (entries: string) => `未定价:${entries}。`,
- modeledUnavailable: (entries: string) => `实测加建模估算不可用:${entries}。`,
- skipReason: {
- 'outside-measured-range': '未在该交互性下实测',
- 'no-power': '缺少全电源配置功率数据',
- 'no-measured-power': '同一组基准测试数据点缺少有效功耗',
- 'no-cpu-power': '同一组基准测试数据点缺少完整的 Grace 或 module 功耗',
- 'incompatible-power-basis': '插值两端的功耗测量口径不同',
- 'unsupported-power-hardware': '该硬件暂无适用的系统功耗模型',
- 'unsupported-power-topology': 'GPU 数量、物理主机或各角色功耗不满足系统模型要求',
- 'outside-power-model': '这些基准测试数据点超出功耗模型的适用范围',
- 'no-cost': '该层级无 TCO 数据',
- 'no-token-mix': '未记录输入/输出 token 比例',
- } satisfies Record,
compareHistory: '对比历史趋势',
gpuConfig: '芯片配置',
gpuConfigTooltip: `最多选择 ${PROFIT_HISTORY_MAX_GPUS} 个芯片配置,对比其收入与利润估算随时间的变化。每个配置都会用当日实测的运行结果,在每个对比日期(日期范围的起止两端,以及从下方配置变更日志中添加的日期或运行)重新估价,软件更新带来的差异会直接体现在柱形上。`,
@@ -1371,29 +1339,6 @@ function ProfitEstimatorInner({
historyCurrentRunIds,
]);
- const powerUnavailable = useMemo(() => {
- const unpriced: string[] = [];
- const measuredUnavailable: string[] = [];
- for (const row of fullEstimate.skipped) {
- const label = rowLabel(
- { ...row, dateLabel: row.date ? historyEntryLabel(row.date) : undefined },
- hardwareConfig,
- );
- const entries =
- powerBasis === 'compare' &&
- fullEstimate.rows.some((priced) => priced.resultKey === `${row.resultKey}__provisioned`)
- ? measuredUnavailable
- : unpriced;
- entries.push(`${label}: ${t.skipReason[row.reason]}`);
- }
- return [
- unpriced.length > 0 ? t.skipped(unpriced.join('; ')) : '',
- measuredUnavailable.length > 0 ? t.modeledUnavailable(measuredUnavailable.join('; ')) : '',
- ]
- .filter(Boolean)
- .join(' ');
- }, [fullEstimate, hardwareConfig, historyEntryLabel, powerBasis, t]);
-
const powerBasisNotes = useMemo(() => {
const notes = new Map();
for (const row of estimate.rows) {
@@ -1434,20 +1379,6 @@ function ProfitEstimatorInner({
{t.powerLabel}: {t.powerOptions[powerBasis]}
)}
- {basis === 'gw-year' && powerBasis !== 'provisioned' && fullEstimate.skipped.length > 0 && (
-
- track('profit_estimator_power_unavailable_toggled')}
- >
- {t.unavailableEstimates(fullEstimate.skipped.length)}
-
- {powerUnavailable}
-
- )}
- {basis === 'gw-year' &&
- estimate.rows.length === 0 &&
- powerBasis !== 'provisioned' && (
-
- {t.powerPreview} {powerUnavailable}
-
- )}
Date: Thu, 1 Oct 2026 15:07:29 -0700
Subject: [PATCH 22/22] fix: select power-valid curves for profit estimates
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
中文:利润估算先选择功耗有效的曲线。
---
docs/dashboard-readonly-views.md | 13 ++
docs/powerx-system-power.md | 21 +-
docs/powerx-system-power.zh.md | 16 +-
.../app/cypress/e2e/profit-power-curves.cy.ts | 126 +++++++++++
.../src/app/api/v1/views/extensions.test.ts | 70 ++++++
.../calculator/ProfitEstimatorDisplay.tsx | 16 +-
.../calculator/profit-history.test.ts | 46 ++++
.../components/calculator/profit-history.ts | 6 +-
.../calculator/profit-power.test.ts | 209 ++++++++++++++++++
.../src/components/calculator/profit-power.ts | 35 ++-
.../calculator/useThroughputData.ts | 7 +-
packages/app/src/lib/api-route-catalog.ts | 10 +-
.../lib/views-api/calculator-extensions.ts | 9 +-
.../app/src/lib/views-api/docs/extensions.ts | 4 +-
packages/app/timings.json | 4 +
.../skills/inferencex-api/integrity.json | 2 +-
.../references/dashboard-views.md | 12 +-
17 files changed, 570 insertions(+), 36 deletions(-)
create mode 100644 packages/app/cypress/e2e/profit-power-curves.cy.ts
diff --git a/docs/dashboard-readonly-views.md b/docs/dashboard-readonly-views.md
index 6ae649aa3..dd99cf859 100644
--- a/docs/dashboard-readonly-views.md
+++ b/docs/dashboard-readonly-views.md
@@ -81,6 +81,13 @@ Missing system estimates use `y: null`, `status: "unavailable"`, and
with blank missing values. Numeric `series` and `count` are unchanged. Each date
comparison and unofficial overlay has its own `tableRows`; latest does not pool history.
+For GW-year profit, `modeled` selects valid system-power points before building the
+curve at the same target, without extrapolation or another snapshot. `compare` uses
+the same valid-curve throughput for both power budgets, retaining the original
+provisioned estimate when no valid curve covers the target. Provisioned-only keeps
+the original performance curve. Official, comparison and unofficial scopes remain
+independent; CPU/module telemetry and compatible sensor-basis requirements still apply.
+
Dense profit charts reserve readable space per bar and scroll within the plot on narrow
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
@@ -256,6 +263,12 @@ those properties.
`unavailableReason` 给出原因;`measuredGpuWatts` 保留实测 GPU 功耗。CSV 导出同一组行,缺失值留空。
`series` 和 `count` 保持不变,仍只包含可绘制的数值点;各日期对比和非官方叠加分别返回自己的 `tableRows`,
Latest 不会合并历史数据。
+
+按 GW 年估算利润时,modeled 先筛选满足系统功耗要求的数据点,再在原目标值上构建曲线,
+不外推,也不借用其他快照。compare 的两种功耗方案使用同一条有效曲线的吞吐量;
+没有有效曲线覆盖目标时,保留原曲线的预配估算。provisioned 单独使用时沿用原性能曲线。
+官方数据、日期对比和非官方叠加各自独立计算;CPU/模块遥测要求和传感器口径兼容性要求同样不变。
+
私有上传、密钥、提示词、反馈及管理操作不作为公开读取接口。
OperatorX 的入口受功能开关控制,页面使用专属的 `/api/v1/operatorx/*`
接口;目前没有发布 `/api/v1/views/operatorx` 契约。
diff --git a/docs/powerx-system-power.md b/docs/powerx-system-power.md
index c64bcf3f7..1b5be58cd 100644
--- a/docs/powerx-system-power.md
+++ b/docs/powerx-system-power.md
@@ -52,9 +52,9 @@ experimental power controls with ↑↑↓↓ if they are hidden.
the selected scope, including B200/H200 multi-node deployments. It shows measured
GPU power even when an all-in estimate is unavailable; the estimate displays
`—` with a reason and stays blank in CSV. The graph plots numeric estimates only.
-The Profit Estimator adds target-range and financial requirements and uses only
-official frontier points. Inference charts and tables also support unofficial-run
-overlays.
+The Profit Estimator adds target-range and financial requirements. All in
+Measured builds its performance frontier from power-valid measurements in the
+selected scope. Inference charts and tables also support unofficial-run overlays.
## Hardware and telemetry requirements
@@ -148,11 +148,14 @@ total, not the fixture's rounded per-GPU display value.
above the 264 kW installed shelf capacity is outside the model domain.
- **Planning reserve:** facility kW/GPU × 1.10 is a separate capacity buffer.
Average power plus this reserve is not a validated electrical peak limit.
-- **Matched comparison:** profit modes keep the same original performance
- frontier, target, token prices, utilization, license share and per-GPU-hour
- costs. Between frontier points, planning power is interpolated only between
- those original points, with compatible model revision, PUE, topology and sensor
- basis. No extrapolation or replacement by another power-valid point occurs.
+- **Measured curves:** All in Measured selects power-valid points before building
+ the performance frontier. Throughput and power use that curve at the requested
+ target; the curve uses a compatible model, PUE, topology and sensor basis.
+ No target extrapolation or borrowing from unselected history occurs.
+- **Matched comparison:** Compare both uses the same power-valid curve, target
+ and financial inputs for its paired bars. If no measured estimate is available,
+ the ordinary provisioned result remains. All in Provisioned keeps the ordinary
+ performance frontier.
Lower planning power increases GPU capacity per GW. Revenue, compute cost and
license fees scale with that capacity under the fixed per-GPU assumptions; profit
@@ -170,7 +173,7 @@ Chart/table rows and target-based profit estimates answer different questions.
| B200/H200 multi-node system estimate is unavailable (`topology`, `role-power`, `gpu-count`) | Inspect physical GPU count, host placement and total/role watts. Use the original producer topology; do not infer chassis placement from a display label or sum TP/EP aliases. |
| NVL72 reports `cpu-telemetry` / `no-cpu-power` | Inspect the same-window CPU audit, sensor kind and complete socket coverage. GPU validity remains independent. |
| `telemetry` / `no-measured-power` | Check the original validation audit and raw samples. Reprocess only when the retained evidence supports the original window; otherwise collect replacement performance and power together. |
-| `outside-measured-range` or power-invalid target bracket | Choose a target supported by the selected serving curve. Both original bounding points need valid power; another valid point elsewhere on the curve cannot fill the gap. |
+| `outside-measured-range` or power-invalid target bracket | Choose a target within the selected power-valid curve. Both bounding points need compatible valid power; points outside the selected scope cannot fill the gap. |
| `incompatible-power-basis` | Do not interpolate between module and GPU-plus-Grace readings, or different model/PUE bases. |
| No cost, token mix or provisioned power | Inspect the financial inputs. This can prevent both profit estimates even when power is valid. |
| `workload`, `hardware`, `model-domain` | Use a supported workload/profile and in-domain input; do not replace the missing estimate with zero or TDP. |
diff --git a/docs/powerx-system-power.zh.md b/docs/powerx-system-power.zh.md
index ec76d8c66..84e7cb700 100644
--- a/docs/powerx-system-power.zh.md
+++ b/docs/powerx-system-power.zh.md
@@ -42,8 +42,8 @@ AgentX 估算属于容量规划预览,不代表已完成 AgentX 校准,也
**预期结果:**All in Measured 表格保留当前筛选范围内所有 GPU 功耗有效的记录,
包括 B200/H200 多节点部署。即使无法计算整体功耗估算,仍会显示实测 GPU 功率;
缺失的估算显示为 `—` 并注明原因,CSV 中对应数值留空。图表只绘制有数值的估算。
-利润估算器另有目标范围和财务输入要求,且只使用官方性能前沿上的数据点;推理
-图表和表格同时支持非官方运行叠加。
+利润估算器另有目标范围和财务输入要求。All in Measured 在当前筛选范围内,
+先选出功耗有效的实测点,再构建性能前沿。推理图表和表格也支持非官方运行叠加。
## 硬件与遥测要求
@@ -128,10 +128,12 @@ NVL72 记录。利润规划始终要求 schema 2。
264 kW 时,超出模型定义域。
- **规划余量:**设施 kW/GPU × 1.10 是独立的容量缓冲。平均功率加此余量不等于
经过验证的供电峰值上限。
-- **匹配比较:**各利润模式保持原性能前沿、目标、token 价格、利用率、授权分成
- 和每 GPU 小时成本一致。目标位于两个前沿数据点之间时,只在这两个原始点之间
- 插值规划功率,且模型版本、PUE、拓扑和传感器边界须兼容;不做范围外推,也不
- 换用其他功耗有效的数据点。
+- **实测曲线:**All in Measured 先选出功耗有效的数据点,再构建性能前沿。
+ 吞吐量和功耗都按这条曲线在所选目标处计算;整条曲线的模型版本、PUE、拓扑
+ 和传感器口径须兼容。不做目标范围外推,也不借用未选中历史曲线上的数据点。
+- **匹配比较:**Compare both 的实测与预配柱形使用同一条功耗有效曲线、同一目标
+ 和相同财务输入。没有可用实测估算时,仍保留常规预配结果。All in Provisioned
+ 继续使用常规性能前沿。
较低的规划功率会提高每 GW 可容纳的 GPU 数量。在固定单卡假设下,收入、计算成本
和授权费用都随容量变化;利润率和每芯片小时的经济指标不会改善,也不会另行重算电费。
@@ -147,7 +149,7 @@ NVL72 记录。利润规划始终要求 schema 2。
| B200/H200 多节点系统估算不可用(`topology`、`role-power`、`gpu-count`) | 检查物理 GPU 数量、主机分布和总功率/角色功率。使用原生产端拓扑,不从展示名称推断机箱分布,也不累加 TP/EP 别名。 |
| NVL72 返回 `cpu-telemetry` / `no-cpu-power` | 检查同一窗口的 CPU 审计、传感器类型和完整 socket 覆盖;GPU 有效性独立判断。 |
| `telemetry` / `no-measured-power` | 检查原验证审计及原始样本。只有保留证据足以支持原窗口时才重处理,否则须重新采集匹配的性能和功耗。 |
-| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前所选曲线支持的目标。两个原始端点的功耗都须有效,曲线上其他位置的有效点不能补齐这个缺口。 |
+| `outside-measured-range`,或目标区间端点功耗无效 | 选择当前功耗有效曲线范围内的目标。插值两端的功耗须有效且口径兼容,不能用筛选范围外的数据点补齐缺口。 |
| `incompatible-power-basis` | 不在 module 读数与 GPU 加 Grace 读数之间插值,也不在不同模型版本或 PUE 取值之间插值。 |
| 缺少成本、token 组成或预配功率 | 检查财务输入;即使功耗有效,也可能无法计算两种利润结果。 |
| `workload`、`hardware`、`model-domain` | 使用受支持的工作负载/profile 和定义域内的输入,不以零值或 TDP 替代缺失估算。 |
diff --git a/packages/app/cypress/e2e/profit-power-curves.cy.ts b/packages/app/cypress/e2e/profit-power-curves.cy.ts
new file mode 100644
index 000000000..d06d53559
--- /dev/null
+++ b/packages/app/cypress/e2e/profit-power-curves.cy.ts
@@ -0,0 +1,126 @@
+import {
+ interceptProfitData,
+ profitBenchmarkRows,
+ PROFIT_DATE,
+ PROFIT_HISTORY_DATE,
+} from '../support/profit-fixtures';
+
+// Invalid MI355X knots dominate the valid curve in the performance frontier.
+// Power selection must recover the remaining curve before interpolation.
+function rowsFor(date = PROFIT_DATE) {
+ const rows = profitBenchmarkRows('kimik3', date).map((row) => ({
+ ...row,
+ metrics: {
+ ...row.metrics,
+ power_valid: Number(
+ row.hardware !== 'b300' && !(row.hardware === 'mi355x' && row.conc === 16),
+ ),
+ power_metric_schema_version: 2,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ },
+ }));
+ return [
+ ...rows,
+ ...rows
+ .filter((row) => row.hardware === 'mi355x')
+ .map((row) => ({
+ ...row,
+ id: row.id + 100_000,
+ framework: 'atom',
+ metrics: { ...row.metrics, power_valid: 1 },
+ })),
+ ];
+}
+
+function setup() {
+ interceptProfitData();
+ cy.intercept('GET', 'https://openrouter.ai/api/v1/models', { data: [] });
+ cy.intercept('GET', '/api/v1/benchmarks*', (req) => {
+ req.reply({
+ body: rowsFor(req.query['date'] === PROFIT_HISTORY_DATE ? PROFIT_HISTORY_DATE : PROFIT_DATE),
+ });
+ });
+}
+
+function unlock(win: Cypress.AUTWindow) {
+ win.localStorage.setItem('inferencex-feature-gate', '1');
+ win.localStorage.setItem('inferencex-star-modal-dismissed', String(Date.now()));
+ win.sessionStorage.setItem('inferencex-reproducibility-nudge-shown', '1');
+}
+
+const chart = () => cy.get('[data-testid="profit-estimator-chart"]');
+const barCount = (count: number) => chart().find('text.revenue-label').should('have.length', count);
+
+describe('Profit power-valid curves', { testIsolation: true }, () => {
+ for (const locale of ['en', 'zh'] as const) {
+ it(`keeps the target and pairs the selected valid curve (${locale})`, () => {
+ setup();
+ const width = locale === 'en' ? 1280 : 390;
+ cy.viewport(width, 900);
+ let csv: Blob | undefined;
+ cy.visit(`${locale === 'zh' ? '/zh' : ''}/profit-estimator-per-gigawatt?c_power=modeled`, {
+ onBeforeLoad: (win) => {
+ unlock(win);
+ win.URL.createObjectURL = (blob) => {
+ if (blob instanceof win.Blob) csv = blob;
+ return 'blob:profit-test';
+ };
+ win.HTMLAnchorElement.prototype.click = () => {};
+ },
+ });
+ cy.get('#profit-target').should('have.value', '45');
+ barCount(3);
+ chart()
+ .find('.x-axis')
+ .should('contain', 'MI355X')
+ .and('not.contain', 'H200')
+ .and('not.contain', 'GB300')
+ .and('not.contain', 'B300');
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
+ cy.get('body').then(($body) => expect($body[0].scrollWidth).to.be.at.most(width));
+ chart().scrollIntoView();
+ cy.screenshot(`profit-valid-curves-${locale}`, { capture: 'viewport', overwrite: true });
+ cy.get('#profit-power').click();
+ cy.get('[role="option"]')
+ .contains(locale === 'en' ? 'Compare both' : '对比两种估算方式')
+ .click();
+ barCount(8);
+ cy.get('#profit-target').should('have.value', '45');
+ cy.get('[data-testid="export-button"]').first().click();
+ cy.get('[data-testid="export-csv-button"]').click();
+ cy.then(() => csv!.text()).then((text) => {
+ const rows = text
+ .split('\n')
+ .filter((line) => /^(?:B200|B300|GB300|MI355X)/.test(line))
+ .map((line) => line.split(','));
+ expect(rows).to.have.length(8);
+ const paired = rows.filter((row) => row[0].includes('MI355X'));
+ expect(paired).to.have.length(4);
+ const revenue = paired.map((row) => row[9]);
+ expect(new Set(revenue).size).to.equal(2);
+ for (const value of new Set(revenue))
+ expect(revenue.filter((r) => r === value)).to.have.length(2);
+ expect(text).not.to.contain('NaN');
+ });
+ cy.get('#profit-power').click();
+ cy.get('[role="option"]')
+ .contains(locale === 'en' ? 'All in Provisioned' : '整体预配功耗')
+ .click();
+ barCount(5);
+ });
+ }
+
+ it('uses valid curves independently for historical comparisons', () => {
+ setup();
+ cy.viewport(1280, 900);
+ cy.visit(
+ `/profit-estimator-per-gigawatt?c_power=compare&i_gpus=mi355x_vllm&i_dstart=${PROFIT_HISTORY_DATE}&i_dend=${PROFIT_HISTORY_DATE}`,
+ { onBeforeLoad: unlock },
+ );
+ cy.get('#profit-target').should('have.value', '45');
+ barCount(4);
+ chart().should('contain', PROFIT_HISTORY_DATE);
+ cy.get('[data-testid="profit-power-unavailable"]').should('not.exist');
+ });
+});
diff --git a/packages/app/src/app/api/v1/views/extensions.test.ts b/packages/app/src/app/api/v1/views/extensions.test.ts
index 069795f57..344ea8e68 100644
--- a/packages/app/src/app/api/v1/views/extensions.test.ts
+++ b/packages/app/src/app/api/v1/views/extensions.test.ts
@@ -390,6 +390,76 @@ describe('new dashboard projections', () => {
expect(body.comparisons[0].data.rows[0].revenuePerGpuHour).toBeCloseTo(5.8536, 4);
},
);
+ it.each(['modeled', 'compare'])(
+ 'uses valid power curves at the same target in official, historical and overlay %s estimates',
+ async (powerBasis) => {
+ const curve = (scale: number, date: string) =>
+ [
+ [20, 9000, 1],
+ [40, 8000, 0],
+ [60, 3000, 1],
+ ].map(([interactivity, throughput, powerValid], index) =>
+ agenticRow({
+ id: scale * 100 + index,
+ conc: 3 - index,
+ // The exact logical snapshot includes a retained endpoint from an
+ // older producer run; power selection must not split that curve.
+ date: scale === 2 && index === 0 ? '2026-09-08' : date,
+ run_url:
+ scale === 2 && index === 0
+ ? 'https://github.com/SemiAnalysisAI/InferenceX/actions/runs/122'
+ : agenticRow().run_url,
+ curve_workflow_run_id: scale * 1000,
+ curve_date: date,
+ metrics: {
+ ...agenticRow().metrics,
+ p90_itl: 1 / interactivity,
+ tput_per_gpu: throughput * scale,
+ input_tput_per_gpu: throughput * scale * 0.9,
+ output_tput_per_gpu: throughput * scale * 0.1,
+ power_valid: powerValid,
+ },
+ }),
+ );
+ mocks.benchmarks.mockImplementation((request: NextRequest) =>
+ Response.json(
+ request.nextUrl.searchParams.get('exact') === 'true'
+ ? curve(2, '2026-09-09')
+ : curve(1, '2026-09-10'),
+ ),
+ );
+ mocks.unofficial.mockImplementation(() =>
+ Response.json({ benchmarks: curve(3, '2026-09-10'), evaluations: [] }),
+ );
+ const query = `model=DeepSeek-V4-Pro&precisions=fp4&target=45&priceSource=custom&inputPrice=1&cachedInputPrice=1&outputPrice=1&powerBasis=${powerBasis}&dates=2026-09-09&unofficialrun=456`;
+ const response = await gw(req('profit-estimator-per-gigawatt', query));
+ expect(response.status).toBe(200);
+ const body = await response.json();
+ for (const [output, scale] of [
+ [body.data, 1],
+ [body.comparisons[0].data, 2],
+ [body.overlays, 3],
+ ]) {
+ expect(output.skipped).toEqual([]);
+ expect(output.rows).toHaveLength(powerBasis === 'compare' ? 2 : 1);
+ // The existing Steffen curve over valid endpoints at 20/60 yields
+ // 4766.6015625 tok/s/GPU at 45.
+ // Both power budgets use that throughput, even though the full performance
+ // frontier includes the faster, power-invalid knot at 40.
+ for (const row of output.rows)
+ expect(row.revenuePerGpuHour).toBeCloseTo(17.159765625 * scale);
+ }
+ const provisioned = await gw(
+ req(
+ 'profit-estimator-per-gigawatt',
+ query.replace(`powerBasis=${powerBasis}`, 'powerBasis=provisioned'),
+ ),
+ );
+ const baseline = await provisioned.json();
+ expect(baseline.data.rows).toHaveLength(1);
+ expect(baseline.data.rows[0].revenuePerGpuHour).toBeGreaterThan(17.159765625);
+ },
+ );
it('returns NVL72 measured basis and matching capacity for official and overlay estimates', async () => {
const tray = agenticRow({
hardware: 'gb200',
diff --git a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
index c59d3ae10..e737a393d 100644
--- a/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
+++ b/packages/app/src/components/calculator/ProfitEstimatorDisplay.tsx
@@ -205,7 +205,7 @@ const STRINGS = {
benchmarkGroup: 'Benchmark Config',
powerLabel: 'Power Estimation',
powerTooltip:
- 'Change only the power budget used to scale the same benchmark result to one GW. Pricing, throughput, utilization and unit costs stay the same.',
+ 'All in Measured uses the best power-valid curve at the selected target. Compare both uses that same curve for both bars when available. Pricing, utilization and unit costs stay the same.',
powerOptions: {
provisioned: POWER_BASIS_LABELS['utility-provisioned'].en,
modeled: POWER_BASIS_LABELS['utility-modeled'].en,
@@ -319,7 +319,7 @@ const STRINGS = {
benchmarkGroup: '基准测试配置',
powerLabel: '功耗估算方式',
powerTooltip:
- '仅更改将同一基准测试结果换算为每 GW 收益时采用的功耗预算。价格、吞吐量、利用率和单位成本保持不变。',
+ '整体实测功耗在选定目标下采用功耗有效的最优曲线。对比两种估算方式时,若实测估算可用,两根柱子采用同一条曲线。价格、利用率和单位成本保持不变。',
powerOptions: {
provisioned: POWER_BASIS_LABELS['utility-provisioned'].zh,
modeled: POWER_BASIS_LABELS['utility-modeled'].zh,
@@ -983,7 +983,15 @@ function ProfitEstimatorInner({
// from that date's run with the same target, prices, and TCO tier.
const fullEstimate = useMemo(() => {
if (!hasData || !pricing) return { rows: [], skipped: [] };
- const current = getResults(targetValue, mode, interpolationCostProvider);
+ const curvePowerBasis = basis === 'gw-year' ? powerBasis : 'provisioned';
+ const current = getResults(
+ targetValue,
+ mode,
+ interpolationCostProvider,
+ undefined,
+ false,
+ curvePowerBasis,
+ );
const results = historyActive
? [
...current.filter((r) => selectedGPUs.includes(r.hwKey)),
@@ -994,6 +1002,7 @@ function ProfitEstimatorInner({
targetValue,
mode,
costProvider: interpolationCostProvider,
+ powerBasis: curvePowerBasis,
currentRunIds: historyCurrentRunIds,
}),
]
@@ -1021,6 +1030,7 @@ function ProfitEstimatorInner({
hasData,
pricing,
getResults,
+ basis,
powerBasis,
t.powerBarLabels,
targetValue,
diff --git a/packages/app/src/components/calculator/profit-history.test.ts b/packages/app/src/components/calculator/profit-history.test.ts
index 63027737d..29643db2c 100644
--- a/packages/app/src/components/calculator/profit-history.test.ts
+++ b/packages/app/src/components/calculator/profit-history.test.ts
@@ -2,6 +2,7 @@ import { describe, expect, it } from 'vitest';
import type { AvailabilityRow, BenchmarkRow, RunConfigRow } from '@/lib/api';
import { Percentile } from '@/lib/data-mappings';
+import { modeledPowerAtTarget } from './profit-power';
import {
buildProfitHistoryResults,
@@ -175,6 +176,51 @@ describe('buildProfitHistoryResults', () => {
costProvider: 'costh' as const,
};
+ it('selects valid knots within each historical run without borrowing from older runs', () => {
+ const row = (x: number, throughput: number, valid: boolean, run = 200) => {
+ const point = agenticRow('2026-06-14', 'b200', x, throughput, {
+ workflow_run_id: run,
+ run_started_at: `2026-06-14T${run === 200 ? '12' : '06'}:00:00Z`,
+ });
+ return {
+ ...point,
+ metrics: {
+ ...point.metrics,
+ power_valid: Number(valid),
+ power_metric_schema_version: 2,
+ avg_power_w: 500,
+ avg_total_gpu_power_w: 4000,
+ },
+ };
+ };
+ const rows = [row(20, 8000, true), row(40, 9000, false), row(80, 3000, true)];
+ const selected = buildProfitHistoryResults([{ date: '2026-06-14', rows }], {
+ ...options,
+ powerBasis: 'modeled',
+ });
+ expect(selected).toHaveLength(1);
+ expect(selected[0].nearestPoints.map((p) => p.interactivity)).toEqual([20, 80]);
+ expect(modeledPowerAtTarget(selected[0], 60)).toHaveProperty('kwPerGpu');
+ const noLatestPower = buildProfitHistoryResults(
+ [
+ {
+ date: '2026-06-14',
+ rows: [
+ row(20, 8000, true, 100),
+ row(80, 3000, true, 100),
+ row(20, 8000, false),
+ row(80, 3000, false),
+ ],
+ },
+ ],
+ { ...options, powerBasis: 'compare' },
+ );
+ expect(noLatestPower[0].nearestPoints.every((p) => p.sourceRow?.workflow_run_id === 200)).toBe(
+ true,
+ );
+ expect(modeledPowerAtTarget(noLatestPower[0], 60)).toEqual({ reason: 'no-measured-power' });
+ });
+
it('interpolates the selected chip at the target on each date and stamps the date', () => {
const rowsByDate = [
{
diff --git a/packages/app/src/components/calculator/profit-history.ts b/packages/app/src/components/calculator/profit-history.ts
index cf1c821a7..dc4513aa9 100644
--- a/packages/app/src/components/calculator/profit-history.ts
+++ b/packages/app/src/components/calculator/profit-history.ts
@@ -29,7 +29,7 @@ import { getHardwareConfig, getModelSortIndex, isKnownGpu } from '@/lib/constant
import { type Percentile, Sequence } from '@/lib/data-mappings';
import { getDisplayLabel } from '@/lib/utils';
-import { interpolateForGPU } from './interpolation';
+import { interpolateProfitForGPU, type ProfitPowerBasis } from './profit-power';
import type { ProfitEstimatorRow } from './profit-estimator';
import { buildGpuGroups, type GroupMeta } from './throughput-data';
import type { CalculatorMode, CostProvider, InterpolatedResult } from './types';
@@ -168,6 +168,7 @@ export function buildProfitHistoryResults(
targetValue: number;
mode: CalculatorMode;
costProvider: CostProvider;
+ powerBasis?: ProfitPowerBasis;
/**
* Per chip, the run its current bar is built from
* (`profitHistoryCurrentRunIds`); a pinned run entry for that same run is
@@ -183,6 +184,7 @@ export function buildProfitHistoryResults(
targetValue,
mode,
costProvider,
+ powerBasis = 'provisioned',
currentRunIds = {},
} = options;
if (selectedGPUs.length === 0) return [];
@@ -211,7 +213,7 @@ export function buildProfitHistoryResults(
for (const [groupKey, points] of Object.entries(grouped)) {
const meta = groupMeta[groupKey];
if (!meta) continue;
- const result = interpolateForGPU(points, targetValue, mode, costProvider);
+ const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis);
if (!result || !(result.value > 0)) continue;
results.push({
...result,
diff --git a/packages/app/src/components/calculator/profit-power.test.ts b/packages/app/src/components/calculator/profit-power.test.ts
index a408619f2..f297608f3 100644
--- a/packages/app/src/components/calculator/profit-power.test.ts
+++ b/packages/app/src/components/calculator/profit-power.test.ts
@@ -8,6 +8,7 @@ import { buildGpuGroups, interpolateForGPU } from './useThroughputData';
import { estimateProfitRows } from './profit-estimator';
import {
estimateProfitByPower,
+ interpolateProfitForGPU,
modeledPowerAtTarget,
type ProfitPowerSource,
} from './profit-power';
@@ -56,6 +57,8 @@ const point: GPUDataPoint = {
throughput: 6000,
inputThroughput: 5940,
outputThroughput: 60,
+ inputTokenShare: 0.99,
+ cacheHitRate: 0.9,
concurrency: 16,
tp: 8,
precision: 'fp4',
@@ -136,6 +139,212 @@ const withPoints = (base: InterpolatedResult, points: GPUDataPoint[]): Interpola
});
describe('profit power basis preview', () => {
+ it('retains the existing Steffen result for the Kimi K3 MI355X valid knots at 45', () => {
+ const knots = [
+ { ...point, interactivity: 14.39677512237259, throughput: 12090.08106 },
+ { ...point, interactivity: 53.85029617662897, throughput: 7227.86065 },
+ ];
+ expect(
+ interpolateProfitForGPU(knots, 45, 'interactivity_to_throughput', 'costh', 'modeled')?.value,
+ ).toBeCloseTo(8059.609379677125, 8);
+ });
+
+ it('selects power-valid raw knots before the frontier at the unchanged target', () => {
+ const low = {
+ ...point,
+ interactivity: 14.396775,
+ throughput: 8000,
+ sourceRow: { ...source, id: 443687, conc: 48 },
+ };
+ const high = {
+ ...point,
+ interactivity: 53.850296,
+ throughput: 4000,
+ sourceRow: { ...source, id: 443686, conc: 14 },
+ };
+ const invalid = {
+ ...point,
+ interactivity: 15.951507,
+ throughput: 9000,
+ sourceRow: {
+ ...source,
+ id: 443692,
+ conc: 44,
+ metrics: { ...source.metrics, power_valid: 0 },
+ },
+ };
+ const points = [low, invalid, high];
+ const original = interpolateForGPU(points, 45, 'interactivity_to_throughput', 'costh')!;
+ expect(original.nearestPoints.map((p) => p.sourceRow?.id)).toEqual([443692, 443686]);
+ expect(modeledPowerAtTarget(original, 45)).toEqual({ reason: 'no-measured-power' });
+ const eligible = interpolateForGPU([low, high], 45, 'interactivity_to_throughput', 'costh')!;
+ for (const basis of ['modeled', 'compare'] as const) {
+ const selected = interpolateProfitForGPU(
+ points,
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ basis,
+ )!;
+ expect(selected).toEqual(eligible);
+ expect(modeledPowerAtTarget(selected, 45)).toHaveProperty('kwPerGpu');
+ const output = estimateProfitByPower(
+ [selected],
+ specs,
+ pricing,
+ assumptions,
+ basis,
+ 45,
+ labels,
+ );
+ expect(output.skipped).toEqual([]);
+ expect(output.rows).toHaveLength(basis === 'compare' ? 2 : 1);
+ if (basis === 'compare') {
+ expect(output.rows[0].revenue / output.rows[0].gpuHours).toBeCloseTo(
+ output.rows[1].revenue / output.rows[1].gpuHours,
+ 10,
+ );
+ }
+ }
+ expect(
+ interpolateProfitForGPU(points, 45, 'interactivity_to_throughput', 'costh', 'provisioned'),
+ ).toEqual(original);
+ });
+
+ it('keeps inherited producer rows in the selected logical curve', () => {
+ const points = [20, 60].map((interactivity, i) => ({
+ ...point,
+ interactivity,
+ throughput: 8000 - i * 4000,
+ sourceRow: {
+ ...source,
+ date: `2026-09-${10 + i}`,
+ workflow_run_id: 100 + i,
+ run_url: `https://github.com/SemiAnalysisAI/InferenceX/actions/runs/${100 + i}`,
+ curve_date: '2026-09-12',
+ curve_workflow_run_id: 102,
+ },
+ }));
+ const selected = interpolateProfitForGPU(
+ points,
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'modeled',
+ )!;
+ expect(selected.clamped).toBe(false);
+ expect(selected.nearestPoints).toEqual(points);
+ expect(modeledPowerAtTarget(selected, 45)).toHaveProperty('kwPerGpu');
+ });
+
+ it('accepts an exact valid point and preserves provisioned fallback outside the valid range', () => {
+ const exact = interpolateProfitForGPU(
+ [point],
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'modeled',
+ )!;
+ expect(modeledPowerAtTarget(exact, 45)).toHaveProperty('kwPerGpu');
+ const invalid = {
+ ...point,
+ interactivity: 60,
+ throughput: 3000,
+ sourceRow: { ...source, metrics: { ...source.metrics, power_valid: 0 } },
+ };
+ const points = [{ ...point, interactivity: 20 }, invalid];
+ const original = interpolateForGPU(points, 45, 'interactivity_to_throughput', 'costh')!;
+ const fallback = interpolateProfitForGPU(
+ points,
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'compare',
+ )!;
+ expect(fallback).toEqual(original);
+ const estimate = estimateProfitByPower(
+ [fallback],
+ specs,
+ pricing,
+ assumptions,
+ 'compare',
+ 45,
+ labels,
+ );
+ expect(estimate.rows).toHaveLength(1);
+ expect(estimate.rows[0].powerLabel).toBe(labels.provisioned);
+ expect(estimate.skipped[0].reason).toBe('no-measured-power');
+ const clamped = interpolateProfitForGPU(
+ [{ ...point, interactivity: 38 }],
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'modeled',
+ )!;
+ expect(modeledPowerAtTarget(clamped, 45)).toEqual({ reason: 'outside-measured-range' });
+ const missingCpu = {
+ ...trayPoint,
+ sourceRow: { ...traySource, metrics: { ...traySource.metrics, cpu_power_valid: 0 } },
+ };
+ const excluded = interpolateProfitForGPU(
+ [missingCpu],
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'modeled',
+ )!;
+ expect(modeledPowerAtTarget(excluded, 45)).toEqual({ reason: 'no-cpu-power' });
+ });
+
+ it('partitions the whole frontier by sensor basis before interpolation', () => {
+ const graceSource: BenchmarkRow = {
+ ...traySource,
+ power_audit: {
+ cpu: { sensor_kind: 'grace_socket', expected_sockets: 2, observed_sockets: 2 },
+ },
+ metrics: { ...traySource.metrics },
+ };
+ delete graceSource.metrics.avg_total_module_power_w;
+ const grace = [30, 60].map((interactivity, i) => ({
+ ...trayPoint,
+ sourceRow: graceSource,
+ interactivity,
+ throughput: 9000 - i * 5000,
+ cacheHitRate: 0.5 + i * 0.2,
+ inputTokenShare: 0.8 + i * 0.1,
+ }));
+ const modules = [10, 20].map((interactivity, i) => ({
+ ...trayPoint,
+ interactivity,
+ throughput: 12000 - i * 1000,
+ cacheHitRate: 0.99,
+ inputTokenShare: 0.99,
+ }));
+ const expected = interpolateForGPU(grace, 45, 'interactivity_to_throughput', 'costh');
+ expect(
+ interpolateForGPU([...modules, ...grace], 45, 'interactivity_to_throughput', 'costh')?.value,
+ ).not.toBe(expected?.value);
+ for (const points of [[...modules, ...grace], [...grace, ...modules].toReversed()]) {
+ expect(
+ interpolateProfitForGPU(points, 45, 'interactivity_to_throughput', 'costh', 'compare'),
+ ).toEqual(expected);
+ }
+ const fasterModules = modules.map((p, i) => ({
+ ...p,
+ interactivity: 30 + i * 30,
+ throughput: 11000 - i * 5000,
+ }));
+ expect(
+ interpolateProfitForGPU(
+ [...fasterModules, ...grace],
+ 45,
+ 'interactivity_to_throughput',
+ 'costh',
+ 'modeled',
+ ),
+ ).toEqual(interpolateForGPU(fasterModules, 45, 'interactivity_to_throughput', 'costh'));
+ });
+
it.each([8])(
'keeps raw power attached through official and run-keyed %i-GPU frontiers',
(gpus) => {
diff --git a/packages/app/src/components/calculator/profit-power.ts b/packages/app/src/components/calculator/profit-power.ts
index 48ed51813..bbaa1e01c 100644
--- a/packages/app/src/components/calculator/profit-power.ts
+++ b/packages/app/src/components/calculator/profit-power.ts
@@ -11,7 +11,8 @@ import {
type ProfitEstimatorSkipReason,
type ProfitEstimatorSpecs,
} from './profit-estimator';
-import type { GPUDataPoint, InterpolatedResult } from './types';
+import { interpolateForGPU } from './interpolation';
+import type { CalculatorMode, CostProvider, GPUDataPoint, InterpolatedResult } from './types';
export type ProfitPowerBasis = 'provisioned' | 'modeled' | 'compare';
@@ -110,7 +111,37 @@ function planningPower(point: GPUDataPoint): PlanningPower {
};
}
-/** Reusing the original frontier prevents the power choice from changing throughput. */
+/** Select a compatible power-valid frontier within the caller's selected curve. */
+export function interpolateProfitForGPU(
+ points: GPUDataPoint[],
+ target: number,
+ mode: CalculatorMode,
+ costProvider: CostProvider,
+ powerBasis: ProfitPowerBasis,
+): InterpolatedResult | null {
+ const original = interpolateForGPU(points, target, mode, costProvider);
+ if (powerBasis === 'provisioned' || mode !== 'interactivity_to_throughput') return original;
+ const cohorts = new Map();
+ for (const point of points) {
+ const power = planningPower(point);
+ if ('reason' in power) continue;
+ const key = powerSourceKey(power.source);
+ const cohort = cohorts.get(key) ?? [];
+ cohort.push(point);
+ cohorts.set(key, cohort);
+ }
+ let best: InterpolatedResult | null = null;
+ // Source ordering makes equal-throughput selection independent of input order.
+ for (const key of [...cohorts.keys()].toSorted()) {
+ const candidate = interpolateForGPU(cohorts.get(key)!, target, mode, costProvider);
+ if (!candidate || 'reason' in modeledPowerAtTarget(candidate, target)) continue;
+ if (!best || candidate.value > best.value) best = candidate;
+ }
+ // Preserve provisioned-only fallback and the existing unavailability reason.
+ return best ?? original;
+}
+
+/** Read power from the same selected knots as throughput, without extrapolation. */
export function modeledPowerAtTarget(result: InterpolatedResult, target: number): PlanningPower {
if (result.clamped) return { reason: 'outside-measured-range' };
const exact = result.nearestPoints.find((p) => Math.abs(p.interactivity - target) < 1e-9);
diff --git a/packages/app/src/components/calculator/useThroughputData.ts b/packages/app/src/components/calculator/useThroughputData.ts
index 0f87a2afe..41dfaeab5 100644
--- a/packages/app/src/components/calculator/useThroughputData.ts
+++ b/packages/app/src/components/calculator/useThroughputData.ts
@@ -24,6 +24,7 @@ import {
sign,
} from './interpolation';
import type { CostProvider, CostType, GPUDataPoint, InterpolatedResult } from './types';
+import { interpolateProfitForGPU, type ProfitPowerBasis } from './profit-power';
// Re-export pure functions so existing imports from this module keep working.
export {
@@ -261,6 +262,7 @@ export function useThroughputData(
costProvider: CostProvider,
visibleHwKeys?: Set,
hideSkuAboveConfigLimit = false,
+ powerBasis: ProfitPowerBasis = 'provisioned',
): InterpolatedResult[] => {
const results: InterpolatedResult[] = [];
@@ -270,7 +272,7 @@ export function useThroughputData(
// Skip GPUs that are not visible (legend filters by hwKey)
if (visibleHwKeys && !visibleHwKeys.has(hwKey)) continue;
- const result = interpolateForGPU(points, targetValue, mode, costProvider);
+ const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis);
if (result && result.value > 0 && !(hideSkuAboveConfigLimit && result.clampedAbove)) {
results.push({
...result,
@@ -305,6 +307,7 @@ export function useThroughputData(
visibleHwKeys?: Set,
runInfoByIndex?: Record,
hideSkuAboveConfigLimit = false,
+ powerBasis: ProfitPowerBasis = 'provisioned',
): InterpolatedResult[] => {
const results: InterpolatedResult[] = [];
@@ -313,7 +316,7 @@ export function useThroughputData(
if (!meta) continue;
if (visibleHwKeys && !visibleHwKeys.has(meta.hwKey)) continue;
- const result = interpolateForGPU(points, targetValue, mode, costProvider);
+ const result = interpolateProfitForGPU(points, targetValue, mode, costProvider, powerBasis);
if (result && result.value > 0 && !(hideSkuAboveConfigLimit && result.clampedAbove)) {
results.push({
...result,
diff --git a/packages/app/src/lib/api-route-catalog.ts b/packages/app/src/lib/api-route-catalog.ts
index 826b2f081..5f51eaa46 100644
--- a/packages/app/src/lib/api-route-catalog.ts
+++ b/packages/app/src/lib/api-route-catalog.ts
@@ -843,6 +843,14 @@ export interface ApiContractSourceDigest {
* touching a route module. Digest changes require an explicit documentation review.
*/
export const apiContractSourceDigests = [
+ {
+ source: 'src/components/calculator/profit-power.ts',
+ sourceSha256: 'ff92b0a954afb78769ab8a29ca6ca7bdc52f83bb49379e2f2f274ef4c179330a',
+ reviewArea: {
+ en: 'Power-valid curve selection at fixed targets, compatible power bases, paired throughput and provisioned fallback shared by Profit UI and API.',
+ zh: '利润界面与 API 共用的有效功耗曲线选择、固定目标值、功耗口径兼容性、配对吞吐量与预配估算回退。',
+ },
+ },
{
source: 'src/components/inference/utils/inference-table-data.ts',
sourceSha256: '816eb0dbc416f6c322fd3955af13fc03f8c6988704d82f36a9bfd8ec9abc9669',
@@ -1023,7 +1031,7 @@ export const apiContractSourceDigests = [
{
source: 'src/lib/views-api/calculator-extensions.ts',
- sourceSha256: '2bcd27b5f5fbea9dcd75b32f7ee88e1c92343bed3ff65a293fc9a8ff5ecf04c3',
+ sourceSha256: '0c571acc23940e65290a5421c0872eb63f2531b289699c520209910b50dee784',
reviewArea: {
en: 'Dashboard read-only selector and calculation parity.',
zh: '仪表板只读接口的选择项与计算一致性。',
diff --git a/packages/app/src/lib/views-api/calculator-extensions.ts b/packages/app/src/lib/views-api/calculator-extensions.ts
index ba33414f7..1373907e9 100644
--- a/packages/app/src/lib/views-api/calculator-extensions.ts
+++ b/packages/app/src/lib/views-api/calculator-extensions.ts
@@ -4,13 +4,15 @@ import {
parseFirstTokenCaps,
selectFirstTokenWinners,
} from '@/components/calculator/first-token-limits';
-import { interpolateForGPU } from '@/components/calculator/interpolation';
import {
DEFAULT_UTILIZATION_PCT,
listPricingToTokenRevenuePricing,
profitModelDefaults,
} from '@/components/calculator/profit-estimator';
-import { estimateProfitByPower } from '@/components/calculator/profit-power';
+import {
+ estimateProfitByPower,
+ interpolateProfitForGPU,
+} from '@/components/calculator/profit-power';
import type { TokenRevenuePricing } from '@/components/inference/types';
import { fetchOpenRouterPricing } from '@/hooks/api/use-openrouter-pricing';
import { cachedJson } from '@/lib/api-cache';
@@ -183,11 +185,12 @@ export function calculatorExtension(view: CalculatorExtension, request: NextRequ
} as const;
function estimate(group: typeof groups.official) {
const results = Object.entries(group.grouped).flatMap(([key, points]) => {
- const result = interpolateForGPU(
+ const result = interpolateProfitForGPU(
points,
target,
'interactivity_to_throughput',
costProvider === 'custom' ? 'costh' : costProvider,
+ basis === 'gw-year' ? powerBasis : 'provisioned',
);
return result && result.value > 0
? [{ ...result, ...group.groupMeta[key], resultKey: key }]
diff --git a/packages/app/src/lib/views-api/docs/extensions.ts b/packages/app/src/lib/views-api/docs/extensions.ts
index 80ee85a20..aa0bd9d98 100644
--- a/packages/app/src/lib/views-api/docs/extensions.ts
+++ b/packages/app/src/lib/views-api/docs/extensions.ts
@@ -122,8 +122,8 @@ const PARAMETER_NOTES: Record = {
'模型许可或收入分成百分比,范围 0 至 100,默认值随模型变化。',
],
powerBasis: [
- 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. All in Measured uses measured GPU power plus modeled unmeasured components and PUE; it is not measured wall power. Eligible measured source rows are required; powerLabel identifies paired estimates and full-chassis extrapolation. NVL72 requires validated GPU and Grace/module telemetry with complete socket coverage; CPU rail alone is insufficient. powerSource records topology, measured basis, sensor, PUE and the app-owned model content revision, TypeScript source path and source hash. compare retains provisioned estimates when measured estimates are unavailable; skipped reasons distinguish missing CPU power and incompatible sensor bases. Missing coverage is not zero.',
- 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。整体实测功耗采用 GPU 实测值,加上未实测组件的功耗估算和 PUE,并非墙上电表读数。该估算需要符合条件的实测数据行;powerLabel 标明对比方式和整机外推。NVL72 需要通过验证的 GPU 与 Grace/模块遥测,并完整覆盖所有 socket;仅 CPU rail 读数不满足要求。powerSource 记录拓扑、实测口径、传感器、PUE 及由应用维护的模型内容版本、TypeScript 源码路径和源码哈希。compare 在实测估算不可用时保留预配估算;skipped 原因区分 CPU 功耗缺失与传感器口径不兼容。缺失数据不按零处理。',
+ 'provisioned (default, All in Provisioned), modeled (All in Measured) or compare. For GW-year estimates, modeled filters source points to valid system-power inputs before building the curve at the same target, without extrapolation or historical substitution. compare uses identical valid-curve throughput for paired budgets; when no valid curve reaches the target, it retains the original provisioned estimate. provisioned-only keeps the original performance curve. All in Measured adds modeled components and PUE to measured inputs; it is not measured wall power. NVL72 requires complete validated GPU and Grace/module telemetry; CPU rail alone is insufficient, and sensor bases must be compatible. powerSource records topology, measured basis, sensor, PUE, model revision and source hash. powerLabel marks paired estimates and full-chassis extrapolation; skipped reasons preserve missing coverage.',
+ 'provisioned(默认,整体预配功耗)、modeled(整体实测功耗)或 compare。按 GW 年估算时,modeled 先筛选满足系统功耗要求的数据点,再在同一目标值上构建曲线,不外推,也不借用历史数据。compare 的两种功耗方案使用同一条有效曲线的吞吐量;没有有效曲线覆盖目标时,保留原曲线的预配估算。provisioned 单独使用时仍沿用原性能曲线。整体实测功耗在实测输入上叠加组件估算和 PUE,并非墙上电表读数。NVL72 需要完整且通过验证的 GPU 与 Grace/模块遥测,仅 CPU rail 读数不足,传感器口径也必须兼容。powerSource 记录拓扑、实测口径、传感器、PUE、模型版本和源码哈希;powerLabel 标明配对估算和整机外推,skipped 保留覆盖缺失原因。',
],
power: [
'Comma-separated certified and/or legacy power tiers. Omit for all tiers.',
diff --git a/packages/app/timings.json b/packages/app/timings.json
index df735201a..976d0b235 100644
--- a/packages/app/timings.json
+++ b/packages/app/timings.json
@@ -248,6 +248,10 @@
"spec": "cypress/e2e/profit-estimator.cy.ts",
"duration": 39450
},
+ {
+ "spec": "cypress/e2e/profit-power-curves.cy.ts",
+ "duration": 3568
+ },
{
"spec": "cypress/e2e/reliability-chart.cy.ts",
"duration": 5511
diff --git a/packages/skills/skills/inferencex-api/integrity.json b/packages/skills/skills/inferencex-api/integrity.json
index a0640e448..4eeb07c27 100644
--- a/packages/skills/skills/inferencex-api/integrity.json
+++ b/packages/skills/skills/inferencex-api/integrity.json
@@ -8,7 +8,7 @@
"references/cli-contract.md": "fa44faa38d889b4fbdee5ba42752b47758ee3706fab150513bd4e6db87a86cb0",
"references/cli.md": "96b228f34cb3600f4548f4dc84df8506531747286763c56ca834c77bee05e1eb",
"references/collectivex.md": "eb794f9c28d4a27bec4db80c42c4685ff3b1d204fd9b12a6788511fb258a789f",
- "references/dashboard-views.md": "61d995291e9aa7280a5ced5af1a672ad5e0b001336db7c56a7655aa1a024ec64",
+ "references/dashboard-views.md": "47fb1611effb6d3c12467e693e1659ff1a91891c7979d36898fb7194fa7caf19",
"references/offline-exports.md": "95aa565dcd4c9e592159baa560a9a58217bc61e395de1c6daf0576795c7f4ec9",
"references/pareto.md": "1b4d2d163f982e3f2d97addce310789eae9c1e5badd82dd7501708c9e0385027",
"references/powerx.md": "cf8adcc5dfe395c5fb659908819980a872475f862f22b723815af43d363cedad",
diff --git a/packages/skills/skills/inferencex-api/references/dashboard-views.md b/packages/skills/skills/inferencex-api/references/dashboard-views.md
index b2765ff2c..7ea5896d6 100644
--- a/packages/skills/skills/inferencex-api/references/dashboard-views.md
+++ b/packages/skills/skills/inferencex-api/references/dashboard-views.md
@@ -116,10 +116,14 @@ watts again. Read `powerSource` for topology, measured basis, sensor, PUE, model
content revision, app TypeScript source path and source hash. Equations and parameters
are maintained in InferenceX-app; the revision is a content digest, not a private-repository
Git commit. Model-only updates recalculate retained valid measurements after deployment;
-they do not require telemetry backfill. `compare` preserves provisioned rows when a measured
-estimate is unavailable; `skipped.reason` distinguishes `no-cpu-power` from
-`incompatible-power-basis`. Modeled-only estimates never substitute provisioned
-watts, and neither mode selects a different serving frontier to fill missing power.
+they do not require telemetry backfill. For GW-year estimates, `modeled` first selects
+points with valid system-power inputs, then builds the curve at the requested target.
+It does not extrapolate or substitute historical snapshots. `compare` uses the same
+valid-curve throughput for both budgets; if no valid curve covers the target, it keeps
+the original provisioned estimate. `provisioned` alone retains the original performance
+curve. Official, comparison and unofficial scopes are evaluated independently.
+`skipped.reason` distinguishes missing CPU power and incompatible sensor bases;
+modeled estimates never substitute provisioned watts.
Prefer equal-service comparisons for article-facing hardware analysis. Use
`xstat=mean` only for fixed-sequence service axes when that statistic is intended: