Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
374b4e1
feat: add NVL72 rack profiles and tray-based measured system power es…
edwingao28 Sep 18, 2026
0bf119f
feat: accept fully measured NVL72 trays in the profit planning gate a…
edwingao28 Sep 18, 2026
e4d95e1
feat: register NVL72 CPU-side power keys and power_audit.cpu across c…
edwingao28 Sep 18, 2026
02cff94
test: cover the GB200 NVL72 basis control in the profit estimator Cyp…
edwingao28 Sep 18, 2026
7a654f4
docs: add the NVL72 rack estimate section to the PowerX system-power doc
edwingao28 Sep 18, 2026
3f0c53e
fix: scrub CPU-side power on its own verdict and model NVL72 trays as…
edwingao28 Sep 18, 2026
a1b21e6
chore: merge feat/powerx-db-ingest and reconcile NVL72 tray planning …
edwingao28 Sep 19, 2026
1745e70
feat: infer NVL72 trays for aggregate multinode rows without a worker…
edwingao28 Sep 19, 2026
f2de031
refactor: cite the producer contract instead of a local ticket in cpu…
edwingao28 Sep 19, 2026
128605b
Merge branch 'feat/powerx-db-ingest' into feat/powerx-nvl72-smart-pro…
edwingao28 Sep 21, 2026
3e7672f
fix: reconcile partial chassis and NVL72 profit estimates
edwingao28 Sep 21, 2026
45aa4be
feat: integrate NVL72 planning with consolidated PowerX UI
edwingao28 Sep 30, 2026
4360563
fix: synchronize NVL72 views with PowerX ruler and tooltip updates
edwingao28 Sep 30, 2026
cc86afd
test: synchronize NVL72 views with Chrome coverage
edwingao28 Sep 30, 2026
d46b18e
fix: pin latest NVL72 reference revision
edwingao28 Sep 30, 2026
0861bdc
docs: split PowerX modeling guide into English and Chinese
edwingao28 Sep 30, 2026
8610123
docs: clarify telemetry contracts and model assumptions
edwingao28 Sep 30, 2026
989bb69
refactor: own PowerX models in InferenceX-app
edwingao28 Sep 30, 2026
35d15b9
test: expect the app-owned system power model in the NVL72 CSV
edwingao28 Sep 30, 2026
03d76fd
chore: merge #1220 comparison fixes into NVL72 modeling
edwingao28 Sep 30, 2026
68e51c0
chore: sync NVL72 modeling with the rebased telemetry parent
edwingao28 Sep 30, 2026
55d4331
test: trim redundant PowerX modeling cases
edwingao28 Sep 30, 2026
78bbbe7
chore: sync PowerX comparison fix into NVL72 modeling
edwingao28 Sep 30, 2026
b41c0e4
fix: align NVL72 chart explanations with measured inputs
edwingao28 Sep 30, 2026
517145b
fix: synchronize NVL72 dashboard with updated parent
edwingao28 Oct 1, 2026
3278b48
fix: keep dense profit comparison charts readable
edwingao28 Oct 1, 2026
84b9d53
fix: distinguish unpriced profit estimates from missing power
edwingao28 Oct 1, 2026
c80f7f3
fix: retain AgentX estimates in all-in power views
edwingao28 Oct 1, 2026
78585ab
fix: retain measured rows without all-in estimates
edwingao28 Oct 1, 2026
cd07394
docs: clarify unavailable multi-node system estimates
edwingao28 Oct 1, 2026
9a62745
fix: remove unavailable estimate lists (#1249)
edwingao28 Oct 1, 2026
c5a86f7
fix: select power-valid curves for profit estimates
edwingao28 Oct 1, 2026
a3b4e1e
fix: sync NVL72 CPU ingestion contract
edwingao28 Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 39 additions & 3 deletions docs/dashboard-readonly-views.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,8 +66,34 @@ state. GPU interactive downsampling does not alter returned raw data or statisti
Power boundary labels are GPU Level Measured, GPU Level Provisioned (TDP), All in
Provisioned, and All in Measured. The last combines measured GPU power with modeled
unmeasured components and PUE; it is not a wall-meter measurement. These labels and
collapsed power-assumption/availability notes do not change metric IDs, API selectors,
power-assumption notes do not change metric IDs, API selectors,
or calculations. Profit comparison `powerLabel` display text follows the same names.
All in Measured watts and energy accept validated 8K/1K and AgentX rows through the
shared chart/API transform, including historical and unofficial rows. AgentX reuses
the chassis or rack model without independent workload calibration. Telemetry and
topology gates still apply; NVL72 needs complete Grace or module power. The standalone
Modeled Chassis AC metric and 8K/1K offline export retain their 8K/1K scope.

For All in Measured, `tableRows` retains every GPU-valid observation in the selected
scope and best-series selection, including axis-clipped and non-frontier points.
Missing system estimates use `y: null`, `status: "unavailable"`, and
`unavailableReason`; `measuredGpuWatts` remains available. CSV exports these rows
with blank missing values. Numeric `series` and `count` are unchanged. Each date
comparison and unofficial overlay has its own `tableRows`; latest does not pool history.

For GW-year profit, `modeled` selects valid system-power points before building the
curve at the same target, without extrapolation or another snapshot. `compare` uses
the same valid-curve throughput for both power budgets, retaining the original
provisioned estimate when no valid curve covers the target. Provisioned-only keeps
the original performance curve. Official, comparison and unofficial scopes remain
independent; CPU/module telemetry and compatible sensor-basis requirements still apply.

Dense profit charts reserve readable space per bar and scroll within the plot on narrow
screens; captions and controls stay fixed. This is presentation-only: API selectors,
calculations, source identities and CSV rows are unchanged. PNG export includes the full
plot regardless of its current scroll position, so no API or skills contract change is needed.
Profit charts omit the per-configuration unavailable-estimate list. This is presentation-only;
the API retains skipped rows and reasons, and calculations and CSV exports are unchanged.
The GPU statistics table includes startup and warmup for all chips in the selected
series, regardless of chip visibility. It is separate from serving-window power,
J/token and selected-time-window calculations. Run telemetry is DB-first with an
Expand Down Expand Up @@ -201,8 +227,7 @@ point identity. Fewer than three distinct output rates return `fit: null` with
power, and R² is null when power did not vary.

These analytical results are JSON-only: `format=csv` with any analysis enabled returns
400, rather than silently exporting only the primary chart. Ordinary CSV retains
its existing plotted-point contract.
400, rather than silently exporting only the primary chart. CSV exports plotted points except for All in Measured, which exports `tableRows`.

| Surface | Dashboard control / share parameter | Read-only API coverage |
| ---------------------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------- |
Expand Down Expand Up @@ -233,6 +258,17 @@ those properties.
@semianalysisai/inferencex-skills 包。所有仪表板路由(含隐藏和功能开关控制的
视图)均在覆盖表中登记;上表列出各只读接口接受的全部查询参数名。
接口复用现有计算函数,公开运行与非官方叠加数据保留各自来源。
整体实测指标的 `tableRows` 保留当前筛选范围和最优曲线选择内所有 GPU 遥测有效的观测点,
不按前沿或坐标轴显示范围裁剪。估算不可用时 `y` 为 null,`status` 为 `unavailable`,
`unavailableReason` 给出原因;`measuredGpuWatts` 保留实测 GPU 功耗。CSV 导出同一组行,缺失值留空。
`series` 和 `count` 保持不变,仍只包含可绘制的数值点;各日期对比和非官方叠加分别返回自己的 `tableRows`,
Latest 不会合并历史数据。

按 GW 年估算利润时,modeled 先筛选满足系统功耗要求的数据点,再在原目标值上构建曲线,
不外推,也不借用其他快照。compare 的两种功耗方案使用同一条有效曲线的吞吐量;
没有有效曲线覆盖目标时,保留原曲线的预配估算。provisioned 单独使用时沿用原性能曲线。
官方数据、日期对比和非官方叠加各自独立计算;CPU/模块遥测要求和传感器口径兼容性要求同样不变。

私有上传、密钥、提示词、反馈及管理操作不作为公开读取接口。
OperatorX 的入口受功能开关控制,页面使用专属的 `/api/v1/operatorx/*`
接口;目前没有发布 `/api/v1/views/operatorx` 契约。
Expand Down
4 changes: 2 additions & 2 deletions docs/data-transforms.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,8 +79,8 @@ Returns `{ chartData: InferenceData[][], hardwareConfig: HardwareConfig }`.
| B4 utility modeled (measured → chassis → PUE) | `modeledSystemPower.deploymentFacilityWatts ÷ modeledSystemPower.gpuCount` | `B1 J/out × (B4 W ÷ B1 W)` | `utilityModeledWatts`, `utilityModeledJPerOutputToken` |

- **N_alloc and total throughput** (`powerBasisNormalization`). Aggregate rows report output per allocated GPU, so `N_alloc` cancels and the builder uses `W ÷ output_tput_per_gpu` without trusting display counts (legacy ingest can encode TP × EP twice). Fixed-sequence disaggregated rows (`disagg && benchmark_type === 'single_turn'`) report output per decode GPU while the deployment also powers the prefill pool, so `total output tok/s = output_tput_per_gpu × num_decode_gpu` and `N_alloc = num_prefill_gpu + num_decode_gpu`; the article's 4P+4D counts eight GPUs, and B3 energy is `(P + D) / D` times `jOutput` on such rows. Other disaggregated benchmark types emit no provisioned energy because it is not verifiable in-app whether AgentX throughput already divides by all GPUs.
- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. `deploymentFacilityWatts` is chassis AC × PUE with PUE applied exactly once inside `estimateChassisPower`, divided by the physical measured GPU count (not `modeledGpuCount`, which over-counts partially allocated chassis: a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8, and the unit test pins that divisor). This is a different quantity from the existing `modeledChassisPowerPerGpu` (chassis AC ÷ modeled GPU count, no PUE). Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator).
- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware such as GB200/GB300 NVL72 (`reason: 'hardware'`), and telemetry/topology failures withhold it. The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here).
- **B4 source.** Reuses the `SystemPowerEstimate` that `rowToAggDataEntry` attached as `entry.modeledSystemPower`; nothing re-runs `modelSystemPower`. B4 W/GPU is `deploymentFacilityWatts ÷ gpuCount`. PUE is applied exactly once to rounded AC power inside `estimateChassisPower` or `estimateRackPower`. `modelSystemPower` combines the resulting chassis or tray shares and attributes partially allocated units to their measured GPUs; NVL72 tray shares come from one rack evaluated at the measured trays’ mean input. Divide the deployment total by the physical measured GPU count, not `modeledGpuCount` (a 4P+4D deployment on two worker hosts models 16 GPUs while measuring 8). The existing `modeledChassisPowerPerGpu` instead uses `chassisAcWatts ÷ modeledGpuCount`, without PUE. Modeled energy scales the producer's same-window `joules_per_output_token` by `B4 W ÷ avg_power_w`, so it inherits B1's token denominator and survives rows whose `output_tput_per_gpu` is missing (the provisioned energies do not; a per-basis point count cannot assume one shared denominator).
- **Null rules.** A value is emitted only when finite and positive; otherwise the key is omitted (never `{ y: 0 }`), because the metric filters drop points by `metricKey in point` and `remapInferencePoint` falls back to raw throughput when a key exists with an unusable value. B2/B3 W are absent for hardware without registry specs (`getGpuSpecs` returns zeros); their energies are also absent without output throughput or disaggregated counts. B4 requires `modeledSystemPower.status === 'supported'` and B1 watts (`avg_power_w` after `rowToAggDataEntry`'s `power_valid !== 0` admission); B4 energy additionally needs `joules_per_output_token`. Telemetry admission belongs to `modelSystemPower`, which accepts `power_valid === 1` with schema v2 or the validated unversioned single-node producer (`telemetryBasis: 'validated-unversioned-single-node'`), so B4 renders on exactly the rows that show B1 and `modeledChassisPowerPerGpu`; off-8k/1k workloads, unsupported hardware, and telemetry/topology failures withhold it. GB200/GB300 NVL72 are supported when validated GPU telemetry is accompanied by complete, matching Grace or module telemetry (`reason: 'cpu-telemetry'` when that CPU-side evidence is missing or invalid). The public API's `strictV2` row filter is not re-applied in the chart for B1, so it is not re-applied for B4 either (plan §3.1 describes B1 with that filter; the app's chart path is the authority here).
- **Ordering invariant** (unit-tested on real B200 telemetry): B3 ≥ B4 ≥ B1 and B2 ≥ B1 for both W/GPU and J/out on the same point.
- **Reconstructed prefill energy** (`reconstructedPrefillJPerOutputToken`, `utils/role-energy.ts`). For validated (`power_valid === 1`, schema 2) disaggregated rows, `prefill_joules_per_input_token × (joules_per_output_token ÷ joules_per_input_token)` carries the prefill pool's energy onto the output-token axis: the ratio is the served input:output token count because schema-2 aggregate energy has one numerator. With `decode_joules_per_output_token` it sums back to the deployment's J/out. It feeds only the `i_pcompare=roles` comparison on the energy axis (PowerX Figure 7) and is never a y-axis of its own; aggregate rows and rows missing any of the four inputs omit it.
- Historical Trends substitutes `output_tput_per_gpu := tput_per_gpu` for legacy rows lacking output throughput; provisioned energies in trends inherit that fallback.
Expand Down
2 changes: 1 addition & 1 deletion docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ Design rationale and non-obvious conventions. See [CLAUDE.md](../CLAUDE.md) for
- [Pareto Boundary API](./pareto-api.md): Query frontier and hinterland observations, preserve provenance, and distinguish API scope from chart highlights.

- [API Skill Examples](./inferencex-api-examples.md) — Install the public skill, query benchmarks, export measured PowerX data, and explain empty results
- [PowerX System Power](./powerx-system-power.md) — Pinned chassis model, measured-input guards, assumptions, and reproducible article exports
- [PowerX System Power](./powerx-system-power.md) / [简体中文](./powerx-system-power.zh.md) — Measured-curve steps, chassis/NVL72 requirements, a worked planning example, missing-data diagnosis, and reproducible exports
- [PowerX Permanent View](./powerx-permanent-view.md) — Power boundaries as gated Measured Energy metrics, `i_metric`/`i_rulers` share links, missing-value states
- [PowerX Persistence and Recovery](./powerx-persistence-recovery.md) — Telemetry receipts, migration prerequisites, and targeted repair
- [API Skill Releases](./inferencex-skills-release.md) — Prepare an immutable package, verify clean installations and agent exports, and publish through the package-specific workflow
Expand Down
16 changes: 8 additions & 8 deletions docs/powerx-permanent-view.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,17 +23,17 @@ group stays out of the selector otherwise. The boundary metrics are members of t

## Boundaries

| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source |
| --------------------- | ---------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------- |
| `gpu-measured` | GPU measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged |
| `gpu-provisioned` | GPU provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` |
| `utility-provisioned` | Utility provisioned (all-in) | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) |
| `utility-modeled` | Utility modeled (PUE) | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` chassis AC × PUE 1.3 (applied once), ÷ measured GPUs |
| Basis (`PowerBasis`) | Selector label | W / GPU metric | J / output token metric | Source |
| --------------------- | --------------------------- | -------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| `gpu-measured` | GPU Level Measured | `y_measuredAvgPower` (+P75/P90, roles, %TDP) | `y_measuredJPerOutputToken` (+ input/total/query) | runner telemetry; existing metrics, unchanged |
| `gpu-provisioned` | GPU Level Provisioned (TDP) | `y_gpuProvisionedWatts` | `y_gpuProvisionedJPerOutputToken` | `HW_REGISTRY.tdp` |
| `utility-provisioned` | All in Provisioned | `y_utilityProvisionedWatts` | `y_utilityProvisionedJPerOutputToken` | `HW_REGISTRY.power` (all-in kW per GPU) |
| `utility-modeled` | All in Measured | `y_utilityModeledWatts` | `y_utilityModeledJPerOutputToken` | `modelSystemPower` deployment AC ÷ measured GPUs × PUE (1.3 air, 1.1 NVL72; applied once) |

Formulas, the all-GPU normalization (`N_alloc` = prefill + decode GPUs for disaggregated
rows) and the null rules are specified in
[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis model
itself in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its
[Data Transforms → Power boundaries](./data-transforms.md#power-boundaries); the chassis and NVL72 rack models
in [PowerX System Power](./powerx-system-power.md). The ungated `jOutput` keeps its
per-decode-GPU normalization; its labels and the boundary metric's `all GPUs` label keep the
two distinguishable in the selector, the availability list and CSV headers.

Expand Down
Loading
Loading