[PowerX] enable NVL72 smart provisioning on measured compute-module power / 为 NVL72 启用基于实测计算模块功耗的智能预配 - #1190
edwingao28 wants to merge 32 commits into
Conversation
…timate Pin inferencex_power_model 963ead8b (feat/gb200-nvl72-rack-model) and emit gb200/gb300 rack profiles plus 204 Python parity cases; the x86 chassis profiles and their 292 cases regenerate byte-identically. Port the rack evaluation to TypeScript with the source rounding order (rack AC rounded before PUE). modelSystemPower admits NVL72 rows as compute trays (one host, four GPUs, two Grace sockets) whose module or GPU-board + Grace-socket watts are measured: cpu_power_valid=1 and the Grace-side keys are required, the Grace side is never modelled, DLC PUE 1.1 applies once at rack AC, a partial tray extrapolates only the GPU-board share, and the result carries topologyBasis 'nvl72-trays', measuredBasis, sensorKind and the new cpu-telemetry reason. x86 rows are unchanged. 中文:固定 inferencex_power_model 963ead8b(feat/gb200-nvl72-rack-model), 生成 gb200/gb300 机架 profile 和 204 条 Python 对照用例,x86 机箱 profile 及 其 292 条用例逐字节不变。TypeScript 按原模型的舍入顺序移植机架计算(机架交流 功率先舍入再乘 PUE)。modelSystemPower 将 NVL72 行按计算 tray(单主机、4 张 GPU、2 个 Grace socket)接纳,输入为实测模块功耗或 GPU 板卡 + Grace socket 功耗:要求 cpu_power_valid=1 及 Grace 侧指标,Grace 侧从不建模,DLC PUE 1.1 在机架交流侧只应用一次,部分分配的 tray 仅外推 GPU 板卡份额,结果新增 topologyBasis 'nvl72-trays'、measuredBasis、sensorKind 和 cpu-telemetry 原因。x86 行为不变。
…nd name the power basis Planning kW/GPU now admits nvl72-trays estimates whose trays are all fully measured next to the eight-GPU single-node chassis; partial trays stay rejected. Two frontier knots must share the same measured basis and sensor kind or the bracket stays unavailable. Modeled rows carry powerSource (topology, measured basis, sensor kind, PUE, pinned profile path, revision and source SHA-256); the bar tooltip, a caption line and three CSV columns name it per row. English strings for x86 rows are byte-identical. 中文:规划 kW/GPU 除八卡单节点机箱外,接受全部 tray 完整实测的 NVL72 估算, 部分 tray 仍被拒绝;两个前沿数据点的实测口径或传感器类型不同时保持不可用。 估算行携带 powerSource(拓扑、实测口径、传感器类型、PUE、固定 profile 路径、 版本和源文件 SHA-256),柱形提示、标注行和三列 CSV 逐行标出;x86 行英文字符串 逐字节不变。
…onstants, ingest and the API reference avg_cpu_socket_power_w, avg_total_cpu_power_w, total_cpu_energy_j, avg_total_module_power_w and total_module_energy_j join MEASURED_POWER_METRIC_KEY_LIST (withheld with power_valid=0 like every measured key); cpu_power_valid joins the contract discriminators and is normalized as a verdict independent of power_valid. extractPowerAudit keeps a bounded power_audit.cpu block matching the consumer's audit_summary. The API registry documents the six keys, the cpu audit schema, and a bilingual measured-power note and example; no route catalog digest changes. 中文:五个 CPU 侧实测指标加入 MEASURED_POWER_METRIC_KEY_LIST(与其他实测键一样 在 power_valid=0 时移除);cpu_power_valid 加入契约字段并按独立于 power_valid 的验证结论归一化;extractPowerAudit 保留与消费端 audit_summary 一致的有界 power_audit.cpu。API 文档新增六个键的说明、cpu 审计 schema 以及中英文 measured-power 说明与示例;路由目录摘要无需刷新。
…ress spec profit-fixtures gains an NVL72 SKU (one four-GPU tray with CPU-side keys and the module sensor) kept out of PROFIT_SKUS so existing bar counts hold. The new case checks the control stays hidden while the gate is locked, then prices the tray on Measured + modeled, names the measured module basis and DLC PUE 1.1 in the caption, and leaves GB300 unavailable. 中文:profit-fixtures 新增 NVL72 SKU(单个四卡 tray,含 CPU 侧指标与模块传感器), 不加入 PROFIT_SKUS 以保持现有柱形数量。新用例验证功能开关锁定时控件隐藏,解锁后 按实测 + 估算为该 tray 定价,标注行写明实测模块口径与液冷 PUE 1.1,GB300 保持不可用。
Measured input and admission rules, the modeled residual table with the UNVERIFIED parameters and their ranges, the shelf overflow bound, and the Profit Estimator gate rules, in English and in the 中文说明 section. The Profit Estimator paragraphs now name the per-hardware PUE policy, the same-basis knot rule and the CSV provenance columns in both languages. 中文:新增 NVL72 机架估算一节:实测输入与接纳条件、含 UNVERIFIED 参数及范围的 建模残差表、电源架容量上限和利润估算器门槛规则,中英文同步;利润估算器段落 双语补充按硬件取值的 PUE 策略、同口径数据点规则和 CSV 出处列。
… one rack - ingest withholds the CPU-side keys on cpu_power_valid != 1 and GPU-side keys on power_valid = 0; mapper tests cover all four verdict combinations; the power manifest carries cpu_power_valid - exporter shares defaultSystemPue with the dashboard (1.3 chassis, 1.1 DLC NVL72); NVL72 rows carry the rack profile, measured basis, sensor kind and CPU-side inputs; x86 rows byte-identical - chart tooltip names NVL72 compute trays and the measured basis instead of eight-GPU chassis (en/zh) - heterogeneous measured trays fold into one rack at their mean; the shelf curve is evaluated once at rack DC load, matching gb200_nvl72_rack_power; estimateRackPower and parity cases unchanged 中文:摄取按 cpu_power_valid 独立清除 CPU 侧指标,GPU 侧仍按 power_valid,mapper 测试覆盖四种组合,功耗清单附带 cpu_power_valid;导出器与仪表板共用 defaultSystemPue(机箱 1.3、液冷 NVL72 1.1),NVL72 行补充机架 profile、实测口径、传感器类型与 CPU 侧输入,x86 行逐字节不变;图表提示改用 NVL72 计算 tray 措辞并标出实测口径(中英文);多 tray 按均值折算为整机架,电源架曲线只在机架直流负载处求值一次,与 Python 模型一致,estimateRackPower 与对照用例不变。
…with multinode chassis planning Brings in the base's deleted-run skip for artifact backfills, the `uniform-hosts` topology basis (aggregate multinode chassis modeled at the deployment mean) and the per-host `power_audit_` telemetry ingest, and reconciles them with the NVL72 tray path: - `modelSystemPower`: the `uniform-hosts` branch keeps the base's x86 semantics unchanged and now fills the shared `MeasuredUnit` list; it is chassis-only (`!rack`), so an aggregate multinode NVL72 row without a per-worker array stays `topology`-unavailable rather than inferring a tray count from the GPU total. The supported-estimate union carries `'single-node' | 'worker-hosts' | 'uniform-hosts'` beside `'nvl72-trays'`. - Profit planning gate: every fully measured estimate is admitted (`chassisBasis: 'full'`); `nvl72-trays` estimates keep the measured-basis and sensor-kind source label, every chassis topology is labeled `chassis`. - Tooltip: the uniform-hosts note renders in the same position as on the base, alongside the NVL72 tray copy; English and Chinese x86 strings are unchanged. - Docs and tests from both sides kept; new tests pin the chassis-only uniform-hosts decision and the `chassis` label for multi-host sources. 中文:合并 feat/powerx-db-ingest(跳过 GitHub 已删除的 run、聚合多节点机箱按 部署平均功耗建模的 `uniform-hosts` 拓扑基准、按主机拆分的 `power_audit_` telemetry 入库),并与 NVL72 tray 路径对齐:`uniform-hosts` 分支保持 x86 语义不变、仅对机箱硬件生效,无 worker 数组的 NVL72 聚合多节点行仍按 `topology` 不可用;Profit Estimator 规划闸门接受所有完整实测的估算, NVL72 tray 保留实测基准与传感器标签,机箱拓扑统一标为 `chassis`;tooltip 的 uniform-hosts 说明位置与基线一致,x86 中英文文案不变;两侧文档与测试 均保留,并新增测试固定上述决策。
… array The Kimi K3 GB200 aggregate producer (dynamo-vLLM TP16, sixteen GPUs on four trays) emits no `workers` array, so `modelSystemPower` rejected every such row as `topology` and the tray path never reached the Profit Estimator. Generalize the base's `uniform-hosts` branch from eight-GPU chassis to the unit size of the hardware: on GB200/GB300, a non-disaggregated multinode row without a worker array is `gpuCount / 4` compute trays, each fed the deployment mean (module total per tray when the module keys are present, otherwise GPU-board plus Grace-socket watts per tray, exactly as on the worker-hosts tray path), with `chassisBasis: 'full'`, `topologyBasis: 'nvl72-trays'` and the usual measured basis and sensor kind. A GPU total that does not fill whole trays is `gpu-count`; the tray count must agree with the Grace-socket count recovered from the CPU-side keys and, when the CPU leg recorded it, with `power_audit.cpu.observed_sockets`, otherwise `cpu-telemetry`. The x86 uniform-hosts path, the worker-hosts tray path and every English string are unchanged. Tests cover module and Grace-socket bases, both socket mismatches, the uneven count and the Profit Estimator source label; the system-power doc records the rule in English and Chinese. 中文:Kimi K3 GB200 聚合采集端(dynamo-vLLM TP16,16 张 GPU 分布在 4 个 tray) 不输出 `workers` 数组,`modelSystemPower` 一律按 `topology` 拒绝,tray 路径无法进入 Profit Estimator。现将基线的 `uniform-hosts` 分支从八卡机箱推广到硬件的单元大小: GB200/GB300 上无 worker 数组的非 disagg 多节点行按 GPU 总数 ÷ 4 推算 tray 数,每个 tray 取部署平均值(有模块指标时用模块总功耗 ÷ tray 数,否则用 GPU 板卡 + Grace socket 每 tray 功耗,与 worker-hosts tray 路径一致),`chassisBasis: 'full'`、 `topologyBasis: 'nvl72-trays'`,实测口径与传感器类型不变。GPU 总数无法填满整数个 tray 时为 `gpu-count`;tray 数须与 CPU 侧指标推算的 Grace socket 数一致,且在 CPU 采集记录了 `power_audit.cpu.observed_sockets` 时与之一致,否则为 `cpu-telemetry`。 x86 uniform-hosts 路径、worker-hosts tray 路径及所有英文文案不变。测试覆盖模块与 Grace socket 两种口径、两类 socket 不一致、非整数 tray 数以及 Profit Estimator 的 来源标签;系统功耗文档中英文同步。
…-side comments Three comments referenced a task-tracker ticket that does not exist in this repository; point them at the InferenceX producer contract instead. No code change. 中文:三处注释引用了仓库中不存在的本地工单,改为指向 InferenceX 的生产端契约文档;无代码变更。
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Claude finished @edwingao28's task in 4m 14s —— View job Review complete — 🟢 LGTM, no blocking issues foundI reviewed the full diff against Todo
What I verified
Chinese copyThe diff touches user-visible zh strings in Notes (non-blocking)
This is a carefully structured, well-tested change. No changes requested. |
|
Claude finished @edwingao28's task in 4m 31s —— View job Claude Code Review — finalizing
|
Integrate the released partial-chassis rules, retain provisioned comparison rows when measured power is unavailable, and validate complete Grace or module provenance. 整合已发布的部分卡机箱规则;实测功耗缺失时保留预配功耗对比结果,并校验完整 Grace 或 module 遥测来源。
Preserve measured module and Grace provenance, retain provisioned comparisons when measured power is unavailable, and align API documentation, exports and collapsed assumptions. 中文:将 NVL72 规划整合到统一 PowerX 界面,保留模块及 Grace 实测来源;实测功耗不可用时仍显示预配对比,并同步 API 文档、导出与折叠说明。
|
Claude finished @edwingao28's task in 3m 1s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What I verified
Chinese copyThe Notes (non-blocking)
Carefully structured, well-tested change. No changes requested. |
|
Found 3 test failures on Blacksmith runners: Failures
|
中文:将运行级标尺分享与移动端提示框修复同步到 NVL72 视图,保留实测传感器校验和模型边界。
中文:同步整合 UI 的 Chrome 功耗菜单测试修正,保留 NVL72 模型与实测传感器校验。
|
Claude finished @edwingao28's task in 4m 36s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
The 3 Blacksmith failures are not attributable to this diffI traced each failing test to the file it lives in and the component it exercises, then diffed those against the base:
What I verified in the diff
Chinese copyThe touched Carefully structured, well-tested change. No changes requested. |
中文:将 NVL72 参考模型固定到最新修订,重新生成来源哈希与版本信息;所有机箱、机架参数及参考用例数值保持不变。
中文:将 PowerX 系统功耗与智能容量规划说明拆分为完整的中英文文档,补充 Kimi K3 缺失结果、NVL72 计算示例和模型更新流程,并明确参考源码尚未发布及校准边界。
Include the parent snapshot identity repair without changing the child scope. 中文:将父分支的 PowerX 曲线对比来源修复同步至NVL72 建模,保持当前 PR 的改动范围。
|
Claude finished @edwingao28's task in 1m 14s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last reviewOnly one new commit: the merge The only net delta the merge introduces into the PR's own file set is a docs/skills sync:
No production code, no user-visible Chinese strings, and no UI surface are touched by this commit, so the Playwright runtime check and the Verified
Notes (non-blocking)
No new issues introduced by the sync. No changes requested. |
Correct measured-power eligibility, energy scaling, and cooling-specific PUE in chart copy and model documentation. Show the app model digest and cover English/Chinese desktop and mobile states. 中文:修正 NVL72 图表说明,使其与实测输入一致。更新实测功耗准入、能耗换算及不同冷却方式的 PUE 说明,显示模型摘要,并覆盖中英文桌面和手机状态。
|
Claude finished @edwingao28's task in 1m 42s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last reviewOnly one new commit:
Verified
Chinese copyThe reworded Notes (non-blocking)
No new issues introduced by the copy alignment. No changes requested. |
中文:同步 NVL72 仪表板父分支,保留对比控件移除结果与现有功耗模型,并更新 API 技能完整性清单。修复新版 Chrome 中滚动操作改变 Cypress 断言对象的问题。
Reserve readable bar spacing, retain the SVG across scrolling thresholds, and export the full plot without moving the live chart. 中文:为密集利润对比图保留足够的柱形间距,在窄屏中启用图内横向滚动,并保持完整 PNG 导出与固定说明区域。切换滚动状态时保留 SVG,避免临界宽度下图形消失。
|
Claude finished @edwingao28's task in 2m 31s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last reviewThree new commits; only one touches production code:
What I verified in
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 3278b48. Configure here.
Group unavailable notices by the exact retained provisioned result, preserving configuration and history identity without changing estimates or API data. 中文:按实际保留的预配置功耗估算区分未定价 SKU 与缺少实测加建模结果的提示,避免将没有基准估算的配置误标为仅缺少实测功耗。保留现有 GPU 与 Grace 功耗回退行为。
|
Claude finished @edwingao28's task in 2m 18s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
The Cursor Bugbot finding is already fixed by
|
Enable the existing AgentX model option in shared chart transforms while keeping the standalone chassis AC metric scoped to 8K/1K. Preserve valid multi-node rows across charts, tables, history, overlays and the public view. 简体中文:整体实测功耗复用现有 AgentX 模型选项,保留符合条件的多节点 曲线与表格记录;独立机箱交流功耗指标仍仅适用于 8K/1K。精简中英文指南, 说明操作步骤、遥测要求和模型假设,移除 PR 与交付历史。
|
Claude finished @edwingao28's task in 5m 30s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last reviewOne production-code commit,
Verified in the diff
Runtime verificationLoaded Chinese copyThe touched Notes (non-blocking)
No new issues introduced by the AgentX All-in-Measured change. No changes requested. |
中文:整体功耗估算不可用时保留实测记录。
|
Claude finished @edwingao28's task in 4m 9s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last reviewThis commit makes the All in Measured table retain every GPU-valid observation even when the all-in (modeled) estimate can't be computed, instead of dropping those rows. The modeled estimate renders as Verified in the diff
Runtime verificationLoaded Test coverageComprehensive and colocated: Chinese copyThe new Notes (non-blocking)
No new issues introduced by the retain-measured-rows change. No changes requested. |
中文:明确多节点系统估算不可用的含义。
中文:移除无法估算配置列表。
|
Claude finished @edwingao28's task in 3m 7s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last review
This also fully supersedes the earlier Cursor Bugbot finding (compare-mode skip mislabeling at Verified in the diff
Runtime verificationLoaded Chinese copyThe touched Notes (non-blocking)
No new issues introduced by the unavailable-list removal. No changes requested. |
|
Claude finished @edwingao28's task in 3m 54s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last review
Verified in
|
中文:利润估算先选择功耗有效的曲线。
|
Claude finished @edwingao28's task in 3m 34s —— View job Re-review — 🟢 LGTM, no blocking issues foundRe-reviewed against
What changed since the last review
Verified in the diff
Runtime verificationLoaded Chinese copyThe one new user-visible Notes (non-blocking)
No new issues introduced by the power-valid curve selection. No changes requested. |


Summary
flowchart TD A[Performance + GPU and Grace/module measurements] B[Validate power and select compatible curve at target] C[Model rack overhead and PUE] D[Apply 10% planning reserve] E[Calculate capacity and profit per GW] F[Keep provisioned fallback in Compare both] A --> B B -->|Valid| C --> D --> E B -->|Unavailable| FNVL72 uses measured module or GPU+Grace power. Measured tables retain GPU-valid rows; profit selects power-valid curves before interpolation, with matched Compare both bars. Guide.
Depends on #1220; collection and calibration pending.
AI model disclosure
GPT-6 (variant unverified): implementation/review/tests; claude-fable-5-1: Chinese review; earlier versions unverified.
Validation
c5a86f7d: 7,383 unit tests (4 skipped) and 195 browser cases pass (EN/ZH, desktop/mobile, history, overlays); typecheck/lint/format pass. No GPU qualification.中文说明
NVL72 采用实测模块功耗,或 GPU 与 Grace 实测功耗。实测表格保留 GPU 功耗有效的记录;利润估算先筛选功耗有效的曲线再插值,Compare both 两根柱子使用同一条曲线。指南。
依赖 #1220;功耗采集和硬件校准仍待完成。
GPT-6 负责实现、审核及测试,精确变体无法核实;claude-fable-5-1 负责中文审核。早期版本无法核实。
c5a86f7d:7,383 项单元测试通过(4 项跳过),195 项浏览器用例通过,覆盖中英文、桌面/手机、历史数据和叠加运行;类型检查、lint 和格式检查通过。未做 GPU 资格验证。Note
High Risk
Changes admission and math for NVL72 system power, GW-year profit capacity, and public inference view/CSV contracts—areas where telemetry gaps or curve selection bugs directly misstate planning numbers.
Overview
NVL72 “All in Measured” now estimates facility power from validated GPU telemetry plus complete Grace or compute-module readings in the same window, with rack overhead, DLC PUE 1.1, and the existing planning reserve. GB200/GB300 are no longer blanket-unsupported in docs and transforms when CPU-side evidence is present; missing telemetry surfaces as
cpu-telemetry/ table reasons instead of hiding rows.Inference / views API: For All in Measured, responses add
tableRowsthat keep every GPU-valid point in scope (including axis-clipped and non-frontier), withy: null,status, andunavailableReasonwhen the system estimate fails;measuredGpuWattsstays populated. CSV for this metric exportstableRows(blank missing values), not only plottedseries. Official, date comparison, and unofficial overlays each carry their owntableRows.Per-GW profit estimator: modeled / compare build throughput and power from a power-valid performance curve at the selected target (no extrapolation across invalid knots); compare pairs bars on the same valid curve and keeps provisioned when measured is unavailable. Charts gain horizontal scroll on narrow viewports (PNG export still full width); the per-SKU unavailable-estimate disclosure list is removed (skipped reasons remain in API/CSV). Tooltips and CSV add power basis / sensor / profile metadata; NVL72 help text moves into option help and captions.
Docs & ops: Power boundary labels align with GPU Level Measured / All in Measured;
powerx-system-poweris rewritten with a Chinese twin. Offline export uses per-hardware default PUE and NVL72 columns; Python profile generation is dropped in favor of app-owned provenance (update-system-power-provenance.ts).NEXT_PUBLIC_APP_SOURCE_REFpins GitHub doc/model links in tooltips.Reviewed by Cursor Bugbot for commit c5a86f7. Bugbot is set up for automated code reviews on this repo. Configure here.