Skip to content

[PowerX] drive GB200 AgentX power from resolved recipes / 根据实际配方接入 GB200 AgentX 功耗 - #3358

Open
edwingao28 wants to merge 2 commits into
mainfrom
feat/gb200-agentx-recipe-power
Open

edwingao28 wants to merge 2 commits into
mainfrom
feat/gb200-agentx-recipe-power

Conversation

@edwingao28

Copy link
Copy Markdown
Collaborator

Description

Replace GB200 AgentX's model whitelist with resolved recipe routing. Enable DSV4 through existing telemetry fields, preserving serving differences and required-power semantics.

Testing: 108 local tests passed; 155 offline combinations checked. Unchanged-code evidence reused after rebase (regressions).

Blocker: GPU sweep/evals and collection qualification pending; page delivery depends on InferenceX-app #1167.

中文

将 GB200 AgentX 的模型白名单改为按实际解析的配方接线。DSV4 使用现有遥测字段接入,保留 serving 差异和 required-power 语义。

测试: 108 项本地测试通过,155 个离线组合已检查。rebase 后相关代码不变,复用原验收证据;回归测试见上方链接。

待办: GPU sweep/evals 和真实采集资格尚未验证;页面交付依赖 InferenceX-app #1167。

AI 模型: GPT-6;无法核实精确运行时版本。用于实现、委派审查和草拟。

无关联 issue。改动类型为功能、配置和文档;清单在下方填写一次。

AI model disclosure

  • Model/version: GPT-6; exact runtime version unavailable.
  • Role: Implementation, delegated review, and drafting.

Related Issue

None.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

通过实际生效的 SRT 配方接入 GB200 AgentX 功耗,保留 required-power、测量窗口和模型运行差异,并为 DSV4 添加普通配方接入与离线行为回归。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

将本次 GB200 配方功耗改动追加条目的占位链接更新为 PR #3358,保留历史条目和运行逻辑。
@edwingao28
edwingao28 marked this pull request as ready for review September 22, 2026 05:44
@edwingao28
edwingao28 requested a review from a team September 22, 2026 05:44
@claude

claude Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 5m 8s —— View job


Review of PR #3358

  • Gather context (read diff against origin/main)
  • Validate master config / perf-changelog rules
  • Review runners/slurm_utils.sh and runners/launch_gb200-nv.sh power routing
  • Review infx/srt_slurm/synthetic_acceptance.py and recipe YAMLs
  • Run utils/test_gb200_recipe_power.py (could not run, see below)
  • Post findings

LGTM - no blocking issues found.

Traced the new flow end to end: setup_srt_slurm still appends the eval passthrough to SRTCTL_EVAL_ARGS before prepare_gb200_srt_power copies it into SRTCTL_RECIPE_ARGS, so the inspection and the srtctl apply see identical overrides. The power-producer-sha.txt and exporter squash provenance moved but are still written before write_srt_cluster_config consumes DCGM_EXPORTER_SQSH, and the MODEL_PATH override for the non-AgentX DSV4 sglang lane still lands before the cluster config is generated. The one GB200 recipe currently selected without an indexed selector (dsr1/sglang/gb200-fp8/8k1k/disagg-tp8-stp-variants.yaml) has no telemetry block and is non-agentic, so it hits the preserved multi-variant none path rather than the new "exactly one selected recipe" error. configs/nvidia-master.yaml only gained a comment, and the perf-changelog.yaml entry is appended at the end of the file without append-only.

Not verified in this environment: the utils/srt-slurm submodule is not checked out on the CI runner and network fetches were declined, so I could not run utils/test_gb200_recipe_power.py or read srtctl's parse_overrides / SrtConfig schema. Two behaviors therefore rest on the author's local test run: (1) the shell validation in runners/slurm_utils.sh:141 no longer rejects duplicate concurrencies the way inject_srt_power_concurrencies.py did, so test_invalid_matrix_concurrencies_fail_before_submission["1 1"] depends on the srtctl schema rejecting [1, 1]; (2) benchmark.concurrencies is now injected for every AgentX GB200 job, not only power-enabled ones, which is fine as long as srtctl's custom-benchmark runner invokes agentic_srt.sh once regardless of that list (the existing kimik3 power recipes already rely on this).

for role in before:
original_spec = json.loads(before[role]["args"].pop("speculative-config"))
generated_spec = json.loads(after[role]["args"].pop("speculative-config"))
assert generated_spec.pop("rejection_sample_method") == "synthetic"
original_spec = json.loads(before[role]["args"].pop("speculative-config"))
generated_spec = json.loads(after[role]["args"].pop("speculative-config"))
assert generated_spec.pop("rejection_sample_method") == "synthetic"
assert generated_spec.pop("synthetic_acceptance_length") > 1
Comment on lines +66 to +72
'source "$1" || exit $?; config="$2"; shift 2; '
'SRTCTL_EVAL_ARGS+=("$@"); '
'prepare_gb200_srt_power "$config" dynamo-vllm || exit $?; '
'printf \'{"dcgm":%s,"agentx":%s}\\n\' '
'"$USES_DCGM_POWER" "$USES_AGENTX_POWER" > "$LANE"; '
'apply_srt_recipe "$config" dynamo-vllm '
'-f "$config" --tags "ordinary AgentX submission" "${SRTCTL_RECIPE_ARGS[@]}"',
Comment thread runners/slurm_utils.sh
Comment on lines +136 to +145
if [[ "$IS_AGENTIC" == "1" ]]; then
check_env_vars CONC_LIST
local -a power_concurrencies
read -r -a power_concurrencies <<< "$CONC_LIST"
for concurrency in "${power_concurrencies[@]}"; do
[[ "$concurrency" =~ ^[1-9][0-9]*$ ]] || {
echo "Error: invalid AgentX concurrency: $concurrency" >&2
return 1
}
done

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 GB200 AgentX submissions can now silently run duplicate concurrency points, wasting a multi-node allocation, where the base branch rejected them before submission. prepare_gb200_srt_power (runners/slurm_utils.sh:136-145) only regex-checks each CONC_LIST entry is a positive integer; it never checks uniqueness, unlike the dedupe check it replaces (inject_srt_power_concurrencies.py's _validate_concurrencies, which raises "concurrencies must be positive unique integers" on duplicates). That script is still used, with its uniqueness check intact, by launch_gb300-nv.sh, launch_b200-nscale-slurm.sh and launch_h200-dgxc-slurm.sh, so this is a regression specific to the GB200 path. …

Why this was flagged

…Fix: prepare_gb200_srt_power must also reject duplicate values in CONC_LIST (e.g. check len(power_concurrencies) unique) before building benchmark.concurrencies/CONC_LIST, matching the sibling launchers' dedupe guard.

Trigger: an operator edits a master config's conc-list (configs/nvidia-master.yaml, e.g. the dsv4 AgentX entries added at line ~5709/5718) to include a duplicate value, or any future CONC_LIST producer emits one; infx/matrix/validation.py only checks conc_list entries are >0, never uniqueness. That value flows unchanged into CONC_LIST for a GB200 AgentX job and reaches prepare_gb200_srt_power (runners/slurm_utils.sh:136-151), which validates format only. benchmark.concurrencies and CONC_LIST are set with the duplicate, and benchmarks/multi_node/agentic_srt.sh iterates CONCURRENCIES as given, re-running and overwriting the same conc_{N} result path with no error. On the base branch this same duplicate would be caught before the srtctl submission by inject_srt_power_concurrencies.py's _validate_concurrencies, aborting the launch instead of burning a multi-node allocation.

Verification: nit. Real but low-severity regression. Base branch's GB200 AgentX power path (removed hunk in runners/launch_gb200-nv.sh) ran inject_srt_power_concurrencies.py, whose _validate_concurrencies (runners/inject_srt_power_concurrencies.py:20-22) rejects duplicates via len(set(concurrencies)) != len(concurrencies) and exits before submission. The new prepare_gb200_srt_power… | nit. Factually…

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants