[PowerX] drive GB200 AgentX power from resolved recipes / 根据实际配方接入 GB200 AgentX 功耗 - #3358
edwingao28 wants to merge 2 commits into
Conversation
通过实际生效的 SRT 配方接入 GB200 AgentX 功耗,保留 required-power、测量窗口和模型运行差异,并为 DSV4 添加普通配方接入与离线行为回归。
|
Thanks for the contribution!
中文感谢你的贡献!
|
将本次 GB200 配方功耗改动追加条目的占位链接更新为 PR #3358,保留历史条目和运行逻辑。
|
Claude finished @edwingao28's task in 5m 8s —— View job Review of PR #3358
LGTM - no blocking issues found. Traced the new flow end to end: Not verified in this environment: the |
| for role in before: | ||
| original_spec = json.loads(before[role]["args"].pop("speculative-config")) | ||
| generated_spec = json.loads(after[role]["args"].pop("speculative-config")) | ||
| assert generated_spec.pop("rejection_sample_method") == "synthetic" |
| original_spec = json.loads(before[role]["args"].pop("speculative-config")) | ||
| generated_spec = json.loads(after[role]["args"].pop("speculative-config")) | ||
| assert generated_spec.pop("rejection_sample_method") == "synthetic" | ||
| assert generated_spec.pop("synthetic_acceptance_length") > 1 |
| 'source "$1" || exit $?; config="$2"; shift 2; ' | ||
| 'SRTCTL_EVAL_ARGS+=("$@"); ' | ||
| 'prepare_gb200_srt_power "$config" dynamo-vllm || exit $?; ' | ||
| 'printf \'{"dcgm":%s,"agentx":%s}\\n\' ' | ||
| '"$USES_DCGM_POWER" "$USES_AGENTX_POWER" > "$LANE"; ' | ||
| 'apply_srt_recipe "$config" dynamo-vllm ' | ||
| '-f "$config" --tags "ordinary AgentX submission" "${SRTCTL_RECIPE_ARGS[@]}"', |
| if [[ "$IS_AGENTIC" == "1" ]]; then | ||
| check_env_vars CONC_LIST | ||
| local -a power_concurrencies | ||
| read -r -a power_concurrencies <<< "$CONC_LIST" | ||
| for concurrency in "${power_concurrencies[@]}"; do | ||
| [[ "$concurrency" =~ ^[1-9][0-9]*$ ]] || { | ||
| echo "Error: invalid AgentX concurrency: $concurrency" >&2 | ||
| return 1 | ||
| } | ||
| done |
There was a problem hiding this comment.
🔴 GB200 AgentX submissions can now silently run duplicate concurrency points, wasting a multi-node allocation, where the base branch rejected them before submission. prepare_gb200_srt_power (runners/slurm_utils.sh:136-145) only regex-checks each CONC_LIST entry is a positive integer; it never checks uniqueness, unlike the dedupe check it replaces (inject_srt_power_concurrencies.py's _validate_concurrencies, which raises "concurrencies must be positive unique integers" on duplicates). That script is still used, with its uniqueness check intact, by launch_gb300-nv.sh, launch_b200-nscale-slurm.sh and launch_h200-dgxc-slurm.sh, so this is a regression specific to the GB200 path. …
Why this was flagged
…Fix: prepare_gb200_srt_power must also reject duplicate values in CONC_LIST (e.g. check len(power_concurrencies) unique) before building benchmark.concurrencies/CONC_LIST, matching the sibling launchers' dedupe guard.
Trigger: an operator edits a master config's conc-list (configs/nvidia-master.yaml, e.g. the dsv4 AgentX entries added at line ~5709/5718) to include a duplicate value, or any future CONC_LIST producer emits one; infx/matrix/validation.py only checks conc_list entries are >0, never uniqueness. That value flows unchanged into CONC_LIST for a GB200 AgentX job and reaches prepare_gb200_srt_power (runners/slurm_utils.sh:136-151), which validates format only. benchmark.concurrencies and CONC_LIST are set with the duplicate, and benchmarks/multi_node/agentic_srt.sh iterates CONCURRENCIES as given, re-running and overwriting the same conc_{N} result path with no error. On the base branch this same duplicate would be caught before the srtctl submission by inject_srt_power_concurrencies.py's _validate_concurrencies, aborting the launch instead of burning a multi-node allocation.
Verification: nit. Real but low-severity regression. Base branch's GB200 AgentX power path (removed hunk in runners/launch_gb200-nv.sh) ran inject_srt_power_concurrencies.py, whose _validate_concurrencies (runners/inject_srt_power_concurrencies.py:20-22) rejects duplicates via len(set(concurrencies)) != len(concurrencies) and exits before submission. The new prepare_gb200_srt_power… | nit. Factually…
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Description
Replace GB200 AgentX's model whitelist with resolved recipe routing. Enable DSV4 through existing telemetry fields, preserving serving differences and required-power semantics.
Testing: 108 local tests passed; 155 offline combinations checked. Unchanged-code evidence reused after rebase (regressions).
Blocker: GPU sweep/evals and collection qualification pending; page delivery depends on InferenceX-app #1167.
中文
将 GB200 AgentX 的模型白名单改为按实际解析的配方接线。DSV4 使用现有遥测字段接入,保留 serving 差异和 required-power 语义。
测试: 108 项本地测试通过,155 个离线组合已检查。rebase 后相关代码不变,复用原验收证据;回归测试见上方链接。
待办: GPU sweep/evals 和真实采集资格尚未验证;页面交付依赖 InferenceX-app #1167。
AI 模型: GPT-6;无法核实精确运行时版本。用于实现、委派审查和草拟。
无关联 issue。改动类型为功能、配置和文档;清单在下方填写一次。
AI model disclosure
Related Issue
None.
Type of Change
Checklist
perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.