[Klaud Cold] Update GB200 DSV41flash vLLM image and retune TP4/DEP2/DEP4 / [Klaud Cold] 更新 GB200 DSV41flash vLLM 镜像并调优 TP4/DEP2/DEP4 - #3687
functionstackx wants to merge 2 commits into
Conversation
Mirror the B200 retune from #3686 on GB200: bump to nightly-dev-arm64-cu130-ac9126e58aa7 with FlashInfer autotuning, run TP4 at concurrency 1-128, and replace TP2 with DEP2 (8-32) and DEP4 (64-128) behind a consistent-hash vLLM Router. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
| description: | ||
| - "Update the GB200 DeepSeek-V4.1-Flash vLLM AgentX image from nightly cd10ed6f to nightly-dev-arm64-cu130-ac9126e58aa7 and enable FlashInfer autotuning." | ||
| - "Run TP4 at concurrency 1-128 and replace TP2 with DEP2 (TP1 x DP2 + EP2, DeepGEMM MegaMoE) at concurrency 8-32 and DEP4 (TP1 x DP4 + EP4, MegaMoE) at concurrency 64-128, both behind a consistent-hash vLLM Router. All points set gpu-memory-utilization 0.97." | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD |
There was a problem hiding this comment.
🔴 This new changelog entry's pr-link uses pull/TBD, which will fail CI's perf-changelog validation for every reviewer/CI run on this PR. validate_perf_changelog.py's validate_added_pr_link only accepts the literal XXX placeholder (or the real pull/<pr_number>) via PR_LINK_PLACEHOLDERS/CANONICAL_PR_LINK; pull/TBD matches neither, so it raises ChangelogValidationError("new PR entry must use ... or an XXX placeholder; found ..."). Fix: use https://github.com/SemiAnalysisAI/InferenceX/pull/XXX (the placeholder other entries use, e.g. line 5598) until the real PR number is known.
Why this was flagged
The entry appended at inferencex-e2e/perf-changelog.yaml:9286 sets pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD. validate_added_pr_link in inferencex-e2e/infx/workflows/validate_perf_changelog.py:139-150 checks the link against PR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"} and CANONICAL_PR_LINK (requires pull/\d+); TBD satisfies neither, so validation raises ChangelogValidationError. On the base branch this same validator is satisfied by using the accepted pull/XXX placeholder (precedent at line 5598); this PR substitutes an unrecognized token instead, breaking the automated changelog-append check that gates merge.
Verification: CORRECTION to the candidate's trigger: this does NOT fail "every reviewer/CI run on this PR." The only pull_request-triggered validator never calls validate_added_pr_link, and matrix/validation.py:999 types pr_link as a bare str with no format check, so the PR check-changelog step PASSES with pull/TBD. validate_added_pr_link is only reached at merge time, via canonicalize_appended_links -> compare_entries (prepare_perf_changelog_merge.py:88,519, calling validate_added_pr_link at line 192-193).
| - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb200-fp4-mtp/agentic.yaml } | ||
| # DEP2 is native TP1 x DP2 + EP2 with MegaMoE behind a consistent-hash vLLM Router. | ||
| - { tp: 2, ep: 2, dp-attn: true, kv-offloading: none, spec-decoding: mtp, conc-list: [8, 16, 32], router: { name: vllm-router, version: "0.1.14" }, srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb200-fp4-mtp/agentic.yaml } | ||
| # DEP4 is native TP1 x DP4 + EP4 with MegaMoE behind the same router. | ||
| - { tp: 4, ep: 4, dp-attn: true, kv-offloading: none, spec-decoding: mtp, conc-list: [64, 128], router: { name: vllm-router, version: "0.1.14" }, srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb200-fp4-mtp/agentic.yaml } |
There was a problem hiding this comment.
🟡 (optional) DEP4 (tp:4, ep:4, dp-attn:true, conc-list [64,128]) at nvidia-master.yaml:8449 shares tp and concurrency with the TP4 arm at :8445, so both get the identical exp-name dsv41flash_tp4_conc64_kvnone_spec-mtp for two different deployments. generate.py's _agentic_entries exp-name formula (infx/matrix/generate.py:1188-1192) encodes model, tp, conc, kv-offload and spec-decoding but never ep/dp-attn. This breaks the uniqueness convention documented at nvidia-master.yaml:1544 and makes filter_exp_names raise 'matched multiple rows' for that name. Fix: fold ep/dp-attn into exp-name for any arm that can share tp+conc with another.
Why this was flagged
Trigger: generating the agentic matrix for dsv41flash-fp4-gb200-vllm-agentic-dspark (full-sweep, generate_config_matrix, or test-config --exp-names) via infx/matrix/generate.py's _agentic_entries. The TP4 row at nvidia-master.yaml:8445 (tp:4, conc-list [1,4,8,16,32,64,128]) and the new DEP4 row at :8449 (tp:4, ep:4, dp-attn:true, conc-list [64,128]) both produce exp-name dsv41flash_tp4_conc64_kvnone_spec-mtp (and the conc128 variant) because generate.py:1188-1192 omits ep and dp-attn. On base no two rows for this key share tp+conc: GB300's DEP2 sibling uses tp:2 and the sglang EP arms use a different tp than their TP arm, so exp-names were unique; this is the first collision. filter_exp_names then raises 'Experiment name(s) matched multiple rows' for that name, blocking targeted selection, and no test in infx/tests/matrix checks exp-name uniqueness across search-space rows, so the claimed 528 passing tests do not catch it.
Verification: nit. The exp-name collision is real and reachable, but its serious consequences are prevented downstream, so this is a convention/tooling issue, not a correctness regression. generate.py:1188-1192 builds the exp-name from tp, conc, kv-offload and spec-decoding, omitting ep and dp-attn, so TP4 (:8445) and DEP4 (:8449) both emit dsv41flash_tp4_conc64_kvnone_spec-mtp, violating the uniqueness convention at :1544 and making filter_exp_names raise "matched multiple rows".
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=37076187881 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=37076187881 |
Description
Applies the B200 retune from #3686 to the GB200 DeepSeek-V4.1-Flash vLLM AgentX arm (
dsv41flash-fp4-gb200-vllm-agentic-dspark).cd10ed6ftovllm/vllm-openai:nightly-dev-arm64-cu130-ac9126e58aa7, the arm64 build of the commit that Update B200 DSV41flash vLLM image and retune TP4/DEP2/DEP4 / 更新 B200 DSV41flash vLLM 镜像并调优 TP4/DEP2/DEP4 #3686 (B200) uses and Update GB300 DSV41flash vLLM image and retune TP4/DEP2 / 更新 GB300 DSV41flash vLLM 镜像并调优 TP4/DEP2 #3652 (GB300) shipped.FULL_AND_PIECEWISECUDA graphs with capture sizes in multiples of the six-token DSpark verification block, an MXFP4 indexer KV cache with sparse logits, an fp8 KV cache andgpu-memory-utilization0.97.Validation
infx/tests/srt_slurmandinfx/tests/matrixpass (528 tests).infx.matrix.generate test-configemits 12 points (TP4 ×7, DEP2 ×3, DEP4 ×2). Each matches exactly one zip variant, and every zip list has consistent lengths.full-sweep-fail-fastlabel is pending. This PR claims no performance results yet.AI model disclosure
claude-opus-5-5) in Claude Code; no delegated agents.Related Issue
None. Companion to #3686 (B200).
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesfull-sweep-fail-fast(recommended),full-sweep-enabled, ornon-canary-full-sweep-enabled. Optional modifiersall-evals,evals-only, andagentx-fastrequire a primary label; the last two block reuse while applied.OWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.中文
改动说明
将 #3686 中 B200 的调优移植到 GB200 DeepSeek-V4.1-Flash vLLM AgentX 配置(
dsv41flash-fp4-gb200-vllm-agentic-dspark)。cd10ed6f更新为vllm/vllm-openai:nightly-dev-arm64-cu130-ac9126e58aa7,即 Update B200 DSV41flash vLLM image and retune TP4/DEP2/DEP4 / 更新 B200 DSV41flash vLLM 镜像并调优 TP4/DEP2/DEP4 #3686(B200)所用 commit 的 arm64 构建,也是 Update GB300 DSV41flash vLLM image and retune TP4/DEP2 / 更新 GB300 DSV41flash vLLM 镜像并调优 TP4/DEP2 #3652(GB300)已合入的镜像。FULL_AND_PIECEWISECUDA graph,捕获尺寸为六 token DSpark 验证块的倍数,使用带稀疏 logits 的 MXFP4 indexer KV cache、fp8 KV cache,并设置gpu-memory-utilization0.97。验证
infx/tests/srt_slurm和infx/tests/matrix通过(528 个测试)。infx.matrix.generate test-config生成 12 个点(TP4 ×7、DEP2 ×3、DEP4 ×2),每个点唯一匹配一个 zip variant,且各 zip 列表长度一致。full-sweep-fail-fastlabel 触发 GPU sweep,结果待出,暂不声明性能结果。AI 模型使用说明
claude-opus-5-5);没有委派其他 agent。关联 issue
无。与 #3686(B200)配套。
改动类型
配置变更、文档更新。
🤖 Generated with Claude Code