Update GB300 DSV41flash image and tune TP/DEP / 更新 GB300 DSV41flash 镜像并调优 TP/DEP - #3575
Juntian777 wants to merge 1 commit into
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
中文:将 perf-changelog 条目链接到 #3575。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked the CONC=192 tail of zip_override_dep2 where max-num-seqs is 384 but max-cudagraph-capture-size stays at 2046 — that gap is intentional (batches beyond the captured graph size just fall back to eager/piecewise execution) and matches the author's own CONC>=96 tier in the PR description, so it's not a bug.
Extended reasoning...
The diff bumps the vLLM nightly image and retunes GB300 DSpark benchmark parallelism (TP4/TP2/new DEP2 zip variants) across the srt-slurm recipe, master config search-space rows, and an appended perf-changelog entry; no auth/crypto/permissions surface is touched. I independently confirmed the already-reported exp_name-collision finding by reading infx/matrix/generate.py:993-997, where the single-node exp_name formula only encodes tp/conc/kv_offloading/spec_decoding (not ep or dp-attn), so the new DEP2 row (tp:2, ep:2, dp-attn:true) and the existing plain TP2 row do produce identical exp_name values at CONC 16/32/64, with no dedup check in validation.py. I also checked the CONC=192 max-num-seqs(384) vs max-cudagraph-capture-size(2046) gap in zip_override_dep2 and ruled it out as intentional, matching the author's stated design.
| - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 4, 8, 16], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml } | ||
| # TP2 covers the middle range with the same attention and graph tiers. | ||
| - { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [2, 4, 8, 16, 32, 64], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml } | ||
| # DEP2 is native TP1 x DP2 + EP2, with MegaAttention/MegaMoE. | ||
| - { tp: 2, ep: 2, dp-attn: true, kv-offloading: none, spec-decoding: mtp, conc-list: [16, 32, 48, 64, 96, 128, 192], router: { name: vllm-router, version: "0.1.14" }, srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml } |
There was a problem hiding this comment.
🟡 (optional) Adding the DEP2 row (line 8402: tp:2, ep:2, dp-attn:true, conc-list [16,32,48,64,96,128,192]) alongside the existing plain TP2 row (line 8398: tp:2, conc-list [...,16,32,64]) makes matrix generation emit duplicate exp-name values for CONC 16, 32 and 64. The single-node exp_name formula in infx/matrix/generate.py:993 (f"{model_code}tp{tp}conc{conc}...") only encodes tp+conc, not ep/dp-attn, unlike multinode_agentic_exp_name which deliberately encodes ep/dpa to avoid exactly this collision (generate.py:490). Both runs at CONC=16 become 'dsv41flash_tp2_conc16_kvnone_spec-mtp', so infx/results/collect_results.py writes both into the same agg<exp_name>.json, silently overwriting one benchmark's results with the other's. …
Why this was flagged
…Fix: make the single-node exp_name include ep/dp-attn (or router) so any two single-node rows sharing tp+conc get distinct names, covering this and future dp-attn arms added at an existing tp value.
When infx/matrix generate.py builds the agentic matrix for dsv41flash-fp4-gb300-vllm-agentic-dspark, it iterates both search-space rows at nvidia-master.yaml:8398 (plain tp:2) and :8402 (tp:2, ep:2, dp-attn:true, DEP2). For CONC in {16,32,64}, both rows pass through generate.py:993's exp_name formula, which only interpolates tp and conc, producing the identical string 'dsv41flash_tp2_conc16_kvnone_spec-mtp' for two structurally different deployments (pure TP2 vs TP1xDP2+EP2 behind a vllm-router). No uniqueness check runs over the generated matrix (filter_exp_names at generate.py:1357 only checks user-supplied --exp-names, not generated rows). Downstream, infx/results/collect_results.py:8-16 writes agg_{exp_name}.json keyed by this string, so one run's results file silently overwrites the other's. On the base branch this row didn't exist, so no gb300 vLLM AgentX arm ever collided this…
Verification: normal. The new DEP2 row at inferencex-e2e/configs/nvidia-master.yaml:8402 ({tp:2, ep:2, dp-attn:true, kv-offloading:none, spec-decoding:mtp, conc-list:[16,32,48,64,96,128,192]}) shares tp:2/kv-none/mtp with the plain TP2 row at :8400 ({tp:2, conc-list:[2,4,8,16,32,64]}), and their conc-lists overlap at 16, 32, 64. The config sets multinode:false (line 8391), so both rows take the single-node…
中文:将 perf-changelog 条目链接到 #3575。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
b1b4461 to
3cfc7db
Compare
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36646378405 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36646378405 |
Move the GB300 DeepSeek-V4.1-Flash vLLM AgentX recipe to nightly 36768d1b with FlashInfer sparse attention, MXFP4 indexer KV, sparse logits, fp8 KV and FlashInfer autotuning disabled. Run TP4 at CONC 1, 4, 8, 16 and TP2 at CONC 2, 4, 8, 16, 32 with huge-page Engram tables and max-num-seqs 2 x CONC (at least 8), and add DEP2 (TP1 x DP2 + EP2, MegaAttention/MegaMoE) at CONC 16, 32, 48, 64, 128 behind consistent-hash vLLM Router 0.1.14. CUDA graphs use FULL_AND_PIECEWISE with DSpark-aligned capture sizes: 2046/2048 at CONC <= 4, 8190/8192 otherwise, and 2046 capture at DEP2 CONC 128. Append the perf-changelog entry and document the GB300 settings. 将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 配方更新到 nightly 36768d1b, 使用 FlashInfer 稀疏 attention、MXFP4 indexer KV、稀疏 logits、fp8 KV, 并关闭 FlashInfer autotune。TP4 覆盖 CONC 1、4、8、16,TP2 覆盖 CONC 2、4、 8、16、32,Engram 表使用透明大页,max-num-seqs 为 2 x CONC(至少 8);新增 DEP2(TP1 x DP2 + EP2,MegaAttention/MegaMoE),覆盖 CONC 16、32、48、64、 128,前置一致性哈希 vLLM Router 0.1.14。CUDA graph 使用 FULL_AND_PIECEWISE 和按 DSpark 对齐的捕获尺寸:CONC <= 4 为 2046/2048,其余为 8190/8192,DEP2 CONC 128 捕获到 2046。追加 perf-changelog 条目,并补充 GB300 配置文档。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
d391a09 to
1a37262
Compare
Update the GB300 DeepSeek-V4.1-Flash vLLM AgentX image to
nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89and retune the parallelism layout: TP4 for low concurrency, TP2 for the middle range and a new DEP2 (TP1 x DP2 + EP2, MegaAttention/MegaMoE) arm behind consistent-hash vLLM Router 0.1.14. All points keep FlashInfer autotuning disabled and use FULL_AND_PIECEWISE CUDA graphs with DSpark-aligned capture sizes from #3396.cpu_offload,use_thpcpu_offload,use_thpcpu_offload,embedding_across_dpValidation: ARM64 image exists; local matrix generation and variant selection passed (14 configurations);
infx/tests/srt_slurmandinfx/tests/matrixpassed. max-num-seqs is sized from measured AgentX in-flight peaks (up to 3 x CONC at CONC 1, about 0.7-1.0 x CONC above CONC 8). Tuning was informed by local GB300 AgentX runs onnightly-dev-arm64-cu130-52cb02dee06e. On this image, a local GB300 DEP2 CONC 64 AgentX run with the PR's engine flags (1800 s, 4,438 requests, 0 errors) reached 87.8 tok/s/user p90 interactivity and 256.0 M tokens/$, versus 84.1 and 262.0 on the earlier image; the TP arms have not been GPU-run on this image yet.AI model disclosure
claude-opus-5-5) in Claude Code; no delegated agents.Related Issue
None. Follows #3514 and #3396.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Pending a final all-green full sweep with evals passing.中文
将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 镜像更新为
nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89,并调整并行布局:低并发用 TP4,中等并发用 TP2,新增 DEP2(TP1 x DP2 + EP2,MegaAttention/MegaMoE),前置 consistent_hash vLLM Router 0.1.14。所有点继续禁用 FlashInfer 自动调优,并沿用 #3396 中按 DSpark 对齐的 FULL_AND_PIECEWISE CUDA graph。cpu_offload、use_thpcpu_offload、use_thpcpu_offload、embedding_across_dp验证:ARM64 镜像已发布;本地矩阵生成与 variant 选择通过(14 个配置);
infx/tests/srt_slurm与infx/tests/matrix通过。max-num-seqs 依据实测的 AgentX 在途请求峰值设置(CONC 1 时最多 3 x CONC,CONC 8 以上约 0.7-1.0 x CONC)。调优依据为本地 GB300 上基于nightly-dev-arm64-cu130-52cb02dee06e的 AgentX 测试。在本镜像上,用本 PR 的引擎参数在本地 GB300 跑了 DEP2 CONC 64 AgentX(1800 s,4,438 个请求,0 个错误),p90 interactivity 为 87.8 tok/s/user、256.0 M tokens/$,旧镜像为 84.1 和 262.0;TP 配置尚未在本镜像上用 GPU 运行。AI 模型披露:使用 Claude Code 中的 Claude Opus 5.5(
claude-opus-5-5),未委派其他 agent;负责起草配方、master 配置、perf-changelog 和文档改动,运行本地 GB300 基准测试与验证,并撰写本说明。关联 issue:无,延续 #3514 和 #3396。改动类型:配置变更、文档更新。文档:在
configuration-procedures.md及其中文版中补充 GB300 的镜像、并发点、CUDA graph 档位、Engram 大页和 max-num-seqs 设置。🤖 Generated with Claude Code