Skip to content

Update GB300 DSV41flash image and tune TP/DEP / 更新 GB300 DSV41flash 镜像并调优 TP/DEP - #3575

Closed
Juntian777 wants to merge 1 commit into
mainfrom
config/dsv41flash-gb300-vllm-af7f9488
Closed

Juntian777 wants to merge 1 commit into
mainfrom
config/dsv41flash-gb300-vllm-af7f9488

Conversation

@Juntian777

@Juntian777 Juntian777 commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Update the GB300 DeepSeek-V4.1-Flash vLLM AgentX image to nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89 and retune the parallelism layout: TP4 for low concurrency, TP2 for the middle range and a new DEP2 (TP1 x DP2 + EP2, MegaAttention/MegaMoE) arm behind consistent-hash vLLM Router 0.1.14. All points keep FlashInfer autotuning disabled and use FULL_AND_PIECEWISE CUDA graphs with DSpark-aligned capture sizes from #3396.

Arm CONC max-num-seqs Graph capture / batched tokens Engram
TP4 1, 4, 8, 16 2 x CONC, at least 8 2046 / 2048 at CONC <= 4, else 8190 / 8192 cpu_offload, use_thp
TP2 2, 4, 8, 16, 32 2 x CONC, at least 8 2046 / 2048 at CONC <= 4, else 8190 / 8192 cpu_offload, use_thp
DEP2 16, 32, 48, 64, 128 CONC per DP rank 8190 / 8192 at CONC <= 64; 2046 / 8192 at CONC 128 cpu_offload, embedding_across_dp

Validation: ARM64 image exists; local matrix generation and variant selection passed (14 configurations); infx/tests/srt_slurm and infx/tests/matrix passed. max-num-seqs is sized from measured AgentX in-flight peaks (up to 3 x CONC at CONC 1, about 0.7-1.0 x CONC above CONC 8). Tuning was informed by local GB300 AgentX runs on nightly-dev-arm64-cu130-52cb02dee06e. On this image, a local GB300 DEP2 CONC 64 AgentX run with the PR's engine flags (1800 s, 4,438 requests, 0 errors) reached 87.8 tok/s/user p90 interactivity and 256.0 M tokens/$, versus 84.1 and 262.0 on the earlier image; the TP arms have not been GPU-run on this image yet.

AI model disclosure

  • Model/version: Claude Opus 5.5 (claude-opus-5-5) in Claude Code; no delegated agents.
  • Role: drafted the recipe, master config, perf-changelog and documentation changes, ran the local GB300 benchmarks and validation, and wrote this description.

Related Issue

None. Follows #3514 and #3396.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Pending a final all-green full sweep with evals passing.
中文

将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 镜像更新为 nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89,并调整并行布局:低并发用 TP4,中等并发用 TP2,新增 DEP2(TP1 x DP2 + EP2,MegaAttention/MegaMoE),前置 consistent_hash vLLM Router 0.1.14。所有点继续禁用 FlashInfer 自动调优,并沿用 #3396 中按 DSpark 对齐的 FULL_AND_PIECEWISE CUDA graph。

配置 CONC max-num-seqs Graph capture / batched tokens Engram
TP4 1, 4, 8, 16 2 x CONC,至少 8 CONC <= 4 为 2046 / 2048,其余 8190 / 8192 cpu_offload、use_thp
TP2 2, 4, 8, 16, 32 2 x CONC,至少 8 CONC <= 4 为 2046 / 2048,其余 8190 / 8192 cpu_offload、use_thp
DEP2 16, 32, 48, 64, 128 每个 DP rank CONC CONC <= 64 为 8190 / 8192;CONC 128 为 2046 / 8192 cpu_offload、embedding_across_dp

验证:ARM64 镜像已发布;本地矩阵生成与 variant 选择通过(14 个配置);infx/tests/srt_slurm 与 infx/tests/matrix 通过。max-num-seqs 依据实测的 AgentX 在途请求峰值设置(CONC 1 时最多 3 x CONC,CONC 8 以上约 0.7-1.0 x CONC)。调优依据为本地 GB300 上基于 nightly-dev-arm64-cu130-52cb02dee06e 的 AgentX 测试。在本镜像上,用本 PR 的引擎参数在本地 GB300 跑了 DEP2 CONC 64 AgentX(1800 s,4,438 个请求,0 个错误),p90 interactivity 为 87.8 tok/s/user、256.0 M tokens/$,旧镜像为 84.1 和 262.0;TP 配置尚未在本镜像上用 GPU 运行。

AI 模型披露:使用 Claude Code 中的 Claude Opus 5.5(claude-opus-5-5),未委派其他 agent;负责起草配方、master 配置、perf-changelog 和文档改动,运行本地 GB300 基准测试与验证,并撰写本说明。

关联 issue:无,延续 #3514 和 #3396。改动类型:配置变更、文档更新。文档:在 configuration-procedures.md 及其中文版中补充 GB300 的镜像、并发点、CUDA graph 档位、Engram 大页和 max-num-seqs 设置。

🤖 Generated with Claude Code

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

Juntian777 added a commit that referenced this pull request Sep 29, 2026
中文:将 perf-changelog 条目链接到 #3575。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@Juntian777 Juntian777 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 29, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the CONC=192 tail of zip_override_dep2 where max-num-seqs is 384 but max-cudagraph-capture-size stays at 2046 — that gap is intentional (batches beyond the captured graph size just fall back to eager/piecewise execution) and matches the author's own CONC>=96 tier in the PR description, so it's not a bug.

Extended reasoning...

The diff bumps the vLLM nightly image and retunes GB300 DSpark benchmark parallelism (TP4/TP2/new DEP2 zip variants) across the srt-slurm recipe, master config search-space rows, and an appended perf-changelog entry; no auth/crypto/permissions surface is touched. I independently confirmed the already-reported exp_name-collision finding by reading infx/matrix/generate.py:993-997, where the single-node exp_name formula only encodes tp/conc/kv_offloading/spec_decoding (not ep or dp-attn), so the new DEP2 row (tp:2, ep:2, dp-attn:true) and the existing plain TP2 row do produce identical exp_name values at CONC 16/32/64, with no dedup check in validation.py. I also checked the CONC=192 max-num-seqs(384) vs max-cudagraph-capture-size(2046) gap in zip_override_dep2 and ruled it out as intentional, matching the author's stated design.

Comment on lines +8398 to +8402
- { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 4, 8, 16], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml }
# TP2 covers the middle range with the same attention and graph tiers.
- { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [2, 4, 8, 16, 32, 64], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml }
# DEP2 is native TP1 x DP2 + EP2, with MegaAttention/MegaMoE.
- { tp: 2, ep: 2, dp-attn: true, kv-offloading: none, spec-decoding: mtp, conc-list: [16, 32, 48, 64, 96, 128, 192], router: { name: vllm-router, version: "0.1.14" }, srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/gb300-fp4-mtp/agentic.yaml }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Adding the DEP2 row (line 8402: tp:2, ep:2, dp-attn:true, conc-list [16,32,48,64,96,128,192]) alongside the existing plain TP2 row (line 8398: tp:2, conc-list [...,16,32,64]) makes matrix generation emit duplicate exp-name values for CONC 16, 32 and 64. The single-node exp_name formula in infx/matrix/generate.py:993 (f"{model_code}tp{tp}conc{conc}...") only encodes tp+conc, not ep/dp-attn, unlike multinode_agentic_exp_name which deliberately encodes ep/dpa to avoid exactly this collision (generate.py:490). Both runs at CONC=16 become 'dsv41flash_tp2_conc16_kvnone_spec-mtp', so infx/results/collect_results.py writes both into the same agg<exp_name>.json, silently overwriting one benchmark's results with the other's. …

Why this was flagged

…Fix: make the single-node exp_name include ep/dp-attn (or router) so any two single-node rows sharing tp+conc get distinct names, covering this and future dp-attn arms added at an existing tp value.

When infx/matrix generate.py builds the agentic matrix for dsv41flash-fp4-gb300-vllm-agentic-dspark, it iterates both search-space rows at nvidia-master.yaml:8398 (plain tp:2) and :8402 (tp:2, ep:2, dp-attn:true, DEP2). For CONC in {16,32,64}, both rows pass through generate.py:993's exp_name formula, which only interpolates tp and conc, producing the identical string 'dsv41flash_tp2_conc16_kvnone_spec-mtp' for two structurally different deployments (pure TP2 vs TP1xDP2+EP2 behind a vllm-router). No uniqueness check runs over the generated matrix (filter_exp_names at generate.py:1357 only checks user-supplied --exp-names, not generated rows). Downstream, infx/results/collect_results.py:8-16 writes agg_{exp_name}.json keyed by this string, so one run's results file silently overwrites the other's. On the base branch this row didn't exist, so no gb300 vLLM AgentX arm ever collided this…

Verification: normal. The new DEP2 row at inferencex-e2e/configs/nvidia-master.yaml:8402 ({tp:2, ep:2, dp-attn:true, kv-offloading:none, spec-decoding:mtp, conc-list:[16,32,48,64,96,128,192]}) shares tp:2/kv-none/mtp with the plain TP2 row at :8400 ({tp:2, conc-list:[2,4,8,16,32,64]}), and their conc-lists overlap at 16, 32, 64. The config sets multinode:false (line 8391), so both rows take the single-node…

@Juntian777 Juntian777 added agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures and removed full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) labels Sep 29, 2026
Juntian777 added a commit that referenced this pull request Sep 29, 2026
中文:将 perf-changelog 条目链接到 #3575。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@Juntian777
Juntian777 force-pushed the config/dsv41flash-gb300-vllm-af7f9488 branch from b1b4461 to 3cfc7db Compare September 29, 2026 17:51
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Move the GB300 DeepSeek-V4.1-Flash vLLM AgentX recipe to nightly
36768d1b with FlashInfer sparse attention, MXFP4 indexer KV, sparse
logits, fp8 KV and FlashInfer autotuning disabled. Run TP4 at CONC 1, 4,
8, 16 and TP2 at CONC 2, 4, 8, 16, 32 with huge-page Engram tables and
max-num-seqs 2 x CONC (at least 8), and add DEP2 (TP1 x DP2 + EP2,
MegaAttention/MegaMoE) at CONC 16, 32, 48, 64, 128 behind consistent-hash
vLLM Router 0.1.14. CUDA graphs use FULL_AND_PIECEWISE with DSpark-aligned
capture sizes: 2046/2048 at CONC <= 4, 8190/8192 otherwise, and 2046
capture at DEP2 CONC 128. Append the perf-changelog entry and document the
GB300 settings.

将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 配方更新到 nightly 36768d1b,
使用 FlashInfer 稀疏 attention、MXFP4 indexer KV、稀疏 logits、fp8 KV,
并关闭 FlashInfer autotune。TP4 覆盖 CONC 1、4、8、16,TP2 覆盖 CONC 2、4、
8、16、32,Engram 表使用透明大页,max-num-seqs 为 2 x CONC(至少 8);新增
DEP2(TP1 x DP2 + EP2,MegaAttention/MegaMoE),覆盖 CONC 16、32、48、64、
128,前置一致性哈希 vLLM Router 0.1.14。CUDA graph 使用 FULL_AND_PIECEWISE
和按 DSpark 对齐的捕获尺寸:CONC <= 4 为 2046/2048,其余为 8190/8192,DEP2
CONC 128 捕获到 2046。追加 perf-changelog 条目,并补充 GB300 配置文档。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants