Skip to content

Update GB200 DeepSeek V4.1 Flash vLLM nightly / 更新 GB200 DeepSeek V4.1 Flash vLLM nightly - #3395

Open
xinli-sw wants to merge 1 commit into
mainfrom
config/dsv41flash-gb200-vllm-nightly
Open

xinli-sw wants to merge 1 commit into
mainfrom
config/dsv41flash-gb200-vllm-nightly

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Pin GB200 DeepSeek-V4.1-Flash vLLM to nightly ddd6fbca148a867aad1fcab7ec72f582b9977db4 on the native srt-slurm AgentX recipe. Use FlashInfer sparse attention at TP2 and TP4 with MXFP4 indexer KV, sparse logits, fp8 KV cache, and Engram CPU offload. The engine readiness timeout is 7200 s. Topology and concurrency remain unchanged.

中文

在原生 srt-slurm AgentX 配方中,将 GB200 DeepSeek-V4.1-Flash vLLM 固定到 nightly ddd6fbca148a867aad1fcab7ec72f582b9977db4。TP2 和 TP4 均使用 FlashInfer 稀疏注意力,并启用 MXFP4 索引器 KV、稀疏 logits、fp8 KV 缓存及 Engram CPU 卸载。 引擎就绪超时为 7200 秒。 拓扑与并发设置不变。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@xinli-sw
xinli-sw marked this pull request as ready for review September 23, 2026 17:53
@xinli-sw
xinli-sw requested a review from a team September 23, 2026 17:53
@claude

claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @xinli-sw's task in 3m 48s —— View job


Review of PR #3395

  • Gather context (diff, AGENTS.md, touched files)
  • Validate master config, perf-changelog, and image accessibility
  • Review benchmark script and launcher changes
  • Post review

Summary: One blocking issue, posted inline. The new ATTENTION_ARGS block in the shared dsv41flash_fp4_vllm_mtp.sh is gated on GPU memory, so it also activates on the B200 and GB300 vLLM arms, which are symlinked to the same script but stay pinned to older nightlies (cd10ed6f, af1c0149) and are not listed in the perf-changelog entry. That is an undocumented recipe change on two other configs and may fail engine start if those nightlies predate the indexer_sparse_logits option or the DSV41 backend names. Suggested fix is to gate on a launcher-exported variable from runners/launch_gb200-nv.sh instead of GPU memory.

Everything else checks out: the image is a public Docker Hub tag, the changelog entry is appended at the tail, --use-chat-template flows through the shared replay helper, and VLLM_ENGINE_READY_TIMEOUT_S=7200 in the launcher correctly overrides the 3600 s runtime default (the workflow sources runtime_settings.sh before the launcher). Two notes: I could not verify the older vLLM SHAs against upstream because external access and the vLLM source MCP server were unavailable from this runner, and the PR body is missing the AI model disclosure section required by AGENTS.md.

中文

摘要:发现一个阻塞问题,已内联评论。共享脚本 dsv41flash_fp4_vllm_mtp.sh 中新增的 ATTENTION_ARGS 块以显存大小作为条件,因此也会在 B200 和 GB300 的 vLLM 配置上生效;这两个配置通过符号链接使用同一脚本,但仍固定在较旧的 nightly(cd10ed6f、af1c0149),且未列入 perf-changelog 条目。这属于对另外两个配置的未记录改动,若旧 nightly 不支持 indexer_sparse_logits 或 DSV41 后端名称,引擎启动会失败。建议改为由 runners/launch_gb200-nv.sh 导出变量来控制,而非依据显存判断。

其余检查均通过:镜像为公开 Docker Hub 标签,changelog 条目追加在文件末尾,--use-chat-template 由共享回放函数传入,启动脚本中的 VLLM_ENGINE_READY_TIMEOUT_S=7200 能正确覆盖 3600 秒的默认值。另注:本次无法从 runner 访问外部资源核实旧版 vLLM SHA;PR 描述缺少 AGENTS.md 要求的 AI model disclosure 部分。

· [Branch](https://github.com/SemiAnalysisAI/InferenceX/tree/config/dsv41flash-gb200-vllm-nightly)

Comment thread benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh Outdated
Comment thread benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh Outdated
@xinli-sw xinli-sw added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 23, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@xinli-sw
xinli-sw marked this pull request as draft September 24, 2026 15:49
@xinli-sw
xinli-sw marked this pull request as ready for review September 24, 2026 22:20

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline findings, I also checked whether doubling VLLM_ENGINE_READY_TIMEOUT_S to 7200s in runners/launch_gb200-nv.sh:113 could cost a sweep its Slurm wall-clock instead of a clean readiness-timeout error — the default SALLOC_TIME_LIMIT for these single-node sweeps is 480 minutes (8h, from the benchmark workflow), comfortably longer than the new 2h engine-ready timeout, so this isn't a practical risk.

Extended reasoning...

Confirmed via runners/runtime_settings.sh and .github/workflows/.yml that SALLOC_TIME_LIMIT for these sweeps defaults to 480 minutes, well above the new 7200s (120 min) VLLM_ENGINE_READY_TIMEOUT_S, so the candidate wall-clock-loss concern is not realistic under current defaults. Separately verified via git diff and runners/.sh that the new DSV41_BLACKWELL_ATTENTION gate is exported only by launch_gb200-nv.sh (not by launch_b200-nscale-slurm.sh, launch_gb300-nv.sh, or launch_h200-dgxc-slurm.sh), so the attention-config args no longer key off GPU_MEM_MIB alone.

Comment thread perf-changelog.yaml
Comment on lines +8850 to +8852

- config-keys:
- dsv41flash-fp4-b200-vllm-agentic-dspark

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) This entry's config-keys list only dsv41flash-fp4-b200-vllm-agentic-dspark and dsv41flash-fp4-gb200-vllm-agentic-dspark, but the fixed script (benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh) is also symlinked from dsv41flash_fp4_gb300_vllm_mtp.sh and dsv41flash_fp4_h200_vllm_mtp.sh, so dsv41flash-fp4-gb300-vllm-agentic-dspark and dsv41flash-fp4-h200-vllm-agentic-dspark get the same SIGPIPE fix but aren't listed. Anyone auditing perf-changelog.yaml by config-key for GB300 or H200 will miss this fix's context even though it changed their behavior too. Fix: include every config that resolves to the shared script (b200, gb200, gb300, h200) in config-keys, not just the two SKUs named in the PR title.

Why this was flagged

configs/nvidia-master.yaml:8210 (dsv41flash-fp4-gb300-vllm-agentic-dspark) and configs/nvidia-master.yaml:8239 (dsv41flash-fp4-h200-vllm-agentic-dspark) both point at benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh via symlink, the same script whose GPU_MEM_MIB probe is fixed by this diff. perf-changelog.yaml:8851-8852 lists only b200 and gb200 in config-keys for that fix entry (perf-changelog.yaml:8850-8858). AGENTS.md requires perf-changelog entries to reference the affected config keys as the source of truth for what changed. An engineer or tool filtering perf-changelog.yaml by dsv41flash-fp4-gb300-vllm-agentic-dspark or dsv41flash-fp4-h200-vllm-agentic-dspark would not find this SIGPIPE fix despite it changing their runs' behavior, unlike on base where the entry didn't exist for anyone.

Verification: nit. The SIGPIPE fix lives in the shared script benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh (diff replaces nvidia-smi ... | head -n1 | tr -d ' ' with GPU_MEM_MIB=$(nvidia-smi ...) + %%$'\n'* + numeric validation). ls -la confirms four vLLM symlinks point at this one file: dsv41flash_fp4_b200_vllm_mtp.sh, gb200, gb300, and h200 all -> dsv41flash_fp4_vllm_mtp.sh. All…

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch 2 times, most recently from e8c965a to e08edaa Compare September 25, 2026 02:03
@functionstackx

Copy link
Copy Markdown
Collaborator

InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest main into this PR: we have migrated single-node AgentX onto native srt-slurm (#3428), so AgentX configs are now declarative YAML recipes, not per-config 1000+ line bash slop scripts. Please also delete the old benchmarks/single_node/** scripts (see this recipe for the new format).

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch 2 times, most recently from 6b659e9 to 0485f33 Compare September 27, 2026 22:13
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch from 0485f33 to 4d49128 Compare September 28, 2026 12:40
@xinli-sw xinli-sw added full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures and removed full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) labels Sep 30, 2026
@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch 2 times, most recently from f42289c to 5a0af7e Compare October 1, 2026 02:18

@wzhao18 wzhao18 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Assessed commit: 5a0af7e0ef64fe17fd5a9559634417e16b2e9c42.
  • PR validation: CI and full sweep, attempt 2 passed on the assessed commit. The completed sweep covers all 16 TP2/TP4 performance jobs at concurrency 1–128. The latest collected artifact contains 16 usable throughput/latency measurements plus one stale empty TP4 c128 record; that empty record is excluded, and the successful TP4 c128 rerun is used. Power data is valid for all eight TP4 measurements; all eight TP2 records report power_valid: 0, so no validated TP2 power claim is made.
  • Evals: the same sweep and image passed GSM8K at TP4 c128, with em_strict = 0.9727065959059894, n_eff = 1319, and infrastructure_success: true. Evidence: eval_results_all. This is one evaluated configuration, not an eval at every performance point.
  • Reuse: /use 36805208558 was posted by CODEOWNER wzhao18 on this PR to select the passing sweep for publication.
  • Recipe documentation: published DeepSeek-V4.1-Flash recipe, with the Blackwell sparse-attention settings documented in vllm-project/recipes#995, merged September 20, and GB200 TP2 Engram offload in vllm-project/recipes#1002, merged September 21.
  • Image and serving stack: upstream vllm/vllm-openai:nightly-ac68c3087215e0a4f3cdfa218508c6aada57235d, matching the recipe and master config. Performance and eval logs report vLLM 0.30.1rc1.dev396+gac68c3087. No engine patch or replacement wheel is introduced.
  • Draft precision: inspected deepseek-ai/DeepSeek-V4.1-Flash@2cba9e42aa026125f3ed06c6d98c1db82f7ca027. Safetensors headers contain 2,401 embedded mtp.* tensors across shards 44–46: packed MXFP4 expert weights, FP8 dense weights and E8M0 scales, plus BF16/F32 auxiliary tensors. The pinned upstream DSpark loader loads these embedded weights through the checkpoint’s default quantization handling. The effective recipe adds no draft checkpoint substitution, dtype override, or quantization override. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent and inapplicable to this vLLM submission.
  • AgentX acceptance and chat path: throughput uses five-token probabilistic DSpark drafting with synthetic AL 3.51, matching thinking_on K5 in the committed golden curve. Runtime artifacts confirm synthetic acceptance; eval uses real block rejection without synthetic acceptance. Replay uses /v1/chat/completions.
  • Pareto evidence: run 36805208558, attempt 2, assessed SHA above, results_bmk, AgentX P90 E2EL versus total token throughput/GPU, one GB200 vLLM FP4 visual curve and one image. There are 16 usable measured throughput/latency points and 7 frontier points: TP4 c8/c16/c32/c128 and TP2 c8/c16/c32. The stale empty TP4 c128 record is excluded; power validity is not used as a throughput/E2EL filter, and the invalid TP2 power records are not presented as validated power measurements. The repository’s pareto_coverage helper reports PASS.
  • append-only is N/A. This PR changes existing recipes and the image and uses a regular changelog entry. All historical changelog bytes are preserved, with the new entry appended at the tail.

Signed: wzhao18

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@wzhao18 staged run 36805208558: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-01~r36805208558

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

@SemiAnalysisAI SemiAnalysisAI deleted a comment from wzhao18 Oct 2, 2026
@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch 3 times, most recently from 94a4b92 to 9100452 Compare October 2, 2026 20:56
将 GB200 的 vLLM 镜像和服务参数与已通过的 B200 配置对齐。
@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb200-vllm-nightly branch from 9100452 to 5f0511d Compare October 3, 2026 13:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants