Skip to content

Add DSV4 GB200 Dynamo+SGLang AgentX configs / 添加 DSV4 GB200 Dynamo+SGLang AgentX 配置 - #3182

Open
nvpohanh wants to merge 3 commits into
mainfrom
dsv4-fp4-gb200-dynamo-sglang
Open

nvpohanh wants to merge 3 commits into
mainfrom
dsv4-fp4-gb200-dynamo-sglang

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

[by Codex]

Add GB200 Dynamo+SGLang AgentX configurations for deepseek-ai/DeepSeek-V4-Pro-0813 using its bundled DSpark head and HiCache.

The sweep covers TP8 aggregate concurrency 1 and 4; 1P1D DEP8/DEP16 concurrency 64 and 128; 1P1D DEP16/DEP32 concurrency 256; and 2P1D DEP16/DEP32 concurrency 768, 1024, and 1280. The branch is rebased onto current main, with its GB200 checkpoint and compute-visible workspace behavior migrated to the current Python launcher.

AI model disclosure

  • GPT-5 (Codex): rebased the branch, resolved the launcher and changelog conflicts, validated the recipes and generated matrix, and prepared this PR update.
中文

为 deepseek-ai/DeepSeek-V4-Pro-0813 添加 GB200 Dynamo+SGLang AgentX 配置,使用其内置的 DSpark 草稿头和 HiCache。

本次 sweep 覆盖 TP8 聚合模式并发 1 和 4;1P1D DEP8/DEP16 并发 64 和 128;1P1D DEP16/DEP32 并发 256;以及 2P1D DEP16/DEP32 并发 768、1024 和 1280。分支已 rebase 到当前 main,并将 GB200 检查点及计算节点可见工作区逻辑迁移到当前 Python launcher。

AI 模型披露

  • GPT-5 (Codex):完成分支 rebase、解决 launcher 与 changelog 冲突、验证 recipe 和生成的 matrix,并准备此次 PR 更新。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

nvpohanh added a commit that referenced this pull request Sep 16, 2026
将性能变更日志中的占位链接更新为公开 PR #3182。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good — this is a mechanical addition of a new GB200 Dynamo+SGLang AgentX benchmark family (6 recipe YAMLs + matching master-config entries + changelog), plus a small, additive conditional in launch_gb200-nv.sh that preserves existing behavior for non-agentic/dynamo-vllm dsv4 paths.

What was reviewed: the new IS_AGENTIC/FRAMEWORK branching for MODEL_PATH and uses_watchtower_shared_fs() (additive, existing dsv4/dynamo-vllm behavior unchanged, no defaulted vars); the recipe/master-config/changelog linkage for all 6 new files; decode max-running-requests vs. cuda-graph-max-bs-decode sizing in the disagg c400 recipe (left at 256 despite higher concurrency — flagged as a candidate but ruled out); and the widened shared-fs routing for AgentX dsv4 jobs (consistent with the existing dynamo-vllm precedent, ruled out as a new risk category).

Extended reasoning...

Overview

The PR adds a new GB200 Dynamo+SGLang AgentX recipe family for DSV4-Pro: six new recipe YAMLs (two aggregate TP8, four disaggregated 1P1D/2P1D DEP8/DEP16), two matching nvidia-master.yaml config-key entries, two perf-changelog.yaml append-only entries, and a small change to runners/launch_gb200-nv.sh that (a) points MODEL_PATH at a node-local NVMe checkpoint specifically for MODEL_PREFIX=dsv4, PRECISION=fp4, FRAMEWORK=dynamo-sglang, IS_AGENTIC=1 jobs, and (b) extends uses_watchtower_shared_fs() to route those same jobs onto the compute-visible shared filesystem, matching the existing dynamo-vllm dsv4 precedent.

Security risks

None identified. This is internal benchmark infrastructure config (image tags, model paths, cluster routing flags) with no user input, auth, or crypto surface. Paths and image references are hardcoded, not interpolated from untrusted sources.

Level of scrutiny

Moderate: most of the diff (recipe YAMLs, master-config, changelog) is mechanical and follows established patterns already present elsewhere in the repo (e.g., the existing gb300 dsv4 dynamo-sglang agentic entries). The bash script change is small (18 lines) and additive — it nests new conditionals inside existing elif branches rather than replacing them, so prior behavior for non-agentic and dynamo-vllm dsv4 jobs is preserved, consistent with the repo's no-silent-default-fallback convention. I traced the two candidate concerns already flagged (decode CUDA-graph capture size vs. max-running-requests in the c400 disagg recipe, and the widened shared-fs routing) and confirmed both are either intentional or consistent with pre-existing precedent, which is why they were ruled out.

Other factors

The perf-changelog entries carry placeholder pull/XXX pr-link values, but grepping the file shows this is an established, tolerated pattern already present elsewhere in perf-changelog.yaml (e.g., line 5598) prior to this PR, so it isn't a new deviation worth raising. No CHANGES_REQUESTED reviews or unresolved third-party objections are evident in the timeline, and the bug hunt exited via dry_streak with no findings. Given the additive, pattern-following nature of the change and the narrow, well-scoped bash edits, I have high confidence this does not need further human scrutiny beyond what's already been examined.

This review covers commit 893b871, which is no longer the latest commit on this pull request; later commits are not covered by it.

@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

@kedarpotdar-nv

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

  • PR validation and performance passed in reused sweep run 35278105888 on the current head. All six AgentX points—c1, c4, c64, c128, c400, and c800—completed successfully.
  • The four AgentX eval jobs passed with GSM8K strict-match scores of 97.73% (c64), 97.65% (c128), 97.42% (c400), and 97.57% (c800).
  • The recipes use the bundled DeepSeek-V4-Pro-0813 DSpark head with block size 6 through the chat-completions workload path. Performance jobs automatically received the committed thinking_on golden acceptance length of 3.77 using match-expected and real draft tokens; eval jobs ran without synthetic acceptance.
  • DeepSeek-V4-Pro-0813 remains active for the Agentic scenario in MODELS.md.
  • The submission uses the upstream lmsysorg/sglang image as shipped and contains no engine or serving-stack patches.
  • The official-recipe requirement is not applicable because this PR exclusively adds multi-node recipes.
  • append-only: true is not used.
  • Sweep reuse was authorized in /use 35278105888.

Signed: kedarpotdar-nv

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the TTFT doesn't seem tuned well and it seems like from 20s to 100s e2el there is no points which would be unfair to vllm submissions & other submissions to have no points on 20s to 100s

https://inferencex.semianalysis.com/inference?unofficialRun=35278105888

Image

@github-actions

Copy link
Copy Markdown
Contributor

✅✅✅ Verdict: PASS ✅✅✅

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @kedarpotdar-nv is a named owner of configs/nvidia-master.yaml; the other changed paths (benchmarks/multi_node/**, perf-changelog.yaml, runners/launch_gb200-nv.sh) fall under the * catch-all, which a recognized CODEOWNER satisfies.

✅ Check 1 (Passing sweep on in-PR commit): PASS — on the pinned head 7220982 (PR tip unchanged), run 35278105888 executed all six multi-node agentic / perf jobs (c1, c4, c64, c128, c400, c800) and all four multi-node agentic eval / jobs with conclusion success (none skipped).

✅ Check 2 (Evals pass): PASS — agg_eval_all.json from that run has four GSM8K rows (c64 0.9773, c128 0.9765, c400 0.9742, c800 0.9757 em_strict, n_eff 1319 each), all above the dsv4 floor of 0.91 in infx/evals/thresholds.yaml; the eval jobs ran on lmsysorg/sglang:nightly-dev-20260916-c9a8fba9, the same image as both new master-config entries.

➖ Check 3 (Recipe linked/merged/complete): N/A — disaggregated/multi-node submission (all recipes under benchmarks/multi_node/srt-slurm-recipes/**, both master entries multinode: true); the recipe-link requirement applies to single-node recipes only.

✅ Check 4 (Reuse command): PASS — /use 35278105888 posted as a whole-line comment by kedarpotdar-nv (COLLABORATOR).

✅ Check 5 (Latest checklist template): PASS — every item of the current docs/PR_REVIEW_CHECKLIST.md template (15 items incl. the nested recipe-link item) is present and checked in the sign-off.

✅ Check 6 (Upstream image, engine-first): PASS — both new entries use framework: dynamo-sglang with upstream lmsysorg/sglang:nightly-dev-20260916-c9a8fba9; vLLM coverage for dsv4 on cluster:gb200-nv already exists (dsv4-fp4-gb200-dynamo-vllm-agentic-mtp-agg / -disagg).

✅ Check 7 (No deprecated models/scenarios): PASS — dsv4 Agentic coding is active as of 2026-09-18, and MODELS.md explicitly sanctions the native DSpark head deepseek-ai/DeepSeek-V4-Pro-0813 for Agentic coding.

✅ Check 8 (No architecture hacks): PASS — no --hf-overrides / model-override args; enable-deepseek-v4-fp4-indexer, enable-w4a4-mxfp4-megamoe and moe-runner-backend flashinfer_mxfp4 are precision/kernel selections matching the already-merged GB300 DSV4 SGLang recipes.

✅ Check 9 (Spec-decode via chat template): PASS — DSPARK configs are driven by agentic_srt.sh → build_replay_cmd, which runs aiperf with --endpoint /v1/chat/completions --endpoint-type chat.

✅ Check 10 (No engine patches): PASS — no .patch, sed -i, heredoc rewrites, or engine wheel installs; dynamo: install: true pulls the published Dynamo 1.5.0.dev20260910 router wheel declared under router: in the master config (the same pattern as the existing srt-slurm recipes), and the SGLang image runs as shipped.

✅ Check 11 (Agentic spec-decode golden AL): PASS — recipes carry no hard-coded AL; apply_srt_recipe → infx/srt_slurm/synthetic_acceptance.py selects dsv4-pro-0813-dspark.yaml thinking_on / block size 6 = 3.77, and the run's server logs show SGLANG_SIMULATE_ACC_LEN=3.77 SGLANG_SIMULATE_ACC_METHOD=match-expected SGLANG_SIMULATE_ACC_TOKEN_MODE=real-draft-token on the perf jobs (evals ran with real acceptance).

➖ Check 12 (Append-only): N/A — neither new perf-changelog.yaml entry uses append-only: true.

Assessed commit: 722098267d530f6cfe222d44315e0e88df53533e.

@nvpohanh

Copy link
Copy Markdown
Collaborator Author

I will fix that

@functionstackx

Copy link
Copy Markdown
Collaborator

thanks @nvpohanh ! appreipcate it! <3

nvpohanh added a commit that referenced this pull request Sep 22, 2026
将性能变更日志中的占位链接更新为公开 PR #3182。
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-gb200-dynamo-sglang branch from 09bcd32 to 5163e42 Compare September 22, 2026 03:01
nvpohanh added a commit that referenced this pull request Sep 23, 2026
将性能变更日志中的占位链接更新为公开 PR #3182。
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-gb200-dynamo-sglang branch from 8642cc5 to a0c4ade Compare September 23, 2026 04:24
nvpohanh added a commit that referenced this pull request Sep 23, 2026
将性能变更日志中的占位链接更新为公开 PR #3182。
@nvpohanh
nvpohanh force-pushed the dsv4-fp4-gb200-dynamo-sglang branch from 8d04085 to 956b6bc Compare September 23, 2026 13:50
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

@nvpohanh
nvpohanh force-pushed the dsv4-fp4-gb200-dynamo-sglang branch from 9b39841 to 38ae07c Compare September 30, 2026 08:45
@nvpohanh nvpohanh changed the title Add DSV4 GB200 Dynamo+SGLang AgentX configs Add DSV4 GB200 Dynamo+SGLang AgentX configs / 添加 DSV4 GB200 Dynamo+SGLang AgentX 配置 Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants