Skip to content

Add B300 Dynamo+SGLang AgentX configs for DeepSeek V4 with aggregate DEP8 c384 / 添加含聚合式 DEP8 c384 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 - #3631

Open
nvpohanh wants to merge 9 commits into
mainfrom
dsv4-fp4-b300-dynamo-sglang-agentx-dep8-agg
Open

nvpohanh wants to merge 9 commits into
mainfrom
dsv4-fp4-b300-dynamo-sglang-agentx-dep8-agg

Conversation

@nvpohanh

@nvpohanh nvpohanh commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

[by Claude Code]

Adds DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX recipes. This PR replaces #3190. It carries the same recipes, rebased on current main, with two changes: a one-node aggregate DEP8 c384 point replaces the 2P1D DEP8 c480 point, and no recipe enables the W4A4 MXFP4 Mega-MoE path.

  • Aggregate: TP8 c1, TP4 c4, and DEP8 c384 (new).
  • Disaggregated with Mooncake KV transfer: 1P1D DEP8 c64/c240.
  • enable-w4a4-mxfp4-megamoe is removed from every recipe. Mega-MoE (moe-a2a-backend: megamoe) uses its default FP8xFP4 kernels.
  • Installs only the public mooncake-transfer-engine-efa-cuda13==0.3.13.post1 wheel during disaggregated container setup; no SGLang source patch is applied.
  • Relies on the DSXE-managed fabric injection without explicit /opt/amazon/efa or /opt/amazon/ofi-nccl mounts.
  • Uses the current Python launcher and cluster registry, including persistent AgentX and Hugging Face caches.
  • Serves Dynamo+SGLang DeepSeek-V4-Pro-0813 weights on b300-dsxe from the node-local NVMe copy (DeepSeek-V4-Pro-0813@scratch, as for vLLM).
  • Prefill and decode keep the same DSpark K=6 method.

Aggregate DEP8 c384

Ported from the single-node SGLang DEP8 c384 point (dsv4-fp4-b300-sglang-agentic-hicache-mtp, benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml, override_dep8_c384):

  • Engine settings kept: DP8 attention + EP8 Mega-MoE, FP4 indexer, DP LM head, local control broadcast, prefill delayer with prefill-decode-interval: 20, stream-interval: 20, chunked-prefill-size: 65536 with SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320, max-running-requests: 768 (96 per DP rank, twice concurrency), mem-fraction-static: 0.88, swa-full-tokens-ratio: 0.075, HiCache ratio 3 (write_back, direct, page_first_direct), DSpark K=6.
  • Updated to this PR's conventions: lmsysorg/sglang:nightly-dev-20260916-c9a8fba9, tp-size/dp-size/ep-size, attention-backend: dsv4 (compressed is a deprecated alias in this image), and no SGLANG_ENABLE_UNIFIED_RADIX_TREE (deprecated; the unified radix tree is now the default). Thinking is enabled through SGLANG_DEFAULT_THINKING and SGLANG_DSV4_REASONING_EFFORT, as in the other recipes.
  • Decode graph limit: cuda-graph-max-bs-decode: 96, the per-rank request cap (768 / 8). The single-node recipe used 544, but SGLang clamps captured decode batches to the per-rank request pool, so 544 never captured more than this.
  • Routing: the Dynamo KV router replaces the SGLang router's DP-aware consistent hashing. Sessions stay sticky through X-Dynamo-Session-ID, and each DP rank publishes KV events (kv-events-config), so the router sees per-rank prefixes. One frontend; etcd and NATS run on the worker node. The frontend sets DYN_TCP_REQUEST_TIMEOUT: "60", as the disaggregated recipes do, so the c384 warmup burst does not exceed Dynamo's 5 s request-plane ack timeout.
  • Dropped: settings that only apply to SGLang's own HTTP server and have no effect behind the Dynamo frontend: tokenizer-worker-num, tool-call and reasoning parsers, chat template, skip-server-warmup, SGLANG_TIMEOUT_KEEP_ALIVE.
  • dram-utilization: 0.80 matches the disaggregated config.

Local validation

  • YAML parse and shell syntax.
  • Exact-key matrix generation for both config keys: five points (aggregate TP8 c1, TP4 c4, and DEP8 c384 with DRAM offload on one node; disaggregated 1P1D c64 and c240 on two nodes).
  • validate_perf_changelog against main; the changelog change is append-only.
  • srt-slurm v2.39.1 validate_config_file and srtctl migrate (already current) for all five recipes.
  • Focused tests: infx/tests/launch, infx/tests/clusters, infx/tests/matrix, infx/tests/srt_slurm, and the changelog workflow tests (769 passed); Ruff on models.py.
  • SGLang flag and environment-variable names checked against the image's SGLang revision (c9a8fba9).

AI model disclosure

中文

新增 DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX 配方。本 PR 取代 #3190:包含相同的配方并 rebase 到最新 main,另有两处改动:以单节点聚合式 DEP8 c384 测试点取代 2P1D DEP8 c480 测试点,并且所有配方均不启用 W4A4 MXFP4 Mega-MoE 路径。

  • 聚合式:TP8 c1、TP4 c4,以及新增的 DEP8 c384。
  • 使用 Mooncake KV 传输的分离式:1P1D DEP8 c64/c240。
  • 所有配方移除 enable-w4a4-mxfp4-megamoe,Mega-MoE(moe-a2a-backend: megamoe)使用默认的 FP8xFP4 kernel。
  • 分离式容器初始化时仅安装公开的 mooncake-transfer-engine-efa-cuda13==0.3.13.post1 wheel,不修改 SGLang 源码。
  • 依赖 DSXE 管理的网络栈注入,不显式挂载 /opt/amazon/efa 或 /opt/amazon/ofi-nccl。
  • 使用当前 Python 启动器和集群注册表,并保留持久化 AgentX 与 Hugging Face 缓存。
  • b300-dsxe 上 Dynamo+SGLang 的 DeepSeek-V4-Pro-0813 权重从节点本地 NVMe 副本(DeepSeek-V4-Pro-0813@scratch,与 vLLM 相同)加载。
  • Prefill 与 decode 保持使用相同的 DSpark K=6 方法。

聚合式 DEP8 c384

移植自单节点 SGLang DEP8 c384 测试点(dsv4-fp4-b300-sglang-agentic-hicache-mtp,benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml,override_dep8_c384):

  • 保留的引擎设置: DP8 注意力 + EP8 Mega-MoE、FP4 indexer、DP LM head、local control broadcast、prefill-decode-interval: 20 的 prefill delayer、stream-interval: 20、chunked-prefill-size: 65536 配合 SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320、max-running-requests: 768(每个 DP rank 96,为并发数两倍)、mem-fraction-static: 0.88、swa-full-tokens-ratio: 0.075、HiCache 比例 3(write_back、direct、page_first_direct)、DSpark K=6。
  • 按本 PR 的约定更新: 使用 lmsysorg/sglang:nightly-dev-20260916-c9a8fba9、tp-size/dp-size/ep-size、attention-backend: dsv4(该镜像中 compressed 为已弃用别名),并且不再设置 SGLANG_ENABLE_UNIFIED_RADIX_TREE(已弃用,unified radix tree 现为默认)。与其他配方一样,通过 SGLANG_DEFAULT_THINKING 和 SGLANG_DSV4_REASONING_EFFORT 启用思考模式。
  • Decode CUDA graph 上限: cuda-graph-max-bs-decode: 96,即每个 rank 的请求上限(768 / 8)。单节点配方使用 544,但 SGLang 会把捕获的 decode batch 限制在每个 rank 的请求池以内,因此 544 实际上从未捕获超过该值。
  • 路由: 以 Dynamo KV 路由器取代 SGLang 路由器的 DP 感知一致性哈希。会话通过 X-Dynamo-Session-ID 保持亲和,各 DP rank 发布 KV 事件(kv-events-config),使路由器能看到每个 rank 的前缀。使用单个前端,etcd 与 NATS 运行在 worker 节点上。前端与分离式配方一样设置 DYN_TCP_REQUEST_TIMEOUT: "60",避免 c384 预热突发请求超过 Dynamo 请求平面默认 5 秒的确认超时。
  • 移除: 仅适用于 SGLang 自身 HTTP 服务端、在 Dynamo 前端之后不起作用的设置:tokenizer-worker-num、工具调用与推理解析器、聊天模板、skip-server-warmup、SGLANG_TIMEOUT_KEEP_ALIVE。
  • dram-utilization: 0.80 与分离式配置一致。

本地验证

  • YAML 解析与 Shell 语法检查。
  • 两个配置键的精确配置键矩阵生成:共五个测试点(单节点聚合式 TP8 c1、TP4 c4,以及使用 DRAM 卸载的 DEP8 c384;双节点分离式 1P1D c64 与 c240)。
  • 针对 main 运行 validate_perf_changelog;changelog 改动仅为追加。
  • 对全部五个配方运行 srt-slurm v2.39.1 的 validate_config_file 与 srtctl migrate(已是最新 schema)。
  • 专项测试:infx/tests/launch、infx/tests/clusters、infx/tests/matrix、infx/tests/srt_slurm 及 changelog 工作流测试(769 项通过);对 models.py 运行 Ruff。
  • 已对照镜像的 SGLang 版本(c9a8fba9)核对 SGLang 参数与环境变量名称。

AI 模型披露

🤖 Generated with Claude Code

nvpohanh and others added 6 commits October 1, 2026 01:12
新增基于 Mooncake 的 B300 DeepSeek V4 AgentX 配方,并适配 Python 启动器。
移除配方中与集群级每 GPU CPU 配置冲突的任务级 CPU 参数。
…gentX

The pluggable launcher resolves DeepSeek-V4-Pro-0813 on b300-dsxe to the
shared /data/models copy for every non-vLLM framework, including multi-node
Dynamo+SGLang, which previously read the node-local /scratch/models copy.
With several points loading concurrently from shared storage, the TP8 c1 agg
worker spent ~2.9h in weight loading and lost its etcd lease; the TP4 c4 agg
worker never became healthy within the 4h window on the previous head.

Pin every point of both configs to /scratch/models/DeepSeek-V4-Pro-0813 via
the MODEL_PATH additional-setting, restoring the storage path under which
this recipe set last passed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…oint

A point-level MODEL_PATH is opaque to the launcher, so srtctl's model
preflight (srt-slurm v2.39.1) ran on the runner host and rejected the
node-local /scratch/models path. Instead, add dynamo-sglang to the b300-dsxe
override that already routes vLLM to DeepSeek-V4-Pro-0813@scratch: the
checkpoint is then known to be node-local, model_paths maps the recipe
alias to /scratch/models/DeepSeek-V4-Pro-0813, and preflight is skipped as
it is for vLLM. Single-node SGLang keeps the shared /data/models copy.

Drop the MODEL_PATH additional-settings added in the previous commit.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Replace the disaggregated 2P1D DEP8 c480 point with a one-node
aggregate DEP8 c384 Dynamo+SGLang recipe that follows the single-node
SGLang DEP8 c384 point: attention DP8 + EP8 Mega-MoE, FP4 indexer,
prefill delayer (interval 20), chunked prefill 65536,
max-running-requests 768, cuda-graph-max-bs-decode 544,
mem-fraction-static 0.88, swa-full-tokens-ratio 0.075 and a HiCache
ratio 3 write_back DRAM tier with the page_first_direct layout.

The recipe uses this PR's lmsysorg/sglang:nightly-dev-20260916-c9a8fba9
image and its arg names (tp-size/dp-size/ep-size, attention-backend dsv4
instead of the deprecated compressed alias, no deprecated
SGLANG_ENABLE_UNIFIED_RADIX_TREE). The Dynamo frontend tokenizes and
routes, so SGLang-server-only settings (tokenizer workers, parsers, chat
template, server warmup, keep-alive) are dropped; per-DP-rank KV events
feed the Dynamo KV router.

Drop enable-w4a4-mxfp4-megamoe from every recipe, so Mega-MoE runs its
default FP8xFP4 kernels.

将 B300 分离式 2P1D DEP8 c480 测试点替换为单节点聚合式 DEP8 c384
Dynamo+SGLang 配方,沿用单节点 SGLang DEP8 c384 测试点:注意力 DP8 +
EP8 Mega-MoE、FP4 indexer、prefill delayer(间隔 20)、chunked prefill
65536、max-running-requests 768、cuda-graph-max-bs-decode 544、
mem-fraction-static 0.88、swa-full-tokens-ratio 0.075,以及 HiCache 比例
3、write_back、page_first_direct 布局的 DRAM 层。

配方使用本 PR 的 lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 镜像及其参数
命名(tp-size/dp-size/ep-size、以 attention-backend dsv4 替代已弃用的
compressed 别名、不再设置已弃用的 SGLANG_ENABLE_UNIFIED_RADIX_TREE)。由于
Dynamo 前端负责分词与路由,移除仅适用于 SGLang 服务端的设置(tokenizer
worker、解析器、聊天模板、服务端预热、keep-alive);各 DP rank 的 KV 事件
供 Dynamo KV 路由器使用。

所有配方移除 enable-w4a4-mxfp4-megamoe,Mega-MoE 使用默认的 FP8xFP4 kernel。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

将 B300 Dynamo+SGLang AgentX 变更日志条目的 pr-link 指向 #3631。

Co-Authored-By: Claude Opus 5.5 <[email protected]>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

cpus-per-gpu: 24
salloc-args: [--mem=0]
volumes:
aiperf-cache: {path: /data/home/sa-gha-runner/aiperf-cache}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 The new aiperf-cache volume for b300-dsxe is never mounted into the multi-node jobs this PR adds, so the PR's claimed "persistent AgentX and Hugging Face caches" don't actually persist. SRT_LANES[("b300-dsxe", LaunchPath.SRT_MULTI)] in infx/launch/drivers/srt/lanes.py:70 is SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES, so lane_mounts() (drivers/srt/config.py:187-197) never binds aiperf-cache or hf-hub-cache for this cluster even though the new agentic-coding recipes set AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache. Fix: add mounts=_AGENTIC_CACHES (or equivalent) to the b300-dsxe SRT_MULTI lane so the new volume is actually wired to agentic multi-node jobs. Pre-existing gap in lanes.py, but this PR's 5 new recipes plus the runners.yaml volume addition substantially widen who hits it.

Why this was flagged

Trigger: any of the 5 new dsv4 AgentX recipes launches on b300-dsxe via the SRT_MULTI path with IS_AGENTIC=1 (set because scenario-type is agentic-coding, per infx/matrix/plan.py:131). The job's container env sets AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache (e.g. agg-tp8-c1-mtp.yaml:127-128), but lane_mounts() only mounts volumes listed in SrtLane.mounts, and lanes.py:70 lists none. On the base branch the same gap exists for the one prior qwen3.5 disagg point; this PR adds the aiperf-cache volume definition to runners.yaml (implying it should now work) and 5 more call sites, without updating lanes.py, so the paths stay ephemeral container-local directories instead of host-backed caches, re-downloading/reprocessing data every run with no persistent cache or error to flag it.

Verification: nit. The aiperf-cache volume added at runners.yaml:581 (and pre-existing hf-hub-cache at :583) is never mounted into b300-dsxe multi-node jobs. The b300-dsxe srt-slurm block has no volume-mounts key, and lanes.py:70 is ("b300-dsxe", LaunchPath.SRT_MULTI): SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES. lane_mounts() (config.py:187-194) iterates only lane.mounts.

Set cuda-graph-max-bs-decode to 96, the per-DP-rank request cap
(max-running-requests 768 / dp-size 8). SGLang already clamps captured
decode batch sizes to the per-rank request pool, so 544 never captured
more than this; the explicit value states the intended limit.

将 B300 聚合式 DEP8 c384 配方的 cuda-graph-max-bs-decode 设为 96,即每个
DP rank 的请求上限(max-running-requests 768 / dp-size 8)。SGLang 本就会把
捕获的 decode batch size 限制在每个 rank 的请求池以内,544 实际上从未捕获
超过该值;显式设置可明确预期的上限。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The c384 warmup burst hit Dynamo's 5 s default request-plane ack
timeout while the worker ingested the large AgentX requests; the
frontend then marked the only worker unreachable and returned 500/503
for every warmup request. Set DYN_TCP_REQUEST_TIMEOUT to 60 s, as in
the disaggregated recipes of this PR.

c384 预热突发请求期间,worker 接收大量 AgentX 大请求时触发了 Dynamo 请求平面
默认 5 秒的确认超时;前端随后将唯一的 worker 标记为不可达,所有预热请求均返回
500/503。与本 PR 的分离式配方一致,将 DYN_TCP_REQUEST_TIMEOUT 设为 60 秒。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant