Add B300 Dynamo+SGLang AgentX configs for DeepSeek V4 with aggregate DEP8 c384 / 添加含聚合式 DEP8 c384 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 - #3631
Conversation
新增基于 Mooncake 的 B300 DeepSeek V4 AgentX 配方,并适配 Python 启动器。
移除配方中与集群级每 GPU CPU 配置冲突的任务级 CPU 参数。
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…gentX The pluggable launcher resolves DeepSeek-V4-Pro-0813 on b300-dsxe to the shared /data/models copy for every non-vLLM framework, including multi-node Dynamo+SGLang, which previously read the node-local /scratch/models copy. With several points loading concurrently from shared storage, the TP8 c1 agg worker spent ~2.9h in weight loading and lost its etcd lease; the TP4 c4 agg worker never became healthy within the 4h window on the previous head. Pin every point of both configs to /scratch/models/DeepSeek-V4-Pro-0813 via the MODEL_PATH additional-setting, restoring the storage path under which this recipe set last passed. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…oint A point-level MODEL_PATH is opaque to the launcher, so srtctl's model preflight (srt-slurm v2.39.1) ran on the runner host and rejected the node-local /scratch/models path. Instead, add dynamo-sglang to the b300-dsxe override that already routes vLLM to DeepSeek-V4-Pro-0813@scratch: the checkpoint is then known to be node-local, model_paths maps the recipe alias to /scratch/models/DeepSeek-V4-Pro-0813, and preflight is skipped as it is for vLLM. Single-node SGLang keeps the shared /data/models copy. Drop the MODEL_PATH additional-settings added in the previous commit. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Replace the disaggregated 2P1D DEP8 c480 point with a one-node aggregate DEP8 c384 Dynamo+SGLang recipe that follows the single-node SGLang DEP8 c384 point: attention DP8 + EP8 Mega-MoE, FP4 indexer, prefill delayer (interval 20), chunked prefill 65536, max-running-requests 768, cuda-graph-max-bs-decode 544, mem-fraction-static 0.88, swa-full-tokens-ratio 0.075 and a HiCache ratio 3 write_back DRAM tier with the page_first_direct layout. The recipe uses this PR's lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 image and its arg names (tp-size/dp-size/ep-size, attention-backend dsv4 instead of the deprecated compressed alias, no deprecated SGLANG_ENABLE_UNIFIED_RADIX_TREE). The Dynamo frontend tokenizes and routes, so SGLang-server-only settings (tokenizer workers, parsers, chat template, server warmup, keep-alive) are dropped; per-DP-rank KV events feed the Dynamo KV router. Drop enable-w4a4-mxfp4-megamoe from every recipe, so Mega-MoE runs its default FP8xFP4 kernels. 将 B300 分离式 2P1D DEP8 c480 测试点替换为单节点聚合式 DEP8 c384 Dynamo+SGLang 配方,沿用单节点 SGLang DEP8 c384 测试点:注意力 DP8 + EP8 Mega-MoE、FP4 indexer、prefill delayer(间隔 20)、chunked prefill 65536、max-running-requests 768、cuda-graph-max-bs-decode 544、 mem-fraction-static 0.88、swa-full-tokens-ratio 0.075,以及 HiCache 比例 3、write_back、page_first_direct 布局的 DRAM 层。 配方使用本 PR 的 lmsysorg/sglang:nightly-dev-20260916-c9a8fba9 镜像及其参数 命名(tp-size/dp-size/ep-size、以 attention-backend dsv4 替代已弃用的 compressed 别名、不再设置已弃用的 SGLANG_ENABLE_UNIFIED_RADIX_TREE)。由于 Dynamo 前端负责分词与路由,移除仅适用于 SGLang 服务端的设置(tokenizer worker、解析器、聊天模板、服务端预热、keep-alive);各 DP rank 的 KV 事件 供 Dynamo KV 路由器使用。 所有配方移除 enable-w4a4-mxfp4-megamoe,Mega-MoE 使用默认的 FP8xFP4 kernel。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 B300 Dynamo+SGLang AgentX 变更日志条目的 pr-link 指向 #3631。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
| cpus-per-gpu: 24 | ||
| salloc-args: [--mem=0] | ||
| volumes: | ||
| aiperf-cache: {path: /data/home/sa-gha-runner/aiperf-cache} |
There was a problem hiding this comment.
🟣 The new aiperf-cache volume for b300-dsxe is never mounted into the multi-node jobs this PR adds, so the PR's claimed "persistent AgentX and Hugging Face caches" don't actually persist. SRT_LANES[("b300-dsxe", LaunchPath.SRT_MULTI)] in infx/launch/drivers/srt/lanes.py:70 is SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES, so lane_mounts() (drivers/srt/config.py:187-197) never binds aiperf-cache or hf-hub-cache for this cluster even though the new agentic-coding recipes set AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache. Fix: add mounts=_AGENTIC_CACHES (or equivalent) to the b300-dsxe SRT_MULTI lane so the new volume is actually wired to agentic multi-node jobs. Pre-existing gap in lanes.py, but this PR's 5 new recipes plus the runners.yaml volume addition substantially widen who hits it.
Why this was flagged
Trigger: any of the 5 new dsv4 AgentX recipes launches on b300-dsxe via the SRT_MULTI path with IS_AGENTIC=1 (set because scenario-type is agentic-coding, per infx/matrix/plan.py:131). The job's container env sets AIPERF_DATASET_MMAP_CACHE_DIR=/aiperf_mmap_cache and HF_HUB_CACHE=/hf_hub_cache (e.g. agg-tp8-c1-mtp.yaml:127-128), but lane_mounts() only mounts volumes listed in SrtLane.mounts, and lanes.py:70 lists none. On the base branch the same gap exists for the one prior qwen3.5 disagg point; this PR adds the aiperf-cache volume definition to runners.yaml (implying it should now work) and 5 more call sites, without updating lanes.py, so the paths stay ephemeral container-local directories instead of host-backed caches, re-downloading/reprocessing data every run with no persistent cache or error to flag it.
Verification: nit. The aiperf-cache volume added at runners.yaml:581 (and pre-existing hf-hub-cache at :583) is never mounted into b300-dsxe multi-node jobs. The b300-dsxe srt-slurm block has no volume-mounts key, and lanes.py:70 is ("b300-dsxe", LaunchPath.SRT_MULTI): SrtLane(frameworks=_DYNAMO) with no mounts=_AGENTIC_CACHES. lane_mounts() (config.py:187-194) iterates only lane.mounts.
Set cuda-graph-max-bs-decode to 96, the per-DP-rank request cap (max-running-requests 768 / dp-size 8). SGLang already clamps captured decode batch sizes to the per-rank request pool, so 544 never captured more than this; the explicit value states the intended limit. 将 B300 聚合式 DEP8 c384 配方的 cuda-graph-max-bs-decode 设为 96,即每个 DP rank 的请求上限(max-running-requests 768 / dp-size 8)。SGLang 本就会把 捕获的 decode batch size 限制在每个 rank 的请求池以内,544 实际上从未捕获 超过该值;显式设置可明确预期的上限。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
The c384 warmup burst hit Dynamo's 5 s default request-plane ack timeout while the worker ingested the large AgentX requests; the frontend then marked the only worker unreachable and returned 500/503 for every warmup request. Set DYN_TCP_REQUEST_TIMEOUT to 60 s, as in the disaggregated recipes of this PR. c384 预热突发请求期间,worker 接收大量 AgentX 大请求时触发了 Dynamo 请求平面 默认 5 秒的确认超时;前端随后将唯一的 worker 标记为不可达,所有预热请求均返回 500/503。与本 PR 的分离式配方一致,将 DYN_TCP_REQUEST_TIMEOUT 设为 60 秒。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36845004131 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36845004131 |
[by Claude Code]
Adds DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX recipes. This PR replaces #3190. It carries the same recipes, rebased on current
main, with two changes: a one-node aggregate DEP8 c384 point replaces the 2P1D DEP8 c480 point, and no recipe enables the W4A4 MXFP4 Mega-MoE path.enable-w4a4-mxfp4-megamoeis removed from every recipe. Mega-MoE (moe-a2a-backend: megamoe) uses its default FP8xFP4 kernels.mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel during disaggregated container setup; no SGLang source patch is applied./opt/amazon/efaor/opt/amazon/ofi-ncclmounts.b300-dsxefrom the node-local NVMe copy (DeepSeek-V4-Pro-0813@scratch, as for vLLM).Aggregate DEP8 c384
Ported from the single-node SGLang DEP8 c384 point (
dsv4-fp4-b300-sglang-agentic-hicache-mtp,benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml,override_dep8_c384):prefill-decode-interval: 20,stream-interval: 20,chunked-prefill-size: 65536withSGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320,max-running-requests: 768(96 per DP rank, twice concurrency),mem-fraction-static: 0.88,swa-full-tokens-ratio: 0.075, HiCache ratio 3 (write_back,direct,page_first_direct), DSpark K=6.lmsysorg/sglang:nightly-dev-20260916-c9a8fba9,tp-size/dp-size/ep-size,attention-backend: dsv4(compressedis a deprecated alias in this image), and noSGLANG_ENABLE_UNIFIED_RADIX_TREE(deprecated; the unified radix tree is now the default). Thinking is enabled throughSGLANG_DEFAULT_THINKINGandSGLANG_DSV4_REASONING_EFFORT, as in the other recipes.cuda-graph-max-bs-decode: 96, the per-rank request cap (768 / 8). The single-node recipe used 544, but SGLang clamps captured decode batches to the per-rank request pool, so 544 never captured more than this.X-Dynamo-Session-ID, and each DP rank publishes KV events (kv-events-config), so the router sees per-rank prefixes. One frontend; etcd and NATS run on the worker node. The frontend setsDYN_TCP_REQUEST_TIMEOUT: "60", as the disaggregated recipes do, so the c384 warmup burst does not exceed Dynamo's 5 s request-plane ack timeout.tokenizer-worker-num, tool-call and reasoning parsers, chat template,skip-server-warmup,SGLANG_TIMEOUT_KEEP_ALIVE.dram-utilization: 0.80matches the disaggregated config.Local validation
validate_perf_changelogagainstmain; the changelog change is append-only.validate_config_fileandsrtctl migrate(already current) for all five recipes.infx/tests/launch,infx/tests/clusters,infx/tests/matrix,infx/tests/srt_slurm, and the changelog workflow tests (769 passed); Ruff onmodels.py.c9a8fba9).AI model disclosure
claude-opus-5-5, via Claude Code): rebased the Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 #3190 commits onto currentmain, replaced the 2P1D point with the aggregate DEP8 c384 recipe, removed the W4A4 MXFP4 Mega-MoE flag, ran the validation, and prepared this PR.中文
新增 DeepSeek-V4-Pro-0813 FP4 B300 Dynamo+SGLang AgentX 配方。本 PR 取代 #3190:包含相同的配方并 rebase 到最新
main,另有两处改动:以单节点聚合式 DEP8 c384 测试点取代 2P1D DEP8 c480 测试点,并且所有配方均不启用 W4A4 MXFP4 Mega-MoE 路径。enable-w4a4-mxfp4-megamoe,Mega-MoE(moe-a2a-backend: megamoe)使用默认的 FP8xFP4 kernel。mooncake-transfer-engine-efa-cuda13==0.3.13.post1wheel,不修改 SGLang 源码。/opt/amazon/efa或/opt/amazon/ofi-nccl。b300-dsxe上 Dynamo+SGLang 的 DeepSeek-V4-Pro-0813 权重从节点本地 NVMe 副本(DeepSeek-V4-Pro-0813@scratch,与 vLLM 相同)加载。聚合式 DEP8 c384
移植自单节点 SGLang DEP8 c384 测试点(
dsv4-fp4-b300-sglang-agentic-hicache-mtp,benchmarks/single_node/srt-slurm-recipes/dsv4/sglang/b300-fp4-mtp/agentic.yaml,override_dep8_c384):prefill-decode-interval: 20的 prefill delayer、stream-interval: 20、chunked-prefill-size: 65536配合SGLANG_OPT_DEEPGEMM_MEGA_MOE_NUM_MAX_TOKENS_PER_RANK=8320、max-running-requests: 768(每个 DP rank 96,为并发数两倍)、mem-fraction-static: 0.88、swa-full-tokens-ratio: 0.075、HiCache 比例 3(write_back、direct、page_first_direct)、DSpark K=6。lmsysorg/sglang:nightly-dev-20260916-c9a8fba9、tp-size/dp-size/ep-size、attention-backend: dsv4(该镜像中compressed为已弃用别名),并且不再设置SGLANG_ENABLE_UNIFIED_RADIX_TREE(已弃用,unified radix tree 现为默认)。与其他配方一样,通过SGLANG_DEFAULT_THINKING和SGLANG_DSV4_REASONING_EFFORT启用思考模式。cuda-graph-max-bs-decode: 96,即每个 rank 的请求上限(768 / 8)。单节点配方使用 544,但 SGLang 会把捕获的 decode batch 限制在每个 rank 的请求池以内,因此 544 实际上从未捕获超过该值。X-Dynamo-Session-ID保持亲和,各 DP rank 发布 KV 事件(kv-events-config),使路由器能看到每个 rank 的前缀。使用单个前端,etcd 与 NATS 运行在 worker 节点上。前端与分离式配方一样设置DYN_TCP_REQUEST_TIMEOUT: "60",避免 c384 预热突发请求超过 Dynamo 请求平面默认 5 秒的确认超时。tokenizer-worker-num、工具调用与推理解析器、聊天模板、skip-server-warmup、SGLANG_TIMEOUT_KEEP_ALIVE。dram-utilization: 0.80与分离式配置一致。本地验证
main运行validate_perf_changelog;changelog 改动仅为追加。validate_config_file与srtctl migrate(已是最新 schema)。infx/tests/launch、infx/tests/clusters、infx/tests/matrix、infx/tests/srt_slurm及 changelog 工作流测试(769 项通过);对models.py运行 Ruff。c9a8fba9)核对 SGLang 参数与环境变量名称。AI 模型披露
claude-opus-5-5,通过 Claude Code):将 Add Mooncake B300 Dynamo+SGLang AgentX configs for DeepSeek V4 / 添加基于 Mooncake 的 DeepSeek V4 B300 Dynamo+SGLang AgentX 配置 #3190 的提交 rebase 到最新main,以聚合式 DEP8 c384 配方取代 2P1D 测试点,移除 W4A4 MXFP4 Mega-MoE 参数,完成验证并准备本 PR。🤖 Generated with Claude Code