[NV] Move GB300 DeepSeek-V4-Pro AgentX disagg to the Mooncake external linker / 将 GB300 DeepSeek-V4-Pro AgentX 分离式配置切换到 Mooncake external linker - #3187
Conversation
8d51100 to
50dce57
Compare
50dce57 to
d77ecff
Compare
| - "Prefill 改为按 2048 token 切分单个请求并延迟中间 chunk 的 KV 传输,mem-fraction-static 从 0.85 提升到 0.92,镜像升级到 lmsysorg/sglang:nightly-dev-cu13-20260908-20ca564b。" | ||
| - "在相同 golden acceptance length 下相比原配置实测:1P1D、2P1D、4P1D 的单卡总吞吐分别提升 13.2%、18.4%、17.5%,平均 TTFT 分别下降 21.6%、30.7%、37.3%,三个点的稳态前缀命中率均为 96.0%-96.2%。" | ||
| - "临时方案:服务端改动仍在 sgl-project/sglang#39694 评审中,因此 runner 将该分支克隆到 srt-slurm 的 configs 目录,服务进程通过 PYTHONPATH 导入,同时继续使用容器内编译好的 kernel。配套 setup 脚本在该源码树缺失或版本不符时直接让任务失败,避免静默回退到容器自带的 SGLang。待改动进入固定镜像后应移除克隆、setup 脚本和 recipe 中的 PYTHONPATH。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER |
There was a problem hiding this comment.
🔴 New changelog entry's pr-link uses .../pull/PLACEHOLDER, which the changelog validator rejects for PR runs, so CI will fail and block this PR from merging. validate_added_pr_link in infx/workflows/validate_perf_changelog.py only accepts the exact pull/{pr_number} link or the literal placeholders "XXX" / "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX" (PR_LINK_PLACEHOLDERS, lines 24-27); PLACEHOLDER matches neither. Fix: replace the pr-link value with XXX (or the real PR number once known) so validate_added_pr_link accepts it.
Extended reasoning...
docs/configuration-procedures.md says the model+hardware playbook permits pr-link: TBD before the PR exists, but the actual validator infx/workflows/validate_perf_changelog.py never allows 'TBD' or 'PLACEHOLDER' — it only allows 'XXX' or the exact pull/ URL (PR_LINK_PLACEHOLDERS set, lines 24-27, checked in validate_added_pr_link lines 144-159). This diff's new entry at perf-changelog.yaml:7589 sets pr-link to 'https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER'. When validate_perf_changelog.py runs in CI with pr_number set (a PR run), it computes expected='https://github.com/SemiAnalysisAI/InferenceX/pull/{pr_number}', sees link not in PR_LINK_PLACEHOLDERS and link != expected, and raises ChangelogValidationError('new PR entry must use ... or an XXX placeholder; found ...'). This blocks the PR's changelog check from passing until a maintainer or the author manually edits the line to 'XXX' or the true PR number.
Verification: normal. The new changelog entry at perf-changelog.yaml:7589 sets pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/PLACEHOLDER. The validator infx/workflows/validate_perf_changelog.py rejects this in every mode: (a) on PR runs (pr_number set), validate_added_pr_link (lines 144-160) requires the link be in PR_LINK_PLACEHOLDERS = {"XXX",… | normal. The new changelog entry sets…
| @@ -6664,7 +9063,7 @@ dsv4-fp4-gb300-dynamo-sglang-agentic-disagg: | |||
| - spec-decoding: draft_model | |||
| conc-list: [480] | |||
| kv-offloading: dram | |||
| kv-offload-backend: { name: hicache } | |||
| kv-offload-backend: { name: mooncake } | |||
There was a problem hiding this comment.
🟡 (optional) Operators reading nvidia-master.yaml lose the Mooncake component version that CONFIGS.md requires, making sweep provenance for this backend unrecoverable if the pinned Mooncake build ever needs to be identified after the fact. All four new entries set kv-offload-backend: { name: mooncake } with no version, but configs/CONFIGS.md:136-137 explicitly names Mooncake as an example of an independently-versioned backend that should supply version (unlike framework-native vLLM/HiCache). The Pydantic validator (KVOffloadBackendMetadata in infx/matrix/validation.py:127-132) doesn't enforce this, so nothing blocks the omission. …
Extended reasoning...
…Fix: add a non-image version (e.g. the Mooncake release/commit) for the mooncake kv-offload-backend entries, covering all four sites (configs/nvidia-master.yaml:9066,9084,9102,9120), or update CONFIGS.md if this backend is now treated as framework-native.
configs/CONFIGS.md:133-138 says kv-offload-backend requires non-empty name, version optional only for framework-native backends (vLLM built-in, SGLang HiCache), and explicitly lists Mooncake as needing version since it's independently versioned. The new entries at lines 9066, 9084, 9102, 9120 all set { name: mooncake } with no version key. infx/matrix/validation.py's KVOffloadBackendMetadata makes version fully Optional with no name-based branching (confirmed by reading lines 127-143, and by utils/matrix_logic/test_validation.py tests that never require version for any specific backend name). So CI will not fail, but the documented provenance convention is silently violated: nobody can tell from the master config which Mooncake build/commit these Pareto points were validated against, unlike the equivalent LMCache examples in…
Verification: nit. The deviation is real but the consequence is mild/overstated. CONFIGS.md:135-138 states version is optional only for framework-native backends and explicitly directs "Supply version for independently versioned backends such as LMCache or Mooncake." All four new search-space entries in configs/nvidia-master.yaml (lines 9066, 9084, 9102, 9120) set kv-offload-backend: { name: mooncake }…
| SGLANG_MOONCAKE_OPT_URL="https://github.com/weireweire/sglang.git" | ||
| SGLANG_MOONCAKE_OPT_PIN="d2cf19e69fe7f9d69a7504012c3617fe6669acbd" | ||
| git init configs/sglang-mooncake-opt | ||
| git -C configs/sglang-mooncake-opt fetch --depth 1 \ | ||
| "$SGLANG_MOONCAKE_OPT_URL" "$SGLANG_MOONCAKE_OPT_PIN" | ||
| git -C configs/sglang-mooncake-opt checkout FETCH_HEAD | ||
| cp "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/configs/dsv4-gb300-sglang-mooncake-opt.sh" \ | ||
| configs/dsv4-gb300-sglang-mooncake-opt.sh |
There was a problem hiding this comment.
🟡 (optional) Every gb300 dsv4 dynamo-sglang agentic launch now clones an external, unmerged fork, not just the 4 Mooncake recipes. The elif at runners/launch_gb300-nv.sh:302 copies the whole sglang/deepseek-v4/agentic directory, which also contains agg-gb300-tp4-mtp-lowlatency.yaml and agg-gb300-tp8-mtp-lowlatency.yaml (neither uses hicache/mooncake/setup_script). The new git init/fetch/checkout block at lines 323-330 runs unconditionally in that branch, and set -exo pipefail (line 5) aborts the whole launch if the fetch fails. Fix: gate the clone on the specific recipe(s) that set setup_script: dsv4-gb300-sglang-mooncake-opt.sh (e.g. check MODEL_SUFFIX/recipe filename or grep the copied recipe for that key) so unrelated gb300 dsv4 launches don't depend on github.com/weireweire/sglang.git availability.
Extended reasoning...
runners/launch_gb300-nv.sh:302 elif matches IS_AGENTIC=1, FRAMEWORK=dynamo-sglang, MODEL_PREFIX=dsv4 for ANY recipe under sglang/deepseek-v4/agentic, including the two lowlatency aggregated recipes that keep no hicache/mooncake settings at all. Lines 325-328 unconditionally run git init, git fetch --depth 1 against https://github.com/weireweire/sglang.git pinned to commit d2cf19e69fe7f9d69a7504012c3617fe6669acbd, and git checkout FETCH_HEAD. Script header has set -exo pipefail (line 5), so if that fetch fails — fork renamed/deleted, GH rate limit, network blip, or the author force-pushes over the pinned commit — the whole launch script exits nonzero and the job never gets submitted, even for the lowlatency recipes that never read configs/sglang-mooncake-opt or PYTHONPATH. Before this diff, launching agg-gb300-tp4-mtp-lowlatency.yaml or agg-gb300-tp8-mtp-lowlatency.yaml never touched any third-party GitHub fork; after merging, both silently gain a hard dependency on an unrelated contributor's personal repo staying reachable and unchanged at that SHA.
Verification: Severity: nit. The mechanism is real and reachable; the launch abort is conditional on the external clone failing. runners/launch_gb300-nv.sh:302 elif matches on coarse env vars (IS_AGENTIC==1, FRAMEWORK==dynamo-sglang, MODEL_PREFIX==dsv4), not per-recipe, so it handles EVERY dsv4/dynamo-sglang/agentic launch. Lines 325-328 run git init configs/sglang-mooncake-opt, `git fetch --depth 1… | nit.…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=35494281738 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=35494281738 |
同步 main,使用字节保留流程恢复 PR #3187 的 perf-changelog 追加项,并记录 CUDA graph 参数修复。
同步 main,使用字节保留流程恢复 PR #3187 的 perf-changelog 追加项,并记录 CUDA graph 参数修复。
ab576df to
9fcb8a6
Compare
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
…linker Port of the change onto the native srt-slurm variants layout. The four per-point recipes it used to edit are now one disagg-variants.yaml, so the settings shared by every point live in base and only the per-point values stay in each override. - base: replace hierarchical cache with the Mooncake unified-cache external linker, add the Mooncake store environment, 2048-token per-request prefill chunks with deferred intermediate KV transfer, sparse SWA retention every 32768 tokens, and --mooncake-store-contributor on decode; bump the image to nightly-dev-cu13-20260916-c9a8fba9 and add the mooncake-master service. - overrides: rename cuda-graph-max-bs to cuda-graph-max-bs-decode. The pinned SGLang branch rejects the legacy key as ambiguous, and every override sets it explicitly, so the rename cannot live in base. - TEMPORARY: the runner clones the reviewed SGLang branch into /configs and a setup script aborts the job if that tree is missing or the wrong revision; a sitecustomize shim restores ServerArgs.get_model_config for the pinned Dynamo wheel. Remove all three once the change ships in the pinned image. - The runner passes this cluster's RDMA devices to the Mooncake store. Each variant resolves to exactly the configuration the previous per-point recipe produced, apart from main's move of the benchmark script to benchmarks/srt_agentic.sh.
The block that clones the reviewed SGLang branch and forwards the Mooncake RDMA device list matched every dsv4 dynamo-sglang AgentX run, including the aggregated recipe, which has no prefill/decode roles for the device list to target and never imports the reviewed branch. Key it on the disaggregated variants file instead.
7565d48 to
6317a45
Compare
|
rabased, glad we have utilized the override feature in srt-slurm. We may also able to utilize zip_override in the future |
Summary / 概要
Move the GB300 DeepSeek-V4-Pro Dynamo-SGLang AgentX disaggregated Pareto ladder (1P1D c480, 2P1D c960, 3P1D c1440, 4P1D c1920) from hierarchical cache to the Mooncake unified-cache external linker, and pick up the matching prefill scheduling changes.
将 GB300 DeepSeek-V4-Pro Dynamo-SGLang AgentX 分离式 Pareto 配置(1P1D c480、2P1D c960、3P1D c1440、4P1D c1920)从 hierarchical cache 切换到 Mooncake unified-cache external linker,并同步相应的 prefill 调度改动。
Serve the host KV tier through the unified-cache external linker on Mooncake, with group semantics enabled.
Let decode ranks contribute host DRAM to the store, and retain only a sparse SWA checkpoint per 32768-token interval, so the same DRAM holds more prefixes at the same eviction watermark.
Split a prefill request into 2048-token chunks and defer intermediate chunk KV transfer.
Raise prefill mem-fraction-static from 0.85 to 0.92.
Pin the image to lmsysorg/sglang:nightly-dev-cu13-20260908-20ca564b and record the offload backend as mooncake in the master config.
Use the split CUDA-graph batch-size flag, since the source this configuration runs against no longer carries the deprecated alias.
主机侧 KV 层改由 Mooncake 上的 unified-cache external linker 提供,并启用 group semantics。
decode rank 也贡献主机 DRAM 容量,并按 32768 token 间隔只保留稀疏 SWA checkpoint,使同样的 DRAM 在相同驱逐水位下容纳更多前缀。
prefill 按 2048 token 切分单个请求,并延迟中间 chunk 的 KV 传输。
prefill mem-fraction-static 从 0.85 提升到 0.92。
镜像固定为 lmsysorg/sglang:nightly-dev-cu13-20260908-20ca564b,master config 中的 offload backend 记为 mooncake。
改用拆分后的 CUDA graph batch size 参数,因为该配置所依赖的源码已移除旧别名。
Temporary source override / 临时源码覆盖
The serving change is still in review as sgl-project/sglang#39694 and is not in a released image. Until it ships, the gb300-nv runner clones that branch (pinned by commit) into the srt-slurm configs directory and the servers import it through PYTHONPATH, while still using the container's compiled kernels. A small compatibility shim restores the ServerArgs accessor that the pinned Dynamo wheel still calls. A setup script runs before every server and aborts the job when either piece is missing or is the wrong revision, so a run cannot silently fall back to the container's SGLang.
This is explicitly temporary. The clone block in the runner, the setup script, the compatibility shim, and the recipes' setup_script/PYTHONPATH keys all come out in one commit once the change is in the pinned image.
服务端改动仍在 sgl-project/sglang#39694 评审中,尚未进入发布镜像。在此之前,gb300-nv runner 会将该分支(按 commit 固定)克隆到 srt-slurm 的 configs 目录,服务进程通过 PYTHONPATH 导入,同时继续使用容器内编译好的 kernel。另有一个兼容补丁,用于恢复固定版本 Dynamo wheel 仍在调用的 ServerArgs 访问方法。setup 脚本在每个服务进程启动前运行,当任一部分缺失或版本不符时直接让任务失败,避免静默回退到容器自带的 SGLang。
这是明确的临时方案:改动进入固定镜像后,runner 中的克隆逻辑、setup 脚本、兼容补丁以及 recipe 中的 setup_script/PYTHONPATH 会在同一个提交中移除。
Validation / 验证
该配置已在 GB300 上于 1P1D、2P1D、4P1D 三个点完整跑通,拓扑、负载、3600 秒测量时长、标准的每 lane 十次 AgentX warmup,以及 golden synthetic acceptance length 均与当前配置一致,三个点都跑满测量且请求错误为零;稳态前缀命中率与该负载的理论上限相差约一个百分点以内。utils/matrix_logic 331 项测试通过,recipe YAML 解析及 runner、setup 脚本的 bash 语法检查均通过。c1440 未在本地跑过,其配置改动与其他点完全一致,由本 PR 的 sweep 覆盖。