Conversation
Remove enable-w4a4-mxfp4-megamoe from every prefill and decode worker in the six GB200 Dynamo+SGLang AgentX disaggregated recipes. The MegaMoE all-to-all backend and FP4 indexer settings are unchanged. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked the PR's core claim — that enable-w4a4-mxfp4-megamoe was removed from all prefill/decode workers in the six disaggregated recipes while moe-a2a-backend: megamoe stays — and confirmed no enable-w4a4-mxfp4-megamoe key remains anywhere in the new recipe files, and container.image matches the master-config image field consistently across all eight new scenarios.
Extended reasoning...
The diff adds 7 new srt-slurm recipe YAMLs plus matching master-config entries, a runners.yaml model-path mapping, a lanes.py policy-table entry, and an append-only changelog addition for a new DeepSeek-V4-Pro-0813 GB200 FP4 AgentX workload; no auth/crypto/permission code is touched, only benchmark config and launch policy tables. A confirmed finding already flags a real routing problem in runners.yaml (new model lacks a numa1-local alias, so it is not visible for node-local staging as the changelog implies), which is sufficient reason a human should look before merge.
| deepseek-r1-0528: {root: lustre-models, dir: deepseek-r1-0528} | ||
| deepseek-r1-0528-fp4-v2: {root: lustre-models, dir: deepseek-r1-0528-fp4-v2} | ||
| DeepSeek-V4-Pro: {root: lustre-models, dir: DeepSeek-V4-Pro} | ||
| DeepSeek-V4-Pro-0813: {root: lustre-models, dir: DeepSeek-V4-Pro-0813} |
There was a problem hiding this comment.
🟡 (optional) The changelog promises node-local checkpoint staging for the new gb200-nv dsv4-pro-0813 AgentX workload, but the config actually routes every run through shared Lustre. runners.yaml:611 adds DeepSeek-V4-Pro-0813: {root: lustre-models, ...} with no @ numa1 alias, and lustre-models has no visibility: node-local (runners.yaml:630), so it defaults to shared (infx/clusters/slurm.py:59). The sibling DeepSeek-V4-Pro@ numa1 entry (runners.yaml:619) is how same-cluster non-overridden dsv4 jobs normally get node-local picks via models.py's min(copies, key=lambda name: not node_local(name)). Fix: add a staged DeepSeek-V4-Pro-0813@ numa1 node-local entry, or correct perf-changelog.yaml:9179/9193's claim.
Why this was flagged
Trigger: any of the 8 new gb200-nv dsv4-pro-0813 dynamo-sglang AgentX recipes launches (agg or disagg, all added in this diff) and models.checkpoint() in infx/launch/drivers/srt/models.py resolves MODEL=deepseek-ai/DeepSeek-V4-Pro-0813 with no OVERRIDES entry for gb200-nv, so it falls to the basename scan over cluster.models.entries. Only one copy exists, DeepSeek-V4-Pro-0813 on lustre-models (shared, runners.yaml:611/630), unlike DeepSeek-V4-Pro@ numa1 (runners.yaml:619) which lets the default node-local-preferring logic pick local storage for the sibling model. So every node in every one of these jobs (up to 32 GPUs at concurrency 1280) reads the checkpoint over the shared Lustre mount at job start, contradicting perf-changelog.yaml:9179 and :9193's explicit claim of 'node-local DeepSeek-V4-Pro checkpoint staging', which this PR itself introduces as a factual record of the change.
Verification: The mechanism is real and reachable: basename matches only runners.yaml:611 (DeepSeek-V4-Pro-0813 on lustre-models), lustre-models (runners.yaml:630) sets no visibility, and Volume.visibility defaults to "shared" (clusters/base.py:22), so node_local is False. All 8 new recipes stage from shared Lustre, while perf-changelog.yaml claims "node-local" staging, and there is no DeepSeek-V4-Pro-0813@ numa1 node-local copy.
[by Claude Code]
Add GB200 Dynamo+SGLang AgentX configurations for
deepseek-ai/DeepSeek-V4-Pro-0813using its bundled DSpark head and HiCache, without W4A4 MXFP4 MegaMoE.This is a copy of #3182, rebased onto current
main, with one change:enable-w4a4-mxfp4-megamoe: trueis removed from every prefill and decode worker in the six disaggregated recipes. The MegaMoE all-to-all backend (moe-a2a-backend: megamoe), the FP4 indexer and all other recipe settings are unchanged. The two TP8 aggregate recipes never set the flag and are unchanged.The sweep covers TP8 aggregate concurrency 1 and 4; 1P1D DEP8/DEP16 concurrency 64 and 128; 1P1D DEP16/DEP32 concurrency 256; and 2P1D DEP16/DEP32 concurrency 768, 1024, and 1280.
AI model disclosure
main, resolved the changelog conflict as an append-only tail, removed the W4A4 MegaMoE flag, validated the YAML and the generated matrix for both config keys, and prepared this PR.中文
为
deepseek-ai/DeepSeek-V4-Pro-0813添加 GB200 Dynamo+SGLang AgentX 配置,使用其内置的 DSpark 草稿头和 HiCache,且不启用 W4A4 MXFP4 MegaMoE。本 PR 复制自 #3182,已 rebase 到当前
main,仅有一处改动:从六个分离式(disagg)recipe 的所有 prefill 和 decode worker 中移除enable-w4a4-mxfp4-megamoe: true。MegaMoE all-to-all 后端(moe-a2a-backend: megamoe)、FP4 indexer 及其他 recipe 设置均保持不变。两个 TP8 聚合 recipe 本来就未设置该选项,保持不变。本次 sweep 覆盖 TP8 聚合模式并发 1 和 4;1P1D DEP8/DEP16 并发 64 和 128;1P1D DEP16/DEP32 并发 256;以及 2P1D DEP16/DEP32 并发 768、1024 和 1280。
AI 模型披露
main,以仅追加末尾的方式解决 changelog 冲突,移除 W4A4 MegaMoE 选项,验证 YAML 和两个配置键生成的 matrix,并准备此 PR。🤖 Generated with Claude Code