[Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX 1P3D disagg on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX 1P3D 分离式 - #3522
Conversation
…EFA vLLM image Disaggregated arm split from #3521 so it sweeps in parallel with the aggregate arm. Uses ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (EFA libfabric + Mooncake EFA baked in, no setup script) and carries the b300-dsxe enroot import retry fix. 从 #3521 拆出的分离式配置,与聚合式并行扫描;使用内置 EFA libfabric 与 Mooncake EFA 的 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(无需安装脚本),并包含 b300-dsxe 的 enroot 导入重试修复。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
填写 perf-changelog 的 PR 链接。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
…o srtctl uses the squash file
srtslurm.yaml maps "${IMAGE}" (ghcr.io#semianalysisai/...) to the pre-imported
.sqsh. The recipe said ghcr.io/semianalysisai/..., so the key never matched and
pyxis pulled the 9 GB image from GHCR on every node at job start; on gpu-04 that
outran srtctl's 300 s etcd readiness window (run 36343238250, Slurm 5853).
srtslurm.yaml 以 "${IMAGE}"(ghcr.io#semianalysisai/...)映射到预先导入的 .sqsh;配方写成 ghcr.io/semianalysisai/...,键不匹配,导致每个节点启动时都从 GHCR 拉取 9 GB 镜像,gpu-04 上超出了 srtctl 对 etcd 的 300 秒就绪等待。
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36343944904 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36343944904 |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked the retry-loop change in runners/launch_b300-dsxe.sh: the ENROOT_MAX_CONNECTIONS="${ENROOT_MAX_CONNECTIONS:-4}" default is a defensible exception to AGENTS.md's "no fallback defaults" rule since it's an enroot-native runtime tuning knob rather than caller-supplied benchmark config, and the new loop doesn't introduce any other unset-unsafe expansions in the script.
Extended reasoning...
The diff adds a new multi-node srt-slurm recipe + matching master-config entry for a Kimi-K3 B300 disaggregated benchmark, appends a perf-changelog entry, and patches a launcher script's image-import retry logic; none of it touches auth/crypto/permissions. A CONFIRMED finding (router version pinned in the master config doesn't match the dynamo rev actually pinned in the recipe YAML) will be posted as an inline comment, which is reason enough to withhold approval. I additionally examined the ENROOT_MAX_CONNECTIONS default and the retry loop for nounset-unsafe expansions per AGENTS.md conventions and found no further issue worth raising.
| runner: cluster:b300-dsxe | ||
| precision: fp4 | ||
| framework: dynamo-vllm | ||
| router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } |
There was a problem hiding this comment.
🟡 (optional) The new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg entry's router.version is ba83080ecd31c1ce918559e576d3c5bc9e092ff1, but the recipe it points at (disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml) actually pins dynamo rev cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b at lines 11 and 236. Anyone reading the master config (dashboards, matrix tooling, or a human comparing runs) sees the wrong dynamo-router version for this benchmark, since the base convention is that these two fields must agree. Fix: set router.version to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b to match the recipe's actual pinned dynamo rev, mirroring the container/image consistency already enforced elsewhere in this file.
Why this was flagged
configs/nvidia-master.yaml:1538 declares router version ba83080ecd31c1ce918559e576d3c5bc9e092ff1 for config kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg. That config's srt-recipe, benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml, pins dynamo.source.rev to a different commit, cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b, at lines 11 and 236. Every other kimik3 recipe using rev ba83080... has master-config router.version ba83080... (grep confirms this), so this entry is the sole outlier where the two disagree. Anyone relying on the master config's router field to know which dynamo build actually ran gets stale/incorrect provenance, since no other field in the recipe corrects it.
Verification: nit. The mismatch is real. configs/nvidia-master.yaml:1538 sets router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } for the new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg entry, but the recipe it points at (benchmarks/.../b300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml) pins dynamo at… | nit. configs/nvidia-master.yaml:1538…
|
Closing: superseded by #3523 (1P1D). Mooncake EFA 190 GB segment registration failed with fi_mr_reg ENOMEM on this sweep. |
Summary
The disaggregated arm, split out of #3521 so it sweeps in parallel with the aggregate arm:
kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg, 1P3D, DCP8/DCP8, DSpark4, Mooncake DRAM offload, conc 32. Recipebenchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml.ghcr.io/semianalysisai/vllm-openai:efa_pr_58768([docker] add vllm-openai-efa target vllm-project/vllm#58768vllm-openai-efatarget ata0c6725d8). EFA libfabric2.6.0amzn1.0at/opt/amazon/efa/libandmooncake-transfer-engine-efa-cuda13==0.3.13.post1are built in, so there's nosetup_scriptand nokimik3-b300-efa-setup.sh.LD_LIBRARY_PATHpoints at/opt/amazon/efa/lib.runners/launch_b300-dsxe.shincludes the same fix as [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX #3521.import_squash_imageretriesenroot importup to 4 times (30/60/90 s backoff) withENROOT_MAX_CONNECTIONS=4, after [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX #3521's first sweep lost both multi-node jobs tocurl: (56) Connection reset by peerpulling from GHCR.The recipe, config entry and launcher are byte-identical to #3521 at
8d119b00. Whichever PR merges second will needmainmerged in.ee3ae93d): the recipe now usesghcr.io#semianalysisai/vllm-openai:efa_pr_58768, exactly the master-configimage:.runners/srt-slurm/b300-dsxe.yamlmaps"${IMAGE}"to the pre-imported.sqsh. With the/spelling the key never matched, so pyxis pulled the 9 GB image from GHCR on every node. On gpu-04 that outran srtctl's 300 s etcd readiness window (run 36343238250, Slurm 5853).AI model disclosure
claude-opus-5-5[1m]) via Claude Code: made the change, validated it, and wrote this PR.Test plan
validate_perf_changelogagainst origin/main passes, with no deletions.infx.matrix.generate test-configyields 1 entry (kimik3_p1x8_d3x8_conc32_kvdram-mooncake).bash -n runners/launch_b300-dsxe.sh.中文
摘要
从 #3521 拆出的分离式配置,与聚合式并行扫描:
kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg(1P3D,DCP8/DCP8,DSpark4,Mooncake DRAM 卸载,c32)。ghcr.io/semianalysisai/vllm-openai:efa_pr_58768,内置 EFA libfabric2.6.0amzn1.0与mooncake-transfer-engine-efa-cuda13==0.3.13.post1,因此无需setup_script与kimik3-b300-efa-setup.sh。enroot import最多重试 4 次(30/60/90 秒)并将并行连接限制为 4。配方、配置与启动脚本与 #3521 的
8d119b00完全一致;第二个合并的 PR 需要先合入main。ee3ae93d): 配方改用与主配置image:完全一致的ghcr.io#semianalysisai/vllm-openai:efa_pr_58768,使b300-dsxe.yaml能映射到预先导入的.sqsh;此前键不匹配,每个节点都从 GHCR 拉取 9 GB 镜像,gpu-04 上超出了 etcd 的 300 秒就绪等待。AI 模型披露
claude-opus-5-5[1m]),通过 Claude Code:完成修改与验证并撰写此 PR。测试计划
validate_perf_changelog通过,无删除。🤖 Generated with Claude Code