Skip to content

[Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX 1P3D disagg on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX 1P3D 分离式 - #3522

Closed
functionstackx wants to merge 3 commits into
mainfrom
klaud/kimik3-b300-mooncake-efa-disagg
Closed

functionstackx wants to merge 3 commits into
mainfrom
klaud/kimik3-b300-mooncake-efa-disagg

Conversation

@functionstackx

@functionstackx functionstackx commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The disaggregated arm, split out of #3521 so it sweeps in parallel with the aggregate arm:

The recipe, config entry and launcher are byte-identical to #3521 at 8d119b00. Whichever PR merges second will need main merged in.

  • Image key fix (ee3ae93d): the recipe now uses ghcr.io#semianalysisai/vllm-openai:efa_pr_58768, exactly the master-config image:. runners/srt-slurm/b300-dsxe.yaml maps "${IMAGE}" to the pre-imported .sqsh. With the / spelling the key never matched, so pyxis pulled the 9 GB image from GHCR on every node. On gpu-04 that outran srtctl's 300 s etcd readiness window (run 36343238250, Slurm 5853).

AI model disclosure

  • Claude Opus 5.5 (claude-opus-5-5[1m]) via Claude Code: made the change, validated it, and wrote this PR.

Test plan

  • validate_perf_changelog against origin/main passes, with no deletions.
  • infx.matrix.generate test-config yields 1 entry (kimik3_p1x8_d3x8_conc32_kvdram-mooncake).
  • bash -n runners/launch_b300-dsxe.sh.
  • b300-dsxe sweep passes.
中文

摘要

从 #3521 拆出的分离式配置,与聚合式并行扫描:

配方、配置与启动脚本与 #3521 的 8d119b00 完全一致;第二个合并的 PR 需要先合入 main。

  • 镜像键修复(ee3ae93d): 配方改用与主配置 image: 完全一致的 ghcr.io#semianalysisai/vllm-openai:efa_pr_58768,使 b300-dsxe.yaml 能映射到预先导入的 .sqsh;此前键不匹配,每个节点都从 GHCR 拉取 9 GB 镜像,gpu-04 上超出了 etcd 的 300 秒就绪等待。

AI 模型披露

  • Claude Opus 5.5(claude-opus-5-5[1m]),通过 Claude Code:完成修改与验证并撰写此 PR。

测试计划

  • 针对 origin/main 的 validate_perf_changelog 通过,无删除。
  • 生成 1 个矩阵条目。
  • 启动脚本语法检查通过。
  • b300-dsxe 扫描通过。

🤖 Generated with Claude Code

…EFA vLLM image

Disaggregated arm split from #3521 so it sweeps in parallel with the
aggregate arm. Uses ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (EFA
libfabric + Mooncake EFA baked in, no setup script) and carries the
b300-dsxe enroot import retry fix.

从 #3521 拆出的分离式配置,与聚合式并行扫描;使用内置 EFA libfabric 与 Mooncake EFA 的 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(无需安装脚本),并包含 b300-dsxe 的 enroot 导入重试修复。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
填写 perf-changelog 的 PR 链接。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@functionstackx functionstackx added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 27, 2026
…o srtctl uses the squash file

srtslurm.yaml maps "${IMAGE}" (ghcr.io#semianalysisai/...) to the pre-imported
.sqsh. The recipe said ghcr.io/semianalysisai/..., so the key never matched and
pyxis pulled the 9 GB image from GHCR on every node at job start; on gpu-04 that
outran srtctl's 300 s etcd readiness window (run 36343238250, Slurm 5853).

srtslurm.yaml 以 "${IMAGE}"(ghcr.io#semianalysisai/...)映射到预先导入的 .sqsh;配方写成 ghcr.io/semianalysisai/...,键不匹配,导致每个节点启动时都从 GHCR 拉取 9 GB 镜像,gpu-04 上超出了 srtctl 对 etcd 的 300 秒就绪等待。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the retry-loop change in runners/launch_b300-dsxe.sh: the ENROOT_MAX_CONNECTIONS="${ENROOT_MAX_CONNECTIONS:-4}" default is a defensible exception to AGENTS.md's "no fallback defaults" rule since it's an enroot-native runtime tuning knob rather than caller-supplied benchmark config, and the new loop doesn't introduce any other unset-unsafe expansions in the script.

Extended reasoning...

The diff adds a new multi-node srt-slurm recipe + matching master-config entry for a Kimi-K3 B300 disaggregated benchmark, appends a perf-changelog entry, and patches a launcher script's image-import retry logic; none of it touches auth/crypto/permissions. A CONFIRMED finding (router version pinned in the master config doesn't match the dynamo rev actually pinned in the recipe YAML) will be posted as an inline comment, which is reason enough to withhold approval. I additionally examined the ENROOT_MAX_CONNECTIONS default and the retry loop for nounset-unsafe expansions per AGENTS.md conventions and found no further issue worth raising.

runner: cluster:b300-dsxe
precision: fp4
framework: dynamo-vllm
router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg entry's router.version is ba83080ecd31c1ce918559e576d3c5bc9e092ff1, but the recipe it points at (disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml) actually pins dynamo rev cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b at lines 11 and 236. Anyone reading the master config (dashboards, matrix tooling, or a human comparing runs) sees the wrong dynamo-router version for this benchmark, since the base convention is that these two fields must agree. Fix: set router.version to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b to match the recipe's actual pinned dynamo rev, mirroring the container/image consistency already enforced elsewhere in this file.

Why this was flagged

configs/nvidia-master.yaml:1538 declares router version ba83080ecd31c1ce918559e576d3c5bc9e092ff1 for config kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg. That config's srt-recipe, benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml, pins dynamo.source.rev to a different commit, cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b, at lines 11 and 236. Every other kimik3 recipe using rev ba83080... has master-config router.version ba83080... (grep confirms this), so this entry is the sole outlier where the two disagree. Anyone relying on the master config's router field to know which dynamo build actually ran gets stale/incorrect provenance, since no other field in the recipe corrects it.

Verification: nit. The mismatch is real. configs/nvidia-master.yaml:1538 sets router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } for the new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg entry, but the recipe it points at (benchmarks/.../b300-fp4/agentx/disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml) pins dynamo at… | nit. configs/nvidia-master.yaml:1538…

@functionstackx

Copy link
Copy Markdown
Collaborator Author

Closing: superseded by #3523 (1P1D). Mooncake EFA 190 GB segment registration failed with fi_mr_reg ENOMEM on this sweep.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant