Skip to content

[Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX - #3521

Closed
functionstackx wants to merge 3 commits into
mainfrom
klaud/kimik3-b300-mooncake-efa-image
Closed

functionstackx wants to merge 3 commits into
mainfrom
klaud/kimik3-b300-mooncake-efa-image

Conversation

@functionstackx

@functionstackx functionstackx commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Re-lands two recipes from #3405 on current main, switched to the new EFA vLLM image:

  • Image: ghcr.io/semianalysisai/vllm-openai:efa_pr_58768, replacing vllm/vllm-openai:nightly-0961bbae…. It is built from [docker] add vllm-openai-efa target vllm-project/vllm#58768 (vllm-openai-efa target, head a0c6725d8, reports 0.30.1rc1.dev30+ga0c6725d8). The image already contains AWS EFA libfabric 2.6.0amzn1.0 at /opt/amazon/efa/lib (registered with ldconfig) and mooncake-transfer-engine-efa-cuda13==0.3.13.post1. I checked inside the image that its engine.so links /opt/amazon/efa/lib/libfabric.so.1.

  • Removed benchmarks/multi_node/srt-slurm-recipes/configs/kimik3-b300-efa-setup.sh, which is not carried over. That runtime installer (EFA 1.50 plus Mooncake EFA wheel) is now baked into the image, so both recipes drop setup_script: and LD_LIBRARY_PATH points at /opt/amazon/efa/lib instead of /opt/amazon/efa-1.50/lib.

  • Only two recipes kept, under benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/:

    • agg-dcp8-dspark7-maxseq2-mooncake-c1.yaml, used by kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-agg (conc 1)
    • disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml, used by kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg (1P3D, conc 32)

    The c4/c70 aggregate and 1P1D/1P2D disaggregated variants and the no-spec aggregate key from [dnm]Add Kimi-K3 B300 Mooncake EFA recipes #3405 are dropped.

Recipe bodies are otherwise identical to #3405 (config/kimik3-b300-mooncake-pool-fix @ dd312397).

  • Launcher (runners/launch_b300-dsxe.sh, 8d119b00): the first sweep (run 36340702106) died on both multi-node jobs while importing the GHCR image. The layer download failed with curl: (56) Recv failure: Connection reset by peer on gpu-09 and gpu-11. import_squash_image now retries enroot import up to 4 times with 30/60/90 s backoff and caps parallel layer downloads at ENROOT_MAX_CONNECTIONS=4 (overridable). Enroot's layer cache keeps finished layers across attempts. This applies to every b300-dsxe image import.

Risk

On b300-dsxe the host may bind-mount its own /opt/amazon/efa into the container. #3405's setup script worked around that by installing to /opt/amazon/efa-1.50. If the host mount shadows the image's libfabric, Mooncake EFA may fail to initialize. This first sweep will show it.

AI model disclosure

  • Claude Opus 5.5 (claude-opus-5-5[1m]) via Claude Code: built and pushed the image, made the changes, validated, and wrote this PR.

Test plan

  • validate_perf_changelog against origin/main passes, with no deletions.
  • infx.matrix.generate test-config for both keys yields 2 entries on the new image.
  • Inside the image: fi_info --version reports libfabric 2.6.0amzn1.0, and Mooncake engine.so resolves libfabric.so.1 from /opt/amazon/efa/lib.
  • b300-dsxe sweep passes.
中文

摘要

在当前 main 上重新提交 #3405 中的两个配方,并改用新的 EFA vLLM 镜像:

  • 镜像: 由 vllm/vllm-openai:nightly-0961bbae… 改为 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768,基于 [docker] add vllm-openai-efa target vllm-project/vllm#58768 的 vllm-openai-efa 目标(head a0c6725d8)构建。镜像内已包含 /opt/amazon/efa/lib 下的 AWS EFA libfabric 2.6.0amzn1.0 以及 mooncake-transfer-engine-efa-cuda13==0.3.13.post1,其 engine.so 链接到镜像内的 libfabric。

  • 删除 kimik3-b300-efa-setup.sh:镜像已内置 EFA 与 Mooncake EFA,因此两个配方均去掉 setup_script:,LD_LIBRARY_PATH 改为 /opt/amazon/efa/lib。

  • 仅保留两个配方:agg-dcp8-dspark7-maxseq2-mooncake-c1.yaml(聚合,c1)与 disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml(1P3D,c32)。[dnm]Add Kimi-K3 B300 Mooncake EFA recipes #3405 中的其余变体已删除。

  • 启动脚本(runners/launch_b300-dsxe.sh,8d119b00): 首次扫描(run 36340702106)的两个多节点任务在导入 GHCR 镜像时因下载层中途 curl: (56) Connection reset by peer 失败。import_squash_image 现对 enroot import 最多重试 4 次(间隔 30/60/90 秒),并将并行下载限制为 ENROOT_MAX_CONNECTIONS=4;适用于所有 b300-dsxe 镜像导入。

风险

b300-dsxe 主机可能将自身的 /opt/amazon/efa 挂载进容器,从而覆盖镜像内的 libfabric;首次扫描会验证这一点。

AI 模型披露

  • Claude Opus 5.5(claude-opus-5-5[1m]),通过 Claude Code:构建并推送镜像、完成修改与验证并撰写此 PR。

测试计划

  • 针对 origin/main 的 validate_perf_changelog 通过,无删除。
  • 两个配置键生成 2 个矩阵条目,均使用新镜像。
  • 镜像内 libfabric 2.6.0amzn1.0 与 Mooncake EFA 链接验证通过。
  • b300-dsxe 扫描通过。

🤖 Generated with Claude Code

functionstackx and others added 2 commits September 27, 2026 14:16
…g on the EFA vLLM image

Re-land two of #3405's recipes on current main, pointed at
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vllm-project/vllm#58768
vllm-openai-efa target), which already ships the AWS EFA libfabric stack and
mooncake-transfer-engine-efa-cuda13 0.3.13.post1. Drop the
kimik3-b300-efa-setup.sh runtime installer and point LD_LIBRARY_PATH at the
image's /opt/amazon/efa/lib.

在当前 main 上重新提交 #3405 中的两个配方,改用内置 AWS EFA libfabric 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1 的 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 镜像;删除 kimik3-b300-efa-setup.sh 运行时安装脚本,LD_LIBRARY_PATH 改指镜像内的 /opt/amazon/efa/lib。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
填写 perf-changelog 的 PR 链接。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment thread perf-changelog.yaml Outdated
description:
- "Add two Kimi-K3 B300 Mooncake EFA AgentX configurations on ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vLLM PR vllm-project/vllm#58768 vllm-openai-efa target at a0c6725d8, which bakes in the AWS EFA libfabric 2.6.0amzn1.0 stack and mooncake-transfer-engine-efa-cuda13 0.3.13.post1): aggregate DCP8 DSpark7 maxseq2 at c1 and disaggregated 1P3D DCP8/DCP8 DSpark4 at c32. Supersedes #3405, dropping its kimik3-b300-efa-setup.sh runtime installer because the image already provides the EFA stack."
- "新增两个基于 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(vLLM PR vllm-project/vllm#58768 在 a0c6725d8 的 vllm-openai-efa 目标,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1)的 Kimi-K3 B300 Mooncake EFA AgentX 配置:聚合 DCP8 DSpark7 maxseq2(c1)与分离式 1P3D DCP8/DCP8 DSpark4(c32)。取代 #3405,并删除其 kimik3-b300-efa-setup.sh 运行时安装脚本,因为镜像已提供 EFA 组件。"
pr-link: PRLINK_PLACEHOLDER

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The new changelog entry ships with pr-link: PRLINK_PLACEHOLDER instead of a real PR URL, so anyone auditing this perf change later gets a dead reference instead of the PR it came from. perf-changelog.yaml is append-only and byte-sensitive, so once merged this line cannot be silently rewritten later without another append; validate_perf_changelog only checks pr_link is a non-empty str (infx/matrix/validation.py:1016), so the placeholder passes CI unnoticed. Fix: replace PRLINK_PLACEHOLDER with the actual PR URL (https://github.com/SemiAnalysisAI/InferenceX/pull/, per CONTRIBUTING.md:190 and docs/configuration-procedures.md:647) before merge.

Why this was flagged

perf-changelog.yaml is append-only per AGENTS.md/CONTRIBUTING.md convention, so this pr-link value can never be edited in place after merge, only superseded by a later append. The ChangelogEntry.pr_link field (infx/matrix/validation.py:1016) is typed str with no URL/format validation, so 'PRLINK_PLACEHOLDER' passes validate_perf_changelog exactly as claimed in the PR's test plan. Documented convention (CONTRIBUTING.md:190, docs/configuration-procedures.md:647) is a real github.com/.../pull/ URL, or 'TBD' only before the PR exists (docs/configuration-procedures.md:650) -- but this PR already exists with a real number. Anyone later tracing why kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-agg/-disagg changed gets a non-resolving placeholder instead of the originating PR, permanently, since the entry can't be amended after merge.

Verification: nit. perf-changelog.yaml:8998 really ships pr-link: PRLINK_PLACEHOLDER (confirmed in the diff hunk at @@ -8986,3 +8986,13 @@ and in the file body). This is a placeholder, not the real PR URL, and should be filled in before merge — so the core observation is true and worth flagging as a review comment. However, the finding is nit, not normal, because its supporting reasoning is partly wrong and…

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Both #3521 multi-node jobs died importing the GHCR image with
"curl: (56) Recv failure: Connection reset by peer" mid layer download.
Retry enroot import up to 4 times with 30/60/90 s backoff and cap parallel
layer downloads at ENROOT_MAX_CONNECTIONS=4 (overridable); enroot's layer
cache keeps finished layers across attempts.

两个 #3521 多节点任务在导入 GHCR 镜像时因下载层中途 "curl: (56) Connection reset by peer" 失败。现对 enroot import 最多重试 4 次(间隔 30/60/90 秒),并将并行下载连接数限制为 ENROOT_MAX_CONNECTIONS=4(可覆盖);enroot 的层缓存会在重试间保留已完成的层。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@functionstackx

Copy link
Copy Markdown
Collaborator Author

Closing: superseded by #3523 (1P1D). Mooncake EFA 190 GB segment registration failed with fi_mr_reg ENOMEM on this sweep.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant