[Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX - #3521
Conversation
…g on the EFA vLLM image Re-land two of #3405's recipes on current main, pointed at ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vllm-project/vllm#58768 vllm-openai-efa target), which already ships the AWS EFA libfabric stack and mooncake-transfer-engine-efa-cuda13 0.3.13.post1. Drop the kimik3-b300-efa-setup.sh runtime installer and point LD_LIBRARY_PATH at the image's /opt/amazon/efa/lib. 在当前 main 上重新提交 #3405 中的两个配方,改用内置 AWS EFA libfabric 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1 的 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 镜像;删除 kimik3-b300-efa-setup.sh 运行时安装脚本,LD_LIBRARY_PATH 改指镜像内的 /opt/amazon/efa/lib。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
填写 perf-changelog 的 PR 链接。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
| description: | ||
| - "Add two Kimi-K3 B300 Mooncake EFA AgentX configurations on ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vLLM PR vllm-project/vllm#58768 vllm-openai-efa target at a0c6725d8, which bakes in the AWS EFA libfabric 2.6.0amzn1.0 stack and mooncake-transfer-engine-efa-cuda13 0.3.13.post1): aggregate DCP8 DSpark7 maxseq2 at c1 and disaggregated 1P3D DCP8/DCP8 DSpark4 at c32. Supersedes #3405, dropping its kimik3-b300-efa-setup.sh runtime installer because the image already provides the EFA stack." | ||
| - "新增两个基于 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(vLLM PR vllm-project/vllm#58768 在 a0c6725d8 的 vllm-openai-efa 目标,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1)的 Kimi-K3 B300 Mooncake EFA AgentX 配置:聚合 DCP8 DSpark7 maxseq2(c1)与分离式 1P3D DCP8/DCP8 DSpark4(c32)。取代 #3405,并删除其 kimik3-b300-efa-setup.sh 运行时安装脚本,因为镜像已提供 EFA 组件。" | ||
| pr-link: PRLINK_PLACEHOLDER |
There was a problem hiding this comment.
🟡 (optional) The new changelog entry ships with pr-link: PRLINK_PLACEHOLDER instead of a real PR URL, so anyone auditing this perf change later gets a dead reference instead of the PR it came from. perf-changelog.yaml is append-only and byte-sensitive, so once merged this line cannot be silently rewritten later without another append; validate_perf_changelog only checks pr_link is a non-empty str (infx/matrix/validation.py:1016), so the placeholder passes CI unnoticed. Fix: replace PRLINK_PLACEHOLDER with the actual PR URL (https://github.com/SemiAnalysisAI/InferenceX/pull/, per CONTRIBUTING.md:190 and docs/configuration-procedures.md:647) before merge.
Why this was flagged
perf-changelog.yaml is append-only per AGENTS.md/CONTRIBUTING.md convention, so this pr-link value can never be edited in place after merge, only superseded by a later append. The ChangelogEntry.pr_link field (infx/matrix/validation.py:1016) is typed str with no URL/format validation, so 'PRLINK_PLACEHOLDER' passes validate_perf_changelog exactly as claimed in the PR's test plan. Documented convention (CONTRIBUTING.md:190, docs/configuration-procedures.md:647) is a real github.com/.../pull/ URL, or 'TBD' only before the PR exists (docs/configuration-procedures.md:650) -- but this PR already exists with a real number. Anyone later tracing why kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-agg/-disagg changed gets a non-resolving placeholder instead of the originating PR, permanently, since the entry can't be amended after merge.
Verification: nit. perf-changelog.yaml:8998 really ships pr-link: PRLINK_PLACEHOLDER (confirmed in the diff hunk at @@ -8986,3 +8986,13 @@ and in the file body). This is a placeholder, not the real PR URL, and should be filled in before merge — so the core observation is true and worth flagging as a review comment. However, the finding is nit, not normal, because its supporting reasoning is partly wrong and…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36341858202 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36341858202 |
Both #3521 multi-node jobs died importing the GHCR image with "curl: (56) Recv failure: Connection reset by peer" mid layer download. Retry enroot import up to 4 times with 30/60/90 s backoff and cap parallel layer downloads at ENROOT_MAX_CONNECTIONS=4 (overridable); enroot's layer cache keeps finished layers across attempts. 两个 #3521 多节点任务在导入 GHCR 镜像时因下载层中途 "curl: (56) Connection reset by peer" 失败。现对 enroot import 最多重试 4 次(间隔 30/60/90 秒),并将并行下载连接数限制为 ENROOT_MAX_CONNECTIONS=4(可覆盖);enroot 的层缓存会在重试间保留已完成的层。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
Closing: superseded by #3523 (1P1D). Mooncake EFA 190 GB segment registration failed with fi_mr_reg ENOMEM on this sweep. |
Summary
Re-lands two recipes from #3405 on current
main, switched to the new EFA vLLM image:Image:
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768, replacingvllm/vllm-openai:nightly-0961bbae…. It is built from [docker] add vllm-openai-efa target vllm-project/vllm#58768 (vllm-openai-efatarget, heada0c6725d8, reports0.30.1rc1.dev30+ga0c6725d8). The image already contains AWS EFA libfabric2.6.0amzn1.0at/opt/amazon/efa/lib(registered with ldconfig) andmooncake-transfer-engine-efa-cuda13==0.3.13.post1. I checked inside the image that itsengine.solinks/opt/amazon/efa/lib/libfabric.so.1.Removed
benchmarks/multi_node/srt-slurm-recipes/configs/kimik3-b300-efa-setup.sh, which is not carried over. That runtime installer (EFA 1.50 plus Mooncake EFA wheel) is now baked into the image, so both recipes dropsetup_script:andLD_LIBRARY_PATHpoints at/opt/amazon/efa/libinstead of/opt/amazon/efa-1.50/lib.Only two recipes kept, under
benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/:agg-dcp8-dspark7-maxseq2-mooncake-c1.yaml, used bykimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-agg(conc 1)disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml, used bykimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg(1P3D, conc 32)The c4/c70 aggregate and 1P1D/1P2D disaggregated variants and the no-spec aggregate key from [dnm]Add Kimi-K3 B300 Mooncake EFA recipes #3405 are dropped.
Recipe bodies are otherwise identical to #3405 (
config/kimik3-b300-mooncake-pool-fix@dd312397).runners/launch_b300-dsxe.sh,8d119b00): the first sweep (run 36340702106) died on both multi-node jobs while importing the GHCR image. The layer download failed withcurl: (56) Recv failure: Connection reset by peeron gpu-09 and gpu-11.import_squash_imagenow retriesenroot importup to 4 times with 30/60/90 s backoff and caps parallel layer downloads atENROOT_MAX_CONNECTIONS=4(overridable). Enroot's layer cache keeps finished layers across attempts. This applies to every b300-dsxe image import.Risk
On b300-dsxe the host may bind-mount its own
/opt/amazon/efainto the container. #3405's setup script worked around that by installing to/opt/amazon/efa-1.50. If the host mount shadows the image's libfabric, Mooncake EFA may fail to initialize. This first sweep will show it.AI model disclosure
claude-opus-5-5[1m]) via Claude Code: built and pushed the image, made the changes, validated, and wrote this PR.Test plan
validate_perf_changelogagainst origin/main passes, with no deletions.infx.matrix.generate test-configfor both keys yields 2 entries on the new image.fi_info --versionreports libfabric 2.6.0amzn1.0, and Mooncakeengine.soresolveslibfabric.so.1from/opt/amazon/efa/lib.中文
摘要
在当前
main上重新提交 #3405 中的两个配方,并改用新的 EFA vLLM 镜像:镜像: 由
vllm/vllm-openai:nightly-0961bbae…改为ghcr.io/semianalysisai/vllm-openai:efa_pr_58768,基于 [docker] add vllm-openai-efa target vllm-project/vllm#58768 的vllm-openai-efa目标(heada0c6725d8)构建。镜像内已包含/opt/amazon/efa/lib下的 AWS EFA libfabric2.6.0amzn1.0以及mooncake-transfer-engine-efa-cuda13==0.3.13.post1,其engine.so链接到镜像内的 libfabric。删除
kimik3-b300-efa-setup.sh:镜像已内置 EFA 与 Mooncake EFA,因此两个配方均去掉setup_script:,LD_LIBRARY_PATH改为/opt/amazon/efa/lib。仅保留两个配方:
agg-dcp8-dspark7-maxseq2-mooncake-c1.yaml(聚合,c1)与disagg-1p3d-dcp8-dcp8-dspark4-mooncake-c32.yaml(1P3D,c32)。[dnm]Add Kimi-K3 B300 Mooncake EFA recipes #3405 中的其余变体已删除。启动脚本(
runners/launch_b300-dsxe.sh,8d119b00): 首次扫描(run 36340702106)的两个多节点任务在导入 GHCR 镜像时因下载层中途curl: (56) Connection reset by peer失败。import_squash_image现对enroot import最多重试 4 次(间隔 30/60/90 秒),并将并行下载限制为ENROOT_MAX_CONNECTIONS=4;适用于所有 b300-dsxe 镜像导入。风险
b300-dsxe 主机可能将自身的
/opt/amazon/efa挂载进容器,从而覆盖镜像内的 libfabric;首次扫描会验证这一点。AI 模型披露
claude-opus-5-5[1m]),通过 Claude Code:构建并推送镜像、完成修改与验证并撰写此 PR。测试计划
validate_perf_changelog通过,无删除。🤖 Generated with Claude Code