[Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 / [Klaud Cold] 新增 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 - #3523
Conversation
…park4 c48 on efa_pr_58768 image Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
7bb1635 to
13de142
Compare
|
Thanks for the contribution!
中文感谢你的贡献!
|
| description: | ||
| - "Add the Kimi-K3 B300 Mooncake EFA AgentX disaggregated 1P1D DCP8/DCP8 DSpark4 configuration at c48 on ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vllm-project/vllm#58768 vllm-openai-efa target, with AWS EFA libfabric 2.6.0amzn1.0 and mooncake-transfer-engine-efa-cuda13 0.3.13.post1 built in). Taken from #3405's 1P1D arm; drops its kimik3-b300-efa-setup.sh runtime installer." | ||
| - "新增基于 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(vllm-project/vllm#58768 的 vllm-openai-efa 目标,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1)的 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8/DCP8 DSpark4 配置(c48)。取自 #3405 的 1P1D 配置;删除其 kimik3-b300-efa-setup.sh 运行时安装脚本。" | ||
| pr-link: PRLINK_PLACEHOLDER |
There was a problem hiding this comment.
🔴 Maintainers and CI get a changelog entry that will fail validation because pr-link: PRLINK_PLACEHOLDER isn't a value the validator accepts. infx/workflows/validate_perf_changelog.py's validate_added_pr_link only allows the exact expected .../pull/<pr_number> URL or a value in PR_LINK_PLACEHOLDERS ({"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"}); PRLINK_PLACEHOLDER matches neither, so it raises ChangelogValidationError("new PR entry must use ... or an XXX placeholder; found 'PRLINK_PLACEHOLDER'"). Fix: use the literal XXX placeholder (or the full .../pull/XXX form) as documented, not an invented string, so the automated canonicalization in prepare_perf_changelog_merge.py can replace it with the real PR link before merge.
Why this was flagged
The new entry appended at perf-changelog.yaml:8989-8997 sets pr-link: PRLINK_PLACEHOLDER. validate_added_pr_link (infx/workflows/validate_perf_changelog.py:133-145) is invoked on PR runs and only tolerates PR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"} or the exact expected pull URL. PRLINK_PLACEHOLDER is none of these, so validation raises ChangelogValidationError, blocking the changelog check and the automated pr-link canonicalization step (infx/workflows/prepare_perf_changelog_merge.py:103-109) that would otherwise fill in the real link on merge. On base branch, entries use XXX, which passes this check.
Verification: normal (mechanism corrected): The appended entry at perf-changelog.yaml:8997 sets pr-link: PRLINK_PLACEHOLDER, which is not an accepted pr-link value. In infx/workflows/validate_perf_changelog.py, PR_LINK_PLACEHOLDERS = {"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"} (lines 21-24); validate_added_pr_link (lines 134-145) raises ChangelogValidationError for any…
| runner: cluster:b300-dsxe | ||
| precision: fp4 | ||
| framework: dynamo-vllm | ||
| router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } |
There was a problem hiding this comment.
🟡 (optional) The new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d entry sets router.version to ba83080ecd31c1ce918559e576d3c5bc9e092ff1, but the recipe it points at (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11,237) builds dynamo from a different rev, cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b. Every sibling entry (e.g. nvidia-master.yaml:7841, kimik3-fp4-gb300-...-disagg) keeps router.version equal to its recipe's dynamo rev; this one doesn't. Per docs/results-and-ingestion.md, router ({name, version}) is stored as benchmark result metadata, so results from this scenario get tagged and grouped under the wrong dynamo-router build, silently mixing them into Pareto comparisons meant for a different dynamo revision. …
Why this was flagged
…Fix: set router.version to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b to match the recipe it references.
configs/nvidia-master.yaml:1538 sets router.version to ba83080ecd31c1ce918559e576d3c5bc9e092ff1 for the new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d key. The recipe it references, benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11 and :237, sets the dynamo rev actually built into the image to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b, a different commit. infx/matrix/validation.py's ComponentMetadata schema for router has no check comparing it to the recipe's dynamo rev, so nothing catches the mismatch. docs/results-and-ingestion.md documents router as {name, version} metadata stored with benchmark results, so results are recorded under the wrong dynamo-router version. Sibling entries like nvidia-master.yaml:7841 keep these values equal, showing this is the convention the new entry breaks.
Verification: Severity: normal — this new entry records wrong dynamo/router provenance. configs/nvidia-master.yaml:1538 sets router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" }, but the recipe it references builds a different dynamo commit: disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11 dynamo.source.rev: cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b (with `install:…
| disagg: true | ||
| scenarios: | ||
| agentic-coding: | ||
| - dram-utilization: 0.75 |
There was a problem hiding this comment.
🟡 (optional) Anyone reading this scenario's recorded allocated_cpu_dram_gb metric gets a value ~12x too large, corrupting DRAM-offload capacity/efficiency comparisons for this benchmark. dram-utilization: 0.75 here computes total-cpu-dram-gb ≈ 2249 GB (full-node fraction for the 8-GPU prefill worker), which becomes TOTAL_CPU_DRAM_GB and is stored verbatim as allocated_cpu_dram_gb, but the recipe's actual centralized Mooncake pool is only global_segment_size: 190GB (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:40). Every sibling entry sizes dram-utilization to match its recipe's segment size (e.g. gb300's 0.1775 -> 160GB at nvidia-master.yaml:7849); this one looks copy-pasted from the single-node b300 kimik3 recipe's 0.75/2249GB pairing. …
Why this was flagged
…Fix: pick dram-utilization so computed total-cpu-dram-gb matches the recipe's real DRAM budget (~0.063 here for 190GB), per the convention every other multinode entry follows.
configs/nvidia-master.yaml:1544 sets dram-utilization: 0.75 under kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d. infx/matrix/generate.py's agentic_dram_offload_gb (lines 429-450) uses the prefill worker's full-node GPU fraction (tp=8 on an 8-GPU node -> fraction 1) to turn that into total_cpu_dram_gb ≈ 2249, coincidentally equal to the unrelated single-node recipe's hardcoded TOTAL_CPU_DRAM_GB=2249 (benchmarks/single_node/srt-slurm-recipes/kimik3/vllm/b300-fp4-mtp/agentic.yaml:109). benchmark-multinode-tmpl.yml:178 wires total-cpu-dram-gb into TOTAL_CPU_DRAM_GB; benchmark_lib.sh:212/406 only checks it's a positive integer, so the mismatch is never caught. infx/results/agentic/init.py:191 records it as allocated_cpu_dram_gb, while the recipe's real capacity is global_segment_size: 190GB (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:40) - about 12x smaller than what gets recorded.
Verification: normal (metric/data-correctness defect in newly added scenario; no serving impact). configs/nvidia-master.yaml:1544 sets dram-utilization: 0.75 for kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d. infx/matrix/generate.py agentic_dram_offload_gb (lines 446-470) computes min(runner DRAM, MAX_AGENTIC_AVAILABLE_CPU_DRAM_MIB=2,861,022 MiB ≈ 2999 GB) × 0.75 × (prefill gpu_count…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36345723679 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36345723679 |
|
Hi, I maintain Mooncake on EFA. I looked into the slow registration and reproduced it on 2× p5en with the same setup (8 ranks × 190GB Store segment per node). Root cause EFA itself isn't slow. The problem is how the Store segment gets its huge pages.
Fix (config only): pre-reserve 2MB hugetlb pages on the host and let Mooncake use them env:
MC_STORE_USE_HUGEPAGE: '1' # 2MB pages by default
# GLIBC_TUNABLES is no longer needed for the Store segment# on each node, as root, before launching (host sysctl; containers inherit it)
sudo sysctl -w vm.nr_hugepages=850000 # ~1.66 TiBSizing: ranks per node × ( Reserve the pages early, ideally via Result: with this setup, 8 × 190GB Store setup took 26–28 s end-to-end on p5en (populate + registration), with zero compaction. Investigation and write-up done with AI assistance (Claude Code); numbers are from real runs on p5en. |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
|
Quick update: the fix for the slow EFA segment registration is merged in Mooncake (kvcache-ai/Mooncake#4393). The Store segment is now allocated as 2MB-aligned On p5en (8 ranks × 190GB, node0 fragmented), registration went from 1306s to 40s. The fix will ship in the next Mooncake release. The current image's The fix requires THP not set to I've only verified this on p5en so far. If you try it on B300, please share the registration time and any |
Summary
Adds the 1P1D arm of #3405 (Kimi-K3 FP4, B300 Dynamo-vLLM, 1P (TP8/DCP8) × 1D (TP8/DCP8), MTP/DSpark4, DRAM KV offload via Mooncake, c48) on the new EFA image
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(vllm-project/vllm#58768vllm-openai-efatarget, with AWS EFA libfabric 2.6.0amzn1.0 and mooncake-transfer-engine-efa-cuda13 0.3.13.post1 built in).kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d→recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yamlghcr.io#…spelling so it matches the pre-imported squash key (otherwise every node pulls 9 GB from GHCR);LD_LIBRARY_PATHpoints at the image's/opt/amazon/efa/lib;setup_script: kimik3-b300-efa-setup.shis removed.runners/launch_b300-dsxe.sh: enroot import retry, identical to [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX #3521 / [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX 1P3D disagg on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX 1P3D 分离式 #3522.Sibling PRs: #3521 (agg), #3522 (1P3D c32).
Test plan
full-sweep-enabledsweep on b300-dsxe: 1P+1D healthy, Mooncake EFA transport up, AgentX benchmark + eval complete.中文
新增 #3405 的 1P1D 配置:Kimi-K3 FP4,B300 Dynamo-vLLM,1P (TP8/DCP8) × 1D (TP8/DCP8),MTP/DSpark4,通过 Mooncake 进行 DRAM KV 卸载,并发 48。使用新的 EFA 镜像
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1。ghcr.io#…写法以匹配预导入的 squash 镜像;LD_LIBRARY_PATH指向镜像内的/opt/amazon/efa/lib;删除kimik3-b300-efa-setup.sh。runners/launch_b300-dsxe.sh增加 enroot 导入重试,与 [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX #3521 / [Klaud Cold] Kimi-K3 B300 Mooncake EFA AgentX 1P3D disagg on the EFA vLLM image / [Klaud Cold] 在 EFA vLLM 镜像上运行 Kimi-K3 B300 Mooncake EFA AgentX 1P3D 分离式 #3522 相同。AI disclosure: authored with Claude Code (Claude Opus 5.5).
🤖 Generated with Claude Code