Skip to content

[Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 / [Klaud Cold] 新增 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 - #3523

Open
functionstackx wants to merge 1 commit into
mainfrom
klaud/kimik3-b300-mooncake-efa-disagg-1p1d
Open

functionstackx wants to merge 1 commit into
mainfrom
klaud/kimik3-b300-mooncake-efa-disagg-1p1d

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Summary

Adds the 1P1D arm of #3405 (Kimi-K3 FP4, B300 Dynamo-vLLM, 1P (TP8/DCP8) × 1D (TP8/DCP8), MTP/DSpark4, DRAM KV offload via Mooncake, c48) on the new EFA image ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vllm-project/vllm#58768 vllm-openai-efa target, with AWS EFA libfabric 2.6.0amzn1.0 and mooncake-transfer-engine-efa-cuda13 0.3.13.post1 built in).

Sibling PRs: #3521 (agg), #3522 (1P3D c32).

Test plan

  • full-sweep-enabled sweep on b300-dsxe: 1P+1D healthy, Mooncake EFA transport up, AgentX benchmark + eval complete.
中文

新增 #3405 的 1P1D 配置:Kimi-K3 FP4,B300 Dynamo-vLLM,1P (TP8/DCP8) × 1D (TP8/DCP8),MTP/DSpark4,通过 Mooncake 进行 DRAM KV 卸载,并发 48。使用新的 EFA 镜像 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1。

AI disclosure: authored with Claude Code (Claude Opus 5.5).

🤖 Generated with Claude Code

…park4 c48 on efa_pr_58768 image

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@functionstackx
functionstackx force-pushed the klaud/kimik3-b300-mooncake-efa-disagg-1p1d branch from 7bb1635 to 13de142 Compare September 27, 2026 19:48
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment thread perf-changelog.yaml Outdated
description:
- "Add the Kimi-K3 B300 Mooncake EFA AgentX disaggregated 1P1D DCP8/DCP8 DSpark4 configuration at c48 on ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 (vllm-project/vllm#58768 vllm-openai-efa target, with AWS EFA libfabric 2.6.0amzn1.0 and mooncake-transfer-engine-efa-cuda13 0.3.13.post1 built in). Taken from #3405's 1P1D arm; drops its kimik3-b300-efa-setup.sh runtime installer."
- "新增基于 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(vllm-project/vllm#58768 的 vllm-openai-efa 目标,内置 AWS EFA libfabric 2.6.0amzn1.0 与 mooncake-transfer-engine-efa-cuda13 0.3.13.post1)的 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8/DCP8 DSpark4 配置(c48)。取自 #3405 的 1P1D 配置;删除其 kimik3-b300-efa-setup.sh 运行时安装脚本。"
pr-link: PRLINK_PLACEHOLDER

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Maintainers and CI get a changelog entry that will fail validation because pr-link: PRLINK_PLACEHOLDER isn't a value the validator accepts. infx/workflows/validate_perf_changelog.py's validate_added_pr_link only allows the exact expected .../pull/<pr_number> URL or a value in PR_LINK_PLACEHOLDERS ({"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"}); PRLINK_PLACEHOLDER matches neither, so it raises ChangelogValidationError("new PR entry must use ... or an XXX placeholder; found 'PRLINK_PLACEHOLDER'"). Fix: use the literal XXX placeholder (or the full .../pull/XXX form) as documented, not an invented string, so the automated canonicalization in prepare_perf_changelog_merge.py can replace it with the real PR link before merge.

Why this was flagged

The new entry appended at perf-changelog.yaml:8989-8997 sets pr-link: PRLINK_PLACEHOLDER. validate_added_pr_link (infx/workflows/validate_perf_changelog.py:133-145) is invoked on PR runs and only tolerates PR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"} or the exact expected pull URL. PRLINK_PLACEHOLDER is none of these, so validation raises ChangelogValidationError, blocking the changelog check and the automated pr-link canonicalization step (infx/workflows/prepare_perf_changelog_merge.py:103-109) that would otherwise fill in the real link on merge. On base branch, entries use XXX, which passes this check.

Verification: normal (mechanism corrected): The appended entry at perf-changelog.yaml:8997 sets pr-link: PRLINK_PLACEHOLDER, which is not an accepted pr-link value. In infx/workflows/validate_perf_changelog.py, PR_LINK_PLACEHOLDERS = {"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"} (lines 21-24); validate_added_pr_link (lines 134-145) raises ChangelogValidationError for any…

runner: cluster:b300-dsxe
precision: fp4
framework: dynamo-vllm
router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d entry sets router.version to ba83080ecd31c1ce918559e576d3c5bc9e092ff1, but the recipe it points at (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11,237) builds dynamo from a different rev, cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b. Every sibling entry (e.g. nvidia-master.yaml:7841, kimik3-fp4-gb300-...-disagg) keeps router.version equal to its recipe's dynamo rev; this one doesn't. Per docs/results-and-ingestion.md, router ({name, version}) is stored as benchmark result metadata, so results from this scenario get tagged and grouped under the wrong dynamo-router build, silently mixing them into Pareto comparisons meant for a different dynamo revision. …

Why this was flagged

…Fix: set router.version to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b to match the recipe it references.

configs/nvidia-master.yaml:1538 sets router.version to ba83080ecd31c1ce918559e576d3c5bc9e092ff1 for the new kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d key. The recipe it references, benchmarks/multi_node/srt-slurm-recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11 and :237, sets the dynamo rev actually built into the image to cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b, a different commit. infx/matrix/validation.py's ComponentMetadata schema for router has no check comparing it to the recipe's dynamo rev, so nothing catches the mismatch. docs/results-and-ingestion.md documents router as {name, version} metadata stored with benchmark results, so results are recorded under the wrong dynamo-router version. Sibling entries like nvidia-master.yaml:7841 keep these values equal, showing this is the convention the new entry breaks.

Verification: Severity: normal — this new entry records wrong dynamo/router provenance. configs/nvidia-master.yaml:1538 sets router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" }, but the recipe it references builds a different dynamo commit: disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:11 dynamo.source.rev: cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b (with `install:…

disagg: true
scenarios:
agentic-coding:
- dram-utilization: 0.75

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Anyone reading this scenario's recorded allocated_cpu_dram_gb metric gets a value ~12x too large, corrupting DRAM-offload capacity/efficiency comparisons for this benchmark. dram-utilization: 0.75 here computes total-cpu-dram-gb ≈ 2249 GB (full-node fraction for the 8-GPU prefill worker), which becomes TOTAL_CPU_DRAM_GB and is stored verbatim as allocated_cpu_dram_gb, but the recipe's actual centralized Mooncake pool is only global_segment_size: 190GB (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:40). Every sibling entry sizes dram-utilization to match its recipe's segment size (e.g. gb300's 0.1775 -> 160GB at nvidia-master.yaml:7849); this one looks copy-pasted from the single-node b300 kimik3 recipe's 0.75/2249GB pairing. …

Why this was flagged

…Fix: pick dram-utilization so computed total-cpu-dram-gb matches the recipe's real DRAM budget (~0.063 here for 190GB), per the convention every other multinode entry follows.

configs/nvidia-master.yaml:1544 sets dram-utilization: 0.75 under kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d. infx/matrix/generate.py's agentic_dram_offload_gb (lines 429-450) uses the prefill worker's full-node GPU fraction (tp=8 on an 8-GPU node -> fraction 1) to turn that into total_cpu_dram_gb ≈ 2249, coincidentally equal to the unrelated single-node recipe's hardcoded TOTAL_CPU_DRAM_GB=2249 (benchmarks/single_node/srt-slurm-recipes/kimik3/vllm/b300-fp4-mtp/agentic.yaml:109). benchmark-multinode-tmpl.yml:178 wires total-cpu-dram-gb into TOTAL_CPU_DRAM_GB; benchmark_lib.sh:212/406 only checks it's a positive integer, so the mismatch is never caught. infx/results/agentic/init.py:191 records it as allocated_cpu_dram_gb, while the recipe's real capacity is global_segment_size: 190GB (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-c48.yaml:40) - about 12x smaller than what gets recorded.

Verification: normal (metric/data-correctness defect in newly added scenario; no serving impact). configs/nvidia-master.yaml:1544 sets dram-utilization: 0.75 for kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d. infx/matrix/generate.py agentic_dram_offload_gb (lines 446-470) computes min(runner DRAM, MAX_AGENTIC_AVAILABLE_CPU_DRAM_MIB=2,861,022 MiB ≈ 2999 GB) × 0.75 × (prefill gpu_count…

@github-actions

Copy link
Copy Markdown
Contributor

@whn09

whn09 commented Sep 29, 2026

Copy link
Copy Markdown

Hi, I maintain Mooncake on EFA. I looked into the slow registration and reproduced it on 2× p5en with the same setup (8 ranks × 190GB Store segment per node).

Root cause

EFA itself isn't slow. The problem is how the Store segment gets its huge pages.

  • EFA has a per-NIC page-table budget (~24M entries), so a 190GB segment has to be backed by 2MB pages. With 4KB pages, fi_mr_reg fails with ENOMEM.
  • With protocol: efa and no MC_STORE_USE_HUGEPAGE, the segment is a plain aligned_alloc. It relies on GLIBC_TUNABLES=glibc.malloc.hugetlb=1 to get THP at fault time, during ibv_reg_mr.
  • Once a NUMA node fills up, every 2MB fault triggers direct compaction, and about 99% of those fail. On p5en this pushed registration from about 35 s to 38 min, which matches the 12–23+ min seen here.
  • RDMA (mlx5) doesn't hit this because it can register 4KB pages.

Fix (config only): pre-reserve 2MB hugetlb pages on the host and let Mooncake use them

env:
  MC_STORE_USE_HUGEPAGE: '1'   # 2MB pages by default
  # GLIBC_TUNABLES is no longer needed for the Store segment
# on each node, as root, before launching (host sysctl; containers inherit it)
sudo sysctl -w vm.nr_hugepages=850000   # ~1.66 TiB

Sizing: ranks per node × (global_segment_size + local_buffer_size) / 2MB, plus at least 5–10% headroom. For this recipe that's 8 × 194GiB ≈ 795k pages. 800k was just exhausted in my run, so use about 850k. If the node can't spare ~1.66 TiB, lower global_segment_size.

Reserve the pages early, ideally via /etc/sysctl.d or the hugepages= boot parameter. A fragmented node may not be able to hand out that many 2MB pages later.

Result: with this setup, 8 × 190GB Store setup took 26–28 s end-to-end on p5en (populate + registration), with zero compaction.

Investigation and write-up done with AI assistance (Claude Code); numbers are from real runs on p5en.

@functionstackx

Copy link
Copy Markdown
Collaborator Author

thanks @whn09

@nvpohanh & @xinli-sw any thoughts on changing this cluster to have ~1.66TB of huge pages? sudo sysctl -w vm.nr_hugepages=850000

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

@whn09

whn09 commented Sep 30, 2026

Copy link
Copy Markdown

Quick update: the fix for the slow EFA segment registration is merged in Mooncake (kvcache-ai/Mooncake#4393). The Store segment is now allocated as 2MB-aligned mmap + MADV_HUGEPAGE + MPOL_INTERLEAVE, so registration no longer waits on direct compaction when a NUMA node is fragmented. It also no longer needs GLIBC_TUNABLES=glibc.malloc.hugetlb=1.

On p5en (8 ranks × 190GB, node0 fragmented), registration went from 1306s to 40s.

The fix will ship in the next Mooncake release. The current image's mooncake-transfer-engine-efa-cuda13 0.3.13.post1 doesn't include it. If you need it sooner, build the latest Mooncake main with -DUSE_EFA=ON -DUSE_CUDA=ON.

The fix requires THP not set to never (/sys/kernel/mm/transparent_hugepage/enabled). The hugetlb reservation I suggested earlier (MC_STORE_USE_HUGEPAGE=1 + vm.nr_hugepages) still works as an alternative.

I've only verified this on p5en so far. If you try it on B300, please share the registration time and any fi_mr_reg errors.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants