Skip to content

[Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 with Mooncake store hugepages / [Klaud Cold] 新增使用 Mooncake store 大页的 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 - #3557

Open
functionstackx wants to merge 3 commits into
mainfrom
klaud/kimik3-b300-mooncake-efa-disagg-1p1d-hugepage

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Summary

Hugepage variant of #3523's 1P1D arm (Kimi-K3 FP4, B300 Dynamo-vLLM, 1P (TP8/DCP8) × 1D (TP8/DCP8), DSpark4, Mooncake DRAM KV offload over EFA, c48, ghcr.io/semianalysisai/vllm-openai:efa_pr_58768 with mooncake 0.3.13.post1). It is ported onto the current inferencex-e2e/ layout and starts the Mooncake store with 2 MB hugepages.

⚠️ Expected risk

In Mooncake 0.3.13.post1, MC_STORE_USE_HUGEPAGE is strict: if the segment can't be mapped from HugeTLB, store setup fails with no fallback to regular pages. On 2026-09-28, dsxe-sa-b300-prd0-gpu-00 had vm.nr_hugepages = 21121 × 2 MB ≈ 41 GiB reserved, well below the ~194 GB this recipe asks for. Unless the B300 nodes' hugepage reservation is raised (needs root on the host), expect worker startup to fail at Mooncake store init. This sweep is meant to confirm that.

Test plan

  • full-sweep-enabled sweep on b300-dsxe: Mooncake store segment allocated from hugepages (Using huge pages in the worker logs), 1P+1D healthy, AgentX benchmark + eval complete.
中文

摘要

#3523 中 1P1D 配置的大页变体:Kimi-K3 FP4,B300 Dynamo-vLLM,1P (TP8/DCP8) × 1D (TP8/DCP8),DSpark4,通过 EFA 进行 Mooncake DRAM KV 卸载,并发 48,镜像为 ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(mooncake 0.3.13.post1)。已迁移到当前的 inferencex-e2e/ 目录布局,并以 2 MB 大页启动 Mooncake store。

⚠️ 预期风险

Mooncake 0.3.13.post1 中 MC_STORE_USE_HUGEPAGE 为严格模式:若无法从 HugeTLB 分配,store 初始化直接失败,不会回退到普通页。2026-09-28 时 dsxe-sa-b300-prd0-gpu-00 只预留了约 41 GiB 大页(21121 × 2 MB),远低于本配方所需的约 194 GB。除非提高 B300 节点的大页预留(需要主机 root 权限),否则 worker 预计会在 Mooncake store 初始化时失败,本次 sweep 用于确认这一点。

AI disclosure: authored with Claude Code (Claude Opus 5.5).

🤖 Generated with Claude Code

Port #3523's 1P1D EFA arm onto the inferencex-e2e layout and run the
Mooncake store with MC_STORE_USE_HUGEPAGE=1 / MC_STORE_HUGEPAGE_SIZE=2097152
on the mooncake-master service and the prefill/decode workers.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, I checked two other candidates around the enroot retry-loop change: the ${ENROOT_MAX_CONNECTIONS:-4} fallback is an internal retry-tuning knob rather than caller-supplied config, so AGENTS.md's no-fallback-default rule doesn't clearly apply; and the added backoff sleeps (up to 180s total) are negligible against the existing 120-minute SLURM --time limit on the import step, so they don't meaningfully change the mid-retry timeout or per-file flock-contention risk for other callers.

Extended reasoning...

The diff adds a new Kimi-K3 B300 recipe/master-config pair, a perf-changelog entry, and a retry loop in the enroot-import helper; none of it touches auth, crypto, or permission logic. Two confirmed findings (perf-changelog pr-link left as "TBD" instead of the CI-required placeholder, and a router-version mismatch between the new master-config entry and the recipe it points at) are being posted as inline comments and are substantive enough to block approval. I additionally traced the SLURM --time budget and flock-wait interaction with the new retry loop, and the ENROOT_MAX_CONNECTIONS fallback-default question, and found neither introduces a new bug beyond what's already flagged.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🔴 inferencex-e2e/perf-changelog.yaml — The new changelog entry sets pr-link: TBD instead of the required XXX placeholder, so this PR's own check-changelog CI job (validate_perf_changelog.py) fails and blocks the automated merge path. validate_added_pr_link in inferencex-e2e/infx/workflows/validate_perf_changelog.py:139-150 only accepts the canonical .../pull/<n> link or a value in PR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"}; TBD is neither, so compare_entries raises ChangelogValidationError("new PR entry must use ... or an XXX placeholder; found 'TBD'"). This also breaks merge_with_reuse.py's automatic canonicalization, which only rewrites appended XXX links (per docs/ci-procedures.md:422). …

    Why this was flagged

    …Fix: change the entry's pr-link to XXX (matching every other pending entry's convention) so CI validation and the reuse-merge canonicalization step both succeed.

    Trigger: this PR appends a changelog entry to inferencex-e2e/perf-changelog.yaml with pr-link: TBD (added at the tail of the file, around line 9024). Entry point: the check-changelog job in .github/workflows/run-sweep.yml runs python -m infx.workflows.validate_perf_changelog with --base-ref/--head-ref, which calls compare_entries -> validate_added_pr_link (inferencex-e2e/infx/workflows/validate_perf_changelog.py:139-150). That function only accepts the canonical PR link or a value in PR_LINK_PLACEHOLDERS ({"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"}); TBD matches neither, so it raises ChangelogValidationError and the check fails. On the base branch every other pending/new entry uses XXX, which this same check accepts and merge_with_reuse.py's canonicalize_appended_links rewrites to the real PR link before merge; TBD gets neither treatment, so the PR is stuck failing CI or…

    Verification: normal (with a corrected attribution): The new changelog entry ends with pr-link: TBD (perf-changelog.yaml tail, ~line 9024). The rejecting logic is real and reachable. PR_LINK_PLACEHOLDERS (validate_perf_changelog.py:21-25) contains only "XXX" and ".../pull/XXX"; validate_added_pr_link (lines 139-150) rejects any pr-link that is neither the canonical .../pull/<n> link nor a…

runner: cluster:b300-dsxe
precision: fp4
framework: dynamo-vllm
router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) nvidia-master.yaml:1540 records router version "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" for the new kimik3-fp4-b300-...-hugepage config key, but the recipe it points at pins a different dynamo build. disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml:11 and :244 set dynamo.source.rev / identity.frameworks.dynamo to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b" instead. Every sibling kimik3 vllm agentx recipe (gb300/gb200 disagg and agg variants) uses "ba83080ec..." in both the recipe and the matching master-config router.version, so this is the only entry where the two diverge. Anyone using the master config to attribute a sweep's results to a dynamo build gets pointed at the wrong commit. …

Why this was flagged

…Fix: set router.version here to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b" to match the recipe's actual dynamo.source.rev (or revert the recipe's rev if the bump was unintentional).

The new master-config entry kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d-hugepage at nvidia-master.yaml:1531-1560 declares router: { version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } at line 1540, reached via the CONFIG_FILE additional-setting at line 1559 pointing to the new recipe. That recipe (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml:8-11,244) actually pins dynamo.source.rev to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b", a different commit not used by any other kimik3 recipe in the repo. All comparable sibling entries (e.g. kimik3-fp4-gb300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg at nvidia-master.yaml:7832-7860, paired with gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml:14-21) keep router.version and dynamo.source.rev identical, showing the master config's router.version field is meant to mirror the recipe's actual dynamo build.…

Verification: nit. Factually accurate discrepancy: nvidia-master.yaml:1540 records router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } for the new key, while the recipe it points at (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml) pins the dynamo build to a different commit in both places: line 11 rev: cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b and line 244…

@github-actions

Copy link
Copy Markdown
Contributor

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants