[Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 with Mooncake store hugepages / [Klaud Cold] 新增使用 Mooncake store 大页的 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 - #3557
functionstackx wants to merge 3 commits into
Conversation
Port #3523's 1P1D EFA arm onto the inferencex-e2e layout and run the Mooncake store with MC_STORE_USE_HUGEPAGE=1 / MC_STORE_HUGEPAGE_SIZE=2097152 on the mooncake-master service and the prefill/decode workers. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
…epage Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
There was a problem hiding this comment.
Beyond the inline findings, I checked two other candidates around the enroot retry-loop change: the ${ENROOT_MAX_CONNECTIONS:-4} fallback is an internal retry-tuning knob rather than caller-supplied config, so AGENTS.md's no-fallback-default rule doesn't clearly apply; and the added backoff sleeps (up to 180s total) are negligible against the existing 120-minute SLURM --time limit on the import step, so they don't meaningfully change the mid-retry timeout or per-file flock-contention risk for other callers.
Extended reasoning...
The diff adds a new Kimi-K3 B300 recipe/master-config pair, a perf-changelog entry, and a retry loop in the enroot-import helper; none of it touches auth, crypto, or permission logic. Two confirmed findings (perf-changelog pr-link left as "TBD" instead of the CI-required placeholder, and a router-version mismatch between the new master-config entry and the recipe it points at) are being posted as inline comments and are substantive enough to block approval. I additionally traced the SLURM --time budget and flock-wait interaction with the new retry loop, and the ENROOT_MAX_CONNECTIONS fallback-default question, and found neither introduces a new bug beyond what's already flagged.
Findings marked 🟡 are optional suggestions and need no follow-up push.
Additional findings (outside the current diff — GitHub can't attach inline comments there):
-
🔴
inferencex-e2e/perf-changelog.yaml— The new changelog entry setspr-link: TBDinstead of the requiredXXXplaceholder, so this PR's owncheck-changelogCI job (validate_perf_changelog.py) fails and blocks the automated merge path.validate_added_pr_linkin inferencex-e2e/infx/workflows/validate_perf_changelog.py:139-150 only accepts the canonical.../pull/<n>link or a value inPR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"};TBDis neither, socompare_entriesraisesChangelogValidationError("new PR entry must use ... or an XXX placeholder; found 'TBD'"). This also breaksmerge_with_reuse.py's automatic canonicalization, which only rewrites appendedXXXlinks (per docs/ci-procedures.md:422). …Why this was flagged
…Fix: change the entry's
pr-linktoXXX(matching every other pending entry's convention) so CI validation and the reuse-merge canonicalization step both succeed.Trigger: this PR appends a changelog entry to inferencex-e2e/perf-changelog.yaml with
pr-link: TBD(added at the tail of the file, around line 9024). Entry point: thecheck-changelogjob in .github/workflows/run-sweep.yml runspython -m infx.workflows.validate_perf_changelogwith--base-ref/--head-ref, which callscompare_entries->validate_added_pr_link(inferencex-e2e/infx/workflows/validate_perf_changelog.py:139-150). That function only accepts the canonical PR link or a value inPR_LINK_PLACEHOLDERS({"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"});TBDmatches neither, so it raises ChangelogValidationError and the check fails. On the base branch every other pending/new entry usesXXX, which this same check accepts andmerge_with_reuse.py'scanonicalize_appended_linksrewrites to the real PR link before merge;TBDgets neither treatment, so the PR is stuck failing CI or…Verification: normal (with a corrected attribution): The new changelog entry ends with
pr-link: TBD(perf-changelog.yaml tail, ~line 9024). The rejecting logic is real and reachable.PR_LINK_PLACEHOLDERS(validate_perf_changelog.py:21-25) contains only"XXX"and".../pull/XXX";validate_added_pr_link(lines 139-150) rejects any pr-link that is neither the canonical.../pull/<n>link nor a…
| runner: cluster:b300-dsxe | ||
| precision: fp4 | ||
| framework: dynamo-vllm | ||
| router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } |
There was a problem hiding this comment.
🟡 (optional) nvidia-master.yaml:1540 records router version "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" for the new kimik3-fp4-b300-...-hugepage config key, but the recipe it points at pins a different dynamo build. disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml:11 and :244 set dynamo.source.rev / identity.frameworks.dynamo to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b" instead. Every sibling kimik3 vllm agentx recipe (gb300/gb200 disagg and agg variants) uses "ba83080ec..." in both the recipe and the matching master-config router.version, so this is the only entry where the two diverge. Anyone using the master config to attribute a sweep's results to a dynamo build gets pointed at the wrong commit. …
Why this was flagged
…Fix: set router.version here to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b" to match the recipe's actual dynamo.source.rev (or revert the recipe's rev if the bump was unintentional).
The new master-config entry kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d-hugepage at nvidia-master.yaml:1531-1560 declares router: { version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } at line 1540, reached via the CONFIG_FILE additional-setting at line 1559 pointing to the new recipe. That recipe (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml:8-11,244) actually pins dynamo.source.rev to "cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b", a different commit not used by any other kimik3 recipe in the repo. All comparable sibling entries (e.g. kimik3-fp4-gb300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg at nvidia-master.yaml:7832-7860, paired with gb300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake.yaml:14-21) keep router.version and dynamo.source.rev identical, showing the master config's router.version field is meant to mirror the recipe's actual dynamo build.…
Verification: nit. Factually accurate discrepancy: nvidia-master.yaml:1540 records router: { name: dynamo-router, version: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" } for the new key, while the recipe it points at (disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml) pins the dynamo build to a different commit in both places: line 11 rev: cfada2fd9d17bfa6bb68dbee9d2f455e12577b8b and line 244…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36503454568 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36503454568 |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Summary
Hugepage variant of #3523's 1P1D arm (Kimi-K3 FP4, B300 Dynamo-vLLM, 1P (TP8/DCP8) × 1D (TP8/DCP8), DSpark4, Mooncake DRAM KV offload over EFA, c48,
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768with mooncake 0.3.13.post1). It is ported onto the currentinferencex-e2e/layout and starts the Mooncake store with 2 MB hugepages.kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d-hugepage→recipes/kimik3/vllm/b300-fp4/agentx/disagg-1p1d-dcp8-dcp8-dspark4-mooncake-hugepage-c48.yaml.MC_STORE_USE_HUGEPAGE: '1'andMC_STORE_HUGEPAGE_SIZE: '2097152'are set on themooncake-masterserviceenvand onroles.prefill.env/roles.decode.env. Withmode: embeddedthe 190 GB global segment and 4 GB local buffer are allocated inside the vLLM workers (srt-slurm rendersstore_configinto the workers'MOONCAKE_CONFIG_PATHJSON), so the worker env is what backs the segment with hugepages.inferencex-e2e/runners/launch_b300-dsxe.sh: [Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 / [Klaud Cold] 新增 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 #3523's enroot import retry.In Mooncake 0.3.13.post1,
MC_STORE_USE_HUGEPAGEis strict: if the segment can't be mapped from HugeTLB, store setup fails with no fallback to regular pages. On 2026-09-28,dsxe-sa-b300-prd0-gpu-00hadvm.nr_hugepages = 21121× 2 MB ≈ 41 GiB reserved, well below the ~194 GB this recipe asks for. Unless the B300 nodes' hugepage reservation is raised (needs root on the host), expect worker startup to fail at Mooncake store init. This sweep is meant to confirm that.Test plan
full-sweep-enabledsweep on b300-dsxe: Mooncake store segment allocated from hugepages (Using huge pagesin the worker logs), 1P+1D healthy, AgentX benchmark + eval complete.中文
摘要
#3523 中 1P1D 配置的大页变体:Kimi-K3 FP4,B300 Dynamo-vLLM,1P (TP8/DCP8) × 1D (TP8/DCP8),DSpark4,通过 EFA 进行 Mooncake DRAM KV 卸载,并发 48,镜像为
ghcr.io/semianalysisai/vllm-openai:efa_pr_58768(mooncake 0.3.13.post1)。已迁移到当前的inferencex-e2e/目录布局,并以 2 MB 大页启动 Mooncake store。kimik3-fp4-b300-dynamo-vllm-agentic-dspark-mooncake-dcp8-disagg-1p1d-hugepage。mooncake-master服务的env以及roles.prefill.env/roles.decode.env中设置MC_STORE_USE_HUGEPAGE: '1'与MC_STORE_HUGEPAGE_SIZE: '2097152'。mode: embedded下,190 GB 全局段与 4 GB 本地缓冲由 vLLM worker 进程分配,因此真正让该段使用大页的是 worker 的环境变量。launch_b300-dsxe.sh沿用 [Klaud Cold] Add Kimi-K3 B300 Mooncake EFA AgentX disagg 1P1D DCP8 DSpark4 c48 / [Klaud Cold] 新增 Kimi-K3 B300 Mooncake EFA AgentX 分离式 1P1D DCP8 DSpark4 c48 配置 #3523 的 enroot 导入重试。Mooncake 0.3.13.post1 中
MC_STORE_USE_HUGEPAGE为严格模式:若无法从 HugeTLB 分配,store 初始化直接失败,不会回退到普通页。2026-09-28 时dsxe-sa-b300-prd0-gpu-00只预留了约 41 GiB 大页(21121 × 2 MB),远低于本配方所需的约 194 GB。除非提高 B300 节点的大页预留(需要主机 root 权限),否则 worker 预计会在 Mooncake store 初始化时失败,本次 sweep 用于确认这一点。AI disclosure: authored with Claude Code (Claude Opus 5.5).
🤖 Generated with Claude Code