Skip to content

[Debug] Native Kimi-K3 FP32 PD on official nightly / [Debug] 基于官方 nightly 的原生 Kimi-K3 FP32 PD - #3582

Draft
YukioZzz wants to merge 6 commits into
mainfrom
yichaozhu/k3-pd-native-fp32-0929
Draft

YukioZzz wants to merge 6 commits into
mainfrom
yichaozhu/k3-pd-native-fp32-0929

Conversation

@YukioZzz

@YukioZzz YukioZzz commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Integrate native Kimi-K3 prefill/decode disaggregation with the current InferenceX Python launcher, including the launcher migration in #3576 and the staged-workspace correction in #3600. Platform resources live in configs/runners.yaml; image preparation uses infx/launch/. There is no restored Bash launcher or separate submission/lifecycle implementation.

The configuration uses MXFP4 target weights, FP32 SSM state, FP8 KV and GMU0.90. Existing topology and concurrency variants are retained. Throughput uses automatic golden acceptance selection; eval uses real verification. The high-concurrency prefill configuration retains HSA_NO_SCRATCH_RECLAIM=0.

Commit structure

  • Framework: declare the required model, draft, fabric and memlock resources; validate the selected recipe's worker image against the matrix and stage its router image through the shared backend. Reuse main's setup, submission, cancellation and artifact collection.
  • Configuration: pair the master config and native recipe with the same immutable official ROCm nightly; include synchronized documentation and append-only performance changelog entries.
  • Temporary debug integration: apply the missing upstream [frontend] data point hover modals "stick" on click #508/#57700 prerequisites, only the independent-heartbeat subset of #58968, and the scoped Ionic provider compatibility mount described below. No shared-MR patch, BF16 development patch, READ-credit mitigation or router-algorithm change is added. Dependency repository URLs and the AIPerf source are unchanged.

Official nightly and removable prerequisites

The worker is vllm/vllm-openai-rocm:nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89, pinned to amd64 digest sha256:50b4f2aec13ffcb11ed846fec049d72250a9c2421eb1ad24c95a5210d46cee4c. Registry and source checks on 2026-09-30 confirm that it includes vLLM#56861 and vLLM#58814, but predates merged vLLM#57700.

  • srt-slurm#508: fetch commit 51cee8904a0b402a834887a26008adb79b8cd26b and its parent, then git cherry-pick --no-commit into this workload's disposable job clone. The checked-in dependency remains unchanged.
  • vLLM#57700: download the exact merged commit 30ae8248b663ea56398d44d30c0d8276a598c4a0 diff, verify its pinned SHA256, check applicability and apply only vllm/* through the recipe's temporary setup script. The official image contains an installed wheel rather than a Git checkout, so this is a Python-only backport, without replacing compiled extensions. Native srt-slurm also runs this setup in the frontend container; only the standalone router image is exempt from engine patching. Worker package lookup, download, checksum or application failure stops startup.

The setup also applies the two-file independent-heartbeat subset from vLLM#58968, pinned at 9cb82832079ede6638b5a5f35ba0d60d203f57a0. The checked-in subset has SHA256 5eaf72c56f3dc7e5bc85b76df11c0c335f7935b492b3271cde28fcd1dfd9683f; it moves rank-zero registration sends into an independent process and adds bounded shutdown/parent-death cleanup. It preserves payloads and intervals and excludes that PR's request-completion, shared-MR, parser and transport changes. It is intentionally not a full cherry-pick of the mixed PR.

Once the official nightly contains #57700 and the independent-heartbeat fix, update both worker-image references and remove the corresponding backports, setup script, patch, associated script regression tests and recipe setting. Independently advance the srt-slurm pin to include #508 and remove its job-local cherry-pick. A worker-image update cannot update srt-slurm. These removals leave the framework/configuration changes; normal runtime qualification and review still apply.

The pinned image's Ionic provider rejects kernel ABI4. The temporary layer therefore restores a single-file read-only mount from the cluster-declared ionic-provider volume, selected only by the Kimi-K3 vLLM PD lane. It leaves the image's libibverbs core and all unrelated providers intact. Remove this mount only when the replacement image opens every expected RDMA device without it. This compatibility requirement is independent of shared-MR and is not an engine patch.

Validation

  • The local CI-equivalent CPU suite passes: 2361 passed, 1 skipped. Ruff lint/formatting and setup-script syntax checks pass.
  • The actual matrix, native schema, worker/router command generation, Python launch entrypoint, result staging and rejected-submission path were exercised locally. Only external scheduler, image import and installation operations were stubbed. Launcher regressions now use separate job-local Python environments without exposing srtctl to the launcher. An additional check with a real job-local srt-slurm installation resolves the requested worker and frontend images successfully.
  • Both online prerequisites were applied to their exact source baselines. The #57700-modified Python files match the merged upstream blobs; [frontend] data point hover modals "stick" on click #508's dedicated discovery-template tests pass. Historical performance-changelog bytes are preserved.
  • A separate CPU-only cluster job used a cold private cache and the exact pinned router image. The old registry URI reproduced the JSON error; the corrected URI imported successfully. Inside the actual container, the old setup reproduced the missing-vLLM failure and the corrected setup passed. Router HTTP and discovery listeners started; /health returned the expected 503 waiting for absent workers. SIGTERM produced exit code 0, both ports closed, and the allocation completed successfully. No P/D workers or benchmark traffic were started.
  • Native make setup ARCH=x86_64 also completed locally with real dependency downloads into a fresh private directory, separate from the mocked installation acceptance.
  • A current-image CPU A/B reproduced the kernel ABI rejection without the provider mount; with only the read-only file mount, all eight RDMA devices opened and closed successfully. The image's core library hash was unchanged, the mounted provider hash matched its host source, and the supplied fabric environment survived the container boundary. The idle router also started and exited successfully with the same mount. The owned allocation completed in 53 seconds with no GPUs. This qualifies device/provider compatibility, not MR registration, traffic, throughput or loaded-worker shutdown.

The first manual 1P2D c48 run stopped before image import or server startup because recipe resolution imported srtctl in the launcher interpreter instead of the job-local environment. The framework commit now resolves images through the job-local Python, following the existing binder/submission pattern. The isolated-interpreter regression reproduces the original failure and passes with the correction, including image-mismatch rejection before import.

The subsequent run imported the worker image but stopped before benchmark submission when the shared resolver sent the router manifest request to the Docker Hub website alias. Explicit Docker Hub hosts now resolve to the registry API without changing image digests or cache keys. The setup-scope correction prevents the next startup failure in the standalone router. Neither failure was a P/D engine crash, and neither run produced throughput results. The cluster check above validates these preparation boundaries, not a full GPU sweep or its shutdown path.

This remains a draft debug integration; GPU qualification is pending. The earlier c48 result used a different harness revision and custom image and is not qualification of this candidate. Full sweep, applicable evals, upstream-image policy and review remain required before merge.

The later GPU attempt loaded the model but failed on all workers at Ionic provider initialization, before traffic. The nightly migration had removed the previously validated provider mount before qualifying the replacement image's ABI. That omission is corrected by the scoped mount and current-image A/B above. The failed run produced no throughput result and released its allocation; one decode step needed teardown escalation. End-to-end qualification remains pending.

The next attempt passed provider initialization and worker health, but prefill discovery expired during warmup, producing HTTP 503 before profiling. The allocation was released. The narrow heartbeat backport addresses that observed failure mechanism: CPU A/B with a 3-second send interval and an 8-second worker GIL hold measured maximum gaps of 8.009 seconds for the old thread and 3.003 seconds for the independent process. Real-socket lifecycle checks cover both role payloads, repeated shutdown and abrupt owner death. The production setup was also executed against a clean pinned-nightly source export, and its installed output matched the qualified subset. This is not complete GPU worker construction or an end-to-end pass; replacement 1P2D c48 fast qualification is pending.

中文

将原生 Kimi-K3 prefill/decode 分离接入当前 InferenceX Python launcher,包含 #3576 的入口迁移和 #3600 的 workspace 打包修复。平台资源声明放在 configs/runners.yaml,镜像准备使用 infx/launch/;不恢复旧 Bash launcher,也不另建提交或生命周期管理路径。

配置使用 MXFP4 主模型权重、FP32 SSM 状态、FP8 KV 和 GMU0.90,保留原有拓扑与并发配置。吞吐测试自动选择 golden acceptance,eval 使用真实校验。高并发 prefill 保留 HSA_NO_SCRATCH_RECLAIM=0。

提交分为框架、配置、临时 debug 集成三层。框架声明模型、draft、fabric、memlock 等资源,导入前校验 worker 镜像与矩阵一致,并通过共享 backend 准备 recipe 指定的 router 镜像;其余安装、提交、取消和结果收集沿用 main。配置同步 master/recipe 的官方镜像标识、双语文档及只追加的性能记录。临时层还保留仅作用于 Kimi-K3 vLLM PD 的 Ionic provider 兼容挂载,不新增 shared-MR、BF16 开发补丁、READ-credit 缓解或 router 算法修改;依赖仓库地址及 AIPerf 源码不变。

官方 ROCm nightly 的 tag/digest 如上。截至 2026-09-30,已确认它包含 vLLM#56861/#58814,但尚不含已合入的 #57700。临时层补齐缺失的上游前置能力:#508 按固定 commit 在任务私有 srt-slurm clone 中执行 git cherry-pick --no-commit;#57700 按固定合入 commit 下载并校验 diff,仅应用 vllm/* 的 Python 修改,不替换编译扩展。官方镜像安装的是 wheel,因此后者不是对镜像内 Git 仓库执行 cherry-pick。原生 srt-slurm 也会在 frontend 容器执行 setup,仅独立 router 镜像跳过 engine 补丁;worker 包查找、下载、校验或应用失败仍终止启动。

同一 setup 还应用 vLLM#58968 在 9cb82832079ede6638b5a5f35ba0d60d203f57a0 的两文件独立心跳子集,固定 SHA256 为 5eaf72c56f3dc7e5bc85b76df11c0c335f7935b492b3271cde28fcd1dfd9683f。仅把 rank-zero 注册发送放入独立进程,并接入有界退出和父进程死亡清理;payload 与间隔不变,不携带该 PR 的请求完成、shared-MR、parser 或传输修改,不对整个混合 PR 做 cherry-pick。

nightly 包含 #57700 和独立心跳修复后,同步更新两个 worker 镜像引用并删除对应 backport、setup script、patch、脚本回归测试及 recipe 引用。srt-slurm 需要独立升级 pin,包含 #508 后再删除任务内 cherry-pick;换 worker 镜像不会更新 srt-slurm。清理后保留框架和配置层,但仍须完成常规运行验证和审查。

与 CI 对齐的本地 CPU 回归为 2361 passed、1 skipped,Ruff、格式和 setup script 语法检查通过。其他本地验证覆盖真实矩阵、native schema、worker/router 命令生成、Python 入口、结果收集及调度拒绝路径;仅调度器、镜像导入和安装等外部操作使用模拟。Launcher 回归现使用独立的任务 Python 环境,不向 launcher 暴露 srtctl;另以实际安装的任务私有 srt-slurm 验证了 worker 和 frontend 镜像解析。两个在线补丁均在精确源码基线上应用成功,#57700 修改后的文件与上游合入版本一致,#508 专项测试通过,历史 changelog 字节不变。

独立 CPU 集群作业使用私有空缓存和相同的固定 router 镜像:旧 registry URI 复现 JSON 错误,修正后成功导入;在真实容器中,旧 setup 复现缺少 vLLM 的错误,修正后通过。Router 的 HTTP 和 discovery 监听均启动,/health 返回等待未注册 worker 的预期 503;SIGTERM 后返回码为 0,两个端口关闭,allocation 正常结束。没有启动 P/D 或发送 benchmark 流量。另在全新本地私有目录中完成原生 make setup ARCH=x86_64 的真实依赖下载,与模拟安装验收分开记录。

首次手动 1P2D c48 作业在镜像导入和 server 启动前失败:recipe 解析错误地在 launcher 解释器中导入任务私有环境的 srtctl。框架 commit 已沿用现有 binder/submission 模式,改为通过任务私有 Python 解析。独立解释器回归能复现旧故障,并在修复后通过,包含镜像不匹配时在导入前拒绝的路径。

后续作业导入了 worker 镜像,但共享解析器将 router manifest 请求发到 Docker Hub 网站别名,因而在 benchmark 提交前失败。现将显式 Docker Hub 主机名规范化为 registry API,镜像 digest 和缓存键不变;setup 作用范围修正也避免了独立 router 紧接着会遇到的启动错误。两次失败均不是 P/D engine 崩溃,均未产生吞吐结果。上述集群检查只证明准备阶段和 router 启动正常,不代表完整 GPU sweep 或其收尾已通过。

本 PR 保持 draft debug 状态,GPU 验收仍待完成。之前的 c48 结果使用不同 harness 和自定义镜像,不能作为本版本验收。合入前仍需完整 sweep、适用 eval、上游镜像政策检查和人工 review。

后续 GPU 作业完成模型加载,但所有 worker 都在 Ionic provider 初始化时拒绝内核 ABI4,尚未产生实际传输或吞吐结果。迁移 nightly 时,在验证新镜像 ABI 前过早删除了之前已验证的 provider 挂载;现在通过集群声明的 ionic-provider 卷,仅为 Kimi-K3 vLLM PD 恢复单文件只读挂载,不替换镜像内的 libibverbs 核心库或其他 provider。这与 shared-MR 无关,也不是 engine 补丁。只有替换镜像在不挂载的情况下成功打开全部预期 RDMA 设备,才可删除。

当前镜像的 CPU A/B 已复现无挂载时的 ABI 错误,并验证单文件挂载后 8 个设备全部打开、关闭成功。核心库哈希不变,挂载后的 provider 哈希与主机源文件一致,显式传入的 fabric 环境在容器内保持正确;带同一挂载的 idle router 也正常启动、退出。自有作业在 53 秒内完成,未申请 GPU。这不是 MR、实际流量、吞吐或已加载 worker 收尾的验证。之前失败的 GPU allocation 已释放,其中一个 decode step 需要升级终止;端到端验收仍待完成。

再下一次作业通过 provider 初始化和 worker 健康检查,但 warmup 时 prefill discovery 过期,返回 HTTP 503,尚未进入 profile,allocation 已释放。此次最小心跳 backport 针对这一已观察到的故障机制:CPU A/B 采用实际 3 秒间隔和 8 秒 GIL 阻塞,旧线程最大空档 8.009 秒,独立进程 3.003 秒。真实 socket 生命周期验证覆盖两种角色的 payload、重复退出及父进程意外死亡;生产 setup 也在干净的固定 nightly 源码导出上实际执行,安装结果与已验证子集一致。这不代表完整 GPU worker 构造或端到端通过,替换版本的 1P2D c48 fast 验收仍待完成。

@YukioZzz
YukioZzz force-pushed the yichaozhu/k3-pd-native-fp32-0929 branch 2 times, most recently from e44d65b to 5856c97 Compare September 29, 2026 14:23
@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

@YukioZzz
YukioZzz force-pushed the yichaozhu/k3-pd-native-fp32-0929 branch from 5856c97 to 4a8520f Compare September 30, 2026 02:13
@YukioZzz YukioZzz changed the title Native Kimi-K3 FP32 PD harness and five points / 原生 Kimi-K3 FP32 PD 框架与五点配置 [Debug] Native Kimi-K3 FP32 PD on official nightly / [Debug] 基于官方 nightly 的原生 Kimi-K3 FP32 PD Sep 30, 2026
@YukioZzz

YukioZzz commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Thanks for the heads-up. Rebased onto current main (44d6f8ad, including #3576 and #3600) and moved the integration to configs/runners.yaml and infx/launch/. The branch uses the shared Python launch, image-import, submission, cancellation and artifact paths; none of the deleted Bash launchers are restored.

The changes are separated into framework, configuration and temporary upstream-integration commits. The custom worker image is replaced by a digest-pinned official ROCm nightly, without shared-MR or the candidate-specific provider bind. The debug layer applies pinned srt-slurm#508 in the job-local checkout and the Python-only merged vLLM#57700 diff in disposable worker containers; each can be removed when its respective upstream dependency ships.

Updated framework at f9313f45: the first manual c48 run exposed an interpreter-isolation bug before image import or server startup. Recipe-image resolution now runs in the job-local srt-slurm Python, following the existing binder/submission pattern. A fresh-launcher regression reproduces the old failure and passes after the fix, including rejection before import on a worker-image mismatch. The test launcher no longer receives srtctl through PYTHONPATH.

The local CI-equivalent CPU suite passes (2332 passed, 1 skipped), as do Ruff lint/formatting and setup-script syntax checks. Actual matrix/native-schema/worker-command checks and Python-launcher-to-native-srtctl submission, result staging and rejected-submission acceptance also pass, with external scheduler, import and installer operations stubbed. An additional real job-local installation passes image resolution without exposing srtctl to the launcher. Both online prerequisites were applied to their exact source baselines. This remains Draft; GPU qualification of the migrated official-nightly stack is still pending, and the earlier custom-image result is not reused as evidence for this revision.

中文

谢谢提醒。已 rebase 到当前 main(44d6f8ad,包含 #3576 和 #3600),并将集成迁移到 configs/runners.yaml 与 infx/launch/。镜像导入、启动、提交、取消及产物收集均使用共享 Python 路径,没有恢复已删除的 Bash launcher。

提交分为框架、配置和临时上游集成三层。自定义 worker 镜像已替换为固定 digest 的官方 ROCm nightly,不再携带 shared-MR 或候选 provider bind。debug 层仅在任务私有 checkout 中应用固定版本的 srt-slurm#508,并在临时 worker 容器中应用已合入的 vLLM#57700 纯 Python diff;各自的上游依赖发布后即可分别删除。

f9313f45 更新了框架:首次手动 c48 在镜像导入及 server 启动前暴露了解释器隔离问题。Recipe 镜像现沿用 binder/submission 模式,通过任务私有 srt-slurm Python 解析。新的独立 launcher 回归可复现旧故障,修复后通过,并验证了镜像不匹配时在导入前拒绝。测试不再通过 PYTHONPATH 向 launcher 注入 srtctl。

本地 CI 对齐 CPU 回归为 2332 passed、1 skipped,Ruff、格式和 setup script 语法检查通过。真实矩阵、native schema、worker 命令,以及 Python launcher→native srtctl 的提交、结果收集和拒绝提交验收也通过;调度、导入和安装等外部操作使用模拟。另使用实际安装的任务私有环境验证了镜像解析,launcher 本身仍不可导入 srtctl。两个在线补丁均在精确源码基线上应用成功。PR 保持 Draft,迁移后的官方 nightly 栈仍待 GPU 验收,不将旧自定义镜像的结果当作本版本证据。

@YukioZzz
YukioZzz force-pushed the yichaozhu/k3-pd-native-fp32-0929 branch 2 times, most recently from f9313f4 to 6624ff5 Compare September 30, 2026 05:32
迁移 Kimi-K3 PD 所需的集群资源声明和 recipe 镜像准备,沿用 main 的任务提交、取消与产物收集。
使用固定 digest 的官方 ROCm nightly,保留角色调优并移除自定义 shared-MR 镜像及 provider bind;说明临时依赖与后续清理条件。
@YukioZzz
YukioZzz force-pushed the yichaozhu/k3-pd-native-fp32-0929 branch from 6624ff5 to 58ffe72 Compare September 30, 2026 06:29
单独保存可删除的上游 backport:任务内 cherry-pick srt-slurm#508,worker 容器校验并应用 vLLM#57700,以及 #58968 的独立心跳 helper 和生命周期子集;不带该混合 PR 的其余修改。保留已验证的单文件只读 Ionic provider 挂载,直到替换镜像通过无挂载 ABI 验证;FP32、GMU0.90、router 与数据面设置不变。
@YukioZzz
YukioZzz force-pushed the yichaozhu/k3-pd-native-fp32-0929 branch from 58ffe72 to 07eb0cd Compare September 30, 2026 07:15
Backport only the FULL pre-forward context hunk from vLLM #58968 at 7595667abf90c03b546c92189d05e5d73392910e. Keep the existing #57700 and heartbeat backports, serving configuration and benchmark acceptance criteria unchanged.

Local c48 completed 1266 requests with 6 errors (0.474%, below the existing 1% limit), exited 0 and drained both engines. The actual SA setup script produces the same model-runner file as the validated local candidate. Native cloud lifecycle and performance qualification remain separate.

仅恢复 FULL graph replay 前的 KV READ 上下文,保留既有补丁、serving 配置和验收标准;本地完整运行及真实安装检查通过,云侧生命周期与性能仍需独立验证。
Validate the complete pinned #58968 runtime diff with the separately extracted #59164 zeroing fix before reducing the production patch set. Preserve the nightly, merged #57700, existing heartbeat, FP32 SSM and benchmark criteria.

在精简生产补丁前,以固定版本 #58968 完整运行时代码和独立 #59164 zeroing 修复建立实验对照。保留 nightly、已合入的 #57700、既有心跳、FP32 SSM 与 benchmark 判据。
使用最新官方 ROCm nightly 与固定版本 #59164 验证最小 FP32 PD 依赖;移除已合入的 #57700 回补及实验性的 #58968、独立心跳补丁,保留完整对照历史。

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants