Conversation
e44d65b to
5856c97
Compare
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
5856c97 to
4a8520f
Compare
|
Thanks for the heads-up. Rebased onto current main ( The changes are separated into framework, configuration and temporary upstream-integration commits. The custom worker image is replaced by a digest-pinned official ROCm nightly, without shared-MR or the candidate-specific provider bind. The debug layer applies pinned srt-slurm#508 in the job-local checkout and the Python-only merged vLLM#57700 diff in disposable worker containers; each can be removed when its respective upstream dependency ships. Updated framework at The local CI-equivalent CPU suite passes (2332 passed, 1 skipped), as do Ruff lint/formatting and setup-script syntax checks. Actual matrix/native-schema/worker-command checks and Python-launcher-to-native-srtctl submission, result staging and rejected-submission acceptance also pass, with external scheduler, import and installer operations stubbed. An additional real job-local installation passes image resolution without exposing srtctl to the launcher. Both online prerequisites were applied to their exact source baselines. This remains Draft; GPU qualification of the migrated official-nightly stack is still pending, and the earlier custom-image result is not reused as evidence for this revision. 中文谢谢提醒。已 rebase 到当前 main( 提交分为框架、配置和临时上游集成三层。自定义 worker 镜像已替换为固定 digest 的官方 ROCm nightly,不再携带 shared-MR 或候选 provider bind。debug 层仅在任务私有 checkout 中应用固定版本的 srt-slurm#508,并在临时 worker 容器中应用已合入的 vLLM#57700 纯 Python diff;各自的上游依赖发布后即可分别删除。
本地 CI 对齐 CPU 回归为 2332 passed、1 skipped,Ruff、格式和 setup script 语法检查通过。真实矩阵、native schema、worker 命令,以及 Python launcher→native srtctl 的提交、结果收集和拒绝提交验收也通过;调度、导入和安装等外部操作使用模拟。另使用实际安装的任务私有环境验证了镜像解析,launcher 本身仍不可导入 srtctl。两个在线补丁均在精确源码基线上应用成功。PR 保持 Draft,迁移后的官方 nightly 栈仍待 GPU 验收,不将旧自定义镜像的结果当作本版本证据。 |
f9313f4 to
6624ff5
Compare
迁移 Kimi-K3 PD 所需的集群资源声明和 recipe 镜像准备,沿用 main 的任务提交、取消与产物收集。
使用固定 digest 的官方 ROCm nightly,保留角色调优并移除自定义 shared-MR 镜像及 provider bind;说明临时依赖与后续清理条件。
6624ff5 to
58ffe72
Compare
单独保存可删除的上游 backport:任务内 cherry-pick srt-slurm#508,worker 容器校验并应用 vLLM#57700,以及 #58968 的独立心跳 helper 和生命周期子集;不带该混合 PR 的其余修改。保留已验证的单文件只读 Ionic provider 挂载,直到替换镜像通过无挂载 ABI 验证;FP32、GMU0.90、router 与数据面设置不变。
58ffe72 to
07eb0cd
Compare
Backport only the FULL pre-forward context hunk from vLLM #58968 at 7595667abf90c03b546c92189d05e5d73392910e. Keep the existing #57700 and heartbeat backports, serving configuration and benchmark acceptance criteria unchanged. Local c48 completed 1266 requests with 6 errors (0.474%, below the existing 1% limit), exited 0 and drained both engines. The actual SA setup script produces the same model-runner file as the validated local candidate. Native cloud lifecycle and performance qualification remain separate. 仅恢复 FULL graph replay 前的 KV READ 上下文,保留既有补丁、serving 配置和验收标准;本地完整运行及真实安装检查通过,云侧生命周期与性能仍需独立验证。
Validate the complete pinned #58968 runtime diff with the separately extracted #59164 zeroing fix before reducing the production patch set. Preserve the nightly, merged #57700, existing heartbeat, FP32 SSM and benchmark criteria. 在精简生产补丁前,以固定版本 #58968 完整运行时代码和独立 #59164 zeroing 修复建立实验对照。保留 nightly、已合入的 #57700、既有心跳、FP32 SSM 与 benchmark 判据。
使用最新官方 ROCm nightly 与固定版本 #59164 验证最小 FP32 PD 依赖;移除已合入的 #57700 回补及实验性的 #58968、独立心跳补丁,保留完整对照历史。
Summary
Integrate native Kimi-K3 prefill/decode disaggregation with the current InferenceX Python launcher, including the launcher migration in #3576 and the staged-workspace correction in #3600. Platform resources live in
configs/runners.yaml; image preparation usesinfx/launch/. There is no restored Bash launcher or separate submission/lifecycle implementation.The configuration uses MXFP4 target weights, FP32 SSM state, FP8 KV and GMU0.90. Existing topology and concurrency variants are retained. Throughput uses automatic golden acceptance selection; eval uses real verification. The high-concurrency prefill configuration retains
HSA_NO_SCRATCH_RECLAIM=0.Commit structure
Official nightly and removable prerequisites
The worker is
vllm/vllm-openai-rocm:nightly-36768d1bfd39094681cdbc8cb37d4b31c0729c89, pinned to amd64 digestsha256:50b4f2aec13ffcb11ed846fec049d72250a9c2421eb1ad24c95a5210d46cee4c. Registry and source checks on 2026-09-30 confirm that it includes vLLM#56861 and vLLM#58814, but predates merged vLLM#57700.51cee8904a0b402a834887a26008adb79b8cd26band its parent, thengit cherry-pick --no-commitinto this workload's disposable job clone. The checked-in dependency remains unchanged.30ae8248b663ea56398d44d30c0d8276a598c4a0diff, verify its pinned SHA256, check applicability and apply onlyvllm/*through the recipe's temporary setup script. The official image contains an installed wheel rather than a Git checkout, so this is a Python-only backport, without replacing compiled extensions. Native srt-slurm also runs this setup in the frontend container; only the standalone router image is exempt from engine patching. Worker package lookup, download, checksum or application failure stops startup.The setup also applies the two-file independent-heartbeat subset from vLLM#58968, pinned at
9cb82832079ede6638b5a5f35ba0d60d203f57a0. The checked-in subset has SHA2565eaf72c56f3dc7e5bc85b76df11c0c335f7935b492b3271cde28fcd1dfd9683f; it moves rank-zero registration sends into an independent process and adds bounded shutdown/parent-death cleanup. It preserves payloads and intervals and excludes that PR's request-completion, shared-MR, parser and transport changes. It is intentionally not a full cherry-pick of the mixed PR.Once the official nightly contains #57700 and the independent-heartbeat fix, update both worker-image references and remove the corresponding backports, setup script, patch, associated script regression tests and recipe setting. Independently advance the srt-slurm pin to include #508 and remove its job-local cherry-pick. A worker-image update cannot update srt-slurm. These removals leave the framework/configuration changes; normal runtime qualification and review still apply.
The pinned image's Ionic provider rejects kernel ABI4. The temporary layer therefore restores a single-file read-only mount from the cluster-declared
ionic-providervolume, selected only by the Kimi-K3 vLLM PD lane. It leaves the image'slibibverbscore and all unrelated providers intact. Remove this mount only when the replacement image opens every expected RDMA device without it. This compatibility requirement is independent of shared-MR and is not an engine patch.Validation
/healthreturned the expected 503 waiting for absent workers. SIGTERM produced exit code 0, both ports closed, and the allocation completed successfully. No P/D workers or benchmark traffic were started.make setup ARCH=x86_64also completed locally with real dependency downloads into a fresh private directory, separate from the mocked installation acceptance.The first manual 1P2D c48 run stopped before image import or server startup because recipe resolution imported srtctl in the launcher interpreter instead of the job-local environment. The framework commit now resolves images through the job-local Python, following the existing binder/submission pattern. The isolated-interpreter regression reproduces the original failure and passes with the correction, including image-mismatch rejection before import.
The subsequent run imported the worker image but stopped before benchmark submission when the shared resolver sent the router manifest request to the Docker Hub website alias. Explicit Docker Hub hosts now resolve to the registry API without changing image digests or cache keys. The setup-scope correction prevents the next startup failure in the standalone router. Neither failure was a P/D engine crash, and neither run produced throughput results. The cluster check above validates these preparation boundaries, not a full GPU sweep or its shutdown path.
This remains a draft debug integration; GPU qualification is pending. The earlier c48 result used a different harness revision and custom image and is not qualification of this candidate. Full sweep, applicable evals, upstream-image policy and review remain required before merge.
The later GPU attempt loaded the model but failed on all workers at Ionic provider initialization, before traffic. The nightly migration had removed the previously validated provider mount before qualifying the replacement image's ABI. That omission is corrected by the scoped mount and current-image A/B above. The failed run produced no throughput result and released its allocation; one decode step needed teardown escalation. End-to-end qualification remains pending.
The next attempt passed provider initialization and worker health, but prefill discovery expired during warmup, producing HTTP 503 before profiling. The allocation was released. The narrow heartbeat backport addresses that observed failure mechanism: CPU A/B with a 3-second send interval and an 8-second worker GIL hold measured maximum gaps of 8.009 seconds for the old thread and 3.003 seconds for the independent process. Real-socket lifecycle checks cover both role payloads, repeated shutdown and abrupt owner death. The production setup was also executed against a clean pinned-nightly source export, and its installed output matched the qualified subset. This is not complete GPU worker construction or an end-to-end pass; replacement 1P2D c48 fast qualification is pending.
中文
将原生 Kimi-K3 prefill/decode 分离接入当前 InferenceX Python launcher,包含 #3576 的入口迁移和 #3600 的 workspace 打包修复。平台资源声明放在
configs/runners.yaml,镜像准备使用infx/launch/;不恢复旧 Bash launcher,也不另建提交或生命周期管理路径。配置使用 MXFP4 主模型权重、FP32 SSM 状态、FP8 KV 和 GMU0.90,保留原有拓扑与并发配置。吞吐测试自动选择 golden acceptance,eval 使用真实校验。高并发 prefill 保留
HSA_NO_SCRATCH_RECLAIM=0。提交分为框架、配置、临时 debug 集成三层。框架声明模型、draft、fabric、memlock 等资源,导入前校验 worker 镜像与矩阵一致,并通过共享 backend 准备 recipe 指定的 router 镜像;其余安装、提交、取消和结果收集沿用 main。配置同步 master/recipe 的官方镜像标识、双语文档及只追加的性能记录。临时层还保留仅作用于 Kimi-K3 vLLM PD 的 Ionic provider 兼容挂载,不新增 shared-MR、BF16 开发补丁、READ-credit 缓解或 router 算法修改;依赖仓库地址及 AIPerf 源码不变。
官方 ROCm nightly 的 tag/digest 如上。截至 2026-09-30,已确认它包含 vLLM#56861/#58814,但尚不含已合入的 #57700。临时层补齐缺失的上游前置能力:#508 按固定 commit 在任务私有 srt-slurm clone 中执行
git cherry-pick --no-commit;#57700 按固定合入 commit 下载并校验 diff,仅应用vllm/*的 Python 修改,不替换编译扩展。官方镜像安装的是 wheel,因此后者不是对镜像内 Git 仓库执行 cherry-pick。原生 srt-slurm 也会在 frontend 容器执行 setup,仅独立 router 镜像跳过 engine 补丁;worker 包查找、下载、校验或应用失败仍终止启动。同一 setup 还应用 vLLM#58968 在
9cb82832079ede6638b5a5f35ba0d60d203f57a0的两文件独立心跳子集,固定 SHA256 为5eaf72c56f3dc7e5bc85b76df11c0c335f7935b492b3271cde28fcd1dfd9683f。仅把 rank-zero 注册发送放入独立进程,并接入有界退出和父进程死亡清理;payload 与间隔不变,不携带该 PR 的请求完成、shared-MR、parser 或传输修改,不对整个混合 PR 做 cherry-pick。nightly 包含 #57700 和独立心跳修复后,同步更新两个 worker 镜像引用并删除对应 backport、setup script、patch、脚本回归测试及 recipe 引用。srt-slurm 需要独立升级 pin,包含 #508 后再删除任务内 cherry-pick;换 worker 镜像不会更新 srt-slurm。清理后保留框架和配置层,但仍须完成常规运行验证和审查。
与 CI 对齐的本地 CPU 回归为 2361 passed、1 skipped,Ruff、格式和 setup script 语法检查通过。其他本地验证覆盖真实矩阵、native schema、worker/router 命令生成、Python 入口、结果收集及调度拒绝路径;仅调度器、镜像导入和安装等外部操作使用模拟。Launcher 回归现使用独立的任务 Python 环境,不向 launcher 暴露 srtctl;另以实际安装的任务私有 srt-slurm 验证了 worker 和 frontend 镜像解析。两个在线补丁均在精确源码基线上应用成功,#57700 修改后的文件与上游合入版本一致,#508 专项测试通过,历史 changelog 字节不变。
独立 CPU 集群作业使用私有空缓存和相同的固定 router 镜像:旧 registry URI 复现 JSON 错误,修正后成功导入;在真实容器中,旧 setup 复现缺少 vLLM 的错误,修正后通过。Router 的 HTTP 和 discovery 监听均启动,
/health返回等待未注册 worker 的预期 503;SIGTERM 后返回码为 0,两个端口关闭,allocation 正常结束。没有启动 P/D 或发送 benchmark 流量。另在全新本地私有目录中完成原生make setup ARCH=x86_64的真实依赖下载,与模拟安装验收分开记录。首次手动 1P2D c48 作业在镜像导入和 server 启动前失败:recipe 解析错误地在 launcher 解释器中导入任务私有环境的 srtctl。框架 commit 已沿用现有 binder/submission 模式,改为通过任务私有 Python 解析。独立解释器回归能复现旧故障,并在修复后通过,包含镜像不匹配时在导入前拒绝的路径。
后续作业导入了 worker 镜像,但共享解析器将 router manifest 请求发到 Docker Hub 网站别名,因而在 benchmark 提交前失败。现将显式 Docker Hub 主机名规范化为 registry API,镜像 digest 和缓存键不变;setup 作用范围修正也避免了独立 router 紧接着会遇到的启动错误。两次失败均不是 P/D engine 崩溃,均未产生吞吐结果。上述集群检查只证明准备阶段和 router 启动正常,不代表完整 GPU sweep 或其收尾已通过。
本 PR 保持 draft debug 状态,GPU 验收仍待完成。之前的 c48 结果使用不同 harness 和自定义镜像,不能作为本版本验收。合入前仍需完整 sweep、适用 eval、上游镜像政策检查和人工 review。
后续 GPU 作业完成模型加载,但所有 worker 都在 Ionic provider 初始化时拒绝内核 ABI4,尚未产生实际传输或吞吐结果。迁移 nightly 时,在验证新镜像 ABI 前过早删除了之前已验证的 provider 挂载;现在通过集群声明的
ionic-provider卷,仅为 Kimi-K3 vLLM PD 恢复单文件只读挂载,不替换镜像内的 libibverbs 核心库或其他 provider。这与 shared-MR 无关,也不是 engine 补丁。只有替换镜像在不挂载的情况下成功打开全部预期 RDMA 设备,才可删除。当前镜像的 CPU A/B 已复现无挂载时的 ABI 错误,并验证单文件挂载后 8 个设备全部打开、关闭成功。核心库哈希不变,挂载后的 provider 哈希与主机源文件一致,显式传入的 fabric 环境在容器内保持正确;带同一挂载的 idle router 也正常启动、退出。自有作业在 53 秒内完成,未申请 GPU。这不是 MR、实际流量、吞吐或已加载 worker 收尾的验证。之前失败的 GPU allocation 已释放,其中一个 decode step 需要升级终止;端到端验收仍待完成。
再下一次作业通过 provider 初始化和 worker 健康检查,但 warmup 时 prefill discovery 过期,返回 HTTP 503,尚未进入 profile,allocation 已释放。此次最小心跳 backport 针对这一已观察到的故障机制:CPU A/B 采用实际 3 秒间隔和 8 秒 GIL 阻塞,旧线程最大空档 8.009 秒,独立进程 3.003 秒。真实 socket 生命周期验证覆盖两种角色的 payload、重复退出及父进程意外死亡;生产 setup 也在干净的固定 nightly 源码导出上实际执行,安装结果与已验证子集一致。这不代表完整 GPU worker 构造或端到端通过,替换版本的 1P2D c48 fast 验收仍待完成。