Skip to content

feat(launch): re-add llm-d multinode support to infx.launch - #3613

Draft
adibarra wants to merge 4 commits into
mainfrom
feat/llmd-python-launcher
Draft

adibarra wants to merge 4 commits into
mainfrom
feat/llmd-python-launcher

Conversation

@adibarra

@adibarra adibarra commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Re-adds llm-d (llmd-vllm) multinode support on the Python launcher from #3576, which dropped the launch_gb200-nv.sh llm-d branch.

llm-d keeps its own Slurm orchestration (submit.sh → sbatch job.slurm → srun + pyxis per node → server.sh). The launcher owns everything around it:

  • policy.launch_path: multinode framework == llmd-vllm → LaunchPath.LLMD, enabled on gb200-nv only (LLMD_CLUSTERS); the table check requires slurm.squash.
  • drivers/llmd.py: resolves the checkpoint from runners.yaml, imports the image via backend.prepare_image, runs llm-d/submit.sh directly with the point's topology, sequence lengths and concurrency, attaches to the printed job id, streams the log, checks the final Slurm state, and stages result JSONs, agentic and eval artifacts plus the server-log bundle. No per-model wrapper script.
  • srt/models.py: dsv4 / fp4 / llmd-vllm on gb200-nv → DeepSeek-V4-Pro@numa1, served as deepseek-ai/DeepSeek-V4-Pro.
  • dsv4-fp4-gb200-llmd-vllm: re-adds the low-latency 8k1k point (1P DEP8 + 1D TP8, conc 1) from the key removed in [Klaud Cold] Enact the September 8, 2026 DeepSeek-V4-Pro Single-turn 8k1k deprecation / 执行 2026 年 9 月 8 日 DeepSeek-V4-Pro 单轮 8k1k 场景下线 #2921, to exercise the path.

Fix: successful jobs no longer end CANCELLED

job.slurm used to scancel its own allocation once the decode coordinator finished, so every successful run ended CANCELLED 0:0, which the launcher reports as a failure (run 36720223235 finished its benchmark and was still marked failed). job.slurm now stops the srun step and exits 0, so the job ends COMPLETED. A step that dies before the coordinator finishes still fails the job with its exit code.

Driver port by @ilmarkov from #2719.

Testing

  • pytest infx/tests/launch infx/tests/clusters infx/tests/matrix infx/tests/workflows: 1230 passed; ruff clean
  • validate_perf_changelog against origin/main: passes
  • Local harness for the job.slurm step control: hung workers + done marker → exit 0; step failing first → its exit code
  • GB200 e2e: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/36746690132

Add an llm-d driver that submits the job through the llm-d wrapper, attaches
to the Slurm job, streams its log, checks its final state and stages results,
agentic and eval artifacts. Route llmd-vllm multinode requests to it on
gb200-nv and resolve DeepSeek-V4-Pro to the node-local numa1 checkpoint.

Restore the GB200 DeepSeek-V4-Pro disagg wrapper.

job.slurm no longer scancels its own allocation when the coordinator finishes;
it stops the srun step and exits 0, so the job ends COMPLETED instead of
CANCELLED, which infx.launch reports as a failure.

Driver port by Ilya Markov from #2719.
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

…est point

Drop the per-model llm-d wrapper script. The driver passes topology, sequence
lengths and concurrency to submit.sh and sets GPUS_PER_NODE, TIME_LIMIT,
CONTAINER_IMAGE and worker counts itself. llmd-vllm supports only P/D
disaggregated points.

Re-add the dsv4-fp4-gb200-llmd-vllm low-latency 8k1k point (1P DEP8 + 1D TP8,
conc 1) to exercise the path end to end.
Route every multinode llmd-vllm request to the llm-d driver. The driver
requires slurm.squash and a staged checkpoint, which is what the whitelist
stood in for.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant