Conversation
Add an llm-d driver that submits the job through the llm-d wrapper, attaches to the Slurm job, streams its log, checks its final state and stages results, agentic and eval artifacts. Route llmd-vllm multinode requests to it on gb200-nv and resolve DeepSeek-V4-Pro to the node-local numa1 checkpoint. Restore the GB200 DeepSeek-V4-Pro disagg wrapper. job.slurm no longer scancels its own allocation when the coordinator finishes; it stops the srun step and exits 0, so the job ends COMPLETED instead of CANCELLED, which infx.launch reports as a failure. Driver port by Ilya Markov from #2719.
Contributor
|
Thanks for the contribution!
中文感谢你的贡献!
|
…est point Drop the per-model llm-d wrapper script. The driver passes topology, sequence lengths and concurrency to submit.sh and sets GPUS_PER_NODE, TIME_LIMIT, CONTAINER_IMAGE and worker counts itself. llmd-vllm supports only P/D disaggregated points. Re-add the dsv4-fp4-gb200-llmd-vllm low-latency 8k1k point (1P DEP8 + 1D TP8, conc 1) to exercise the path end to end.
Route every multinode llmd-vllm request to the llm-d driver. The driver requires slurm.squash and a staged checkpoint, which is what the whitelist stood in for.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Re-adds llm-d (
llmd-vllm) multinode support on the Python launcher from #3576, which dropped thelaunch_gb200-nv.shllm-d branch.llm-d keeps its own Slurm orchestration (
submit.sh→sbatch job.slurm→srun+ pyxis per node →server.sh). The launcher owns everything around it:policy.launch_path: multinodeframework == llmd-vllm→LaunchPath.LLMD, enabled ongb200-nvonly (LLMD_CLUSTERS); the table check requiresslurm.squash.drivers/llmd.py: resolves the checkpoint fromrunners.yaml, imports the image viabackend.prepare_image, runsllm-d/submit.shdirectly with the point's topology, sequence lengths and concurrency, attaches to the printed job id, streams the log, checks the final Slurm state, and stages result JSONs, agentic and eval artifacts plus the server-log bundle. No per-model wrapper script.srt/models.py:dsv4/fp4/llmd-vllmon gb200-nv →DeepSeek-V4-Pro@numa1, served asdeepseek-ai/DeepSeek-V4-Pro.dsv4-fp4-gb200-llmd-vllm: re-adds the low-latency 8k1k point (1P DEP8 + 1D TP8, conc 1) from the key removed in [Klaud Cold] Enact the September 8, 2026 DeepSeek-V4-Pro Single-turn 8k1k deprecation / 执行 2026 年 9 月 8 日 DeepSeek-V4-Pro 单轮 8k1k 场景下线 #2921, to exercise the path.Fix: successful jobs no longer end CANCELLED
job.slurmused toscancelits own allocation once the decode coordinator finished, so every successful run endedCANCELLED 0:0, which the launcher reports as a failure (run 36720223235 finished its benchmark and was still marked failed).job.slurmnow stops thesrunstep and exits 0, so the job endsCOMPLETED. A step that dies before the coordinator finishes still fails the job with its exit code.Driver port by @ilmarkov from #2719.
Testing
pytest infx/tests/launch infx/tests/clusters infx/tests/matrix infx/tests/workflows: 1230 passed; ruff cleanvalidate_perf_changelogagainstorigin/main: passesjob.slurmstep control: hung workers + done marker → exit 0; step failing first → its exit code