-
Notifications
You must be signed in to change notification settings - Fork 315
feat(launch): re-add llm-d multinode support to infx.launch #3613
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
8 commits
Select commit
Hold shift + click to select a range
cb0d0a7
feat(launch): route llmd-vllm multinode jobs through infx.launch
adibarra 4d0813f
chore(launch): trim llm-d comments
adibarra fcdd311
feat(launch): run llm-d submit.sh directly; re-add GB200 llm-d 8k1k t…
adibarra 4b2ffc9
refactor(launch): drop the llm-d cluster whitelist
adibarra 62cb067
fix(launch): leave node-local llm-d checkpoints to the job
adibarra 34b73b8
Merge remote-tracking branch 'origin/main' into feat/llmd-python-laun…
adibarra f90c7a3
Merge remote-tracking branch 'origin/main' into feat/llmd-python-laun…
adibarra 7fd158a
chore: drop the GB200 llm-d 8k1k test point
adibarra File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,158 @@ | ||
| """llm-d vLLM multinode jobs submitted through benchmarks/multi_node/llm-d/submit.sh.""" | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| import os | ||
| import shutil | ||
| import subprocess | ||
| import sys | ||
| from pathlib import Path | ||
|
|
||
| from infx.launch import artifacts, policy, proc | ||
| from infx.launch.backends.base import BackendError | ||
| from infx.launch.backends.slurm import cli | ||
| from infx.launch.context import Launch, LaunchError | ||
| from infx.launch.drivers.srt import models | ||
| from infx.launch.drivers.srt.run import slurm_backend | ||
| from infx.launch.request import LlmdRequest, RequestError | ||
|
|
||
| CANCEL_TIMEOUT_S = 600.0 | ||
| DEFAULT_TIME_LIMIT = "08:00:00" | ||
| LLMD_DIR = "benchmarks/multi_node/llm-d" | ||
|
|
||
|
|
||
| def _find_eval_dir(logs_dir: Path) -> Path | None: | ||
| for root, dirs, _files in os.walk(logs_dir): | ||
| if "eval_results" in dirs: | ||
| return Path(root) / "eval_results" | ||
| return None | ||
|
|
||
|
|
||
| def _stage_agentic(logs_dir: Path, workspace: Path) -> None: | ||
| agentic = logs_dir / "agentic" | ||
| if not agentic.is_dir(): | ||
| return | ||
| staged = workspace / "LOGS" / "agentic" | ||
| staged.mkdir(parents=True, exist_ok=True) | ||
| for entry in agentic.iterdir(): | ||
| destination = staged / entry.name | ||
| if entry.is_dir(): | ||
| shutil.copytree(entry, destination, dirs_exist_ok=True) | ||
| elif entry.is_file(): | ||
| shutil.copy2(entry, destination) | ||
|
|
||
|
|
||
| def run(launch: Launch) -> int: | ||
| """Submit the llm-d Slurm job, follow its log, and stage benchmark artifacts.""" | ||
| backend = slurm_backend(launch) | ||
| request = LlmdRequest.from_env(launch.request.env) | ||
| if backend.settings.squash is None: | ||
| raise LaunchError(f"llmd-vllm: cluster {launch.cluster.id!r} has no slurm.squash") | ||
| if not request.disagg: | ||
| raise LaunchError("llmd-vllm supports only P/D disaggregated points") | ||
|
|
||
| checkpoint = models.checkpoint(launch.cluster, request) | ||
| if checkpoint is None: | ||
| raise LaunchError( | ||
| f"cluster {launch.cluster.id!r} stages no checkpoint for MODEL={request.model}" | ||
| ) | ||
| model_path = models.host_path(launch.cluster, checkpoint) | ||
| if not checkpoint.node_local and not (model_path / "config.json").is_file(): | ||
| raise LaunchError(f"model checkpoint is unavailable: {model_path / 'config.json'}") | ||
|
|
||
| squash = backend.prepare_image(request.image) | ||
| logs_dir = request.workspace / "benchmark_logs" | ||
| logs_dir.mkdir(parents=True, exist_ok=True) | ||
|
|
||
| account = backend.settings.account or cli.default_account() | ||
| if not account: | ||
| raise RequestError.missing("SLURM_ACCOUNT") | ||
|
|
||
| env = policy.runtime_env( | ||
| launch.cluster, | ||
| request, | ||
| models.job_env(launch.cluster, request, str(model_path)), | ||
| { | ||
| "SLURM_PARTITION": backend.settings.partition, | ||
| "SLURM_ACCOUNT": account, | ||
| "MODEL_PATH": str(model_path), | ||
| "MODEL_NAME": request.model, | ||
| "CONTAINER_IMAGE": request.image, | ||
| "GPUS_PER_NODE": str(launch.cluster.gpus_per_node), | ||
| "TIME_LIMIT": request.env.get("TIME_LIMIT") or DEFAULT_TIME_LIMIT, | ||
| "PREFILL_WORKERS": str(request.prefill_num_workers), | ||
| "DECODE_WORKERS": str(request.decode_num_workers), | ||
| "LLMD_CONTAINER_ENGINE": "pyxis", | ||
| "LLMD_SQUASH_FILE": squash.reference, | ||
| "BENCHMARK_LOGS_DIR": str(logs_dir), | ||
| }, | ||
| ) | ||
|
|
||
| argv = [ | ||
| "bash", | ||
| "submit.sh", | ||
| str(request.prefill_nodes), | ||
| str(request.decode_nodes), | ||
| str(request.isl), | ||
| str(request.osl), | ||
| "x".join(map(str, request.conc_list)), | ||
| "inf", | ||
| request.random_range_ratio, | ||
| ] | ||
| proc.echo(argv, env) | ||
| submitted = subprocess.run( | ||
| argv, | ||
| stdout=subprocess.PIPE, | ||
| stderr=sys.stderr, | ||
| text=True, | ||
| env=env, | ||
| cwd=request.workspace / LLMD_DIR, | ||
| check=False, | ||
| ) | ||
| job_id = submitted.stdout.strip() | ||
| if submitted.returncode != 0 or not job_id: | ||
| print("ERROR: llm-d submit.sh failed before returning a Slurm job id", file=sys.stderr) | ||
| return 1 | ||
| if not (job_id.isascii() and job_id.isdigit()): | ||
| print( | ||
| f"ERROR: llm-d submit.sh printed {job_id!r} instead of a Slurm job id", | ||
| file=sys.stderr, | ||
| ) | ||
| return 1 | ||
|
|
||
| log_file = logs_dir / f"slurm_job-{job_id}.out" | ||
| job = backend.attach(job_id, log=log_file, outputs=logs_dir) | ||
| print(f"Submitted llm-d job: {job_id}", flush=True) | ||
|
|
||
| launch.life.callback( | ||
| artifacts.bundle_server_logs, logs_dir, request.workspace / "multinode_server_logs.tar.gz" | ||
| ) | ||
| launch.life.callback(backend.cancel, job, wait_s=CANCEL_TIMEOUT_S) | ||
|
|
||
| try: | ||
| backend.stream_logs(job) | ||
| except BackendError: | ||
| return 1 | ||
|
|
||
| status = backend.state(job) | ||
| rc = 0 if status.succeeded else 1 | ||
|
|
||
| for result_file in sorted(logs_dir.glob(f"{request.result_filename}*.json")): | ||
| try: | ||
| artifacts.copy_to_workspace(result_file, request.workspace / result_file.name) | ||
| except artifacts.ArtifactError as error: | ||
| print(f"ERROR: {error}", file=sys.stderr) | ||
| rc = 1 | ||
|
|
||
| if request.is_agentic and not request.eval_only: | ||
| _stage_agentic(logs_dir, request.workspace) | ||
|
|
||
| if request.run_eval: | ||
| eval_dir = _find_eval_dir(logs_dir) or logs_dir / "eval_results" | ||
| try: | ||
| artifacts.copy_eval_artifacts(eval_dir, request.workspace) | ||
| except artifacts.ArtifactError as error: | ||
| print(f"ERROR: {error}", file=sys.stderr) | ||
| rc = 1 | ||
|
|
||
| return rc |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔴 Any startup failure in the decode coordinator (EPP/Envoy not ready, prefill health timeout) now makes the whole llm-d job report success instead of failure, with no benchmark results. server.sh:303 has
trap 'touch "$BENCH_DONE_MARKER"' EXIT, which fires on every exit of that branch, including theexit 1failure paths, writing the same marker file as the real success touch at server.sh:573. job.slurm:82[[ -f "$BENCH_DONE_MARKER" ]] && return 0treats marker-exists as unconditional success, ignoring rc. Fix: use a distinct marker/exit path for genuine completion so job.slurm only succeeds when the benchmark actually ran, not merely when the coordinator exited for any reason.Why this was flagged
Trigger: the decode-leader branch in server.sh (ROLE==decode && LWS_WORKER_INDEX==0) hits any of its
exit 1checks - EPP not binding within 60s (server.sh:384/388), Envoy /ready failing within 120s (server.sh:407/412), or prefill /health timing out within 300s (server.sh:509) - reached via job.slurm's run_until_bench_done -> srun -> server.sh. The EXIT trap at server.sh:303 touches $BENCH_DONE_MARKER on that exit, same file as the genuine-success touch at server.sh:573. job.slurm's run_until_bench_done (job.slurm:70-85) only checks file existence at line 82, not the step's rc, so it returns 0. set -eo pipefail (job.slurm:10) then lets the script finish normally, Slurm records COMPLETED, and drivers/llmd.py:137-138 computes rc=0 from that state; the result-JSON glob (llmd.py:140) finds nothing but the loop simply doesn't run, so no error is raised. On the base branch this same failure ends CANCELLED, which the launcher already treats as failure.Verification: server.sh:303's
trap 'touch "$BENCH_DONE_MARKER"' EXITfires on every exit of the decode-coordinator branch, including the startupexit 1failures at server.sh:384, 388, 407, 412, and 509, writing the same marker as the genuine-success touch at server.sh:573. job.slurm's[[ -f "$BENCH_DONE_MARKER" ]] && return 0then discards the failing rc whenever that marker exists.