refactor(launch): port hardware launchers to a pluggable Python launcher - #3576
Merged
Merged
Conversation
Replace inferencex-e2e/runners/launch_*.sh, slurm_utils.sh, runtime_settings.sh and the per-cluster srt-slurm profiles with `python -m infx.launch`: - configs/runners.yaml `clusters:` is the single static record per cluster (scheduler settings, volumes, model entries, image policy, srt-slurm settings) - infx.launch.backends: small Backend interface; Slurm/pyxis implementation with one enroot import (submit-host, compute, all-nodes, pyxis, pre-staged, unchecked) - drivers: srt-slurm (single/multi-node), script (SpeedBench), legacy (amd_utils, tilert, slated for removal) - lifecycle: signal-safe cleanup, exit-code preservation, log snapshot before delete - post-eval env from one declared workload-env pattern list; no host secrets - historical replay (klaud, ingest recovery, changelog dispatch) runs the revision's own generator; klaud regeneration runs in a step without secrets - workflows call the Python entrypoint; dead launcher paths removed 中文:将硬件启动器迁移为可插拔的 Python 启动器 使用 `python -m infx.launch` 取代 inferencex-e2e/runners/launch_*.sh、slurm_utils.sh、 runtime_settings.sh 以及各集群的 srt-slurm 配置:configs/runners.yaml 的 `clusters:` 成为每个集群唯一的静态配置;新增精简的 Backend 接口及 Slurm/pyxis 实现,统一 enroot 镜像导入;提供 srt-slurm、脚本与遗留(amd_utils、tilert,待移除)驱动;信号安全清理并保留 退出码;评估环境变量由统一的工作负载模式列表导出,不再透传主机密钥;历史版本回放改为运行 该版本自身的生成器;工作流改为调用 Python 入口并移除失效的启动路径。
# Conflicts: # inferencex-e2e/runners/launch_gb300-nv.sh # inferencex-e2e/runners/srt-slurm/patches/README.md
Contributor
|
Thanks for the contribution!
中文感谢你的贡献!
|
中文:operatorx 生成器测试夹具改用 clusters 形式的 runner 配置。
…t lanes - srt model_paths map each recipe's own model.path alias to MODEL's staged checkpoint (by basename, node-local first); a short override table covers the real exceptions - single-node points use the same lookup; hub-only clusters are cluster data - srt lanes drop accidental per-cluster divergence: job status verified, logs copied, setup retried and health checks raised uniformly; one job-name scheme; --no-preflight derived from node-local checkpoints - trim comments and docstrings; inline one-caller helpers 中文:srt 模型路径改为读取配方自身的 model.path 别名,并按 MODEL 名称映射到已暂存的检查点 (优先节点本地),仅保留少量例外;单节点共用同一查找逻辑;统一各集群 srt 流程中无意产生的差异 (作业状态校验、日志复制、setup 重试、健康检查、作业命名,并由节点本地检查点推导 --no-preflight); 精简注释与文档字符串。
中文:精简 srt 驱动中的辅助函数、注释与测试。
…late the launcher uv cache srtctl always renders --account and falls back to the literal "default", which CW clusters reject. Clusters without an account now use the submitting user's default account from sacctmgr. The launcher venv step uses a per-job uv cache because some runner homes hold a ~/.cache/uv uv cannot create. 中文:集群未配置账户时,srt 使用提交用户在 sacctmgr 中的默认账户(srtctl 否则会传入 无效的 "default");启动器虚拟环境步骤改用每个作业独立的 uv 缓存目录。
… single-node srt jobs
Real-cluster smoke runs showed srtctl's untyped --gpus-per-node and default
--segment are rejected by the CW partitions ("Requested node configuration is
not available"). Single-node jobs never emit a segment; clusters that set
gpus-per-node-directive: false get their typed slurm.gres instead. The uv
setup steps use a per-runner cache because some runner homes hold a
~/.cache/uv that uv cannot create.
中文:CW 集群改用带类型的 GRES 申请 GPU,单节点 srt 作业不再传入 --segment(真实集群冒烟测试中
二者被 CW 分区拒绝);uv 准备步骤改用每个 runner 独立的缓存目录。
CW single-node submissions still fail with "Requested node configuration is not available" after typed GRES; the old salloc path there never passed --exclusive. 中文:允许集群让单节点 srt 作业不使用 --exclusive(CW 分区上此前的 salloc 路径从未使用该参数)。
…et as multi-node On h100-cw a cached 397B FP8 checkpoint took ~24 min to load from shared storage plus ~12 min of DeepGEMM warmup, past srtctl's 1800 s default. 中文:单节点 srt 作业与多节点使用相同的 2 小时健康检查时限(h100-cw 上模型加载与预热超过 srtctl 默认的 1800 秒)。
…ronment The TileRT smoke run on b200-nscale got the runner host's UCX_NET_DEVICES instead of the cluster's eight HCAs. As with the old runtime_settings.sh, cluster settings now win over the host environment and lose only to a point's additional-settings. 中文:集群运行时设置优先于 runner 主机自身的环境变量,仅由测试点的 additional-settings 覆盖 (与原 runtime_settings.sh 的行为一致;修复 b200-nscale TileRT 使用了主机 UCX_NET_DEVICES 的问题)。
…one env precedence rule - every srtctl job is named inferencex-<runner> and cleanup cancels both names, so other tools' scancel --name=<runner> cannot kill InferenceX jobs - single-node points read the Hub unless the cluster stages them (b200-nscale, b300-dsxe); B300 DeepSeek-V4-Pro-0813 reads NVMe only for vLLM - cluster env sources override the runner host but never a point's additional-settings; the Slurm account resolves once per launch - TileRT fork keeps its own health and recipe behavior; health floor applies only inside health_check blocks - CW clusters drop --exclusive everywhere; uv Pythons also stay per runner - tests use synthetic records instead of checked-in values 中文:所有 srtctl 作业统一命名为 inferencex-<runner>,清理同时取消两种名称;单节点默认使用 Hub, 仅 b200-nscale、b300-dsxe 使用暂存检查点;集群环境变量优先于主机环境但不覆盖测试点设置; Slurm 账户每次启动解析一次;TileRT fork 保留原有行为;CW 集群不再使用 --exclusive;测试改用合成配置。
The cluster renamed compute-0; sbatch rejects the old name. 中文:mi300x-amd 集群分区已改名为 MI300X-UBUNTU,旧名 compute-0 被 sbatch 拒绝。
b300-dsxe requests cpus-per-task=192 (its full node) since #3477, but srtctl runs TRT-LLM with one task per GPU, so dynamo-trt submissions asked for 8x192 CPUs per node and sbatch refused them. 中文:TRT-LLM 的 srt 作业不再申请整节点 CPU(b300 的 cpus-per-task=192 与每 GPU 一个任务冲突, 导致 dynamo-trt 作业被 sbatch 拒绝)。
Instead of dropping cpus-per-task, divide it by the tasks srtctl places on each node, so b300 TRT jobs still reserve the whole node (8 x 24 CPUs). 中文:TRT-LLM 每个 GPU 一个任务时,将节点 CPU 配额按任务数均分(b300 为 8 × 24),而不是不申请。
#3477 set cpus-per-task=192 (a whole b300 node). srtctl runs TRT-LLM with one task per GPU, so dynamo-trt jobs asked for 8x192 CPUs per node and sbatch refused them. cpus-per-gpu=24 reserves the same 192 CPUs per node for every task layout (verified with sbatch --test-only on batch_1). Clusters declare either field; the TRT task-count special case is gone. 中文:b300 改为按 GPU 申请 CPU(cpus-per-gpu=24),替代 #3477 的 cpus-per-task=192。后者在 TRT-LLM 每 GPU 一个任务时申请 8×192 个 CPU 被 sbatch 拒绝;新设置在所有任务布局下均保留整节点 192 个 CPU(已在 batch_1 上用 sbatch --test-only 验证)。
runners.yaml listed gb200-nv_0..3, but the registered runners are gb200-nv_00..17 (no _08). The launcher resolves the cluster from these labels, so every gb200 job failed before submission. The generated sweep matrix is unchanged. 中文:更新 gb200-nv 的 runner 列表为实际注册的 gb200-nv_00..17(无 _08),此前启动器无法解析集群; 生成的扫描矩阵不变。
Carries main's srt-slurm v2.36.0 pin, its per-profile cluster names (now rendered from the cluster id), the MI300X partition rename (already on this branch), and the AgentX one-concurrency-per-deployment check, whose test now drives infx.launch.
python3 -P fails on images older than Python 3.11 (the MI355X sglang-disagg image is one); PYTHONSAFEPATH gives the same isolation where supported and is ignored elsewhere.
Docstrings stay; lint pragmas and action version pins are kept.
…ners.yaml The job tag is the cluster's GPU label (srt-slurm.job-tag); MORI_RDMA_TC is the MI355X fabric's traffic class, now in the cluster env like b200-nb's UCX device.
中文:恢复本分支未改动代码的文件中 main 原有的注释。
This was referenced Sep 30, 2026
Open
Draft
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replaces the per-cluster bash launchers (
inferencex-e2e/runners/launch_*.sh,slurm_utils.sh,runtime_settings.shand the 13runners/srt-slurm/<cluster>.yamlprofiles) with one Python launcher,python -m infx.launch run|cleanup, and one static record per cluster inconfigs/runners.yaml.Draft. Smoke-tested on real hardware on 8 clusters (table below), at
5cccb7cor earlier. Later commits (comment cleanup, moving the srtctl job tag and MI355XMORI_RDMA_TCintorunners.yaml, merges ofmainincluding #3596) are covered by unit tests only. The branch is level withmain.Design
configs/runners.yaml→clusters.<id>owns the cluster facts: scheduler settings (partition, account, exclusions, GPU and CPU requests), named volumes (model roots, caches), model checkpoints, the image cache policy, and srt-slurm settings (including the srtctl job tag). Workload quirks keyed by cluster (accepted models and frameworks, model overrides, power lanes, TileRT's UCX settings) stay in Python tables that are checked against the records at startup. A runner resolves to its cluster through itscluster:<id>label.infx/launch/backends/defines a smallBackendinterface: prepare image, run container, stream logs, state, fetch outputs, cleanup. Slurm + pyxis is the implementation today. A new scheduler (for example k8s or host docker) is new files, a registry entry and a cluster record, and a fake-backend test shows that end to end. Script-driver points run on any backend; srt-slurm points need Slurm.submit-host,compute,all-nodes(node-local storage, one import per node),pyxis,pre-staged,unchecked. Locked, validated, written to a temp file and renamed atomically; digests are kept.srt(srt-slurm single-node and multi-node, the main path),script(SpeedBench),legacy(the remaining TileRT andamd_utilslanes, to be removed once they move to srt-slurm).model_pathsmap each recipe's ownmodel.pathaliases to MODEL's staged checkpoint (by basename, node-local copy first), with a short override table for real exceptions. Single-node points use the Hub unless the cluster stages checkpoints (b200-nscale, b300-dsxe).inferencex-<runner>, and cleanup cancels both that name and the runner name.EVAL_*,SWEBENCH_*,AIPERF_*,AGENTIC_*, …), so host secrets are never forwarded. Required inputs are enforced by per-path request models.infx/matrix/revision.py). klaud regenerates producers in a step that holds no secrets.Removed
launch_*.sh,slurm_utils.sh,runners/runtime_settings.sh,inject_srt_power_concurrencies.py, and the 13 srt profiles.h100-cr(no runner labels), raw fallbacks that ran the no-longer-existingbenchmarks/single_node/{fixed_seq_len,agentic}/*.sh, gb200 llm-d and legacy SGLang branches,SAGEMAKER_SHM_PATHleftovers.Real-cluster smoke runs
e2e-tests.yml/speedbench-al.ymldispatches of this branch at5cccb7cor earlier; nothing is ingested.agentx-fast)MI325-UBUNTUpartition, no runners yet); the current hosts cannot host uv's default directoriesSome matrix rows (for example TP2/EP2, some MTP rows) have no matching recipe variant; they fail recipe selection the same way on
main.Fixes these runs found
cpus-per-task: 192(a whole node, from perf(b300): update vLLM AgentX to DSpark6 with native SRT recipes / 使用原生 SRT 配方更新 B300 vLLM AgentX 至 DSpark6 #3477) combined with TRT-LLM's one task per GPU asked for 8×192 CPUs per node. b300 now requestscpus-per-gpu: 24, the same 192 CPUs per node for every layout (checked withsbatch --test-onlyonbatch_1).mainhas the same bug.runners.yamllistedgb200-nv_0..3; the registered runners aregb200-nv_00..17. Fixed; the generated sweep matrix is byte-identical.compute-0toMI300X-UBUNTU.python3 -P→PYTHONSAFEPATH=1, so images older than Python 3.11 can run it.mainhas the same problem.--exclusive/--segment, for single-node jobs.--account=default).runtime_settings.sh(fixes TileRT'sUCX_NET_DEVICES).Not yet exercised on hardware
amd_utils), b200-nscale TileRT through srt-slurm, the B300 batch-wrapped dsv41flash lane, the H200 salloc time bump, the GB300 8 h lane, DCGM power lanes,profile.yml.merge-ingest.yml, artifact reuse) can only run after merge.Intended behavior differences from the bash launchers
COMPLETED 0:0), not only on some lanes.inferencex-<runner>(previously gb200 only).--no-preflightfollows from whether the checkpoint sits on node-local storage.agentictag on every tagged lane.MORI_RDMA_TC=104is cluster env, so it reaches every MI355X job (its srt recipes already set 104) and a point's additional-settings can override it.clusterstatus name is the cluster id (b200-nscale,h100-dgxc,h200-dgxc) rather than the runner prefixmainuses (*-slurm).Verification (offline)
pytestacross inferencex-e2e, collectivex and operatorx: 2324 passed (1 known macOS-only failure in a Kimi eval test that does not run in CI).ruff checkandruff format --checkare clean, and PR CI is green.