Skip to content

refactor(launch): port hardware launchers to a pluggable Python launcher - #3576

Merged
adibarra merged 22 commits into
mainfrom
feat/python-launchers
Sep 29, 2026
Merged

adibarra merged 22 commits into
mainfrom
feat/python-launchers

Conversation

@adibarra

@adibarra adibarra commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Replaces the per-cluster bash launchers (inferencex-e2e/runners/launch_*.sh, slurm_utils.sh, runtime_settings.sh and the 13 runners/srt-slurm/<cluster>.yaml profiles) with one Python launcher, python -m infx.launch run|cleanup, and one static record per cluster in configs/runners.yaml.

Draft. Smoke-tested on real hardware on 8 clusters (table below), at 5cccb7c or earlier. Later commits (comment cleanup, moving the srtctl job tag and MI355X MORI_RDMA_TC into runners.yaml, merges of main including #3596) are covered by unit tests only. The branch is level with main.

Design

  • Cluster records. configs/runners.yaml → clusters.<id> owns the cluster facts: scheduler settings (partition, account, exclusions, GPU and CPU requests), named volumes (model roots, caches), model checkpoints, the image cache policy, and srt-slurm settings (including the srtctl job tag). Workload quirks keyed by cluster (accepted models and frameworks, model overrides, power lanes, TileRT's UCX settings) stay in Python tables that are checked against the records at startup. A runner resolves to its cluster through its cluster:<id> label.
  • Pluggable backends. infx/launch/backends/ defines a small Backend interface: prepare image, run container, stream logs, state, fetch outputs, cleanup. Slurm + pyxis is the implementation today. A new scheduler (for example k8s or host docker) is new files, a registry entry and a cluster record, and a fake-backend test shows that end to end. Script-driver points run on any backend; srt-slurm points need Slurm.
  • One image import. Explicit modes: submit-host, compute, all-nodes (node-local storage, one import per node), pyxis, pre-staged, unchecked. Locked, validated, written to a temp file and renamed atomically; digests are kept.
  • Drivers. srt (srt-slurm single-node and multi-node, the main path), script (SpeedBench), legacy (the remaining TileRT and amd_utils lanes, to be removed once they move to srt-slurm).
  • Models. srt model_paths map each recipe's own model.path aliases to MODEL's staged checkpoint (by basename, node-local copy first), with a short override table for real exceptions. Single-node points use the Hub unless the cluster stages checkpoints (b200-nscale, b300-dsxe).
  • Lifecycle. SIGINT/SIGTERM/SIGHUP run cleanups (cancel the exact job, save logs before deleting outputs) and exit 130/143/129; the workload's exit code wins over cleanup failures. srtctl jobs are named inferencex-<runner>, and cleanup cancels both that name and the runner name.
  • Environment. Cluster runtime settings override the runner host's environment but never a point's additional-settings. Post-eval passthrough comes from one declared list of workload families (EVAL_*, SWEBENCH_*, AIPERF_*, AGENTIC_*, …), so host secrets are never forwarded. Required inputs are enforced by per-path request models.
  • Historical replay. klaud baselines, ingest recovery and changelog dispatch run the historical revision's own generator (infx/matrix/revision.py). klaud regenerates producers in a step that holds no secrets.

Removed

  • All 14 launch_*.sh, slurm_utils.sh, runners/runtime_settings.sh, inject_srt_power_concurrencies.py, and the 13 srt profiles.
  • Dead paths: h100-cr (no runner labels), raw fallbacks that ran the no-longer-existing benchmarks/single_node/{fixed_seq_len,agentic}/*.sh, gb200 llm-d and legacy SGLang branches, SAGEMAKER_SHM_PATH leftovers.

Real-cluster smoke runs

e2e-tests.yml / speedbench-al.yml dispatches of this branch at 5cccb7c or earlier; nothing is ingested.

Cluster Path Result
h200-dgxc single-node dsr1 fp8 sglang (STP + MTP) ✅
h200-dgxc multi-node dsr1 fp8 dynamo-sglang MTP, 2 nodes ✅
gb300-nv multi-node dsr1 fp8 dynamo-sglang, 2 nodes ✅
b300-dsxe multi-node dsr1 fp8 dynamo-trt, 2 nodes ✅
b300-dsxe single-node qwen3.5 fp4 sglang ✅
b300-dsxe SpeedBench AL collector (script driver) ✅
b200-nscale single-node qwen3.5 fp4 sglang ✅
b200-nscale AgentX single-node (agentx-fast) ✅
b200-nscale eval-only single-node ✅
b200-nscale multi-node dsr1 fp8 dynamo-trt, 2 nodes ✅
b200-nscale cancel mid-run ✅ Slurm job cancelled 22 s after the workflow cancel; logs saved
h100-cw single-node qwen3.5 fp8 sglang ✅
mi300x-amd single-node dsr1 fp8 sglang ✅
b200-nscale legacy TileRT multi-node ⚠️ benchmark completed (16/16 requests); then the TileRT script's own decode drain timed out (exit 137)
gb200-nv multi-node qwen3.5 fp8 dynamo-sglang, 2 nodes ✅
mi355x-amds single-node qwen3.5 fp8 atom ✅
mi355x-amds multi-node qwen3.5 fp8 sglang-disagg, 2 nodes ✅
mi325x-amds single-node blocked: the cluster is being replaced (new MI325-UBUNTU partition, no runners yet); the current hosts cannot host uv's default directories
h100-dgxc, h200-cw none runners offline, or the scheduler marks Slurm unavailable
b200-cw, b200-nb none not run yet

Some matrix rows (for example TP2/EP2, some MTP rows) have no matching recipe variant; they fail recipe selection the same way on main.

Fixes these runs found

  • b300 TRT-LLM: cpus-per-task: 192 (a whole node, from perf(b300): update vLLM AgentX to DSpark6 with native SRT recipes / 使用原生 SRT 配方更新 B300 vLLM AgentX 至 DSpark6 #3477) combined with TRT-LLM's one task per GPU asked for 8×192 CPUs per node. b300 now requests cpus-per-gpu: 24, the same 192 CPUs per node for every layout (checked with sbatch --test-only on batch_1). main has the same bug.
  • gb200: runners.yaml listed gb200-nv_0..3; the registered runners are gb200-nv_00..17. Fixed; the generated sweep matrix is byte-identical.
  • mi300x: the cluster renamed its partition from compute-0 to MI300X-UBUNTU.
  • Multi-node client: python3 -P → PYTHONSAFEPATH=1, so images older than Python 3.11 can run it. main has the same problem.
  • CW clusters: typed GRES, and no --exclusive/--segment, for single-node jobs.
  • Accounts: clusters without a declared account use the submitting user's Slurm default (srtctl otherwise passes --account=default).
  • Health checks: single-node srt jobs get the same 2 h health-check budget as multi-node.
  • Env precedence: cluster runtime settings beat the runner host's environment again, as with the old runtime_settings.sh (fixes TileRT's UCX_NET_DEVICES).
  • uv: the setup steps keep uv's cache and Pythons per runner.

Not yet exercised on hardware

  • MI355X AgentX (amd_utils), b200-nscale TileRT through srt-slurm, the B300 batch-wrapped dsv41flash lane, the H200 salloc time bump, the GB300 8 h lane, DCGM power lanes, profile.yml.
  • The merge-time path (merge-ingest.yml, artifact reuse) can only run after merge.

Intended behavior differences from the bash launchers

  • Every srt job's final Slurm state is checked (COMPLETED 0:0), not only on some lanes.
  • Every recipe health check is raised to at least 720 attempts.
  • srtctl jobs on every cluster are named inferencex-<runner> (previously gb200 only).
  • --no-preflight follows from whether the checkpoint sits on node-local storage.
  • AgentX jobs use the agentic tag on every tagged lane.
  • MI355X MORI_RDMA_TC=104 is cluster env, so it reaches every MI355X job (its srt recipes already set 104) and a point's additional-settings can override it.
  • srtctl's cluster status name is the cluster id (b200-nscale, h100-dgxc, h200-dgxc) rather than the runner prefix main uses (*-slurm).

Verification (offline)

  • pytest across inferencex-e2e, collectivex and operatorx: 2324 passed (1 known macOS-only failure in a Kimi eval test that does not run in CI). ruff check and ruff format --check are clean, and PR CI is green.
  • All 1676 generated matrix rows resolve to a cluster and a launch path, and the policy tables agree with the cluster records.

Replace inferencex-e2e/runners/launch_*.sh, slurm_utils.sh, runtime_settings.sh
and the per-cluster srt-slurm profiles with `python -m infx.launch`:

- configs/runners.yaml `clusters:` is the single static record per cluster
  (scheduler settings, volumes, model entries, image policy, srt-slurm settings)
- infx.launch.backends: small Backend interface; Slurm/pyxis implementation with
  one enroot import (submit-host, compute, all-nodes, pyxis, pre-staged, unchecked)
- drivers: srt-slurm (single/multi-node), script (SpeedBench), legacy (amd_utils,
  tilert, slated for removal)
- lifecycle: signal-safe cleanup, exit-code preservation, log snapshot before delete
- post-eval env from one declared workload-env pattern list; no host secrets
- historical replay (klaud, ingest recovery, changelog dispatch) runs the
  revision's own generator; klaud regeneration runs in a step without secrets
- workflows call the Python entrypoint; dead launcher paths removed

中文:将硬件启动器迁移为可插拔的 Python 启动器

使用 `python -m infx.launch` 取代 inferencex-e2e/runners/launch_*.sh、slurm_utils.sh、
runtime_settings.sh 以及各集群的 srt-slurm 配置:configs/runners.yaml 的 `clusters:`
成为每个集群唯一的静态配置;新增精简的 Backend 接口及 Slurm/pyxis 实现,统一 enroot
镜像导入;提供 srt-slurm、脚本与遗留(amd_utils、tilert,待移除)驱动;信号安全清理并保留
退出码;评估环境变量由统一的工作负载模式列表导出,不再透传主机密钥;历史版本回放改为运行
该版本自身的生成器;工作流改为调用 Python 入口并移除失效的启动路径。
# Conflicts:
#	inferencex-e2e/runners/launch_gb300-nv.sh
#	inferencex-e2e/runners/srt-slurm/patches/README.md
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

中文:operatorx 生成器测试夹具改用 clusters 形式的 runner 配置。
…t lanes

- srt model_paths map each recipe's own model.path alias to MODEL's staged
  checkpoint (by basename, node-local first); a short override table covers
  the real exceptions
- single-node points use the same lookup; hub-only clusters are cluster data
- srt lanes drop accidental per-cluster divergence: job status verified,
  logs copied, setup retried and health checks raised uniformly; one job-name
  scheme; --no-preflight derived from node-local checkpoints
- trim comments and docstrings; inline one-caller helpers

中文:srt 模型路径改为读取配方自身的 model.path 别名,并按 MODEL 名称映射到已暂存的检查点
(优先节点本地),仅保留少量例外;单节点共用同一查找逻辑;统一各集群 srt 流程中无意产生的差异
(作业状态校验、日志复制、setup 重试、健康检查、作业命名,并由节点本地检查点推导 --no-preflight);
精简注释与文档字符串。
中文:精简 srt 驱动中的辅助函数、注释与测试。
…late the launcher uv cache

srtctl always renders --account and falls back to the literal "default",
which CW clusters reject. Clusters without an account now use the submitting
user's default account from sacctmgr. The launcher venv step uses a per-job
uv cache because some runner homes hold a ~/.cache/uv uv cannot create.

中文:集群未配置账户时,srt 使用提交用户在 sacctmgr 中的默认账户(srtctl 否则会传入
无效的 "default");启动器虚拟环境步骤改用每个作业独立的 uv 缓存目录。
… single-node srt jobs

Real-cluster smoke runs showed srtctl's untyped --gpus-per-node and default
--segment are rejected by the CW partitions ("Requested node configuration is
not available"). Single-node jobs never emit a segment; clusters that set
gpus-per-node-directive: false get their typed slurm.gres instead. The uv
setup steps use a per-runner cache because some runner homes hold a
~/.cache/uv that uv cannot create.

中文:CW 集群改用带类型的 GRES 申请 GPU,单节点 srt 作业不再传入 --segment(真实集群冒烟测试中
二者被 CW 分区拒绝);uv 准备步骤改用每个 runner 独立的缓存目录。
CW single-node submissions still fail with "Requested node configuration is not
available" after typed GRES; the old salloc path there never passed --exclusive.

中文:允许集群让单节点 srt 作业不使用 --exclusive(CW 分区上此前的 salloc 路径从未使用该参数)。
…et as multi-node

On h100-cw a cached 397B FP8 checkpoint took ~24 min to load from shared storage
plus ~12 min of DeepGEMM warmup, past srtctl's 1800 s default.

中文:单节点 srt 作业与多节点使用相同的 2 小时健康检查时限(h100-cw 上模型加载与预热超过 srtctl
默认的 1800 秒)。
…ronment

The TileRT smoke run on b200-nscale got the runner host's UCX_NET_DEVICES instead
of the cluster's eight HCAs. As with the old runtime_settings.sh, cluster settings
now win over the host environment and lose only to a point's additional-settings.

中文:集群运行时设置优先于 runner 主机自身的环境变量,仅由测试点的 additional-settings 覆盖
(与原 runtime_settings.sh 的行为一致;修复 b200-nscale TileRT 使用了主机 UCX_NET_DEVICES 的问题)。
…one env precedence rule

- every srtctl job is named inferencex-<runner> and cleanup cancels both names,
  so other tools' scancel --name=<runner> cannot kill InferenceX jobs
- single-node points read the Hub unless the cluster stages them (b200-nscale,
  b300-dsxe); B300 DeepSeek-V4-Pro-0813 reads NVMe only for vLLM
- cluster env sources override the runner host but never a point's
  additional-settings; the Slurm account resolves once per launch
- TileRT fork keeps its own health and recipe behavior; health floor applies
  only inside health_check blocks
- CW clusters drop --exclusive everywhere; uv Pythons also stay per runner
- tests use synthetic records instead of checked-in values

中文:所有 srtctl 作业统一命名为 inferencex-<runner>,清理同时取消两种名称;单节点默认使用 Hub,
仅 b200-nscale、b300-dsxe 使用暂存检查点;集群环境变量优先于主机环境但不覆盖测试点设置;
Slurm 账户每次启动解析一次;TileRT fork 保留原有行为;CW 集群不再使用 --exclusive;测试改用合成配置。
@adibarra adibarra changed the title refactor(launch): port hardware launchers to a pluggable Python launcher / 重构启动器:将硬件启动器迁移为可插拔的 Python 启动器 refactor(launch): port hardware launchers to a pluggable Python launcher Sep 29, 2026
The cluster renamed compute-0; sbatch rejects the old name.

中文:mi300x-amd 集群分区已改名为 MI300X-UBUNTU,旧名 compute-0 被 sbatch 拒绝。
b300-dsxe requests cpus-per-task=192 (its full node) since #3477, but srtctl runs
TRT-LLM with one task per GPU, so dynamo-trt submissions asked for 8x192 CPUs
per node and sbatch refused them.

中文:TRT-LLM 的 srt 作业不再申请整节点 CPU(b300 的 cpus-per-task=192 与每 GPU 一个任务冲突,
导致 dynamo-trt 作业被 sbatch 拒绝)。
Instead of dropping cpus-per-task, divide it by the tasks srtctl places on each
node, so b300 TRT jobs still reserve the whole node (8 x 24 CPUs).

中文:TRT-LLM 每个 GPU 一个任务时,将节点 CPU 配额按任务数均分(b300 为 8 × 24),而不是不申请。
#3477 set cpus-per-task=192 (a whole b300 node). srtctl runs TRT-LLM with one
task per GPU, so dynamo-trt jobs asked for 8x192 CPUs per node and sbatch
refused them. cpus-per-gpu=24 reserves the same 192 CPUs per node for every
task layout (verified with sbatch --test-only on batch_1). Clusters declare
either field; the TRT task-count special case is gone.

中文:b300 改为按 GPU 申请 CPU(cpus-per-gpu=24),替代 #3477 的 cpus-per-task=192。后者在
TRT-LLM 每 GPU 一个任务时申请 8×192 个 CPU 被 sbatch 拒绝;新设置在所有任务布局下均保留整节点
192 个 CPU(已在 batch_1 上用 sbatch --test-only 验证)。
runners.yaml listed gb200-nv_0..3, but the registered runners are gb200-nv_00..17
(no _08). The launcher resolves the cluster from these labels, so every gb200
job failed before submission. The generated sweep matrix is unchanged.

中文:更新 gb200-nv 的 runner 列表为实际注册的 gb200-nv_00..17(无 _08),此前启动器无法解析集群;
生成的扫描矩阵不变。
Carries main's srt-slurm v2.36.0 pin, its per-profile cluster names (now rendered
from the cluster id), the MI300X partition rename (already on this branch), and
the AgentX one-concurrency-per-deployment check, whose test now drives infx.launch.
python3 -P fails on images older than Python 3.11 (the MI355X sglang-disagg
image is one); PYTHONSAFEPATH gives the same isolation where supported and is
ignored elsewhere.
Docstrings stay; lint pragmas and action version pins are kept.
…ners.yaml

The job tag is the cluster's GPU label (srt-slurm.job-tag); MORI_RDMA_TC is the
MI355X fabric's traffic class, now in the cluster env like b200-nb's UCX device.
中文:恢复本分支未改动代码的文件中 main 原有的注释。
@adibarra
adibarra marked this pull request as ready for review September 29, 2026 22:49
@adibarra
adibarra requested a review from a team September 29, 2026 22:49
This was referenced Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant