From 032c3c25d0a0ac6a9d6b4dc6691e01f8168e6c8b Mon Sep 17 00:00:00 2001 From: Cam Quilici Date: Wed, 30 Sep 2026 09:59:03 -0500 Subject: [PATCH 1/4] refactor(tilert): use upstream srt-slurm integration --- .github/AGENT_OPERATIONS.md | 2 +- .../multi_node/srt-slurm-recipes/RECIPES.md | 5 +- .../srt-slurm-recipes/RECIPES_zh.md | 5 +- .../docs/configuration-procedures.md | 4 +- .../docs/configuration-procedures_zh.md | 4 +- .../infx/launch/drivers/srt/__init__.py | 11 +--- .../infx/launch/drivers/srt/checkout.py | 58 ++++--------------- .../infx/launch/drivers/srt/config.py | 9 +-- .../infx/launch/drivers/srt/models.py | 9 +-- .../infx/launch/drivers/srt/submit.py | 58 +++++-------------- .../infx/tests/launch/test_srt_config.py | 2 +- .../infx/tests/launch/test_srt_driver.py | 14 ++--- .../infx/tests/launch/test_srt_policy.py | 5 +- .../runners/srt-slurm/patches/README.md | 2 +- 14 files changed, 52 insertions(+), 136 deletions(-) diff --git a/.github/AGENT_OPERATIONS.md b/.github/AGENT_OPERATIONS.md index 9ad1ea8e57..dfd6500e76 100644 --- a/.github/AGENT_OPERATIONS.md +++ b/.github/AGENT_OPERATIONS.md @@ -94,7 +94,7 @@ Multinode disaggregated results add `prefill_gpu_energy_j`, `decode_gpu_energy_j Every power result — valid or invalid, single-node or multinode — carries `power_metric_schema_version`. Version 2 defines each unprefixed `joules_per_*` field as whole-deployment GPU-board energy over the named denominator; role-scoped energy uses the explicit `prefill_*` / `decode_*` keys. Rows without the field predate the whole-deployment switch and their unprefixed joules are not comparable across topologies. -For srt-slurm recipes, `telemetry.enabled: true` with `telemetry.dcgm_exporter` enables official energy collection. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for the shared srt-slurm commit, used by both power and non-power NVIDIA lanes. TileRT is the single documented fork exception. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated. +For srt-slurm recipes, `telemetry.enabled: true` with `telemetry.dcgm_exporter` enables official energy collection. The Git submodule pointer at `inferencex-e2e/utils/srt-slurm` is the source of truth for every srt-slurm job, including TileRT. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. The aggregate-power and AgentX power tests validate telemetry and provenance. These local tests do not prove hardware power collection. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated. Power audit artifacts are named `power_audit_` and contain `power_validation_.json` for single-node runs or `power_validation__*.json` for multinode runs. They are uploaded even when validation fails. diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md index 3527d72748..42c3c5bb73 100644 --- a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md @@ -27,9 +27,9 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml Shared runtime assets stay under `configs/` beside the model directories; they are not standalone recipes. The four files in `configs/dsv4-moe-load-balancer-configs/` are copied verbatim from NVIDIA/srt-slurm commit `deb1dfd9934398664f92d194169c183e009da83b`, preserving the EPLB initial expert assignments formerly used by the DSV4 TRT recipes; no checked-in recipe currently references them. The srt driver ([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py)) stages them into the job checkout's `configs/` directory for the recipes' bind mounts. Keeping a recipe in this tree does not activate it; the master configs determine the benchmark matrix. -## TileRT exception +## TileRT -For `FRAMEWORK=tilert`, the srt driver checks out the SemiAnalysisAI/srt-slurm fork directly at `6bc3f306bdafa1edfb5dded2fcda8f1ccede1bde` into the job checkout. This is the schema-2 TileRT port in [SemiAnalysisAI/srt-slurm#13](https://github.com/SemiAnalysisAI/srt-slurm/pull/13). It is the only alternate checkout; its pin lives in `SRT_FORKS` in [`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py) because the TileRT backend and router are absent from the NVIDIA pin. TileRT uses the same schema-2 recipe layout and native post-eval dispatch as NVIDIA. TileRT jobs need network access to the fork at setup time. Remove the fork exception once those features are available upstream. +TileRT uses the pinned upstream srt-slurm submodule. Recipes select `roles.prefill.engine: vllm`, `roles.decode.engine: tilert`, and `frontend.type: tilert-router`. ## Schema 2 and master configuration @@ -58,7 +58,6 @@ Install the shared pin in an isolated environment, then use its CLI: srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang # Repeat for the other model/engine directories. -# Use the pinned TileRT fork for tilert/ recipe directories. python -m pytest infx/tests/matrix/ -q python -m infx.matrix.generate full-sweep \ --config-files configs/nvidia-master.yaml \ diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md index de36f6e0ee..41d431845f 100644 --- a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md @@ -27,9 +27,9 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml 共享运行时资源保留在模型目录旁的 `configs/` 中,不属于独立基准测试配置。`configs/dsv4-moe-load-balancer-configs/` 中的四个文件原样取自 NVIDIA/srt-slurm 提交 `deb1dfd9934398664f92d194169c183e009da83b`,保留了此前 DSV4 TRT 配置使用的 EPLB 初始专家分配;目前没有已提交的配置引用这些文件。srt driver([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py))将这些文件复制到作业仓库的 `configs/` 目录,供配置中的绑定挂载使用。将配置文件放入本目录不会启用该配置;实际基准测试矩阵由主配置决定。 -## TileRT 例外 +## TileRT -当 `FRAMEWORK=tilert` 时,srt driver 直接从 SemiAnalysisAI/srt-slurm 分支仓库获取提交 `6bc3f306bdafa1edfb5dded2fcda8f1ccede1bde`,检出到作业目录。该版本为 [SemiAnalysisAI/srt-slurm#13](https://github.com/SemiAnalysisAI/srt-slurm/pull/13) 中支持 schema 2 的 TileRT 移植。这是唯一的备用检出路径;由于统一的 NVIDIA 版本尚未包含 TileRT 后端和路由器,该例外的固定提交在 [`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py) 的 `SRT_FORKS` 中指定。TileRT 使用与 NVIDIA 相同的 schema 2 配置结构和原生评估调度。TileRT 作业在准备阶段需要通过网络访问分支仓库。上游支持这些功能后,应删除此分支仓库例外。 +TileRT 使用固定版本的上游 srt-slurm 子模块。配置指定 `roles.prefill.engine: vllm`、`roles.decode.engine: tilert` 和 `frontend.type: tilert-router`。 ## Schema 2 与主配置 @@ -58,7 +58,6 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang # 对其他模型/引擎目录重复执行。 -# 迁移 tilert/ 配置目录时,使用固定提交的 TileRT 分支仓库。 python -m pytest infx/tests/matrix/ -q python -m infx.matrix.generate full-sweep \ --config-files configs/nvidia-master.yaml \ diff --git a/inferencex-e2e/docs/configuration-procedures.md b/inferencex-e2e/docs/configuration-procedures.md index 9cbb9bd093..7b1f746529 100644 --- a/inferencex-e2e/docs/configuration-procedures.md +++ b/inferencex-e2e/docs/configuration-procedures.md @@ -25,7 +25,7 @@ Delete retired entries from the active master configs; they are not archived. Gi ## Dependency submodules -Git records the exact dependency commits. [`.gitmodules`](../../.gitmodules) defines the repositories: AIPerf at `utils/aiperf`, NVIDIA srt-slurm at `utils/srt-slurm`. TileRT is a documented manual fork checkout in the srt driver ([`infx/launch/drivers/srt/checkout.py`](../infx/launch/drivers/srt/checkout.py)), not a separate submodule. +Git records the exact dependency commits. [`.gitmodules`](../../.gitmodules) defines the repositories: AIPerf at `utils/aiperf`, NVIDIA srt-slurm at `utils/srt-slurm`. All srt-slurm jobs, including TileRT, use the pinned upstream submodule. Initialize them before running benchmarks locally: @@ -33,7 +33,7 @@ Initialize them before running benchmarks locally: git submodule update --init ``` -To upgrade, fetch and check out the desired commit inside the relevant submodule, then commit the updated submodule pointer in InferenceX. Benchmark workflows already initialize submodules. Slurm launchers make a local Git clone for each job so recipe staging and runtime writes do not modify the submodule, and record the actual commit for result provenance. NVIDIA setup clones locally; TileRT setup fetches its pinned fork commit over the network. +To upgrade, fetch and check out the desired commit inside the relevant submodule, then commit the updated submodule pointer in InferenceX. Benchmark workflows already initialize submodules. Slurm launchers make a local Git clone for each job so recipe staging and runtime writes do not modify the submodule, and record the actual commit for result provenance. Single-node fixed-sequence recipes use NVIDIA upstream srt-slurm. ATOM recipes use the native `atomesh` frontend with one aggregate worker and diff --git a/inferencex-e2e/docs/configuration-procedures_zh.md b/inferencex-e2e/docs/configuration-procedures_zh.md index 83a00ee528..e74c9c5d58 100644 --- a/inferencex-e2e/docs/configuration-procedures_zh.md +++ b/inferencex-e2e/docs/configuration-procedures_zh.md @@ -25,7 +25,7 @@ ## 依赖子模块 -Git 记录依赖的精确提交版本。[`.gitmodules`](../../.gitmodules) 定义各仓库:AIPerf 位于 `utils/aiperf`,NVIDIA srt-slurm 位于 `utils/srt-slurm`。TileRT 由 srt 驱动([`infx/launch/drivers/srt/checkout.py`](../infx/launch/drivers/srt/checkout.py))手动检出已记录的分支仓库,不是独立子模块。 +Git 记录依赖的精确提交版本。[`.gitmodules`](../../.gitmodules) 定义各仓库:AIPerf 位于 `utils/aiperf`,NVIDIA srt-slurm 位于 `utils/srt-slurm`。包括 TileRT 在内的所有 srt-slurm 作业均使用固定的上游子模块。 本地运行基准测试前,先初始化子模块: @@ -33,7 +33,7 @@ Git 记录依赖的精确提交版本。[`.gitmodules`](../../.gitmodules) 定 git submodule update --init ``` -升级时,在对应子模块中获取并检出目标提交,再将更新后的子模块指针提交到 InferenceX。基准测试工作流已配置为自动初始化子模块。Slurm 启动器为每个作业创建本地 Git 克隆,避免配方准备和运行时写入修改子模块,并记录实际提交以供结果溯源。NVIDIA 启动器使用本地克隆;TileRT 启动器通过网络获取固定的分支提交。 +升级时,在对应子模块中获取并检出目标提交,再将更新后的子模块指针提交到 InferenceX。基准测试工作流已配置为自动初始化子模块。Slurm 启动器为每个作业创建本地 Git 克隆,避免配方准备和运行时写入修改子模块,并记录实际提交以供结果溯源。 单节点固定序列长度配方使用 NVIDIA 上游 srt-slurm。ATOM 配方使用原生 `atomesh` frontend、一个聚合 worker,并设置 `enable_multiple_frontends: false`。旧版基准 worker diff --git a/inferencex-e2e/infx/launch/drivers/srt/__init__.py b/inferencex-e2e/infx/launch/drivers/srt/__init__.py index 4a3aa0cd3d..4f7099498e 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/__init__.py +++ b/inferencex-e2e/infx/launch/drivers/srt/__init__.py @@ -22,7 +22,6 @@ from infx.launch.context import Launch from infx.launch.drivers.srt import collect, config, lanes, models, power, submit from infx.launch.drivers.srt.checkout import ( - SRT_FORKS, checkout_dir, compute_workspace, install_srtctl, @@ -65,7 +64,6 @@ def run_single_node(launch: Launch) -> int: mounts=[(str(hf_cache), request.hf_hub_cache)], single_node=True, account=run.account, - fork=checkout.fork, ) config.create_volume_mounts(run) config.write(checkout.root / "srtslurm.yaml", config.render(run.cluster, job_config)) @@ -128,8 +126,7 @@ def run_multinode(launch: Launch) -> int: decision = power.resolve_power(launch.cluster.id, launch.path, request) model = models.checkpoint(launch.cluster, request) served = models.served_path(launch.cluster, request, model) - fork = request.framework in SRT_FORKS - model_paths = models.model_paths(launch.cluster, request, config_file, served, fork=fork) + model_paths = models.model_paths(launch.cluster, request, config_file, served) run = SrtRun.create(launch, request, models.job_env(launch.cluster, request, served)) preflight = run.srt.preflight and not (model_paths and model and model.node_local) if request.framework == "tilert": @@ -151,10 +148,8 @@ def run_multinode(launch: Launch) -> int: conc_list = request.env.get("CONC_LIST", "") if decision.dcgm else None job_name = srtctl_job_name(request.runner_name) prepare_recipe(checkout.root, config_file, job_name, run.srt.dist_timeout_s, conc_list) - arguments = submit.multinode_arguments( - run, lane, checkout, config_file, overrides, preflight=preflight - ) - manifest = None if checkout.fork else run.workspace / submit.MULTINODE_SUBMISSION + arguments = submit.multinode_arguments(run, lane, config_file, overrides, preflight=preflight) + manifest = run.workspace / submit.MULTINODE_SUBMISSION submitted = submit.Submitted(manifest=manifest) run.life.callback(submitted.cancel, run.backend) if rc := submit.submit_lane(run, submitted, checkout, config_file, arguments): diff --git a/inferencex-e2e/infx/launch/drivers/srt/checkout.py b/inferencex-e2e/infx/launch/drivers/srt/checkout.py index 605d7080a8..822ef43e3b 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/checkout.py +++ b/inferencex-e2e/infx/launch/drivers/srt/checkout.py @@ -24,34 +24,12 @@ UV_INSTALLER = "https://astral.sh/uv/install.sh" -@dataclass(frozen=True) -class SrtFork: - """A framework's srt-slurm fork, checked out instead of the pinned submodule. - - Forks get no InferenceX patches, predate ``--json``, ``--no-preflight`` and - ``benchmark.stream_output``, and report their job only in prose. Their jobs keep - srtctl's own health-check default, and may run recipes the workspace mirror lacks. - """ - - url: str - commit: str - - -SRT_FORKS: dict[str, SrtFork] = { - "tilert": SrtFork( - "https://github.com/SemiAnalysisAI/srt-slurm.git", - "6bc3f306bdafa1edfb5dded2fcda8f1ccede1bde", - ), -} - - @dataclass(frozen=True) class Checkout: """A job-local srt-slurm checkout with InferenceX recipes staged.""" root: Path commit: str - fork: bool @property def venv(self) -> Path: @@ -95,31 +73,17 @@ def prepare_checkout(run: SrtRun, destination: Path, *, power: bool) -> Checkout if destination.exists(): print(f"Removing existing {destination}...", flush=True) shutil.rmtree(destination) - fork = SRT_FORKS.get(run.request.framework) - if fork is not None: - _git("init", "--quiet", destination) - _git("-C", destination, "remote", "add", "origin", fork.url) - _git("-C", destination, "fetch", "--quiet", "--depth=1", "origin", fork.commit) - _git("-C", destination, "checkout", "--quiet", "--detach", fork.commit) - commit = fork.commit - else: - source = repository_root() / SUBMODULE - if not (source / ".git").exists(): - raise LaunchError( - "Missing srt-slurm submodule; run git submodule update --init before launching." - ) - commit = _git("-C", source, "rev-parse", "HEAD", capture=True) - _git( - "-c", - "advice.detachedHead=false", - "clone", - "--quiet", - "--no-hardlinks", - source, - destination, + source = repository_root() / SUBMODULE + if not (source / ".git").exists(): + raise LaunchError( + "Missing srt-slurm submodule; run git submodule update --init before launching." ) - for patch in sorted((run.workspace / PATCHES).glob("*.patch")): - _git("-C", destination, "apply", patch) + commit = _git("-C", source, "rev-parse", "HEAD", capture=True) + _git( + "-c", "advice.detachedHead=false", "clone", "--quiet", "--no-hardlinks", source, destination + ) + for patch in sorted((run.workspace / PATCHES).glob("*.patch")): + _git("-C", destination, "apply", patch) head = _git("-C", destination, "rev-parse", "HEAD", capture=True) if head != commit: raise LaunchError(f"srt-slurm checkout is at {head}, expected {commit}") @@ -133,7 +97,7 @@ def prepare_checkout(run: SrtRun, destination: Path, *, power: bool) -> Checkout shutil.copytree(recipes, destination / "recipes", symlinks=True, dirs_exist_ok=True) (destination / RECIPES_MIRROR).symlink_to("../../recipes") shutil.copytree(recipes / "configs", destination / "configs", symlinks=True, dirs_exist_ok=True) - return Checkout(destination, commit, fork is not None) + return Checkout(destination, commit) def _uv(run: SrtRun) -> str: diff --git a/inferencex-e2e/infx/launch/drivers/srt/config.py b/inferencex-e2e/infx/launch/drivers/srt/config.py index f44fecbfda..569156c4c2 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/config.py +++ b/inferencex-e2e/infx/launch/drivers/srt/config.py @@ -49,7 +49,6 @@ class SrtJob: mounts: Sequence[tuple[str, str]] = () single_node: bool = False account: str | None = None - fork: bool = False def pyxis_spelling(image: str) -> str: @@ -108,8 +107,7 @@ def render(cluster: Cluster, job: SrtJob) -> dict[str, Any]: config["output_dir"] = str(srt.outputs) if job.model_paths: config["model_paths"] = dict(job.model_paths) - if not job.fork: - config["default_health_check"] = dict(HEALTH_CHECK) + config["default_health_check"] = dict(HEALTH_CHECK) containers = dict.fromkeys(srt.container_aliases, job.container) containers[job.image] = job.container containers[pyxis_spelling(job.image)] = job.container @@ -218,8 +216,8 @@ def write_lane_config( ) containers: dict[str, str] = {} if request.framework == "tilert": - prefill = backend.stage_image(request.env["PREFILL_IMAGE"]).reference - containers = {"tilert-decode": container, "tilert-prefill": prefill} + prefill_image = request.env["PREFILL_IMAGE"] + containers[prefill_image] = backend.stage_image(prefill_image).reference if power.dcgm: exporter = backend.stage_image(DCGM_EXPORTER_IMAGE, helper="dcgm-exporter") (run.workspace / EXPORTER_PROVENANCE).write_text(f"{backend.image_provenance(exporter)}\n") @@ -236,7 +234,6 @@ def write_lane_config( model_paths=model_paths, mounts=lane_mounts(run, lane), account=run.account, - fork=checkout.fork, ) config_yaml = checkout.root / "srtslurm.yaml" write(config_yaml, render(run.cluster, job)) diff --git a/inferencex-e2e/infx/launch/drivers/srt/models.py b/inferencex-e2e/infx/launch/drivers/srt/models.py index a8f1dad3fc..39799fdb0f 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/models.py +++ b/inferencex-e2e/infx/launch/drivers/srt/models.py @@ -172,17 +172,10 @@ def model_paths( request: LaunchRequest, config_file: str, served: str | None, - *, - fork: bool = False, ) -> dict[str, str]: - """srtslurm.yaml ``model_paths``: each alias of ``config_file``'s recipe mapped to ``served``. - - A ``fork`` may run a recipe of its own, which the workspace mirror lacks; it maps none. - """ + """srtslurm.yaml ``model_paths``: each recipe alias mapped to ``served``.""" recipe = recipe_mirror_path(request.workspace, config_file) if not recipe.is_file(): - if fork: - return {} raise LaunchError(f"CONFIG_FILE {config_file} is not in the recipe mirror: {recipe}") aliases = recipe_aliases(recipe) if aliases and served is None: diff --git a/inferencex-e2e/infx/launch/drivers/srt/submit.py b/inferencex-e2e/infx/launch/drivers/srt/submit.py index a6319192f2..9995eebd79 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/submit.py +++ b/inferencex-e2e/infx/launch/drivers/srt/submit.py @@ -7,7 +7,6 @@ import json import re import subprocess -import sys from collections.abc import Mapping from dataclasses import dataclass from datetime import datetime @@ -44,7 +43,6 @@ "SGLANG_TORCH_PROFILER_DIR", "VLLM_TORCH_PROFILER_DIR", ) # fmt: skip _SHELL_NAME = re.compile(r"[A-Za-z_][A-Za-z0-9_]*") -_PROSE_JOB_IDS = (re.compile(r"✅ Job ([0-9]+)"), re.compile(r"Job ([0-9]+)")) def eval_args(env: Mapping[str, str], command: str) -> list[str]: @@ -85,14 +83,13 @@ def apply( config: str, arguments: list[str], *, - stdout: Path | None = None, + stdout: Path, ) -> subprocess.CompletedProcess[str]: """Run ``srtctl apply`` for ``config``, through the golden AgentX acceptance planner. Every container starts in the workspace mount, as the legacy launchers did: PyTorch's generated module imports fail from / with PYTHONPYCACHEPREFIX set. - ``arguments`` still win. ``stdout`` receives srtctl's JSON manifest; without - it stdout and stderr are captured and echoed. + ``arguments`` still win. ``stdout`` receives srtctl's JSON manifest. """ argv = [ str(checkout.venv / "bin/python"), "-m", "infx.srt_slurm.synthetic_acceptance", @@ -100,11 +97,6 @@ def apply( "--set", 'srun_options.container-workdir="/infmax-workspace"', *arguments, ] # fmt: skip env = {**run.env, "RUNNER_NAME": srtctl_job_name(run.request.runner_name)} - if stdout is None: - result = proc.run(argv, env=env, cwd=checkout.root, capture=True) - sys.stdout.write(result.stdout + result.stderr) - sys.stdout.flush() - return result proc.echo(argv, env) with stdout.open("w") as handle: rc = subprocess.run(argv, env=env, cwd=checkout.root, stdout=handle, check=False).returncode @@ -118,15 +110,13 @@ def _job(backend: SlurmBackend, job_id: str, output: Path) -> SlurmJob: @dataclass class Submitted: - """The job a submission created, once known; its ``--json`` manifest, if any.""" + """The job a submission created, once known, and its ``--json`` manifest.""" - manifest: Path | None = None + manifest: Path job: SlurmJob | None = None def read_manifest(self, backend: SlurmBackend) -> SlurmJob: """Adopt the job srtctl reported in its ``--json`` manifest.""" - if self.manifest is None: - raise LaunchError("this submission writes no manifest") try: job_id, output = submission_fields(self.manifest) except (OSError, ValueError, KeyError, TypeError) as error: @@ -138,7 +128,7 @@ def read_manifest(self, backend: SlurmBackend) -> SlurmJob: def recover(self, backend: SlurmBackend) -> SlurmJob | None: """The job, read from the manifest when the submission was interrupted after writing it.""" - if self.job is None and self.manifest is not None and self.manifest.is_file(): + if self.job is None and self.manifest.is_file(): with contextlib.suppress(LaunchError): self.read_manifest(backend) return self.job @@ -154,34 +144,17 @@ def adopted(self) -> SlurmJob: return self.job -def _prose_job(backend: SlurmBackend, output: str, checkout: Checkout) -> SlurmJob: - """Adopt the one job id in srtctl's human-readable output.""" - for pattern in _PROSE_JOB_IDS: - ids = sorted(set(pattern.findall(output))) - if len(ids) > 1: - raise LaunchError(f"srtctl submitted several jobs: {', '.join(ids)}") - if ids: - return _job(backend, ids[0], checkout.root / "outputs" / ids[0]) - raise LaunchError("Failed to extract JOB_ID from srtctl output") - - def submit_lane( run: SrtRun, submitted: Submitted, checkout: Checkout, config_file: str, arguments: list[str] ) -> int: """Submit a multi-node lane job, record it in ``submitted``, and return srtctl's exit code.""" - if submitted.manifest is None: - applied = apply(run, checkout, config_file, arguments) - if applied.returncode: - return applied.returncode - submitted.job = _prose_job(run.backend, applied.stdout + applied.stderr, checkout) - else: - applied = apply( - run, checkout, config_file, [*arguments, "--json", "--yes"], stdout=submitted.manifest - ) - print(applied.stdout, end="", flush=True) - if applied.returncode: - return applied.returncode - submitted.read_manifest(run.backend) + applied = apply( + run, checkout, config_file, [*arguments, "--json", "--yes"], stdout=submitted.manifest + ) + print(applied.stdout, end="", flush=True) + if applied.returncode: + return applied.returncode + submitted.read_manifest(run.backend) print(f"Extracted JOB_ID: {submitted.adopted().id}", flush=True) return 0 @@ -189,7 +162,6 @@ def submit_lane( def multinode_arguments( run: SrtRun, lane: SrtLane, - checkout: Checkout, config_file: str, overrides: list[str], *, @@ -197,15 +169,15 @@ def multinode_arguments( ) -> list[str]: """The ``srtctl apply`` arguments of a multi-node lane submission.""" request = run.request - stream = [] if checkout.fork else ["--set", "benchmark.stream_output=true"] arguments = [ *eval_args(run.env, MULTINODE_EVAL_COMMAND), - *stream, + "--set", + "benchmark.stream_output=true", *overrides, "-f", config_file, ] - if not checkout.fork and not preflight: + if not preflight: arguments.append("--no-preflight") if run.srt.job_tag is not None: isl, osl = request.env.get("ISL", ""), request.env.get("OSL", "") diff --git a/inferencex-e2e/infx/tests/launch/test_srt_config.py b/inferencex-e2e/infx/tests/launch/test_srt_config.py index b7a400ada8..dffd525478 100644 --- a/inferencex-e2e/infx/tests/launch/test_srt_config.py +++ b/inferencex-e2e/infx/tests/launch/test_srt_config.py @@ -205,7 +205,7 @@ def rsync(source: Path, staging: Path, *, exclude: tuple[str, ...]) -> Path: return staging run = SimpleNamespace(workspace=workspace, backend=SimpleNamespace(stage_workspace=rsync)) - checkout = Checkout(tmp_path / "runs/srt-slurm-1-1-abc", "sha", False) + checkout = Checkout(tmp_path / "runs/srt-slurm-1-1-abc", "sha") staged = compute_workspace(run, checkout, shared=True) assert staged == tmp_path / "runs/infmax-workspace-1-1-abc" diff --git a/inferencex-e2e/infx/tests/launch/test_srt_driver.py b/inferencex-e2e/infx/tests/launch/test_srt_driver.py index 1377d2beac..eb884ae51b 100644 --- a/inferencex-e2e/infx/tests/launch/test_srt_driver.py +++ b/inferencex-e2e/infx/tests/launch/test_srt_driver.py @@ -19,7 +19,6 @@ from infx.launch.__main__ import main from infx.launch.drivers.srt import lanes, models -from infx.launch.drivers.srt.checkout import SRT_FORKS from infx.launch.drivers.srt.lanes import LaneMount, SrtLane from infx.launch.drivers.srt.models import Override from infx.launch.policy import LaunchPath, Match @@ -394,22 +393,21 @@ def test_b300_flash_agentx_reenters_inside_a_batch_allocation(harness): assert list(runner_temp.glob("srt-batch.*.sh")) == [] -def test_tilert_native_lane_runs_on_its_fork_without_patches(harness): +def test_tilert_native_lane_uses_upstream_submission_and_role_images(harness): env = lane_env( harness, "b200-nscale", MODEL_PREFIX="glm5.1", PRECISION="fp8", FRAMEWORK="tilert", MODEL="zai-org/GLM-5.1-FP8", SPEC_DECODING="mtp", IS_AGENTIC="1", ISL="0", OSL="0", FAKE_RESULTS="agentic", PREFILL_IMAGE="prefill:tag", - FAKE_SRT_COMMIT=SRT_FORKS["tilert"].commit, ) # fmt: skip assert_ok(launch(env, harness.config, harness.workspace)) - assert not any(" apply " in f" {line} " for line in lines(harness.logs, "git")) + assert any(" apply " in f" {line} " for line in lines(harness.logs, "git")) [call] = srtctl_calls(harness.logs) - assert not {"--json", "--no-preflight", "benchmark.stream_output=true"} & set(call["argv"]) + assert {"--json", "--no-preflight", "benchmark.stream_output=true"} <= set(call["argv"]) config = srtslurm(Path(call["cwd"])) - assert "default_health_check" not in config - assert config["containers"]["tilert-decode"].endswith("/test_tag.sqsh") - assert config["containers"]["tilert-prefill"].endswith("/prefill_tag.sqsh") + assert config["default_health_check"]["max_attempts"] > 0 + assert config["containers"]["test:tag"].endswith("/test_tag.sqsh") + assert config["containers"]["prefill:tag"].endswith("/prefill_tag.sqsh") assert config["default_mounts"][str(harness.workspace)] == "/infmax-workspace" assert json.loads((harness.workspace / "point-identity_conc4.json").read_text()) == {"conc": 4} diff --git a/inferencex-e2e/infx/tests/launch/test_srt_policy.py b/inferencex-e2e/infx/tests/launch/test_srt_policy.py index ce4b3d00d5..1f6a8ad26f 100644 --- a/inferencex-e2e/infx/tests/launch/test_srt_policy.py +++ b/inferencex-e2e/infx/tests/launch/test_srt_policy.py @@ -145,11 +145,10 @@ def test_every_recipe_alias_maps_to_the_checkpoint_and_literals_pass_through(tmp assert resolved == {alias: str(tmp_path / path) for alias, path in paths.items()} -def test_only_a_fork_may_run_a_recipe_the_mirror_lacks(tmp_path): +def test_missing_recipe_fails_before_model_resolution(tmp_path): c, point = cluster(tmp_path), request(MODEL="org/M", GITHUB_WORKSPACE=str(tmp_path)) - assert model_paths(c, point, "recipes/fork-only.yaml", "/m", fork=True) == {} with pytest.raises(LaunchError, match="not in the recipe mirror"): - model_paths(c, point, "recipes/fork-only.yaml", "/m") + model_paths(c, point, "recipes/missing.yaml", "/m") def test_matching_single_node_points_read_the_shared_hub_cache(tmp_path, monkeypatch): diff --git a/inferencex-e2e/runners/srt-slurm/patches/README.md b/inferencex-e2e/runners/srt-slurm/patches/README.md index ba9db355f1..e6ba038f31 100644 --- a/inferencex-e2e/runners/srt-slurm/patches/README.md +++ b/inferencex-e2e/runners/srt-slurm/patches/README.md @@ -2,7 +2,7 @@ As shown in the [CODEOWNERS](../../../../.github/CODEOWNERS) file, InferenceX core maintainers control the patches here, so there is ZERO dependency on upstream srt-slurm NVIDIA maintainers for any srt-slurm patch, ensuring that InferenceX is vendor neutral. srt-slurm allows for declarative YAML launching instead of the previous unmaintainable, low-quality pile of 1000+ bash scripts. -The srt driver ([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py)) applies every `*.patch` here to the job's srt-slurm clone after checking out the pinned submodule. TileRT jobs use the fork checkout and skip these patches. +The srt driver ([`infx/launch/drivers/srt/checkout.py`](../../../infx/launch/drivers/srt/checkout.py)) applies every `*.patch` here to the job's srt-slurm clone after checking out the pinned submodule. Each patch is a temporary fix for an open upstream PR. When the PR merges and the submodule pin includes it, delete the patch and its row. From b42f90df3af597dfae9ff4d270727a62efaefc76 Mon Sep 17 00:00:00 2001 From: Cam Quilici Date: Wed, 30 Sep 2026 14:05:12 -0500 Subject: [PATCH 2/4] fix(eval): initialize native multi-node evaluation clients --- inferencex-e2e/benchmarks/benchmark_lib.sh | 23 +++++++-- inferencex-e2e/docs/eval-agentx-procedures.md | 2 + .../docs/eval-agentx-procedures_zh.md | 2 + .../infx/launch/drivers/srt/submit.py | 3 +- .../tests/evals/test_run_eval_dispatch.py | 36 ++++++++++++++ .../infx/tests/launch/test_srt_driver.py | 47 +++++++++++++++++++ 6 files changed, 108 insertions(+), 5 deletions(-) diff --git a/inferencex-e2e/benchmarks/benchmark_lib.sh b/inferencex-e2e/benchmarks/benchmark_lib.sh index 08fc8a11e0..0bd6c2b707 100644 --- a/inferencex-e2e/benchmarks/benchmark_lib.sh +++ b/inferencex-e2e/benchmarks/benchmark_lib.sh @@ -2003,11 +2003,26 @@ get_native_max_context_length() { if [ -n "${MODEL_PATH:-}" ] && [ -d "${MODEL_PATH}" ]; then model_path="${MODEL_PATH}" fi - python3 -c " + python3 - "$model_path" <<'PY' +import json +import sys +from pathlib import Path + +fields = ['max_position_embeddings', 'max_sequence_length', 'seq_length', 'n_positions'] +try: + config = json.loads((Path(sys.argv[1]) / 'config.json').read_text()) + for field in fields: + value = config.get(field) + if type(value) is int and value > 0: + print(value) + sys.exit(0) +except (OSError, ValueError, AttributeError): + pass + try: from transformers import AutoConfig - config = AutoConfig.from_pretrained('${model_path}', trust_remote_code=True) - for attr in ['max_position_embeddings', 'max_sequence_length', 'seq_length', 'n_positions']: + config = AutoConfig.from_pretrained(sys.argv[1], trust_remote_code=True) + for attr in fields: if hasattr(config, attr): print(getattr(config, attr)) break @@ -2015,7 +2030,7 @@ try: print(0) except Exception: print(0) -" +PY } # Requested benchmark context capped at the model's native max. Sets diff --git a/inferencex-e2e/docs/eval-agentx-procedures.md b/inferencex-e2e/docs/eval-agentx-procedures.md index c93a93cfd3..c78a5ed0f6 100644 --- a/inferencex-e2e/docs/eval-agentx-procedures.md +++ b/inferencex-e2e/docs/eval-agentx-procedures.md @@ -119,6 +119,8 @@ Set `EVAL_ONLY=true` **before server launch**. It is not merely a switch inside Relevant implementation: [context setup](../benchmarks/benchmark_lib.sh#L2016-L2042), [eval dispatch and failure policy](../benchmarks/benchmark_lib.sh#L2893-L3073), and [workflow inputs](../../.github/workflows/benchmark-tmpl.yml#L40-L57). +Native multi-node post-eval reads the mounted checkpoint at `/model` and enables dataset downloads in the eval process, without changing worker environments. Context lookup reads numeric limits from local `config.json` before falling back to Transformers; an explicit `EVAL_MAX_MODEL_LEN` still takes precedence. + Do not toggle `EVAL_ONLY` after a throughput-sized server is already running and assume the context changed. Restart through the recipe. In eval-only mode an eval failure is returned after available artifacts are staged. In a workflow, upload happens with `always()` before score validation so failed evidence survives ([single-node upload and gate](../../.github/workflows/benchmark-tmpl.yml#L467-L494), [multi-node upload and gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L487-L518)). ## 4. Batched eval concurrency diff --git a/inferencex-e2e/docs/eval-agentx-procedures_zh.md b/inferencex-e2e/docs/eval-agentx-procedures_zh.md index e2c2b52a86..3ffd020867 100644 --- a/inferencex-e2e/docs/eval-agentx-procedures_zh.md +++ b/inferencex-e2e/docs/eval-agentx-procedures_zh.md @@ -114,6 +114,8 @@ python3 -m infx.evals.validate_scores --model-prefix "$MODEL_PREFIX" 4. 吞吐量路径立即返回或被跳过。 5. 运行 `run_eval` 和 artifact staging。 +原生多节点 post-eval 从 `/model` 读取挂载的检查点,并仅在评估进程中启用数据集下载,不改变工作进程环境。上下文查询先读取本地 `config.json` 中的数值上限,再回退到 Transformers;显式设置的 `EVAL_MAX_MODEL_LEN` 仍优先。 + 相关实现:[context 设置](../benchmarks/benchmark_lib.sh#L2016-L2042)、[eval 分派与失败策略](../benchmarks/benchmark_lib.sh#L2893-L3073) 和[工作流输入](../../.github/workflows/benchmark-tmpl.yml#L40-L57)。 不要在吞吐量规格的服务已经运行后才切换 `EVAL_ONLY`,并假定 context 会随之变化。应通过 recipe 重启。Eval-only 模式会在暂存已有 artifact 后返回 eval 失败;在工作流中,上传步骤使用 `always()`,并位于分数校验前,因此失败证据仍会保留([单节点上传与 gate](../../.github/workflows/benchmark-tmpl.yml#L467-L494)、[多节点上传与 gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L487-L518))。 diff --git a/inferencex-e2e/infx/launch/drivers/srt/submit.py b/inferencex-e2e/infx/launch/drivers/srt/submit.py index 9995eebd79..e26b8da18f 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/submit.py +++ b/inferencex-e2e/infx/launch/drivers/srt/submit.py @@ -27,7 +27,8 @@ SINGLE_NODE_SUBMISSION = "srt-single-node-submission.json" MULTINODE_SUBMISSION = "srt-submission.json" MULTINODE_EVAL_COMMAND = ( - '["bash", "{infmax_workspace}/benchmarks/multi_node/srt_eval.sh", "{endpoint}", ' + '["env", "HF_HUB_OFFLINE=0", "HF_DATASETS_OFFLINE=0", "TRANSFORMERS_OFFLINE=0", ' + '"MODEL_PATH=/model", "bash", "{infmax_workspace}/benchmarks/multi_node/srt_eval.sh", "{endpoint}", ' '"{infmax_workspace}"]' ) SINGLE_NODE_EVAL_COMMAND = ( diff --git a/inferencex-e2e/infx/tests/evals/test_run_eval_dispatch.py b/inferencex-e2e/infx/tests/evals/test_run_eval_dispatch.py index 48fbcf0809..4399676528 100644 --- a/inferencex-e2e/infx/tests/evals/test_run_eval_dispatch.py +++ b/inferencex-e2e/infx/tests/evals/test_run_eval_dispatch.py @@ -23,6 +23,42 @@ MULTINODE_AGENTIC_SCRIPT = REPO_ROOT / "benchmarks/srt_agentic.sh" +@pytest.mark.parametrize("use_model_path", [False, True]) +def test_local_context_does_not_require_registered_transformers_model(tmp_path, use_model_path): + model = tmp_path / "new model's weights" + model.mkdir() + (model / "config.json").write_text( + json.dumps( + { + "model_type": "not_registered_yet", + "max_position_embeddings": 1048576, + "seq_length": 4096, + } + ) + ) + (tmp_path / "transformers.py").write_text('raise RuntimeError("model not registered")\n') + env = { + **os.environ, + "BENCHMARK_LIB": str(BENCHMARK_LIB), + "PYTHONPATH": str(tmp_path), + "MODEL_PATH": str(model) if use_model_path else "", + "MODEL_ARG": "served-alias" if use_model_path else str(model), + "KV_OFFLOADING": "none", + } + result = subprocess.run( + [ + "bash", + "-c", + 'source "$BENCHMARK_LIB"; get_native_max_context_length "$MODEL_ARG"', + ], + env=env, + capture_output=True, + text=True, + check=True, + ) + assert result.stdout.strip() == "1048576" + + @pytest.fixture(autouse=True) def explicit_runtime_inputs(monkeypatch: pytest.MonkeyPatch) -> None: """Provide explicit caller inputs before each case applies its overrides.""" diff --git a/inferencex-e2e/infx/tests/launch/test_srt_driver.py b/inferencex-e2e/infx/tests/launch/test_srt_driver.py index eb884ae51b..aa8c5e2ab6 100644 --- a/inferencex-e2e/infx/tests/launch/test_srt_driver.py +++ b/inferencex-e2e/infx/tests/launch/test_srt_driver.py @@ -512,3 +512,50 @@ def test_post_eval_is_handed_the_workload_contract_and_no_other_secret(harness): if arg.startswith("post_eval.passthrough_env=")] # fmt: skip assert handed.keys() <= set(names) assert not {*withheld, "PATH", "HF_HUB_CACHE"} & set(names) + + +def test_multinode_eval_overrides_image_offline_mode_and_host_model_path(harness): + env = lane_env( + harness, + "h200-dgxc", + MODEL_PREFIX="dsr1", + PRECISION="fp8", + FRAMEWORK="dynamo-sglang", + MODEL="deepseek-ai/DeepSeek-R1-0528", + ) + assert_ok(launch(env, harness.config, harness.workspace)) + [call] = srtctl_calls(harness.logs) + [command] = [json.loads(arg.split("=", 1)[1]) for arg in call["argv"] + if arg.startswith("post_eval.command=")] # fmt: skip + # Stand in for the external evaluator and inspect the environment it receives. + script = harness.workspace / "benchmarks/multi_node/srt_eval.sh" + script.write_text( + 'printf "%s\\n" "$HF_HUB_OFFLINE" "$HF_DATASETS_OFFLINE" ' + '"$TRANSFORMERS_OFFLINE" "$MODEL_PATH" "$EVAL_MAX_MODEL_LEN" "$1" "$2"\n' + ) + result = subprocess.run( + [ + arg.format(infmax_workspace=harness.workspace, endpoint="http://worker:8000") + for arg in command + ], + env={ + **os.environ, + "HF_HUB_OFFLINE": "1", + "HF_DATASETS_OFFLINE": "1", + "TRANSFORMERS_OFFLINE": "1", + "MODEL_PATH": "/host-only/checkpoint", + "EVAL_MAX_MODEL_LEN": "9472", + }, + capture_output=True, + text=True, + check=True, + ) + assert result.stdout.splitlines() == [ + "0", + "0", + "0", + "/model", + "9472", + "http://worker:8000", + str(harness.workspace), + ] From bc02bd7e3c198400a17841b4ddc78551207046db Mon Sep 17 00:00:00 2001 From: Cam Quilici Date: Wed, 30 Sep 2026 09:59:59 -0500 Subject: [PATCH 3/4] refactor: port MI355X TileRT AgentX to srt-slurm --- .../agentic/glm5.3_fp8_mi355x_tilert.sh | 172 ------ .../benchmarks/multi_node/amd_utils/env.sh | 27 +- .../benchmarks/multi_node/amd_utils/job.slurm | 67 +- .../multi_node/amd_utils/models_tilert.yaml | 6 - .../benchmarks/multi_node/amd_utils/server.sh | 7 +- .../multi_node/amd_utils/server_sglang.sh | 1 - .../multi_node/amd_utils/server_tilert.sh | 572 ------------------ .../multi_node/amd_utils/setup_deps.sh | 136 ----- .../benchmarks/multi_node/runtime_settings.sh | 16 +- .../configs/glm5.3-tilert-rocm.sh | 28 + .../agentx/disagg-1p1d-tp8-mtp.yaml | 102 ++++ inferencex-e2e/configs/amd-master.yaml | 22 +- inferencex-e2e/configs/runners.yaml | 1 + .../infx/launch/drivers/srt/lanes.py | 5 +- .../infx/srt_slurm/synthetic_acceptance.py | 50 +- .../srt_slurm/test_synthetic_acceptance.py | 113 ++++ 16 files changed, 302 insertions(+), 1023 deletions(-) delete mode 100644 inferencex-e2e/benchmarks/multi_node/agentic/glm5.3_fp8_mi355x_tilert.sh delete mode 100644 inferencex-e2e/benchmarks/multi_node/amd_utils/models_tilert.yaml delete mode 100644 inferencex-e2e/benchmarks/multi_node/amd_utils/server_tilert.sh delete mode 100644 inferencex-e2e/benchmarks/multi_node/amd_utils/setup_deps.sh create mode 100644 inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/configs/glm5.3-tilert-rocm.sh create mode 100644 inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml diff --git a/inferencex-e2e/benchmarks/multi_node/agentic/glm5.3_fp8_mi355x_tilert.sh b/inferencex-e2e/benchmarks/multi_node/agentic/glm5.3_fp8_mi355x_tilert.sh deleted file mode 100644 index 82e81aee32..0000000000 --- a/inferencex-e2e/benchmarks/multi_node/agentic/glm5.3_fp8_mi355x_tilert.sh +++ /dev/null @@ -1,172 +0,0 @@ -#!/usr/bin/env bash - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" -source "$SCRIPT_DIR/../../benchmark_lib.sh" - -check_env_vars \ - CONC_LIST \ - ISL \ - OSL \ - IMAGE \ - SPEC_DECODING \ - MODEL_PATH \ - MODEL_NAME \ - PREFILL_NUM_WORKERS \ - PREFILL_TP \ - PREFILL_EP \ - PREFILL_DP_ATTN \ - DECODE_NUM_WORKERS \ - DECODE_TP \ - DECODE_EP \ - DECODE_DP_ATTN \ - PREFILL_NODES \ - DECODE_NODES \ - RANDOM_RANGE_RATIO \ - DURATION \ - MODEL_PREFIX \ - PRECISION \ - RESULT_FILENAME \ - KV_OFFLOADING \ - IS_AGENTIC \ - FRAMEWORK \ - PREFILL_IMAGE - -if [[ -n "$SLURM_JOB_ID" ]]; then - echo "JOB $SLURM_JOB_ID running on $SLURMD_NODENAME" -fi - -set -x - -cd "$GITHUB_WORKSPACE/benchmarks/multi_node/amd_utils" || exit 1 - -export TIME_LIMIT=08:00:00 -export MODEL_PATH=$MODEL_PATH -export MODEL_NAME=$MODEL_NAME -export CONTAINER_IMAGE=$IMAGE -export PREFILL_IMAGE - -export RESULT_FILENAME - -if [[ "$PREFILL_NODES" -ne 1 || "$DECODE_NODES" -ne 1 || \ - "$PREFILL_NUM_WORKERS" -ne 1 || "$DECODE_NUM_WORKERS" -ne 1 ]]; then - echo "Error: tilert supports exactly 1 prefill node/worker + 1 decode node/worker" \ - "(got PREFILL_NODES=$PREFILL_NODES x$PREFILL_NUM_WORKERS, DECODE_NODES=$DECODE_NODES x$DECODE_NUM_WORKERS)" >&2 - exit 1 -fi - -if [[ "$KV_OFFLOADING" != "none" ]]; then - echo "Error: tilert has no KV offload backend; kv-offloading must be 'none' (got '$KV_OFFLOADING')" >&2 - exit 1 -fi - -# TileRT configuration. Every value is explicit here: server_tilert.sh -# validates each one with check_env_vars and supplies no defaults of its own. -export TILERT_VERSION=0.1.6.post2 -export TILERT_PROFILE=glm5_2 # decode_server --model (TileRT model profile) -export TILERT_MODEL_TYPE=glm-5 # weight_converter --model_type (fallback converter) -export TILERT_MODEL_PKG=glm_5_2_rocm # per-model converter package, preferred when importable -export SERVED_MODEL_NAME=glm5_2 -# GLM-5.3's full context window (config.json max_position_embeddings), as every -# in-tree GLM-5.2 recipe serves. (202752 was GLM-5.1's, inherited from the B200 -# TileRT recipe this mirrors.) -# -# Memory at this context, per rank, bf16 wire layout (verified against the -# tilert 0.1.6 and vLLM 0.24.0 sources and the MI355X logs, 287.98 GiB cards; -# the undivided PD buffer sizes are what 0.1.6 allocated on one card): -# decode : weights 90.72 GiB + engine cache window 93.25 GiB -# + PD receive buffer 99.06 GiB (receive_server.py, dense in max_seq_len) -# prefill: weights 90.45 GiB + profiling/non-torch 40.3 GiB + vLLM KV 91.71 GiB -# + PD staging buffer 99.06 GiB (prefill_connector.py, TP rank 0, -# allocated OUTSIDE vLLM's gpu-memory-utilization budget) -# Undivided, neither side starts: the decode rank is node-marginal (~283 of -# 288 GiB) and the prefill rank cannot fit at any utilization (~321 GiB). -# tilert 0.1.6.post1 keeps both buffers on the GPU but shards them by layer -# across the eight devices (layer lid on device lid % 8, TILERT_PD_SHARDS, -# default on), so each card holds 12.54 GiB instead of 99.06 GiB on one. -# convert() dequantises each layer on the device that received it, which spreads -# its transients too (108.6 KiB/token, 82.9 GiB at 800k tokens) instead of -# leaving them on cuda:0. Measured on 2x8 MI350X at this context with bf16 KV: -# decode peaks at 202.1 GiB per card, prefill at 269.2 GiB of 287.69 GiB, and -# the KV path stays device-to-device at 108 GB/s (81 GB in 751 ms, 54% of the -# 4x400 GbE line rate). No host hop and no patch: the wheel runs as shipped. -export TILERT_MAX_MODEL_LEN=1048576 -export TILERT_TRANSPORT=mooncake -export TILERT_PARSER=none -export TILERT_RDMA_STRICT=0 -export TILERT_CONVERT_LOCK_WAIT=21600 -export TILERT_SIMULATE_ACC_METHOD=match-expected -export TILERT_WEIGHTS_DIR="/models/${MODEL_NAME}-tilert-tp${DECODE_TP}" -# bf16 MLA KV on both roles. This is the only layout TileRT 0.1.6 can consume -# from vLLM on ROCm: MlaNsaProfile.classify_layers infers the layout from the -# cache tensor stride and accepts exactly 1152 B/token (bf16) or 656 B/token -# (fp8_ds_mla). vLLM's ROCM_AITER_MLA_SPARSE backend has no fp8_ds_mla; its -# plain "fp8" writes a flat 576 B/token row, which the connector rejects at -# register_kv_caches. Explicit bfloat16 rather than auto so the stride does not -# depend on the model dtype. Never float16: it passes the 1152 B check and is -# then read as bf16. -export PREFILL_KV_DTYPE=bfloat16 -# The ROCm backend supports block sizes [1, 64] and vLLM picks 1, which makes -# the connector's KI plane copy fail and MLA address the wrong rows. -export PREFILL_BLOCK_SIZE=64 -export DECODE_KV_DTYPE=bf16 -# The PD staging shard sits outside vLLM's budget, so vLLM needs 90.45 (weights) -# + 40.3 (profiling) + 91.71 GiB (KV for one 1048576-token request) = 222.5 GiB -# inside it: 0.85 x 287.98 = 244.8 GiB leaves 22 GiB of KV margin and 43 GiB -# outside the budget for the 12.54 GiB staging shard plus the ~6.3 GiB non-torch -# baseline measured on the decode OOM node (287.98 - 95.94 free - 184.17 - 1.58 -# reserved). 0.75 (216 GiB) refuses with "91.71 GiB KV cache is needed ... -# available 85.25 GiB". -export GPU_MEM_UTIL=0.85 -export SKIP_CONTAINER_BARRIER=0 -# Two images, one per rank, ~32 GB each. On a node that has neither cached the -# pull alone outlasts the SGLang path's 300s default and the 1800s this script -# used to hardcode, and the rank that comes up first waits out the whole -# timeout while its peer is still pulling. -export CONTAINER_BARRIER_TIMEOUT=5400 -export ROUTER_PORT=30000 -export PREFILL_PORT=8000 -export DECODE_CTRL_PORT=5556 -export DECODE_HTTP_PORT=5557 -export DECODE_WAIT=7200 # prefill waits for the decode ctrl port -export PREFILL_WAIT=3600 # prefill waits for its own vLLM port -export ROUTER_WAIT=10800 # decode waits for the router port to open - -if [[ "$SPEC_DECODING" == "mtp" ]]; then - # TileRT decode drafts at depth 3 (the only depth the ROCm GLM profile - # builds) and the golden acceptance curve is keyed on it. The vLLM prefill - # rank only has to materialise the MTP layer's KV, so it runs at 1. - export DECODE_MTP_SIZE=3 - export PREFILL_SPEC_TOKENS=1 -else - export DECODE_MTP_SIZE=0 - export PREFILL_SPEC_TOKENS=0 -fi -export TILERT_QUEUE_TIMEOUT=1800 # requests wait on the bs=1 decode engine -export THINKING_MODE=thinking_on - -if [[ "$PREFILL_EP" -ne 1 || "$DECODE_EP" -ne 1 || \ - "$PREFILL_DP_ATTN" == "true" || "$DECODE_DP_ATTN" == "true" ]]; then - echo "Error: tilert runs pure TP8 on both roles; ep must be 1 and dp-attn false" >&2 - exit 1 -fi -export PREFILL_ENABLE_EP=false -export PREFILL_ENABLE_DP=false -export DECODE_ENABLE_EP=false -export DECODE_ENABLE_DP=false - -JOB_ID=$(bash ./submit.sh $PREFILL_NODES \ - $PREFILL_NUM_WORKERS \ - $DECODE_NODES \ - $DECODE_NUM_WORKERS \ - $ISL $OSL "${CONC_LIST// /x}" inf \ - ${PREFILL_ENABLE_EP} ${PREFILL_ENABLE_DP} \ - ${DECODE_ENABLE_EP} ${DECODE_ENABLE_DP} \ - ${PREFILL_TP} ${DECODE_TP} \ - ${RANDOM_RANGE_RATIO}) - -if [[ $? -ne 0 ]]; then - echo "Failed to submit job" >&2 - exit 1 -fi - -echo "$JOB_ID" diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/env.sh b/inferencex-e2e/benchmarks/multi_node/amd_utils/env.sh index ba48879cb4..04adb875ac 100755 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/env.sh +++ b/inferencex-e2e/benchmarks/multi_node/amd_utils/env.sh @@ -2,21 +2,12 @@ source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only check_env_vars ENGINE -# MoRI-IO queue-pair tuning, the UCX RoCE GID index, SGLang router logging and the -# SGLang decode cuda-graph NCCL workaround. Only the SGLang MoRI KV path -# below reads these. ENGINE=tilert moves KV over mooncake and starts no SGLang -# router, so it is neither given nor reads them: validating them there would force -# the recipe to invent MoRI tuning for a transport it never uses. -if [[ "$ENGINE" != "tilert" ]]; then - check_env_vars \ - MORI_IO_SQ_BACKOFF_TIMEOUT_US MORI_IO_QP_MAX_SEND_WR MORI_IO_QP_MAX_CQE MORI_IO_QP_MAX_SGE MORI_IO_TC_DISABLE \ - UCX_IB_GID_INDEX MORI_APP_LOG_LEVEL SGLANG_ROUTER_STDOUT_LOGS TORCH_NCCL_BLOCKING_WAIT NCCL_BLOCKING_WAIT \ - SGLANG_OPT_USE_AITER_INDEXER -fi -# Dual-engine environment setup for multi-node disaggregated serving. -# -# ENGINE=sglang-disagg or tilert selects the engine-specific block. -# +# SGLang MoRI environment. +check_env_vars \ + MORI_IO_SQ_BACKOFF_TIMEOUT_US MORI_IO_QP_MAX_SEND_WR MORI_IO_QP_MAX_CQE MORI_IO_QP_MAX_SGE MORI_IO_TC_DISABLE \ + UCX_IB_GID_INDEX MORI_APP_LOG_LEVEL SGLANG_ROUTER_STDOUT_LOGS TORCH_NCCL_BLOCKING_WAIT NCCL_BLOCKING_WAIT \ + SGLANG_OPT_USE_AITER_INDEXER + # REQUIRED ENVIRONMENT VARIABLES: # IBDEVICES - RDMA/InfiniBand device names (e.g., ionic_0,ionic_1,... or mlx5_0,mlx5_1,...) # Set by runner or auto-detected from hostname. @@ -119,10 +110,6 @@ else fi fi -if [[ "$ENGINE" == "tilert" ]]; then - echo "[INFO] tilert: IBDEVICES=$IBDEVICES NCCL_SOCKET_IFNAME=$NCCL_SOCKET_IFNAME NCCL_IB_HCA=$NCCL_IB_HCA" - -else export SGLANG_USE_AITER=1 export AITER_LOG_LEVEL=ERROR @@ -237,5 +224,3 @@ else export GPU_MAX_HW_QUEUES=2 fi fi - -fi diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/job.slurm b/inferencex-e2e/benchmarks/multi_node/amd_utils/job.slurm index 881ca6a017..40548f07e5 100755 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/job.slurm +++ b/inferencex-e2e/benchmarks/multi_node/amd_utils/job.slurm @@ -34,11 +34,7 @@ echo "" # Use $(pwd) not BASH_SOURCE — sbatch copies the script to /var/spool/slurmd/ # at runtime, but the CWD remains the submit-time directory (amd_utils/). -if [[ "$ENGINE" == "tilert" ]]; then - MODELS_YAML="$(pwd)/models_tilert.yaml" -else - MODELS_YAML="$(pwd)/models.yaml" -fi +MODELS_YAML="$(pwd)/models.yaml" if [[ ! -f "$MODELS_YAML" ]]; then echo "Error: models YAML not found at $MODELS_YAML" @@ -50,16 +46,6 @@ if [[ -z "${DOCKER_IMAGE_NAME:-}" ]]; then exit 1 fi -if [[ "$ENGINE" == "tilert" && -z "${PREFILL_IMAGE:-}" ]]; then - echo "Error: ENGINE=tilert requires PREFILL_IMAGE (e.g. PREFILL_IMAGE=vllm/vllm-openai-rocm:nightly- in prefill.additional-settings)." - exit 1 -fi -if [[ "$ENGINE" == "tilert" ]]; then - # server_tilert.sh takes the container-creation barrier timeout from the - # recipe (no 300s default as on the SGLang path); fail here, before sbatch - # work is done, rather than inside the container. - check_env_vars CONTAINER_BARRIER_TIMEOUT -fi # Resolve the models.yaml entry the same way server_sglang.sh does: agentic runs # (IS_AGENTIC) use the '-AgentX' recipe, non-agentic disaggregated runs use @@ -541,46 +527,9 @@ DOCKER_ENV_COMMON=( ) # Engine-specific env vars -if [[ "$ENGINE" == "tilert" ]]; then - DOCKER_ENV_ENGINE=( - -e MODEL_PATH=$DOCKER_MODEL_PATH - -e PREFILL_IMAGE=${PREFILL_IMAGE} - -e TILERT_VERSION=${TILERT_VERSION} - -e TILERT_PROFILE=${TILERT_PROFILE} - -e TILERT_MODEL_TYPE=${TILERT_MODEL_TYPE} - -e TILERT_MODEL_PKG=${TILERT_MODEL_PKG} - -e TILERT_MAX_MODEL_LEN=${TILERT_MAX_MODEL_LEN} - -e TILERT_TRANSPORT=${TILERT_TRANSPORT} - -e TILERT_PARSER=${TILERT_PARSER} - -e TILERT_QUEUE_TIMEOUT=${TILERT_QUEUE_TIMEOUT} - -e TILERT_WEIGHTS_DIR=${TILERT_WEIGHTS_DIR} - -e TILERT_RDMA_STRICT=${TILERT_RDMA_STRICT} - -e TILERT_CONVERT_LOCK_WAIT=${TILERT_CONVERT_LOCK_WAIT} - -e TILERT_SIMULATE_ACC_METHOD=${TILERT_SIMULATE_ACC_METHOD} - -e \"TILERT_EXTRA_ENV=${TILERT_EXTRA_ENV:-}\" - -e SERVED_MODEL_NAME=${SERVED_MODEL_NAME} - -e PREFILL_KV_DTYPE=${PREFILL_KV_DTYPE} - -e PREFILL_BLOCK_SIZE=${PREFILL_BLOCK_SIZE} - -e PREFILL_SPEC_TOKENS=${PREFILL_SPEC_TOKENS} - -e DECODE_KV_DTYPE=${DECODE_KV_DTYPE} - -e GPU_MEM_UTIL=${GPU_MEM_UTIL} - -e PREFILL_PORT=${PREFILL_PORT} - -e DECODE_CTRL_PORT=${DECODE_CTRL_PORT} - -e DECODE_HTTP_PORT=${DECODE_HTTP_PORT} - -e DECODE_WAIT=${DECODE_WAIT} - -e PREFILL_WAIT=${PREFILL_WAIT} - -e ROUTER_WAIT=${ROUTER_WAIT} - -e SKIP_CONTAINER_BARRIER=${SKIP_CONTAINER_BARRIER} - # Golden-acceptance selection on agentic MTP runs; unset elsewhere. - -e THINKING_MODE=${THINKING_MODE:-} - -e IBDEVICES=${IBDEVICES:-} - -e PYTHONPYCACHEPREFIX=/tmp/pycache - ) -else - DOCKER_ENV_ENGINE=( - -e SGLANG_WS_PATH=${WS_PATH} - ) -fi +DOCKER_ENV_ENGINE=( + -e SGLANG_WS_PATH=${WS_PATH} +) # HiCache / Mooncake settings are delivered via a bind-mounted config file rather # than a long list of docker -e flags. Write it once to the shared benchmark-logs @@ -650,12 +599,6 @@ echo \"Rank \$SLURM_PROCID on \$(hostname)\" eval \"\$DOCKER_CMD_DETECT\" echo \"[docker-detect] rank \$SLURM_PROCID: DOCKER_CMD=\$DOCKER_CMD\" -RANK_IMAGE= -if [[ \"$ENGINE\" == \"tilert\" && \"\$SLURM_PROCID\" -lt \"$xP\" ]]; then - RANK_IMAGE=\"$PREFILL_IMAGE\" - echo \"[tilert] rank \$SLURM_PROCID is a prefill rank; using PREFILL_IMAGE=\$RANK_IMAGE\" -fi - exec \$DOCKER_CMD run \ --init \ --stop-timeout 10 \ @@ -696,7 +639,7 @@ exec \$DOCKER_CMD run \ ${CLIENT_DOCKER_ENV} \ --name \"$DOCKER_CONT_NAME\" \ --entrypoint \"\" \ - \"\${RANK_IMAGE:-$DOCKER_IMAGE_NAME}\" bash -lc ' + \"$DOCKER_IMAGE_NAME\" bash -lc ' set -o pipefail mkdir -p /run_logs/slurm_job-'\"\$SLURM_JOB_ID\"' '"$RUN_FILE_FULL"' 2>&1 | tee /run_logs/slurm_job-'\"\$SLURM_JOB_ID\"'/server_\$(hostname).log diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/models_tilert.yaml b/inferencex-e2e/benchmarks/multi_node/amd_utils/models_tilert.yaml deleted file mode 100644 index 1697059d54..0000000000 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/models_tilert.yaml +++ /dev/null @@ -1,6 +0,0 @@ -# Model-owned engine settings for ENGINE=tilert. Everything the caller owns -# (profile, topology, ports, dtypes, draft depth) is set by the recipe, not here. -GLM-5.3: - prefill_env: "VLLM_ROCM_USE_AITER=1 VLLM_ROCM_USE_AITER_MOE=1 VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1 VLLM_ENGINE_READY_TIMEOUT_S=10800" - prefill_extra_flags: "" - decode_extra_flags: "" diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/server.sh b/inferencex-e2e/benchmarks/multi_node/amd_utils/server.sh index a3830b57aa..345fb13a79 100755 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/server.sh +++ b/inferencex-e2e/benchmarks/multi_node/amd_utils/server.sh @@ -4,7 +4,6 @@ source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only # Multi-Engine Disaggregated Server Dispatcher # Dispatches to the engine-specific server launcher based on ENGINE env var. # ENGINE=sglang-disagg (default) -> server_sglang.sh (SGLang + MoRI) -# ENGINE=tilert -> server_tilert.sh (vLLM prefill + TileRT decode) check_env_vars ENGINE WS_PATH if [[ -f /config/hicache_mc.env ]]; then @@ -16,8 +15,4 @@ export WS_PATH ENGINE echo "[DISPATCHER] ENGINE=$ENGINE WS_PATH=$WS_PATH" -if [[ "$ENGINE" == "tilert" ]]; then - source "$WS_PATH/server_tilert.sh" -else - source "$WS_PATH/server_sglang.sh" -fi +source "$WS_PATH/server_sglang.sh" diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/server_sglang.sh b/inferencex-e2e/benchmarks/multi_node/amd_utils/server_sglang.sh index 3447d0030e..73af40c0d3 100755 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/server_sglang.sh +++ b/inferencex-e2e/benchmarks/multi_node/amd_utils/server_sglang.sh @@ -20,7 +20,6 @@ BENCH_MAX_CONC_VALUE=$(echo "$BENCH_MAX_CONCURRENCY" | tr 'x' '\n' | sort -n | t # can resolve formulas like "BENCH_MAX_CONC_VALUE*2" for max_running_requests. export BENCH_MAX_CONC_VALUE -source $SGLANG_WS_PATH/setup_deps.sh source $SGLANG_WS_PATH/env.sh # Install before starting UMBP or serving processes. Early readiness failures must diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/server_tilert.sh b/inferencex-e2e/benchmarks/multi_node/amd_utils/server_tilert.sh deleted file mode 100644 index 37bb8fa6ee..0000000000 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/server_tilert.sh +++ /dev/null @@ -1,572 +0,0 @@ -#!/bin/bash -# TileRT disaggregated launcher: upstream vLLM ROCm prefill (TileRTConnector, -# kv_producer) + TileRT decode_server + the OpenAI-compatible pd_router. -# Every value below is supplied by the recipe through job.slurm; this script -# validates them and never invents a default for caller-owned configuration. - -source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only -if [[ "${EVAL_ONLY:-}" != true ]]; then - validate_agentic_concurrency "${BENCH_MAX_CONCURRENCY:-}" || exit 1 -fi - -check_env_vars \ - NODE0_ADDR NODE_RANK MODEL_DIR MODEL_NAME MODEL_PATH xP yD IPADDRS \ - DRY_RUN GPUS_PER_NODE PREFILL_TP_SIZE DECODE_TP_SIZE \ - BENCH_INPUT_LEN BENCH_OUTPUT_LEN BENCH_MAX_CONCURRENCY \ - RUN_EVAL EVAL_ONLY EVAL_FRAMEWORK BENCHMARK_LOGS_DIR WS_PATH \ - SLURM_JOB_ID SPEC_DECODING \ - TILERT_PROFILE TILERT_MODEL_TYPE TILERT_MODEL_PKG TILERT_MAX_MODEL_LEN \ - TILERT_TRANSPORT TILERT_PARSER TILERT_QUEUE_TIMEOUT TILERT_WEIGHTS_DIR \ - TILERT_RDMA_STRICT TILERT_CONVERT_LOCK_WAIT TILERT_SIMULATE_ACC_METHOD \ - PREFILL_KV_DTYPE PREFILL_BLOCK_SIZE PREFILL_SPEC_TOKENS DECODE_KV_DTYPE \ - DECODE_MTP_SIZE GPU_MEM_UTIL SERVED_MODEL_NAME \ - DECODE_CTRL_PORT DECODE_HTTP_PORT PREFILL_PORT ROUTER_PORT \ - DECODE_WAIT PREFILL_WAIT ROUTER_WAIT SKIP_CONTAINER_BARRIER \ - CONTAINER_BARRIER_TIMEOUT - -LOG_DIR="/run_logs/slurm_job-${SLURM_JOB_ID}" -SHARED_LOG_DIR="${BENCHMARK_LOGS_DIR}/logs/slurm_job-${SLURM_JOB_ID}" -mkdir -p "$LOG_DIR" - -if [[ "$xP" -ne 1 || "$yD" -ne 1 ]]; then - echo "ERROR: tilert supports exactly 1 prefill + 1 decode worker (got xP=$xP yD=$yD)" >&2 - exit 1 -fi -if [[ "$NODE_RANK" -lt "$xP" ]]; then - TILERT_ROLE=prefill -else - TILERT_ROLE=decode -fi -export TILERT_ROLE - -source "$WS_PATH/setup_deps.sh" -source "$WS_PATH/env.sh" -# benchmark_lib.sh derives AIPERF_DIR from this at source time, so -# it must be set before the library is loaded, not in run_agentic_replay. The -# AgentX replay runs in this container, where the repo is mounted at /workspace. -export INFMAX_CONTAINER_WORKSPACE=/workspace -source /workspace/benchmarks/benchmark_lib.sh - -# Model-specific engine environment (not caller configuration): the prefill -# vLLM env block lives with the model. Everything else is passed in by the recipe. -MODELS_YAML="${WS_PATH}/models_tilert.yaml" -eval "$("$PY" - "$MODELS_YAML" "$MODEL_NAME" <<'PYEOF' -import shlex, sys, yaml -path, name = sys.argv[1], sys.argv[2] -with open(path) as f: - models = yaml.safe_load(f) or {} -if name not in models: - sys.exit(f"model '{name}' is not present in {path}") -m = models[name] or {} -for key, var in (("prefill_env", "TILERT_PREFILL_ENV"), - ("prefill_extra_flags", "TILERT_PREFILL_EXTRA_FLAGS"), - ("decode_extra_flags", "TILERT_DECODE_EXTRA_FLAGS")): - print(f"{var}={shlex.quote(str(m.get(key) or ''))}") -PYEOF -)" || { echo "ERROR: cannot read the tilert model entry for '$MODEL_NAME' from $MODELS_YAML" >&2; exit 1; } -echo "[tilert] model entry '$MODEL_NAME' loaded from $MODELS_YAML" - -export ROUTER_PORT -export SERVED_MODEL_NAME - -PREFILL_SPEC=() -DECODE_MTP=() -if [[ "$SPEC_DECODING" == "mtp" ]]; then - # The prefill rank only has to build the MTP layer's KV; TileRT decode owns - # the draft depth (DECODE_MTP_SIZE), so the two counts differ by design. - PREFILL_SPEC=(--speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":${PREFILL_SPEC_TOKENS}}") - # decode_server only accepts depth 3 today, but pass it explicitly so the - # converter's --num_mtp, the golden-curve key and the engine depth agree by - # data flow rather than by coincidence of defaults. - DECODE_MTP=(--with-mtp --num-mtp "$DECODE_MTP_SIZE") -fi - -# The only TileRT recipe on this cluster is AgentX (agentic-coding). -if [[ "${IS_AGENTIC:-0}" != "1" && "${IS_AGENTIC:-}" != "true" && "${SCENARIO_TYPE:-}" != "agentic-coding" ]]; then - echo "ERROR: server_tilert.sh only runs agentic-coding (IS_AGENTIC=${IS_AGENTIC:-} SCENARIO_TYPE=${SCENARIO_TYPE:-})" >&2 - exit 1 -fi - -IFS=',' read -ra IP_ARRAY <<< "$IPADDRS" -PREFILL_HOST="${IP_ARRAY[0]:-$NODE0_ADDR}" -DECODE_HOST="${IP_ARRAY[$xP]:-}" -if [[ -z "$DECODE_HOST" ]]; then - echo "ERROR: cannot resolve the decode node IP from IPADDRS='$IPADDRS' (xP=$xP)" >&2 - exit 1 -fi -host_ip=$(ip route get 1.1.1.1 2>/dev/null | awk '/src/ {print $7}') -host_name=$(hostname) - -echo "[tilert] ROLE=$TILERT_ROLE rank=$NODE_RANK host=$host_name ($host_ip)" -echo "[tilert] PREFILL_HOST=$PREFILL_HOST:$PREFILL_PORT DECODE_HOST=$DECODE_HOST:$DECODE_CTRL_PORT/$DECODE_HTTP_PORT ROUTER=:$ROUTER_PORT" -echo "[tilert] MODEL_PATH=$MODEL_PATH profile=$TILERT_PROFILE served=$SERVED_MODEL_NAME max_len=$TILERT_MAX_MODEL_LEN transport=$TILERT_TRANSPORT kv=${PREFILL_KV_DTYPE}->${DECODE_KV_DTYPE} mtp=${SPEC_DECODING}" - -# Enable libibverbs fork safety on both ranks before any verbs context exists. -# Without it, ibv_fork_init() can fail in these containers while Mooncake -# initialization still reports success, silently falling back from RDMA to TCP. -# Slower prefill-to-decode KV transfer degrades TTFT; TPOT is unaffected. -export RDMAV_FORK_SAFE=1 - -for env_pair in ${TILERT_EXTRA_ENV}; do - export "${env_pair?}" - echo "[tilert][EXTRA_ENV] $env_pair" -done - -log_and_run_bg() { - local label="$1" logfile="$2"; shift 2 - { printf '===== [%s] %s =====\n' "$label" "$(date '+%F %T')" - printf '[cmd]'; printf ' %q' "$@"; printf '\n' - printf '[cwd] %s\n[host] %s\n\n' "$PWD" "$host_name" - } | tee -a "$logfile" - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: [$label] not started" - LAST_BG_PID="" - return 0 - fi - "$@" >>"$logfile" 2>&1 & - LAST_BG_PID=$! - echo "[$label] pid=$LAST_BG_PID log=$logfile" -} - -rdma_preflight() { - local warn=0 - echo "[rdma] role=$TILERT_ROLE IBDEVICES=${IBDEVICES:-} NCCL_SOCKET_IFNAME=${NCCL_SOCKET_IFNAME:-}" - local uverbs=(/dev/infiniband/uverbs*) - if [[ -e "${uverbs[0]}" ]]; then - echo "[rdma] verbs devices: ${uverbs[*]}" - else - echo "[rdma] WARNING: /dev/infiniband/uverbs* missing -- the container has no RDMA device nodes (job.slurm passes --device /dev/infiniband)" >&2 - warn=1 - fi - local ml; ml="$(ulimit -l 2>/dev/null)" - if [[ "$ml" == "unlimited" ]]; then - echo "[rdma] memlock: unlimited" - else - echo "[rdma] WARNING: memlock=$ml (not unlimited) -- pinning memory for RDMA may fail (job.slurm passes --ulimit memlock=-1)" >&2 - warn=1 - fi - if command -v ibv_devices >/dev/null 2>&1; then - echo "[rdma] ibv_devices:"; ibv_devices 2>&1 | sed 's/^/[rdma] /' - fi - if (( warn )) && [[ "$TILERT_RDMA_STRICT" == "1" ]]; then - echo "[rdma] TILERT_RDMA_STRICT=1 and preflight did not fully pass -- aborting" >&2 - return 1 - fi - return 0 -} - -stage_tokenizer_files() { - local staged=0 f b - for f in "$MODEL_PATH"/*; do - [[ -f "$f" ]] || continue - b="$(basename "$f")" - [[ "$b" == *.safetensors ]] && continue - [[ "$b" == "model.safetensors.index.json" ]] && continue - [[ -e "$TILERT_WEIGHTS_DIR/$b" ]] && continue - cp -p "$f" "$TILERT_WEIGHTS_DIR/$b" && staged=$((staged+1)) - done - echo "[stage_tokenizer] staged $staged auxiliary file(s) from $MODEL_PATH -> $TILERT_WEIGHTS_DIR" - local missing=() - [[ -f "$TILERT_WEIGHTS_DIR/chat_template.jinja" ]] || missing+=(chat_template.jinja) - [[ -f "$TILERT_WEIGHTS_DIR/tokenizer_config.json" || -f "$TILERT_WEIGHTS_DIR/tokenizer.json" ]] \ - || missing+=("tokenizer.json/tokenizer_config.json") - if (( ${#missing[@]} )); then - echo "[stage_tokenizer] ERROR: $TILERT_WEIGHTS_DIR is missing ${missing[*]}; check that MODEL_PATH=$MODEL_PATH is an HF directory with the tokenizer" >&2 - return 1 - fi - return 0 -} - -_tilert_weights_cached() { - local r - for r in $(seq 0 $((DECODE_TP_SIZE - 1))); do - [[ -f "$TILERT_WEIGHTS_DIR/rank${r}/model.safetensors.index.json" ]] || return 1 - done - # The engine refuses a cache converted without the MTP module (end2end.py - # checks tilert_meta.json num_mtp), but only after loading ~90 GiB of - # weights. Check the same field here so a stale non-MTP cache is - # re-converted instead of failing late. - [[ -f "$TILERT_WEIGHTS_DIR/tilert_meta.json" ]] || return 1 - if [[ "$SPEC_DECODING" == "mtp" ]]; then - "$PY" - "$TILERT_WEIGHTS_DIR/tilert_meta.json" <<'PYEOF' || return 1 -import json, sys -sys.exit(0 if int(json.load(open(sys.argv[1])).get("num_mtp", 0)) >= 1 else 1) -PYEOF - fi - return 0 -} - -convert_weights() { - if _tilert_weights_cached; then - echo "[weight_converter] cache hit (${DECODE_TP_SIZE}/${DECODE_TP_SIZE} rank index.json), skipping conversion: $TILERT_WEIGHTS_DIR" - return 0 - fi - mkdir -p "$TILERT_WEIGHTS_DIR" || { echo "[weight_converter] ERROR: cannot create $TILERT_WEIGHTS_DIR (set TILERT_WEIGHTS_DIR to a writable shared path)" >&2; return 1; } - exec 9>"$TILERT_WEIGHTS_DIR/.convert.lock" - flock -w "$TILERT_CONVERT_LOCK_WAIT" 9 || { - echo "[weight_converter] timed out waiting for the conversion lock (another job still converting?)" >&2; return 1; } - if _tilert_weights_cached; then - echo "[weight_converter] cache produced by a concurrent job, skipping conversion"; exec 9>&-; return 0 - fi - if [[ -n "$(ls -A "$TILERT_WEIGHTS_DIR" 2>/dev/null | grep -v '^\.convert\.lock$')" ]]; then - echo "[weight_converter] leftovers without index.json (previous conversion incomplete); cleaning and re-converting" - find "$TILERT_WEIGHTS_DIR" -mindepth 1 ! -name '.convert.lock' -delete - fi - echo "[weight_converter] $MODEL_PATH -> $TILERT_WEIGHTS_DIR (model_type=$TILERT_MODEL_TYPE)" - local conv_mod conv_args - if "$PY" -c "import tilert.models.${TILERT_MODEL_PKG}.weight_converter" 2>/dev/null; then - conv_mod="tilert.models.${TILERT_MODEL_PKG}.weight_converter" - conv_args=(--model_dir "$MODEL_PATH" --save_dir "$TILERT_WEIGHTS_DIR" - --device "cuda:$((GPUS_PER_NODE - 1))") - [[ "$SPEC_DECODING" == "mtp" ]] && conv_args+=(--num_mtp "$DECODE_MTP_SIZE") - else - conv_mod="tilert.models.preprocess.weight_converter" - conv_args=(--model_type "$TILERT_MODEL_TYPE" --model_dir "$MODEL_PATH" --save_dir "$TILERT_WEIGHTS_DIR") - fi - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: $PY -m $conv_mod ${conv_args[*]}" - exec 9>&-; return 0 - fi - echo "[weight_converter] using $conv_mod" - "$PY" -m "$conv_mod" "${conv_args[@]}" \ - 2>&1 | tee "$LOG_DIR/tilert_weight_converter_${host_name}.log" - local rc=${PIPESTATUS[0]} - exec 9>&- - if [[ $rc -ne 0 ]] || ! _tilert_weights_cached; then - echo "[weight_converter] ERROR: conversion failed (rc=$rc, per-rank index.json complete: $(_tilert_weights_cached && echo yes || echo no))" >&2 - return 1 - fi - echo "[weight_converter] conversion done and cached: $TILERT_WEIGHTS_DIR" -} - -start_decode() { - # shellcheck disable=SC2206 - local extra=( ${TILERT_DECODE_EXTRA_FLAGS} ) - if [[ "$SPEC_DECODING" == "mtp" && "$EVAL_ONLY" != "true" && "$RUN_EVAL" != "true" ]]; then - check_env_vars MODEL_PREFIX THINKING_MODE - local curve="${WS_PATH%/benchmarks/*}/infx/golden_al_distribution/${MODEL_PREFIX}_mtp.yaml" - TILERT_SIMULATE_ACC_LEN="$("$PY" - "$curve" "$THINKING_MODE" "$DECODE_MTP_SIZE" <<'PYEOF' -import sys, yaml -path, thinking, tokens = sys.argv[1], sys.argv[2], int(sys.argv[3]) -data = yaml.safe_load(open(path)) -if not isinstance(data, dict) or len(data) != 1: - sys.exit(f"golden curve {path} must hold exactly one model key") -model, modes = next(iter(data.items())) -try: - value = float(modes[thinking][tokens]) -except (KeyError, TypeError, ValueError): - sys.exit(f"no golden acceptance for {model}/{thinking}/{tokens} draft tokens in {path}") -if not 1 <= value <= tokens + 1: - sys.exit(f"golden acceptance {value} out of range for {tokens} draft tokens") -print(f"{value:g}") -PYEOF -)" || { echo "[tilert] ERROR: golden AL lookup failed (curve=$curve)" >&2; exit 1; } - echo "[tilert] golden AL ${TILERT_SIMULATE_ACC_LEN} from $(basename "$curve") ($THINKING_MODE, K=${DECODE_MTP_SIZE})" - fi - - if [[ -n "${TILERT_SIMULATE_ACC_LEN:-}" && "$EVAL_ONLY" != "true" ]]; then - export TILERT_SIMULATE_ACC_LEN - export TILERT_SIMULATE_ACC_METHOD - echo "[decode] simulated acceptance: TILERT_SIMULATE_ACC_LEN=${TILERT_SIMULATE_ACC_LEN}" \ - "method=${TILERT_SIMULATE_ACC_METHOD} (output text is meaningless by design)" - else - unset TILERT_SIMULATE_ACC_LEN TILERT_SIMULATE_ACC_METHOD - echo "[decode] real MTP verification (no simulated acceptance)" - fi - local cmd=("$PY" -m tilert.pd_vllm.decode_server - --engine tilert --model "$TILERT_PROFILE" - --model-weights-dir "$TILERT_WEIGHTS_DIR" - --max-seq-len "$TILERT_MAX_MODEL_LEN" - --kv-cache-dtype "$DECODE_KV_DTYPE" --transport "$TILERT_TRANSPORT" - --ctrl-port "$DECODE_CTRL_PORT" --http-port "$DECODE_HTTP_PORT" - "${DECODE_MTP[@]}" "${extra[@]}") - log_and_run_bg decode "$LOG_DIR/decode_${host_name}.log" "${cmd[@]}" - DECODE_PID=$LAST_BG_PID -} - -start_prefill() { - for env_pair in ${TILERT_PREFILL_ENV}; do - export "${env_pair?}" - echo "[PREFILL_ENV] $env_pair" - done - local served=("$SERVED_MODEL_NAME") - [[ -n "$MODEL_NAME" && "$MODEL_NAME" != "$SERVED_MODEL_NAME" ]] && served+=("$MODEL_NAME") - # shellcheck disable=SC2206 - local extra=( ${TILERT_PREFILL_EXTRA_FLAGS} ) - local kv_cfg - kv_cfg=$(printf '{"kv_connector":"TileRTConnector","kv_connector_module_path":"tilert.pd_vllm.prefill_connector","kv_role":"kv_producer","kv_connector_extra_config":{"tilert_host":"%s","tilert_ctrl_port":%s,"tilert_model":"%s","tilert_max_seq_len":%s,"tilert_transport":"%s"}}' \ - "$DECODE_HOST" "$DECODE_CTRL_PORT" "$TILERT_PROFILE" "$TILERT_MAX_MODEL_LEN" "$TILERT_TRANSPORT") - local cmd=(vllm serve "$MODEL_PATH" - --served-model-name "${served[@]}" --port "$PREFILL_PORT" - --tensor-parallel-size "$PREFILL_TP_SIZE" --max-model-len "$TILERT_MAX_MODEL_LEN" - --enforce-eager --trust-remote-code --return-tokens-as-token-ids - --gpu-memory-utilization "$GPU_MEM_UTIL" --kv-cache-dtype "$PREFILL_KV_DTYPE" - --block-size "$PREFILL_BLOCK_SIZE" - "${PREFILL_SPEC[@]}" - --kv-transfer-config "$kv_cfg" - "${extra[@]}") - log_and_run_bg prefill "$LOG_DIR/prefill_${host_name}.log" "${cmd[@]}" - PREFILL_PID=$LAST_BG_PID -} - -start_router() { - local cmd=(env HIP_VISIBLE_DEVICES= ROCR_VISIBLE_DEVICES= CUDA_VISIBLE_DEVICES= - "$PY" -m tilert.pd_vllm.pd_router - --vllm-url "http://$PREFILL_HOST:$PREFILL_PORT" - --decode "$DECODE_HOST:$DECODE_CTRL_PORT:$DECODE_HTTP_PORT" - --host 0.0.0.0 --port "$ROUTER_PORT" --model-path "$MODEL_PATH" --parser "$TILERT_PARSER" - --queue-timeout "$TILERT_QUEUE_TIMEOUT") - log_and_run_bg router "$LOG_DIR/router_${host_name}.log" "${cmd[@]}" - ROUTER_PID=$LAST_BG_PID -} - -tcp_open() { (exec 3<>"/dev/tcp/$1/$2") 2>/dev/null; } - -wait_for_tcp() { - local host="$1" port="$2" timeout="${3:-600}" pid="${4:-}" - local deadline=$(( SECONDS + timeout )) - until tcp_open "$host" "$port"; do - if [[ -n "$pid" ]] && ! kill -0 "$pid" 2>/dev/null; then - echo "[wait_for_tcp] process $pid exited before $host:$port opened" >&2; return 2 - fi - if [[ $SECONDS -ge $deadline ]]; then - echo "[wait_for_tcp] timeout: $host:$port not open after ${timeout}s" >&2; return 1 - fi - sleep 5 - done - echo "[wait_for_tcp] $host:$port ready" -} - -wait_for_tcp_close() { - local host="$1" port="$2" pid="${3:-}" - while tcp_open "$host" "$port"; do - if [[ -n "$pid" ]] && ! kill -0 "$pid" 2>/dev/null; then - echo "[wait_for_tcp_close] process $pid exited while $host:$port is still open" >&2; return 2 - fi - sleep 10 - done - echo "[wait_for_tcp_close] $host:$port closed" -} - -copy_logs_to_shared() { - [[ "$DRY_RUN" -eq 0 ]] || return 0 - mkdir -p "$SHARED_LOG_DIR" && cp -r "$LOG_DIR"/. "$SHARED_LOG_DIR"/ \ - && echo "Copied $LOG_DIR -> $SHARED_LOG_DIR" \ - || echo "WARNING: failed to copy $LOG_DIR to $SHARED_LOG_DIR" >&2 -} -trap copy_logs_to_shared EXIT - -run_lm_eval_on_router() { - echo "Running lm-eval evaluation on the router..." - local ok=false _attempt - for _attempt in 1 2 3; do - if curl -sf --max-time 10 "http://0.0.0.0:${ROUTER_PORT}/health" >/dev/null 2>&1; then ok=true; break; fi - echo "Eval health check attempt $_attempt failed, retrying in 10s..."; sleep 10 - done - if [[ "$ok" != "true" ]]; then - echo "ERROR: router health check failed after 3 attempts; skipping eval" >&2 - return 1 - fi - local eval_failed=0 - pushd /workspace >/dev/null || return 1 - if [[ -n "${EVAL_CONC:-}" ]]; then - export EVAL_CONCURRENT_REQUESTS="${EVAL_CONC}" - else - export EVAL_CONCURRENT_REQUESTS=$(echo "$BENCH_MAX_CONCURRENCY" | tr 'x' '\n' | sort -n | tail -1) - fi - # run_lm_eval reads the endpoint from PORT (check_env_vars) and names the - # model from MODEL_NAME, which the vLLM prefill also serves next to - # SERVED_MODEL_NAME; MODEL stays the local HF dir for the context lookup. - export PORT="$ROUTER_PORT" - export MODEL="$MODEL_PATH" - export MAX_MODEL_LEN="$TILERT_MAX_MODEL_LEN" - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: run_eval --port $ROUTER_PORT (framework=${EVAL_FRAMEWORK}, conc=${EVAL_CONCURRENT_REQUESTS})" - else - run_eval --port "$ROUTER_PORT" - local eval_rc=$? - if [[ $eval_rc -ne 0 ]]; then - echo "ERROR: run_eval exited rc=$eval_rc; preserving failure artifacts" >&2 - eval_failed=1 - else - export TP="${PREFILL_TP_SIZE}" CONC="${EVAL_CONCURRENT_REQUESTS}" EP_SIZE=1 - export PREFILL_TP="${PREFILL_TP_SIZE}" PREFILL_EP=1 PREFILL_NUM_WORKERS="${xP}" - export DECODE_TP="${DECODE_TP_SIZE}" DECODE_EP=1 DECODE_NUM_WORKERS="${yD}" - export DP_ATTENTION=false PREFILL_DP_ATTENTION=false DECODE_DP_ATTENTION=false - export ISL="${BENCH_INPUT_LEN}" OSL="${BENCH_OUTPUT_LEN}" - # As on the SGLang path: rewrite meta_env.json from the exports above, - # then stage unless run_eval already did (eval-only). - rewrite_lm_eval_meta_env - if [[ "$EVAL_ONLY" != "true" ]]; then - append_lm_eval_summary - fi - fi - local eval_copy_dir="$LOG_DIR/eval_results" - if stage_eval_artifacts "$eval_copy_dir" /workspace "${EVAL_RESULT_DIR:-}"; then - echo "Eval artifacts staged in $eval_copy_dir" - else - echo "ERROR: failed to stage eval artifacts in $eval_copy_dir" >&2 - eval_failed=1 - fi - fi - popd >/dev/null || true - return $eval_failed -} - -run_agentic_replay() { - local rc=0 - wait_for_server_ready --port "$ROUTER_PORT" --server-log "$LOG_DIR/router_${host_name}.log" --server-pid "$ROUTER_PID" - cd /workspace || return 1 - - export PORT="$ROUTER_PORT" - export MODEL="$MODEL_PATH" # aiperf --tokenizer (local HF dir) - export SERVED_MODEL_NAME # aiperf --model (name the router/vLLM serve) - check_env_vars DURATION RESULT_FILENAME INFMAX_CONTAINER_WORKSPACE - export MAX_MODEL_LEN="$TILERT_MAX_MODEL_LEN" - # TileRT decode exposes no /metrics route; only the vLLM prefill is scraped. - export AIPERF_SERVER_METRICS_URLS="http://${PREFILL_HOST}:${PREFILL_PORT}/metrics" - export TRANSFORMERS_VERBOSITY=error TOKENIZERS_PARALLELISM=false - # Keep the trace corpus and aiperf's HF downloads on the node's /tmp mount - # instead of the container's ephemeral ~/.cache, as the SGLang client does. - export HF_HOME=/run_logs/hf_cache - - local result_dir="$LOG_DIR/agentic" - local result_filename_base="$RESULT_FILENAME" - mkdir -p "$result_dir" - - # Neither server.sh nor this script runs with errexit; a failed bootstrap - # must not fall through into replay and its misleading cascade. - resolve_trace_source || return 1 - install_agentic_deps || return 1 - - local conc="$BENCH_MAX_CONCURRENCY" conc_result_dir - validate_agentic_concurrency "$conc" || return 1 - echo "Agentic trace replay: conc=$conc" - conc_result_dir="$result_dir/conc_${conc}" - mkdir -p "$conc_result_dir" - export CONC="$conc" USERS="$conc" - build_replay_cmd "$conc_result_dir" - export RESULT_FILENAME="${result_filename_base}_conc${conc}" - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: $REPLAY_CMD" - elif ! run_agentic_replay_and_write_outputs "$conc_result_dir"; then - echo "WARNING: agentic trace replay for conc=$conc failed (replay or validation) after writing available results" >&2 - rc=1 - fi - export RESULT_FILENAME="$result_filename_base" - return $rc -} - -echo "Waiting at the container creation barrier on $host_name" -if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: skipping container creation barrier" -elif [[ "$SKIP_CONTAINER_BARRIER" == "1" ]]; then - echo "SKIP_CONTAINER_BARRIER=1: caller asserts all containers are up" -else - # --grace 60: after the barrier passes, sync.py keeps the port open for - # max(60, timeout/2) seconds in the foreground so a peer one poll behind - # still sees it. At CONTAINER_BARRIER_TIMEOUT=5400 that is a 45-minute idle - # sleep on every rank (jobs 45373/45374 slept 06:28-07:13). Both ranks pass - # within one 5 s poll of each other and the stages below have their own - # readiness waits, so 60 s is plenty. - "$PY" "$WS_PATH/sync.py" barrier \ - --local-ip "${host_ip}" --local-port 5000 --enable-port \ - --node-ips "${IPADDRS}" --node-ports 5000 \ - --wait-for-all-ports --timeout "$CONTAINER_BARRIER_TIMEOUT" --grace 60 \ - || { echo "ERROR: container creation barrier failed after ${CONTAINER_BARRIER_TIMEOUT}s -- the peer rank never opened port 5000." \ - "A cold image pull is the usual cause: this recipe pulls two ~32 GB images, one per rank, and the rank that" \ - "comes up first waits out the whole timeout while the other is still pulling." >&2; exit 1; } -fi - -case "$TILERT_ROLE" in - decode) - echo "${host_name}:${host_ip} is the TileRT Decode Node (Model: ${MODEL_NAME}, profile: ${TILERT_PROFILE})" - rdma_preflight || exit 1 - convert_weights || exit 1 - stage_tokenizer_files || exit 1 - start_decode - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: decode role complete"; exit 0 - fi - echo "Waiting for the router port ${PREFILL_HOST}:${ROUTER_PORT} to open (timeout ${ROUTER_WAIT}s)..." - wait_for_tcp "$PREFILL_HOST" "$ROUTER_PORT" "$ROUTER_WAIT" "$DECODE_PID"; wrc=$? - if [[ $wrc -eq 2 ]]; then - echo "ERROR: decode_server exited before the router came up (see $LOG_DIR/decode_${host_name}.log)" >&2 - tail -50 "$LOG_DIR/decode_${host_name}.log" >&2 || true - copy_logs_to_shared; exit 1 - elif [[ $wrc -ne 0 ]]; then - echo "WARNING: router never opened within ${ROUTER_WAIT}s; shutting down decode" >&2 - kill "$DECODE_PID" 2>/dev/null || true - copy_logs_to_shared; exit 1 - fi - echo "Waiting until the router port closes..." - wait_for_tcp_close "$PREFILL_HOST" "$ROUTER_PORT" "$DECODE_PID"; wrc=$? - if [[ $wrc -eq 2 ]]; then - echo "ERROR: decode_server died while the benchmark was running (see $LOG_DIR/decode_${host_name}.log)" >&2 - tail -50 "$LOG_DIR/decode_${host_name}.log" >&2 || true - copy_logs_to_shared; exit 1 - fi - echo "Killing the decode server" - kill "$DECODE_PID" 2>/dev/null || true - sleep 2 - copy_logs_to_shared - ;; - prefill) - echo "NODE INFO =======================================" - echo "Node List : ${SLURM_JOB_NODELIST:-}" - echo "Node IPs : ${IPADDRS}" - echo "Model : ${MODEL_NAME}" - echo "${host_name}:${host_ip} is the Prefill Node (vLLM + TileRTConnector) and Router Node" - echo "================================================" - rdma_preflight || exit 1 - echo "Waiting for the decode ctrl port ${DECODE_HOST}:${DECODE_CTRL_PORT} (timeout ${DECODE_WAIT}s)..." - if [[ "$DRY_RUN" -eq 0 ]]; then - wait_for_tcp "$DECODE_HOST" "$DECODE_CTRL_PORT" "$DECODE_WAIT" \ - || echo "WARNING: timed out waiting for the decode ctrl port; starting prefill anyway" >&2 - fi - start_prefill - if [[ "$DRY_RUN" -eq 0 ]]; then - wait_for_tcp "$PREFILL_HOST" "$PREFILL_PORT" "$PREFILL_WAIT" "$PREFILL_PID"; wrc=$? - if [[ $wrc -ne 0 ]]; then - echo "ERROR: vLLM prefill did not open ${PREFILL_HOST}:${PREFILL_PORT} (rc=$wrc, see $LOG_DIR/prefill_${host_name}.log)" >&2 - tail -50 "$LOG_DIR/prefill_${host_name}.log" >&2 || true - kill "$PREFILL_PID" 2>/dev/null || true - copy_logs_to_shared; exit 1 - fi - fi - start_router - if [[ "$DRY_RUN" -eq 1 ]]; then - echo "DRY RUN: prefill/router role complete"; exit 0 - fi - echo "Ready for benchmarking on ${host_name}:${host_ip}" - cd "$WS_PATH" || exit 1 - # EVAL_ONLY skips the AgentX replay and runs GSM8K on the same router; - # RUN_EVAL after a replay runs it once the replay has finished. - if [[ "$EVAL_ONLY" == "true" ]]; then - echo "EVAL_ONLY mode: skipping the AgentX replay" - wait_for_server_ready --port "$ROUTER_PORT" --server-log "$LOG_DIR/router_${host_name}.log" --server-pid "$ROUTER_PID" - export TRANSFORMERS_VERBOSITY=error TOKENIZERS_PARALLELISM=false - run_lm_eval_on_router; BENCH_RC=$? - else - run_agentic_replay; BENCH_RC=$? - if [[ "$RUN_EVAL" == "true" ]]; then - run_lm_eval_on_router || BENCH_RC=1 - fi - fi - copy_logs_to_shared - echo "Killing the router and the prefill server" - kill "$ROUTER_PID" "$PREFILL_PID" 2>/dev/null || true - sleep 2 - pkill -f "tilert.pd_vllm.pd_router" 2>/dev/null || true - pkill -f "vllm serve" 2>/dev/null || true - if [[ "$BENCH_RC" -ne 0 ]]; then - echo "ERROR: benchmark/eval reported rc=$BENCH_RC" >&2 - exit "$BENCH_RC" - fi - ;; - *) - echo "ERROR: unknown TILERT_ROLE='$TILERT_ROLE'" >&2; exit 2 ;; -esac - -echo "Script completed successfully" -exit 0 diff --git a/inferencex-e2e/benchmarks/multi_node/amd_utils/setup_deps.sh b/inferencex-e2e/benchmarks/multi_node/amd_utils/setup_deps.sh deleted file mode 100644 index 1a5c5cfacf..0000000000 --- a/inferencex-e2e/benchmarks/multi_node/amd_utils/setup_deps.sh +++ /dev/null @@ -1,136 +0,0 @@ -#!/bin/bash - -source "$(dirname "${BASH_SOURCE[0]}")/../../benchmark_lib.sh" --validation-only -# Install missing disagg dependencies at container start. Each installer is -# idempotent and gated on $ENGINE. - -_SETUP_START=$(date +%s) -_SETUP_INSTALLED=() - -# Pinned by the recipe (TILERT_VERSION); the rest are fixed properties of the -# TileRT 0.1.x runtime rather than caller configuration. -TILERT_PACKAGE=tilert -TILERT_HTTP_DEPS="fastapi uvicorn httpx" -TILERT_TRANSPORT_DEPS="mooncake-transfer-engine-rocm>=0.3.13" -TILERT_TRANSFORMERS_SPEC="transformers>=4.56" - -_tilert_resolve_python() { - if [[ -n "${PY:-}" ]] && command -v "$PY" >/dev/null 2>&1; then :; else - PY="" - local c - for c in python3 python; do command -v "$c" >/dev/null 2>&1 && { PY="$c"; break; }; done - fi - [[ -n "$PY" ]] || { echo "[SETUP] ERROR: neither python3 nor python found"; exit 1; } - export PY - echo "[SETUP] interpreter PY=$PY ($(command -v "$PY"))" -} - -_tilert_installed_version() { - "$PY" - "$1" <<'PYEOF' 2>/dev/null -import sys -from importlib.metadata import version, PackageNotFoundError -try: - print(version(sys.argv[1])) -except PackageNotFoundError: - pass -PYEOF -} - -_tilert_pip() { - "$PY" -m pip install --quiet --no-cache-dir "$@" -} - -_tilert_install_missing() { - local probe="$1"; shift - [[ $# -gt 0 ]] || return 0 - if "$PY" -c "import $probe" 2>/dev/null; then - echo "[SETUP] $probe already present, skipping ($*)" - return 0 - fi - echo "[SETUP] installing $* (probe module '$probe' missing)" - _tilert_pip "$@" || { echo "[SETUP] ERROR: failed to install: $*"; exit 1; } - _SETUP_INSTALLED+=("$*") -} - -install_tilert_container_tools() { - if command -v ip >/dev/null 2>&1 && command -v curl >/dev/null 2>&1 \ - && command -v ibv_devices >/dev/null 2>&1; then - echo "[SETUP] Container RDMA/net tools already present" - return 0 - fi - echo "[SETUP] Installing iproute2 + curl + ibverbs userspace in container..." - apt-get update -q -y && apt-get install -q -y --no-install-recommends \ - iproute2 curl ibverbs-utils libibverbs1 librdmacm1 ibverbs-providers \ - && rm -rf /var/lib/apt/lists/* - if ! command -v ip >/dev/null 2>&1 || ! command -v curl >/dev/null 2>&1; then - echo "[SETUP] ERROR: failed to install iproute2/curl"; exit 1 - fi - _SETUP_INSTALLED+=("iproute2+curl+ibverbs") -} - -_tilert_install_wheel() { - local mode="$1" # full | no-deps - local have; have="$(_tilert_installed_version tilert)" - if [[ "$have" == "$TILERT_VERSION" ]]; then - echo "[SETUP] tilert $have already installed, skipping" - return 0 - fi - [[ -n "$have" ]] && echo "[SETUP] tilert $have installed, switching to pinned $TILERT_VERSION" - if [[ "$mode" == "no-deps" ]]; then - echo "[SETUP] installing $TILERT_PIP_SPEC --no-deps (connector plugin + router on top of the image's vLLM)" - _tilert_pip --no-deps "$TILERT_PIP_SPEC" || { echo "[SETUP] ERROR: failed to install $TILERT_PIP_SPEC (--no-deps)"; exit 1; } - else - echo "[SETUP] installing $TILERT_PIP_SPEC (TileRT ROCm build, official PyPI wheel)" - _tilert_pip "$TILERT_PIP_SPEC" || { echo "[SETUP] ERROR: failed to install $TILERT_PIP_SPEC"; exit 1; } - fi - have="$(_tilert_installed_version tilert)" - [[ "$have" == "$TILERT_VERSION" ]] || { - echo "[SETUP] ERROR: tilert is ${have:-not installed} after install, expected $TILERT_VERSION"; exit 1; } - _SETUP_INSTALLED+=("$TILERT_PACKAGE==$TILERT_VERSION($mode)") -} - -install_tilert_decode() { - install_tilert_container_tools - _tilert_install_wheel full - _tilert_install_missing uvicorn $TILERT_HTTP_DEPS - _tilert_install_missing mooncake.engine "$TILERT_TRANSPORT_DEPS" - _tilert_install_missing transformers "$TILERT_TRANSFORMERS_SPEC" - "$PY" -c "import tilert.pd_vllm.decode_server" 2>/dev/null || { - echo "[SETUP] ERROR: import tilert.pd_vllm.decode_server failed:" - "$PY" -c "import tilert.pd_vllm.decode_server" 2>&1 | tail -3 - exit 1; } - echo "[SETUP] tilert.pd_vllm.decode_server imports OK" -} - -install_tilert_prefill() { - local vllm_v; vllm_v="$(_tilert_installed_version vllm)" - if [[ -z "$vllm_v" ]]; then - echo "[SETUP] ERROR: no vLLM in the prefill image (PREFILL_IMAGE must be a vllm/vllm-openai-rocm image)." - exit 1 - fi - echo "[SETUP] prefill-side vLLM $vllm_v" - install_tilert_container_tools - _tilert_install_wheel no-deps - _tilert_install_missing mooncake.engine "$TILERT_TRANSPORT_DEPS" - "$PY" -c "import tilert.pd_vllm.prefill_connector" 2>/dev/null || { - echo "[SETUP] WARN: import tilert.pd_vllm.prefill_connector failed (vLLM will report again when loading the connector plugin):" - "$PY" -c "import tilert.pd_vllm.prefill_connector" 2>&1 | tail -3; } -} - -if [[ "$ENGINE" == "tilert" ]]; then - check_env_vars TILERT_VERSION - TILERT_PIP_SPEC="$TILERT_PACKAGE==$TILERT_VERSION" - _tilert_resolve_python - case "${TILERT_ROLE:-}" in - decode) install_tilert_decode ;; - prefill) install_tilert_prefill ;; - *) echo "[SETUP] ERROR: ENGINE=tilert needs TILERT_ROLE=decode|prefill (got '${TILERT_ROLE:-}')"; exit 1 ;; - esac -fi - -_SETUP_END=$(date +%s) -if [[ ${#_SETUP_INSTALLED[@]} -eq 0 ]]; then - echo "[SETUP] All dependencies already present ($(( _SETUP_END - _SETUP_START ))s wallclock)" -else - echo "[SETUP] Installed: ${_SETUP_INSTALLED[*]} in $(( _SETUP_END - _SETUP_START ))s" -fi diff --git a/inferencex-e2e/benchmarks/multi_node/runtime_settings.sh b/inferencex-e2e/benchmarks/multi_node/runtime_settings.sh index 1590aac836..faf40b4a2b 100644 --- a/inferencex-e2e/benchmarks/multi_node/runtime_settings.sh +++ b/inferencex-e2e/benchmarks/multi_node/runtime_settings.sh @@ -41,9 +41,7 @@ case "$FRAMEWORK" in fi ;; tilert) - # RUNNER_TYPE selects the AMD block below, so a missing value must fail - # here rather than silently skip it. - check_env_vars GITHUB_WORKSPACE RUNNER_TYPE + check_env_vars GITHUB_WORKSPACE export BENCHMARK_LOGS_DIR="$GITHUB_WORKSPACE" RESULT_DIR=/workspace export GPU_MEM_UTIL=0.75 DECODE_CTRL_PORT=5556 DECODE_HTTP_PORT=5557 PREFILL_PORT=8000 export DECODE_WAIT=3600 PREFILL_WAIT=3600 TILERT_QUEUE_TIMEOUT=0 @@ -52,18 +50,6 @@ case "$FRAMEWORK" in if [[ "$IS_AGENTIC" == 1 || "$IS_AGENTIC" == true ]]; then export TILERT_QUEUE_TIMEOUT=1800 fi - # The MI355X TileRT recipe runs through the shared amd_utils chain - # (submit.sh -> job.slurm -> server.sh -> setup_deps.sh), which validates - # the same orchestration inputs the AMD SGLang arm receives. - # Without them submit.sh exits before sbatch and the launcher never gets a job id. - if [[ "$RUNNER_TYPE" == *mi355x-amds* ]]; then - export SKIP_RDMA_CHECK=0 SKIP_GPU_SANITY=0 - export ROUTER_TYPE=tilert-pd-router ROUTER_PORT=30000 PROXY_PING_PORT=36367 - export HEADNODE_PORT=20000 SERVER_PORT=2584 PROXY_STREAM_IDLE_TIMEOUT=300 - export ENABLE_METRICS=0 PREFILL_ROUTER_POLICY=random DECODE_ROUTER_POLICY=random - export DECODE_MTP_SIZE=0 - export ROCM_PATH=/opt/rocm UCX_HOME=/usr/local/ucx RIXL_HOME=/usr/local/rixl - fi ;; llmd-vllm) export LLMD_CONTAINER_ENGINE=docker VLLM_RANDOMIZE_DP_DUMMY_INPUTS=1 diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/configs/glm5.3-tilert-rocm.sh b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/configs/glm5.3-tilert-rocm.sh new file mode 100644 index 0000000000..e9906ebb2c --- /dev/null +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/configs/glm5.3-tilert-rocm.sh @@ -0,0 +1,28 @@ +#!/usr/bin/env bash +set -eo pipefail + +source /infmax-workspace/benchmarks/benchmark_lib.sh --validation-only +check_env_vars TILERT_VERSION TILERT_ROLE +case "$TILERT_ROLE" in + prefill|decode|router) ;; + *) echo "Unknown TileRT role: $TILERT_ROLE" >&2; exit 1 ;; +esac + +install_args=(--quiet --no-cache-dir) +if [[ "$TILERT_ROLE" == prefill ]]; then + # Preserve the prefill image's vLLM/Torch dependency set. + install_args+=(--no-deps) +fi +python3 -m pip install "${install_args[@]}" "tilert==$TILERT_VERSION" + +if ! python3 -c 'import mooncake.engine' >/dev/null 2>&1; then + python3 -m pip install --quiet --no-cache-dir 'mooncake-transfer-engine-rocm>=0.3.13' +fi +if [[ "$TILERT_ROLE" != prefill ]]; then + if ! python3 -c 'import uvicorn' >/dev/null 2>&1; then + python3 -m pip install --quiet --no-cache-dir fastapi uvicorn httpx + fi + if ! python3 -c 'import transformers' >/dev/null 2>&1; then + python3 -m pip install --quiet --no-cache-dir 'transformers>=4.56' + fi +fi diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml new file mode 100644 index 0000000000..6ee36f441f --- /dev/null +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml @@ -0,0 +1,102 @@ +schema: 2 +name: glm5.3-mi355x-tilert-agentx +model: + path: GLM-5.3 + container: ghcr.io/tile-ai/tilert-rocm-decode:0.1.6 + precision: fp8 +slurm: + time_limit: "08:00:00" +resources: + gpu_type: mi355x + gpus_per_node: 8 +setup_script: glm5.3-tilert-rocm.sh +environment: + TILERT_VERSION: "0.1.6.post2" + RDMAV_FORK_SAFE: "1" + PYTHONDONTWRITEBYTECODE: "1" + NCCL_SOCKET_IFNAME: eno0 + GLOO_SOCKET_IFNAME: eno0 + NCCL_IB_HCA: rdma0,rdma1,rdma2,rdma3,rdma4,rdma5,rdma6,rdma7 +roles: + prefill: + engine: + type: vllm + set_visible_devices: true + container: ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6 + nodes: 1 + workers: 1 + gpus: 8 + env: + TILERT_ROLE: prefill + VLLM_ROCM_USE_AITER: "1" + VLLM_ROCM_USE_AITER_MOE: "1" + VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "1" + VLLM_ENGINE_READY_TIMEOUT_S: "10800" + args: + served-model-name: glm5_2 + tensor-parallel-size: 8 + max-model-len: 1048576 + enforce-eager: true + trust-remote-code: true + return-tokens-as-token-ids: true + gpu-memory-utilization: 0.85 + kv-cache-dtype: bfloat16 + block-size: 64 + speculative-config: '{"method":"mtp","num_speculative_tokens":1}' + kv-transfer-config: >- + {"kv_connector":"TileRTConnector","kv_connector_module_path":"tilert.pd_vllm.prefill_connector", + "kv_role":"kv_producer","kv_connector_extra_config":{ + "tilert_model":"glm5_2","tilert_max_seq_len":1048576,"tilert_transport":"mooncake"}} + decode: + engine: tilert + container: ghcr.io/tile-ai/tilert-rocm-decode:0.1.6 + nodes: 1 + workers: 1 + gpus: 8 + env: + TILERT_ROLE: decode + GLM5_AR_N: "2" + args: + model: glm5_2 + model-weights-dir: /models/GLM-5.3-tilert-tp8 + max-seq-len: 1048576 + kv-cache-dtype: bf16 + transport: mooncake + with-mtp: true + num-mtp: 3 +frontend: + type: tilert-router + container_image: ghcr.io/tile-ai/tilert-rocm-decode:0.1.6 + enable_multiple_frontends: false + env: + TILERT_ROLE: router + args: + parser: none + model-path: /model + queue-timeout: 1800 +sbatch_directives: + cpus-per-task: "128" + mem: "0" +srun_options: + mem: "0" + container-writable: "" + container-remap-root: "" +health_check: + max_attempts: 2160 + interval_seconds: 5 +benchmark: + type: custom + command: bash /infmax-workspace/benchmarks/srt_agentic.sh + env: + CONC: "1" + MODEL: /model + SERVED_MODEL_NAME: glm5_2 + AIPERF_MAX_CONTEXT_LENGTH: "1048576" + RESULT_DIR: /infmax-workspace/LOGS/agentic + AGENTIC_OUTPUT_DIR: /infmax-workspace + AIPERF_DATASET_MMAP_CACHE_DIR: /aiperf_mmap_cache + AIPERF_REQUIRED_SERVER_METRIC_PREFIX: "vllm:" + HF_HOME: /logs/hf_cache + HF_HUB_OFFLINE: "" + TRANSFORMERS_VERBOSITY: error + TOKENIZERS_PARALLELISM: "false" diff --git a/inferencex-e2e/configs/amd-master.yaml b/inferencex-e2e/configs/amd-master.yaml index 50b957ff30..27d23c28a5 100644 --- a/inferencex-e2e/configs/amd-master.yaml +++ b/inferencex-e2e/configs/amd-master.yaml @@ -1477,21 +1477,8 @@ dsv41flash-fp4-mi355x-vllm-agentic-dspark: - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/mi355x-fp4-mtp/agentic.yaml } - { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/mi355x-fp4-mtp/agentic.yaml } -# Speculative decoding on an agentic scenario must run with simulated -# synthetic acceptance at the committed golden AL for this model, thinking mode -# and draft length (docs/PR_REVIEW_CHECKLIST.md), and a submission may not -# substitute its own target. Nothing about that target is written here: -# AGENTS.md forbids hard-coding an acceptance length in a master config, so -# server_tilert.sh reads infx/golden_al_distribution/glm5.3_mtp.yaml at launch and -# fails the run if the curve or the draft length is missing. -# The curve is consumed in the same units as every other framework here, and -# AgentX replays run with thinking on. -# -# NOTE FOR REVIEWERS: that curve is GLM-5.2's, copied because no SPEED-Bench run -# on GLM-5.3 exists and the two share a base. It is committed as provisional and -# labelled as such in the file. Flagging it rather than letting it read as a -# measured 5.3 curve -- please say if you would rather see a measured curve, a -# non-MTP agentic entry, or a waiver. +# The golden acceptance curve remains the provisional GLM-5.2-derived curve +# documented in infx/golden_al_distribution/glm5.3_mtp.yaml. glm5.3-fp8-mi355x-tilert-agentic: image: ghcr.io/tile-ai/tilert-rocm-decode:0.1.6 model: zai-org/GLM-5.3 @@ -1515,15 +1502,12 @@ glm5.3-fp8-mi355x-tilert-agentic: dp-attn: false additional-settings: - "PREFILL_IMAGE=ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6" - - "PREFILL_NODES=1" + - "CONFIG_FILE=recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml" decode: num-worker: 1 tp: 8 ep: 1 dp-attn: false - additional-settings: - - "DECODE_NODES=1" - - "TILERT_EXTRA_ENV=GLM5_AR_N=2" # User-selected official DeepSeek V4.1 MI35X model preview, pinned by digest. # Official MI350X TP4/EP4 candidate; compare against the separate TP4/EP1 vLLM baseline. diff --git a/inferencex-e2e/configs/runners.yaml b/inferencex-e2e/configs/runners.yaml index cf67501a63..c96b90734d 100644 --- a/inferencex-e2e/configs/runners.yaml +++ b/inferencex-e2e/configs/runners.yaml @@ -720,6 +720,7 @@ clusters: models: entries: DeepSeek-V4-Pro-0813: {root: it-share-data, dir: DeepSeek-V4-Pro-0813} + GLM-5.3: {root: it-share-data, dir: GLM-5.3} Qwen3.5-397B-A17B-FP8: {root: it-share-data, dir: Qwen3.5-397B-A17B-FP8} scheduler: slurm slurm: diff --git a/inferencex-e2e/infx/launch/drivers/srt/lanes.py b/inferencex-e2e/infx/launch/drivers/srt/lanes.py index e8217eb607..7be02aeb4b 100644 --- a/inferencex-e2e/infx/launch/drivers/srt/lanes.py +++ b/inferencex-e2e/infx/launch/drivers/srt/lanes.py @@ -100,7 +100,10 @@ class SrtLane: long_time=Match(any_of("dsv4"), frameworks=any_of("dynamo-sglang"), agentic=True), ), ("mi355x-amds", LaunchPath.SRT_MULTI): SrtLane( - mounts=(LaneMount(Match(), "aiperf-cache", "/aiperf_mmap_cache"),), + mounts=( + LaneMount(Match(), "aiperf-cache", "/aiperf_mmap_cache"), + LaneMount(Match(frameworks=any_of("tilert")), "it-share-data", "/models"), + ), eval_unsets=( "roles.prefill.args.ep-dispatch-algorithm", "roles.decode.args.ep-dispatch-algorithm", diff --git a/inferencex-e2e/infx/srt_slurm/synthetic_acceptance.py b/inferencex-e2e/infx/srt_slurm/synthetic_acceptance.py index 691c86c009..19ba42a60d 100644 --- a/inferencex-e2e/infx/srt_slurm/synthetic_acceptance.py +++ b/inferencex-e2e/infx/srt_slurm/synthetic_acceptance.py @@ -29,6 +29,7 @@ "dynamo-trt": "trtllm", "atom": "atom", "atom-disagg": "atom", + "tilert": "tilert", } SGLANG_VARIABLES = ( "SGLANG_SIMULATE_ACC_LEN", @@ -36,10 +37,15 @@ "SGLANG_SIMULATE_ACC_TOKEN_MODE", ) TRT_VARIABLE = "TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS" +TILERT_VARIABLES = ("TILERT_SIMULATE_ACC_LEN", "TILERT_SIMULATE_ACC_METHOD") def spec_parameters(role: Mapping[str, Any], engine: str) -> dict[str, Any]: args = role.get("args", {}) + if engine == "tilert": + if args.get("with-mtp") is not True: + return {} + return {"method": "mtp", "num_speculative_tokens": args.get("num-mtp")} if engine == "atom": method = args.get("method") if not method: @@ -117,7 +123,11 @@ def build_overrides( environment["MODEL_PREFIX"], spec, environment["THINKING_MODE"], golden_dir ) overrides = [] - variables = {"sglang": SGLANG_VARIABLES, "trtllm": (TRT_VARIABLE,)}.get(engine, ()) + variables = { + "sglang": SGLANG_VARIABLES, + "trtllm": (TRT_VARIABLE,), + "tilert": TILERT_VARIABLES, + }.get(engine, ()) # SRT applies recipe-wide environment after role environment. Keep simulation # role-local so global values cannot override the golden AL or leak into evals. for key in variables: @@ -127,13 +137,27 @@ def build_overrides( if name not in ("agg", "prefill", "decode"): continue prefix = f"roles.{name}" - worker_spec = spec_parameters(role, engine) - if engine == "vllm": + worker_engine = engine + worker_al = al + if engine == "tilert": + selected = role.get("engine", recipe.get("engine", "tilert")) + worker_engine = selected.get("type") if isinstance(selected, Mapping) else selected + # TileRT's first token and draft cache come from real vLLM prefill. + # Only its decode runtime simulates acceptance, using environment + # variables (decode_server has no --simulate-acc-* CLI options). + if name != "decode" or worker_engine != "tilert": + worker_al = None + worker_spec = spec_parameters(role, worker_engine) + if engine == "tilert" and (worker_al is None or not worker_spec): + for key in TILERT_VARIABLES: + if key in (role.get("env") or {}): + overrides += ["--unset", f"{prefix}.env.{key}"] + if worker_engine == "vllm": if not worker_spec: continue - if al is not None: + if worker_al is not None: worker_spec.update( - rejection_sample_method="synthetic", synthetic_acceptance_length=al + rejection_sample_method="synthetic", synthetic_acceptance_length=worker_al ) elif ( worker_spec.get("rejection_sample_method") == "synthetic" @@ -147,22 +171,24 @@ def build_overrides( "--set", f"{prefix}.args.speculative-config={json.dumps(worker_spec)}", ] - elif engine == "atom": + elif worker_engine == "atom": # ATOM forces acceptance with a server flag rather than environment. key = "spec-decode-acceptance-length" - if al is not None and worker_spec: - overrides += ["--set", f"{prefix}.args.{key}={al:g}"] + if worker_al is not None and worker_spec: + overrides += ["--set", f"{prefix}.args.{key}={worker_al:g}"] elif key in (role.get("args") or {}): overrides += ["--unset", f"{prefix}.args.{key}"] - elif al is not None and worker_spec: + elif worker_al is not None and worker_spec: values = ( - (f"{al:g}", "match-expected", "real-draft-token") + (f"{worker_al:g}", "match-expected", "real-draft-token") if engine == "sglang" - else (f"{al - 1:g}",) + else (f"{worker_al:g}", "match-expected") + if engine == "tilert" + else (f"{worker_al - 1:g}",) ) for key, value in zip(variables, values, strict=True): overrides += ["--set", f"{prefix}.env.{key}={json.dumps(value)}"] - else: + elif engine != "tilert": for key in variables: if key in (role.get("env") or {}): overrides += ["--unset", f"{prefix}.env.{key}"] diff --git a/inferencex-e2e/infx/tests/srt_slurm/test_synthetic_acceptance.py b/inferencex-e2e/infx/tests/srt_slurm/test_synthetic_acceptance.py index aa6cbf5be4..6b43d9ac49 100644 --- a/inferencex-e2e/infx/tests/srt_slurm/test_synthetic_acceptance.py +++ b/inferencex-e2e/infx/tests/srt_slurm/test_synthetic_acceptance.py @@ -50,6 +50,7 @@ def golden_dir(tmp_path: Path) -> Path: ), ("minimaxm3_eagle3.yaml", "minimax-m3", 2.5), ("minimaxm3_eagle3_gqa.yaml", "minimax-m3", 2.6), + ("glm5.3_mtp.yaml", "glm-5.3", 3.2), ]: (directory / filename).write_text( yaml.safe_dump( @@ -289,6 +290,118 @@ def test_atom_forces_golden_acceptance_by_server_flag( assert "spec-decode-acceptance-length" not in evaluated["roles"]["agg"]["args"] +def tilert_recipe() -> dict[str, Any]: + return { + "roles": { + "prefill": { + "engine": "vllm", + "args": { + "speculative-config": '{"method":"mtp","num_speculative_tokens":1}', + }, + }, + "decode": { + "engine": {"type": "tilert"}, + "args": {"with-mtp": True, "num-mtp": 3}, + "env": {"GLM5_AR_N": "2"}, + }, + }, + } + + +def test_tilert_plan_uses_caller_decode_depth_and_keeps_prefill_real( + tmp_path: Path, golden_dir: Path +) -> None: + recipe = tilert_recipe() + recipe["roles"]["decode"]["args"]["num-mtp"] = 2 + recipe["environment"] = { + "KEEP": "global", + "TILERT_SIMULATE_ACC_LEN": "99", + "TILERT_SIMULATE_ACC_METHOD": "stale-method", + } + path = tmp_path / "tilert.yaml" + path.write_text(yaml.safe_dump(recipe)) + commands = plan_commands( + str(path), + "tilert", + ["--set", "roles.decode.args.num-mtp=3"], + {**ENV, "MODEL_PREFIX": "glm5.3", "RUN_EVAL": "true"}, + golden_dir=golden_dir, + ) + assert len(commands) == 1 + result = apply_native(recipe, commands[0]) + assert result["roles"]["decode"]["env"] == { + "GLM5_AR_N": "2", + "TILERT_SIMULATE_ACC_LEN": "3.2", + "TILERT_SIMULATE_ACC_METHOD": "match-expected", + } + assert result["roles"]["decode"]["args"] == {"with-mtp": True, "num-mtp": 3} + assert result["environment"] == {"KEEP": "global"} + assert json.loads(result["roles"]["prefill"]["args"]["speculative-config"]) == { + "method": "mtp", + "num_speculative_tokens": 1, + } + + +@pytest.mark.parametrize( + "environment", + [{"EVAL_ONLY": "true"}, {"IS_AGENTIC": "0"}, {"SPEC_DECODING": "none"}, {}], +) +def test_tilert_real_verification_removes_stale_role_and_global_simulation( + tmp_path: Path, environment: dict[str, str] +) -> None: + recipe = tilert_recipe() + if not environment: + recipe["roles"]["decode"]["args"]["with-mtp"] = False + stale = {"TILERT_SIMULATE_ACC_LEN": "99", "TILERT_SIMULATE_ACC_METHOD": "match-expected"} + recipe["environment"] = {"KEEP": "global", **stale} + for role in recipe["roles"].values(): + role.setdefault("env", {}).update(stale) + recipe["roles"]["prefill"]["engine"] = {"type": "vllm"} + recipe["roles"]["prefill"]["args"]["speculative-config"] = json.dumps( + { + "method": "mtp", + "num_speculative_tokens": 1, + "rejection_sample_method": "synthetic", + "synthetic_acceptance_length": 99, + } + ) + result = apply_native( + recipe, + build_overrides( + recipe, + "tilert", + {**ENV, "MODEL_PREFIX": "glm5.3", **environment}, + golden_dir=tmp_path / "missing", + ), + ) + assert result["environment"] == {"KEEP": "global"} + assert result["roles"]["decode"]["env"] == {"GLM5_AR_N": "2"} + assert result["roles"]["prefill"]["env"] == {} + assert json.loads(result["roles"]["prefill"]["args"]["speculative-config"]) == { + "method": "mtp", + "num_speculative_tokens": 1, + "rejection_sample_method": "block", + } + + +@pytest.mark.parametrize("depth", [None, 7]) +def test_tilert_requires_explicit_measured_decode_depth( + golden_dir: Path, depth: int | None +) -> None: + recipe = tilert_recipe() + if depth is None: + del recipe["roles"]["decode"]["args"]["num-mtp"] + else: + recipe["roles"]["decode"]["args"]["num-mtp"] = depth + with pytest.raises(ValueError, match="positive integer draft length|No golden acceptance"): + build_overrides( + recipe, + "tilert", + {**ENV, "MODEL_PREFIX": "glm5.3"}, + golden_dir=golden_dir, + ) + + @pytest.mark.parametrize( "curve", [ From 064fe4452d3aa54d038761a951209ba4d38e82d4 Mon Sep 17 00:00:00 2001 From: CrimsonDump Date: Thu, 1 Oct 2026 03:38:40 +0800 Subject: [PATCH 4/4] tilert: GLM-5.3 FP8 MI355X AgentX on 0.1.6.post3 with a vLLM 0.28 + ATOM prefill MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Stacked on the declarative srt-slurm TileRT recipe (#3552). Bump tilert 0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1 (the master config's PREFILL_IMAGE, which the launcher stages, follows). post3 turns on multi-sender staging, per-layer pipelined KV send and router-side incremental chat tokenization by default. The prefill role drops enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling, fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry. 基于声明式 srt-slurm TileRT 配方(#3552)。tilert 由 0.1.6.post2 升级到 0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像 ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1(master config 中供启动器预导 镜像的 PREFILL_IMAGE 同步更新)。post3 默认开启多发送端暂存、逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager,开启 CUDA graph (FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、prefix caching、 16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方的 AITER 设置。 追加 perf-changelog 条目。 Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL --- .../mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml | 12 +++++++++--- inferencex-e2e/configs/amd-master.yaml | 4 ++-- inferencex-e2e/perf-changelog.yaml | 9 +++++++++ 3 files changed, 20 insertions(+), 5 deletions(-) diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml index 6ee36f441f..e1372c7fb4 100644 --- a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml @@ -11,7 +11,7 @@ resources: gpus_per_node: 8 setup_script: glm5.3-tilert-rocm.sh environment: - TILERT_VERSION: "0.1.6.post2" + TILERT_VERSION: "0.1.6.post3" RDMAV_FORK_SAFE: "1" PYTHONDONTWRITEBYTECODE: "1" NCCL_SOCKET_IFNAME: eno0 @@ -22,7 +22,7 @@ roles: engine: type: vllm set_visible_devices: true - container: ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6 + container: ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1 nodes: 1 workers: 1 gpus: 8 @@ -32,16 +32,22 @@ roles: VLLM_ROCM_USE_AITER_MOE: "1" VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS: "1" VLLM_ENGINE_READY_TIMEOUT_S: "10800" + AITER_QUICK_REDUCE_QUANTIZATION: INT4 + AITER_USE_FLYDSL_MOE_SORTING: "1" args: served-model-name: glm5_2 tensor-parallel-size: 8 max-model-len: 1048576 - enforce-eager: true trust-remote-code: true return-tokens-as-token-ids: true gpu-memory-utilization: 0.85 kv-cache-dtype: bfloat16 block-size: 64 + async-scheduling: true + load-format: fastsafetensors + compilation-config: '{"cudagraph_mode":"FULL_AND_PIECEWISE"}' + max-num-batched-tokens: 16384 + enable-prefix-caching: true speculative-config: '{"method":"mtp","num_speculative_tokens":1}' kv-transfer-config: >- {"kv_connector":"TileRTConnector","kv_connector_module_path":"tilert.pd_vllm.prefill_connector", diff --git a/inferencex-e2e/configs/amd-master.yaml b/inferencex-e2e/configs/amd-master.yaml index 27d23c28a5..421a9feaf9 100644 --- a/inferencex-e2e/configs/amd-master.yaml +++ b/inferencex-e2e/configs/amd-master.yaml @@ -1486,7 +1486,7 @@ glm5.3-fp8-mi355x-tilert-agentic: runner: cluster:mi355x-amds precision: fp8 framework: tilert - router: { name: tilert-pd-router, version: "0.1.6.post2" } + router: { name: tilert-pd-router, version: "0.1.6.post3" } multinode: true disagg: true kv-p2p-transfer: mooncake @@ -1501,7 +1501,7 @@ glm5.3-fp8-mi355x-tilert-agentic: ep: 1 dp-attn: false additional-settings: - - "PREFILL_IMAGE=ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6" + - "PREFILL_IMAGE=ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1" - "CONFIG_FILE=recipes/glm5.3/tilert/mi355x-fp8/agentx/disagg-1p1d-tp8-mtp.yaml" decode: num-worker: 1 diff --git a/inferencex-e2e/perf-changelog.yaml b/inferencex-e2e/perf-changelog.yaml index b415cf56a5..8983198c18 100644 --- a/inferencex-e2e/perf-changelog.yaml +++ b/inferencex-e2e/perf-changelog.yaml @@ -9158,3 +9158,12 @@ - "Port the DeepSeek-V4-Pro-0813 MI355X ATOM 1P1D AgentX config from the legacy amd_utils path to a native srt-slurm recipe (benchmarks/multi_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4/agentx/disagg-lmcache-dspark.yaml), one override variant per point. Server flags, env, AToMesh routing policies and decode CUDA-graph capture sizes match the legacy server_atom.sh for every tier: TP8 at concurrency 1 and 16, DP attention with prefill TBO at 64 and 128, and DP attention with ATOM's in-process LMCache CPU offload (lmcache_offload, 187 GB per prefill rank) at 256. Golden acceptance 3.01 for DSpark with three draft tokens is now injected by the srt-slurm path, the same value the legacy models_atom.yaml hardcoded." - "The LMCache tier uses srt-slurm's extra-kv-connectors (NVIDIA/srt-slurm#507, in v2.36.0) to add lmcache_offload next to the generated Mooncake connector." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3543 + +- config-keys: + - glm5.3-fp8-mi355x-tilert-agentic + scenario-type: + - agentic-coding + description: + - "Bump tilert 0.1.6.post2 -> 0.1.6.post3 (PyPI, 2026-09-28) on both TileRT ranks; router metadata follows. post3 turns on three prefill-to-decode latency features by default: multi-sender staging (TILERT_PD_SENDERS=8, each TP rank extracts and sends its own layers), per-layer pipelined KV send (TILERT_PD_PIPELINE=1, a layer is extracted and RDMA-written as soon as it is computed), and router-side incremental chat tokenization (TILERT_ROUTER_TOKENIZE=1, the router tokenizes only the new turn through a per-segment cache and hands vLLM token ids; a start-up self-check against vLLM /tokenize keeps it off on any mismatch). Move the prefill rank to a vLLM 0.28 + ATOM image (ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built on rocm/atom-dev:vllm-v0.28.0-nightly_20260923) with CUDA graphs, async scheduling, prefix caching and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Decode image, 1M context, bf16 KV and layer-sharded PD buffers unchanged. TileRT reports AgentX TTFT p90 4.35 s -> 1.60 s at concurrency 1 (local 2x8 MI350X against the post2 MI355X run)." + - "两侧 TileRT rank 上 tilert 由 0.1.6.post2 升级到 0.1.6.post3(PyPI,2026-09-28),router 元数据随之更新。post3 默认打开三项 prefill 到 decode 的时延优化:多发送端暂存(TILERT_PD_SENDERS=8,每个 TP rank 抽取并发送自己负责的层)、逐层流水线发送 KV(TILERT_PD_PIPELINE=1,每层算完即抽取并 RDMA 写出)、router 侧增量对话分词(TILERT_ROUTER_TOKENIZE=1,router 经按段缓存只对新一轮分词,把 token id 交给 vLLM;启动自检与 vLLM /tokenize 不一致即保持关闭)。prefill rank 改用 vLLM 0.28 + ATOM 镜像(ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于 rocm/atom-dev:vllm-v0.28.0-nightly_20260923),开启 CUDA graph、异步调度、prefix caching,并沿用 GLM-5.2 ATOM MI355X agentic 配方的 AITER 设置。decode 镜像、1M 上下文、bf16 KV 与按层分片的 PD 缓冲不变。TileRT 报告并发 1 下 AgentX TTFT p90 由 4.35 秒降至 1.60 秒(本地 2x8 MI350X 对比 post2 的 MI355X 官方轮次)。" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3563