Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,10 @@ on:
- 'inferencex-e2e/infx/ruff.toml'
- '**/pytest.ini'
- 'inferencex-e2e/utils/srt-slurm'
# Tests check recipe containers against master images.
- 'inferencex-e2e/configs/*-master.yaml'
- 'inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/**'
- 'inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/**'
push:
branches: [main]
paths: *python-paths
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ base:
name: dsv4-fp4-b200-sglang-agentic
model:
path: hf:deepseek-ai/DeepSeek-V4-Pro-0813
container: lmsysorg/sglang:v0.5.19-cu130
container: lmsysorg/sglang:v0.5.20-cu130
precision: fp4
resources:
gpu_type: b200
Expand Down Expand Up @@ -85,7 +85,7 @@ override_tp8_c1:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 2
cuda-graph-max-bs: 2
cuda-graph-max-bs-decode: 2
benchmark:
env:
CONC: '1'
Expand All @@ -100,7 +100,7 @@ override_tp8_c2:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 4
cuda-graph-max-bs: 4
cuda-graph-max-bs-decode: 4
benchmark:
env:
CONC: '2'
Expand All @@ -115,7 +115,7 @@ override_tp8_c3:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 6
cuda-graph-max-bs: 6
cuda-graph-max-bs-decode: 6
benchmark:
env:
CONC: '3'
Expand All @@ -130,7 +130,7 @@ override_tp8_c4:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 8
cuda-graph-max-bs: 8
cuda-graph-max-bs-decode: 8
benchmark:
env:
CONC: '4'
Expand All @@ -145,7 +145,7 @@ override_tp8_c5:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 10
cuda-graph-max-bs: 10
cuda-graph-max-bs-decode: 10
benchmark:
env:
CONC: '5'
Expand All @@ -160,7 +160,7 @@ override_tp8_hicache_c8:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 16
cuda-graph-max-bs: 16
cuda-graph-max-bs-decode: 16
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand All @@ -181,7 +181,7 @@ override_tp8_hicache_c10:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 20
cuda-graph-max-bs: 20
cuda-graph-max-bs-decode: 20
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand All @@ -202,7 +202,7 @@ override_tp8_hicache_c16:
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
max-running-requests: 32
cuda-graph-max-bs: 32
cuda-graph-max-bs-decode: 32
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand Down Expand Up @@ -246,7 +246,7 @@ override_dep8_hicache_c64:
swa-full-tokens-ratio: 0.02
chunked-prefill-size: 49152
max-running-requests: 128
cuda-graph-max-bs: 32
cuda-graph-max-bs-decode: 32
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand Down Expand Up @@ -294,7 +294,7 @@ override_dep8_hicache_c96:
swa-full-tokens-ratio: 0.02
chunked-prefill-size: 49152
max-running-requests: 192
cuda-graph-max-bs: 32
cuda-graph-max-bs-decode: 32
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand Down Expand Up @@ -342,7 +342,7 @@ override_dep8_hicache_c128:
swa-full-tokens-ratio: 0.02
chunked-prefill-size: 49152
max-running-requests: 256
cuda-graph-max-bs: 32
cuda-graph-max-bs-decode: 32
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand Down Expand Up @@ -391,7 +391,7 @@ override_dep8_hicache_c160:
swa-full-tokens-ratio: 0.02
chunked-prefill-size: 49152
max-running-requests: 320
cuda-graph-max-bs: 32
cuda-graph-max-bs-decode: 32
enable-hierarchical-cache: true
hicache-write-policy: write_through
hicache-io-backend: direct
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ base:
name: dsv41flash-fp4-mi355x-vllm-agentic
model:
path: hf:deepseek-ai/DeepSeek-V4.1-Flash
container: vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423
container: vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098
precision: fp4
resources:
gpu_type: mi355x
Expand Down
94 changes: 94 additions & 0 deletions inferencex-e2e/infx/tests/srt_slurm/test_recipe_images.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,94 @@
"""Keep checked-in srt-slurm recipe images identical to their master-config images."""

import sys
from collections import defaultdict
from pathlib import Path

import yaml

from infx.srt_slurm.synthetic_acceptance import selected_recipes

ROOT = Path(__file__).resolve().parents[3]
sys.path.insert(0, str(ROOT / "utils/srt-slurm/src"))
MULTI_NODE_RECIPES = "benchmarks/multi_node/srt-slurm-recipes/"
# Their recipes still name a stale image on main; remove a key once its recipe is aligned.
KNOWN_STALE_KEYS = {
"glm5.2-fp8-mi325x-sglang-agentic-mtp",
"minimaxm3-fp8-mi300x-vllm-agentic-mtp",
}


def recipe_references(node):
"""Yield (multinode, reference) for every srt-slurm recipe a master entry names."""
if isinstance(node, dict):
for key, value in node.items():
if key == "srt-recipe":
yield False, value
else:
yield from recipe_references(value)
elif isinstance(node, list):
for item in node:
yield from recipe_references(item)
elif isinstance(node, str) and node.startswith("CONFIG_FILE="):
path = node.removeprefix("CONFIG_FILE=")
# Launchers stage multi-node recipes under recipes/ in the srt-slurm checkout.
for staged in ("recipes/", MULTI_NODE_RECIPES):
if path.startswith(staged):
yield True, MULTI_NODE_RECIPES + path.removeprefix(staged)


def master_references():
"""Yield (config key, image, multinode, reference) from every master config."""
for master in sorted((ROOT / "configs").glob("*-master.yaml")):
for key, entry in yaml.safe_load(master.read_text()).items():
for multinode, reference in sorted(set(recipe_references(entry))):
yield key, entry["image"], multinode, reference


def variants(reference):
path, _, selector = reference.partition(":")
recipe = yaml.safe_load((ROOT / path).read_text())
return path, selected_recipes(recipe, selector or None)


def canonical_image(image):
# Enroot spells the nvcr.io registry separator as "#"; launchers key aliases either way.
return image.replace("nvcr.io#", "nvcr.io/", 1)


def test_single_node_recipes_use_their_master_image():
references = [
(key, image, reference)
for key, image, multinode, reference in master_references()
if not multinode and key not in KNOWN_STALE_KEYS
]
images = defaultdict(set)
for _, image, reference in references:
images[reference.partition(":")[0]].add(image)
problems = set()
for key, image, reference in references:
path, selected = variants(reference)
containers = {recipe["model"]["container"] for _, recipe in selected}
if image not in containers:
problems.add(f"{key}: no variant of {reference} uses master image {image}")
for container in containers - images[path]:
problems.add(f"{path}: {container} is not the image of any master key using it")
assert not problems, "\n".join(sorted(problems))


def test_multi_node_recipes_use_their_master_image():
problems = set()
for key, image, multinode, reference in master_references():
if not multinode:
continue
path, selected = variants(reference)
for name, recipe in selected:
label = f"{key}: {path}" + (f":{name}" if name else "")
container = recipe["model"]["container"]
# srtctl pulls a literal missing from the alias map, so a stale one runs silently.
if ":" in container and canonical_image(container) != canonical_image(image):
problems.add(f"{label} model.container {container} != master image {image}")
identity = ((recipe.get("identity") or {}).get("container") or {}).get("image")
if identity is not None and canonical_image(identity) != canonical_image(image):
problems.add(f"{label} identity.container.image {identity} != master image {image}")
assert not problems, "\n".join(sorted(problems))
14 changes: 14 additions & 0 deletions inferencex-e2e/perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9035,3 +9035,17 @@
- "Enable breakable CUDA-graph prefill (BCG) at conc <= 4."
- "Enable EP8 + MegaMoE for DP-attention arms at conc >= 128 (including conc 384)."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3430

- config-keys:
- dsv41flash-fp4-mi355x-vllm-agentic-dspark
- dsv4-fp4-b200-sglang-agentic-hicache-mtp
scenario-type:
- agentic-coding
description:
- "Align the single-node srt-slurm recipe model.container with the unchanged master image for two AgentX keys whose recipes #3428 ported from the legacy scripts as they were before the image bumps in #3420 and #3334; every point of these keys would fail before submission with 'Single-node SRT image: recipe/matrix'."
- "Recipe containers: dsv41flash-fp4-mi355x-vllm-agentic-dspark vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423 -> vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098; dsv4-fp4-b200-sglang-agentic-hicache-mtp lmsysorg/sglang:v0.5.19-cu130 -> lmsysorg/sglang:v0.5.20-cu130."
- "The B200 SGLang v0.5.20 recipe also renames cuda-graph-max-bs to cuda-graph-max-bs-decode, as #3334 did in the legacy script, because v0.5.20 removed the deprecated alias (sgl-project/sglang#38375). All other serving flags and sweep points are unchanged."
- "将两个 AgentX key 的单节点 srt-slurm 配方 model.container 与未改动的主配置镜像对齐。这些配方由 #3428 从旧脚本移植而来,而移植所依据的是 #3420 和 #3334 升级镜像之前的版本;此前这些 key 的每个点都会在提交前以 'Single-node SRT image: recipe/matrix' 失败。"
- "配方镜像:dsv41flash-fp4-mi355x-vllm-agentic-dspark vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423 -> vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098;dsv4-fp4-b200-sglang-agentic-hicache-mtp lmsysorg/sglang:v0.5.19-cu130 -> lmsysorg/sglang:v0.5.20-cu130。"
- "B200 的 SGLang v0.5.20 配方同时将 cuda-graph-max-bs 改为 cuda-graph-max-bs-decode,与 #3334 对旧脚本的修改一致,因为 v0.5.20 已移除该弃用别名(sgl-project/sglang#38375)。其余服务参数和 sweep 点均保持不变。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3567
Loading