fix(srt): align B200 and MI355X recipe images with master configs and add a static check / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置并新增静态检查 - #3567
chunfangamd wants to merge 7 commits into
Conversation
#3428 ported these AgentX recipes from the legacy scripts as they were before the 09-25 image bumps (#3361, #3362, #3334, #3420), which changed only the master images. Every point of these keys now fails before submission with 'Single-node SRT image: recipe/matrix'. The SGLang v0.5.20 recipes also take the --cuda-graph-max-bs-decode rename that #3362 and #3334 applied to the legacy scripts; v0.5.20 no longer accepts the deprecated --cuda-graph-max-bs alias (sgl-project/sglang#38375). Co-authored-by: Cursor <[email protected]>
single_node.py rejects a point whose recipe container differs from the matrix image only once a GPU job starts, and the multi-node path has no such check: srtctl pulls a literal container missing from the alias map. Check both statically, including identity.container.image. Co-authored-by: Cursor <[email protected]>
Co-authored-by: Cursor <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
Co-authored-by: Cursor <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36521758562 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36521758562 |
There was a problem hiding this comment.
Looks good, straightforward config alignment change.
What was reviewed: verified each of the four bumped model.container values against the corresponding (unchanged) master-config image in amd-master.yaml/nvidia-master.yaml — all match exactly; checked that every cuda-graph-max-bs occurrence in the two touched v0.5.20 sglang recipes was renamed to cuda-graph-max-bs-decode, and confirmed no other recipe in the repo still pairs the v0.5.20 sglang image with the old flag name; confirmed the new test drives real recipe expansion (generate_override_configs via selected_recipes) rather than pinning static config, and that the perf-changelog diff only appends a new entry at the file's tail.
Extended reasoning...
The diff is confined to four recipe YAML image/flag edits, a new pure-Python static test, and an append-only changelog entry — no auth, crypto, or data-exposure surface. I cross-checked all four container values against the master configs directly and confirmed they match, verified the flag rename is complete and scoped correctly, and confirmed the new test exercises real config-selection logic rather than pinning literals; the perf-changelog append preserves history. None of the touched paths fall under a CODEOWNERS-restricted pattern in .github/CODEOWNERS. The pull/XXX placeholder in the changelog PR-link is a known pre-merge fill-in matching existing repo convention, not a functional defect.
This review covers commit 75c7220, which is no longer the latest commit on this pull request; later commits are not covered by it.
Restore the MI325X GLM-5.2 and MI300X MiniMax-M3 recipes and drop the static image test from this PR; they will follow separately. The changelog entry now selects only the B200 and MI355X keys. Co-authored-by: Cursor <[email protected]>
…sistency Co-authored-by: Cursor <[email protected]> # Conflicts: # inferencex-e2e/perf-changelog.yaml
Restore the static check with the MI325X GLM-5.2 and MI300X MiniMax-M3 keys exempt until their recipes are aligned, and run CI Tests when master configs or srt-slurm recipes change so YAML-only PRs are checked before any GPU job. Co-authored-by: Cursor <[email protected]>
Description
Two single-node AgentX srt-slurm recipes name a container that differs from their master-config
image.infx/srt_slurm/single_node.pytherefore rejects every point of these keys before Slurm submission (Single-node SRT image: recipe/matrix ...), the same failure #3446 hit in its first sweep.Cause. #3428 ported these recipes from the legacy scripts as they were before the image bumps in #3420 and #3334, which merged about ten hours earlier. Those bumps changed only the master images, which was correct while the keys still ran the legacy scripts. The PRs touched different files, so git saw no conflict, and no check compares the two copies before a GPU job starts.
image(unchanged)model.containeronmaindsv41flash-fp4-mi355x-vllm-agentic-dsparkvllm/vllm-openai-rocm:nightly-rocm100-29468dde…vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398…dsv4-fp4-b200-sglang-agentic-hicache-mtp(NVIDIA)lmsysorg/sglang:v0.5.20-cu130lmsysorg/sglang:v0.5.19-cu130Changes
model.containerto the master image. The B200 SGLang v0.5.20 recipe also renamescuda-graph-max-bstocuda-graph-max-bs-decodein all 12 variants, as [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 did in the legacy script. SGLang v0.5.20 no longer accepts the deprecated alias ([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375), so an image-only fix would fail at server startup. No other serving flag or sweep point changes.inferencex-e2e/infx/tests/srt_slurm/test_recipe_images.pyruns in the Tests job in about 4 s without GPUs.single_node.py.model.containerand anyidentity.container.imagemust equal the master image, treatingnvcr.io#asnvcr.io/. Aliases such asdynamo-sglangresolve through the cluster profile and are skipped. This path has no runtime check: srtctl pulls a literal container missing from the alias map, so drift there would run and mislabel results silently.glm5.2-fp8-mi325x-sglang-agentic-mtpandminimaxm3-fp8-mi300x-vllm-agentic-mtpstill drift onmainand are listed inKNOWN_STALE_KEYSuntil their recipes are aligned.ci.yml: CI Tests previously ran only for Python and tooling changes, so a YAML-only config or recipe PR such as [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 never ran them. The trigger paths now includeconfigs/*-master.yamland both srt-slurm recipe trees, so such PRs are checked within minutes and before any GPU job.perf-changelog.yaml: one entry for the two keys so the sweep re-validates them on the native srt-slurm path.Validation (local)
glm5.2-fp4-mi355xrecipe back to the 20260923 image). Multi-node literal andidentity.container.imagemismatches are also caught.infx.workflows.validate_perf_changelogpasses.infx.matrix.planselects 28 throughput jobs (16 onmi355x-amds, 12 onb200-nscale) and 2 AgentX eval jobs.select_recipe,runtime_arguments,plan_commands) was emulated for all 30 planned jobs with 0 failures. A negative control reproduces the runtime error text.Notes for reviewers
minimaxm3-fp8-mi300x-vllm-agentic-mtpfromKNOWN_STALE_KEYS.mi325x-amdsrunners atuv python install(failed to create directory /home/gharunner-mi325x/.cache/uv: File exists). The same step passed on all elevenb200-nscalerunners, so this is a runner-environment issue, not a recipe issue.reuse-sweep-gateruns only onsynchronize, andcanary-selecthas noalways(), so it inherits the skip.dsv41flashcontainer line (currently a TBD image), so whichever merges second resolves a one-line conflict.dsv4-fp4-b200-sglang-agentic-hicache-mtpis an NVIDIA recipe and needs an NVIDIA CODEOWNER review. Theci.ymlchange needs a workflow owner's review.AI model disclosure
Related Issue
No issue. Related: #3428, #3446, #3545, #3555.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
改动说明
两个单节点 AgentX srt-slurm 配方的 container 与主配置
image不一致,导致infx/srt_slurm/single_node.py在提交 Slurm 之前拒绝这些 key 的每一个点(Single-node SRT image: recipe/matrix ...),与 #3446 第一次 sweep 遇到的失败相同。原因: #3428 移植这些配方时,依据的是 #3420 和 #3334 升级镜像之前的旧脚本,而这两个升级约在 #3428 合入前十小时已经合入。升级 PR 只改了主配置镜像,这在这些 key 仍运行旧脚本时是正确的。两边改的是不同文件,git 没有冲突,而在 GPU 任务开始之前也没有任何检查比较两份拷贝。上表列出了两个 key 的主配置镜像(未改动)和
main上配方的旧镜像。改动:
model.container改为主配置镜像。B200 的 SGLang v0.5.20 配方同时在全部 12 个 variant 中将cuda-graph-max-bs改为cuda-graph-max-bs-decode,与 [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 对旧脚本的修改一致。SGLang v0.5.20 已移除该弃用别名([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375),只改镜像会导致服务启动失败。其余服务参数和 sweep 点均不变。inferencex-e2e/infx/tests/srt_slurm/test_recipe_images.py,在 Tests 任务中运行,约 4 秒,不需要 GPU。single_node.py一样按字符串精确比较。model.container以及identity.container.image必须与主配置镜像一致(nvcr.io#视同nvcr.io/);dynamo-sglang等别名经 cluster profile 解析,不做检查。多节点路径没有运行时检查:srtctl 会直接拉取别名表中不存在的完整镜像名,漂移会静默运行并记错结果标签。glm5.2-fp8-mi325x-sglang-agentic-mtp和minimaxm3-fp8-mi300x-vllm-agentic-mtp在main上仍不一致,暂列在KNOWN_STALE_KEYS中,配方对齐后删除。ci.yml: CI Tests 之前只在改动 Python 和工具配置时运行,像 [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 这样只改 YAML 的配置或配方 PR 从未运行过 Tests。现在触发路径加入了configs/*-master.yaml和两个 srt-slurm 配方目录,这类 PR 会在几分钟内、任何 GPU 任务之前完成检查。perf-changelog.yaml: 为这两个 key 追加一条记录,让 sweep 在新的 srt-slurm 路径上重新验证。本地验证:
glm5.2-fp4-mi355x配方改回 20260923 镜像)时都会失败;多节点完整镜像名不一致和identity.container.image不一致也能被发现。infx.workflows.validate_perf_changelog通过。infx.matrix.plan选出 28 个吞吐任务(mi355x-amds16 个、b200-nscale12 个)和 2 个 AgentX eval 任务。select_recipe、runtime_arguments、plan_commands),失败 0 个;反向对照能复现运行时的报错文本。审阅注意事项:
KNOWN_STALE_KEYS删除minimaxm3-fp8-mi300x-vllm-agentic-mtp。mi325x-amdsrunner 上于uv python install失败(failed to create directory /home/gharunner-mi325x/.cache/uv: File exists);同一步在全部 11 台b200-nscalerunner 上成功,属于 runner 环境问题,与配方无关。reuse-sweep-gate只在synchronize时运行,而canary-select的条件没有always(),会随之被跳过。dsv41flashcontainer(目前是 TBD 镜像),后合入的一方需要解决一行冲突。dsv4-fp4-b200-sglang-agentic-hicache-mtp是 NVIDIA 的配方,需要 NVIDIA CODEOWNER 审阅;ci.yml的改动需要 workflow 负责人审阅。AI 模型使用说明
关联 issue
无。相关 PR:#3428、#3446、#3545、#3555。
改动类型
Bug 修复、配置修改,以及新增配方与主配置镜像一致性的静态测试和 CI 触发路径修改。