Skip to content

fix(srt): align B200 and MI355X recipe images with master configs and add a static check / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置并新增静态检查 - #3567

Open
chunfangamd wants to merge 7 commits into
mainfrom
chun/recipe-image-consistency
Open

chunfangamd wants to merge 7 commits into
mainfrom
chun/recipe-image-consistency

Conversation

@chunfangamd

@chunfangamd chunfangamd commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Two single-node AgentX srt-slurm recipes name a container that differs from their master-config image. infx/srt_slurm/single_node.py therefore rejects every point of these keys before Slurm submission (Single-node SRT image: recipe/matrix ...), the same failure #3446 hit in its first sweep.

Cause. #3428 ported these recipes from the legacy scripts as they were before the image bumps in #3420 and #3334, which merged about ten hours earlier. Those bumps changed only the master images, which was correct while the keys still ran the legacy scripts. The PRs touched different files, so git saw no conflict, and no check compares the two copies before a GPU job starts.

Config key Master image (unchanged) Recipe model.container on main
dsv41flash-fp4-mi355x-vllm-agentic-dspark vllm/vllm-openai-rocm:nightly-rocm100-29468dde… vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398…
dsv4-fp4-b200-sglang-agentic-hicache-mtp (NVIDIA) lmsysorg/sglang:v0.5.20-cu130 lmsysorg/sglang:v0.5.19-cu130

Changes

  1. Recipes: set model.container to the master image. The B200 SGLang v0.5.20 recipe also renames cuda-graph-max-bs to cuda-graph-max-bs-decode in all 12 variants, as [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 did in the legacy script. SGLang v0.5.20 no longer accepts the deprecated alias ([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375), so an image-only fix would fail at server startup. No other serving flag or sweep point changes.
  2. Static check: new inferencex-e2e/infx/tests/srt_slurm/test_recipe_images.py runs in the Tests job in about 4 s without GPUs.
    • Single-node: each master key's image must appear among the recipe variants it selects, and every variant container must be the image of a master key that uses the recipe. The comparison is exact, as in single_node.py.
    • Multi-node: a literal model.container and any identity.container.image must equal the master image, treating nvcr.io# as nvcr.io/. Aliases such as dynamo-sglang resolve through the cluster profile and are skipped. This path has no runtime check: srtctl pulls a literal container missing from the alias map, so drift there would run and mislabel results silently.
    • glm5.2-fp8-mi325x-sglang-agentic-mtp and minimaxm3-fp8-mi300x-vllm-agentic-mtp still drift on main and are listed in KNOWN_STALE_KEYS until their recipes are aligned.
  3. ci.yml: CI Tests previously ran only for Python and tooling changes, so a YAML-only config or recipe PR such as [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 never ran them. The trigger paths now include configs/*-master.yaml and both srt-slurm recipe trees, so such PRs are checked within minutes and before any GPU job.
  4. perf-changelog.yaml: one entry for the two keys so the sweep re-validates them on the native srt-slurm path.

Validation (local)

  • The new test passes on this branch. It fails when the B200 or MI355X fix is reverted, and when the [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 case is reproduced (glm5.2-fp4-mi355x recipe back to the 20260923 image). Multi-node literal and identity.container.image mismatches are also caught.
  • The full CI Tests command with the locked Python 3.12 environment: 2152 passed, 1 skipped.
  • infx.workflows.validate_perf_changelog passes. infx.matrix.plan selects 28 throughput jobs (16 on mi355x-amds, 12 on b200-nscale) and 2 AgentX eval jobs.
  • The runner's pre-submission binding (select_recipe, runtime_arguments, plan_commands) was emulated for all 30 planned jobs with 0 failures. A negative control reproduces the runtime error text.
  • GPU validation is pending this PR's sweep.

Notes for reviewers

AI model disclosure

  • Model/version: Claude Opus 5.5 (the identifier exposed by the Cursor agent runtime). No delegated agents were used.
  • Role: investigated the recipe/master image drift, wrote the recipe fixes, the static test, the CI trigger change and the changelog entry, ran the local validation, diagnosed the first sweep failure, and drafted this description. @chunfangamd reviewed the changes and chose the scope of this PR.

Related Issue

No issue. Related: #3428, #3446, #3545, #3555.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe): static recipe/master image consistency test and CI trigger paths

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.
中文

改动说明

两个单节点 AgentX srt-slurm 配方的 container 与主配置 image 不一致,导致 infx/srt_slurm/single_node.py 在提交 Slurm 之前拒绝这些 key 的每一个点(Single-node SRT image: recipe/matrix ...),与 #3446 第一次 sweep 遇到的失败相同。

原因: #3428 移植这些配方时,依据的是 #3420 和 #3334 升级镜像之前的旧脚本,而这两个升级约在 #3428 合入前十小时已经合入。升级 PR 只改了主配置镜像,这在这些 key 仍运行旧脚本时是正确的。两边改的是不同文件,git 没有冲突,而在 GPU 任务开始之前也没有任何检查比较两份拷贝。上表列出了两个 key 的主配置镜像(未改动)和 main 上配方的旧镜像。

改动:

  1. 配方: 将 model.container 改为主配置镜像。B200 的 SGLang v0.5.20 配方同时在全部 12 个 variant 中将 cuda-graph-max-bs 改为 cuda-graph-max-bs-decode,与 [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 对旧脚本的修改一致。SGLang v0.5.20 已移除该弃用别名([Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement sgl-project/sglang#38375),只改镜像会导致服务启动失败。其余服务参数和 sweep 点均不变。
  2. 静态检查: 新增 inferencex-e2e/infx/tests/srt_slurm/test_recipe_images.py,在 Tests 任务中运行,约 4 秒,不需要 GPU。
    • 单节点:每个主配置 key 的镜像必须出现在它所选的配方 variant 中,且每个 variant 的 container 必须是某个引用该配方的 key 的镜像;与 single_node.py 一样按字符串精确比较。
    • 多节点:配方中写完整镜像名的 model.container 以及 identity.container.image 必须与主配置镜像一致(nvcr.io# 视同 nvcr.io/);dynamo-sglang 等别名经 cluster profile 解析,不做检查。多节点路径没有运行时检查:srtctl 会直接拉取别名表中不存在的完整镜像名,漂移会静默运行并记错结果标签。
    • glm5.2-fp8-mi325x-sglang-agentic-mtp 和 minimaxm3-fp8-mi300x-vllm-agentic-mtp 在 main 上仍不一致,暂列在 KNOWN_STALE_KEYS 中,配方对齐后删除。
  3. ci.yml: CI Tests 之前只在改动 Python 和工具配置时运行,像 [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 这样只改 YAML 的配置或配方 PR 从未运行过 Tests。现在触发路径加入了 configs/*-master.yaml 和两个 srt-slurm 配方目录,这类 PR 会在几分钟内、任何 GPU 任务之前完成检查。
  4. perf-changelog.yaml: 为这两个 key 追加一条记录,让 sweep 在新的 srt-slurm 路径上重新验证。

本地验证:

  • 新测试在本 branch 上通过;撤销 B200 或 MI355X 的修复、或复现 [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 的情况(把 glm5.2-fp4-mi355x 配方改回 20260923 镜像)时都会失败;多节点完整镜像名不一致和 identity.container.image 不一致也能被发现。
  • 使用锁定依赖的 Python 3.12 环境运行 CI 完整 Tests 命令:2152 个通过,1 个跳过。
  • infx.workflows.validate_perf_changelog 通过。infx.matrix.plan 选出 28 个吞吐任务(mi355x-amds 16 个、b200-nscale 12 个)和 2 个 AgentX eval 任务。
  • 对全部 30 个计划任务模拟了 runner 提交前的绑定步骤(select_recipe、runtime_arguments、plan_commands),失败 0 个;反向对照能复现运行时的报错文本。
  • GPU 验证待本 PR 的 sweep 完成。

审阅注意事项:

AI 模型使用说明

  • 模型/版本:Claude Opus 5.5(Cursor agent 运行环境提供的标识),未使用其他委派 agent。
  • 工作内容:排查配方与主配置镜像不一致,编写配方修复、静态测试、CI 触发路径修改和 changelog 记录,运行本地验证,诊断第一次 sweep 的失败,并起草本说明;@chunfangamd 审阅了全部改动并确定了本 PR 的范围。

关联 issue

无。相关 PR:#3428、#3446、#3545、#3555。

改动类型

Bug 修复、配置修改,以及新增配方与主配置镜像一致性的静态测试和 CI 触发路径修改。

chunfangamd and others added 3 commits September 29, 2026 01:26
#3428 ported these AgentX recipes from the legacy scripts as they were before the 09-25 image bumps (#3361, #3362, #3334, #3420), which changed only the master images. Every point of these keys now fails before submission with 'Single-node SRT image: recipe/matrix'.

The SGLang v0.5.20 recipes also take the --cuda-graph-max-bs-decode rename that #3362 and #3334 applied to the legacy scripts; v0.5.20 no longer accepts the deprecated --cuda-graph-max-bs alias (sgl-project/sglang#38375).

Co-authored-by: Cursor <[email protected]>
single_node.py rejects a point whose recipe container differs from the matrix image only once a GPU job starts, and the multi-node path has no such check: srtctl pulls a literal container missing from the alias map. Check both statically, including identity.container.image.

Co-authored-by: Cursor <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward config alignment change.
What was reviewed: verified each of the four bumped model.container values against the corresponding (unchanged) master-config image in amd-master.yaml/nvidia-master.yaml — all match exactly; checked that every cuda-graph-max-bs occurrence in the two touched v0.5.20 sglang recipes was renamed to cuda-graph-max-bs-decode, and confirmed no other recipe in the repo still pairs the v0.5.20 sglang image with the old flag name; confirmed the new test drives real recipe expansion (generate_override_configs via selected_recipes) rather than pinning static config, and that the perf-changelog diff only appends a new entry at the file's tail.

Extended reasoning...

The diff is confined to four recipe YAML image/flag edits, a new pure-Python static test, and an append-only changelog entry — no auth, crypto, or data-exposure surface. I cross-checked all four container values against the master configs directly and confirmed they match, verified the flag rename is complete and scoped correctly, and confirmed the new test exercises real config-selection logic rather than pinning literals; the perf-changelog append preserves history. None of the touched paths fall under a CODEOWNERS-restricted pattern in .github/CODEOWNERS. The pull/XXX placeholder in the changelog PR-link is a known pre-merge fill-in matching existing repo convention, not a functional defect.

This review covers commit 75c7220, which is no longer the latest commit on this pull request; later commits are not covered by it.

Restore the MI325X GLM-5.2 and MI300X MiniMax-M3 recipes and drop the static image test from this PR; they will follow separately. The changelog entry now selects only the B200 and MI355X keys.

Co-authored-by: Cursor <[email protected]>
@chunfangamd chunfangamd changed the title fix(srt): align four single-node recipe images with master configs and add a static check / fix(srt):对齐四个单节点配方镜像与主配置并新增静态检查 fix(srt): align B200 and MI355X single-node recipe images with master configs / fix(srt):对齐 B200 与 MI355X 单节点配方镜像与主配置 Sep 29, 2026
chunfangamd and others added 2 commits September 29, 2026 03:41
…sistency

Co-authored-by: Cursor <[email protected]>

# Conflicts:
#	inferencex-e2e/perf-changelog.yaml
Restore the static check with the MI325X GLM-5.2 and MI300X MiniMax-M3 keys exempt until their recipes are aligned, and run CI Tests when master configs or srt-slurm recipes change so YAML-only PRs are checked before any GPU job.

Co-authored-by: Cursor <[email protected]>
@chunfangamd chunfangamd changed the title fix(srt): align B200 and MI355X single-node recipe images with master configs / fix(srt):对齐 B200 与 MI355X 单节点配方镜像与主配置 fix(srt): align B200 and MI355X recipe images with master configs and add a static check / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置并新增静态检查 Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant