feat(srt): check srt-slurm recipes against master images before sweep dispatch / feat(srt):在派发 sweep 前检查 srt-slurm 配方与主配置镜像是否一致 - #3624
chunfangamd wants to merge 1 commit into
Conversation
…atch run-sweep, e2e-tests and profile now pipe their matrix through infx.srt_slurm.preflight after benchmark_schema. Each single-node point must select exactly one variant through the runtime's select_recipe, and every CONFIG_FILE and EVAL_CONFIG_FILE recipe a multi-node point can launch must name a container the job's srt-slurm config resolves to the point's image. A mismatch now fails setup before any canary or benchmark job is dispatched. Co-authored-by: Cursor <[email protected]>
| if [ -f "$PRIORITY_ROOT/infx/srt_slurm/preflight.py" ] && [ -f "$MEASURED_ROOT/infx/srt_slurm/preflight.py" ]; then | ||
| git -C "$MEASURED_ROOT" submodule update --init utils/srt-slurm | ||
| CONFIG_JSON=$(printf '%s' "$CONFIG_JSON" | env PYTHONPATH="$PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src" \ | ||
| uv run --no-project --exclude-newer PT12H --python 3.12 --with pydantic --with pyyaml \ | ||
| --with marshmallow --with marshmallow-dataclass --with ruamel.yaml --with requests \ | ||
| python -P -m infx.srt_slurm.preflight --root "$MEASURED_ROOT") |
There was a problem hiding this comment.
🔴 This runs the "trusted" preflight check with srtctl imported from the PR's own utils/srt-slurm submodule checkout, letting a malicious PR execute code inside the trusted-tooling process instead of only the sandboxed measured tree. check_multi_node -> selected_recipes does from srtctl.core.config import generate_override_configs, and PYTHONPATH is $PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src (e2e-tests.yml:336) while only $MEASURED_ROOT ever gets git submodule update --init utils/srt-slurm (line 335) — $PRIORITY_ROOT never does, so srtctl can only resolve from the PR-controlled submodule commit. -P doesn't block PYTHONPATH imports. Fix: give PRIORITY_ROOT its own trusted srt-slurm checkout (or vendor/pin the srtctl it imports) so the trusted interpreter never imports a package whose source lives only in the untrusted MEASURED_ROOT tree; same pattern in profile.yml:111-116.
Why this was flagged
The whole point of checking out .ci-priority/PRIORITY_ROOT separately from MEASURED_ROOT in this job is to run decision logic with code the PR cannot modify. This diff defeats that: infx.srt_slurm.preflight (loaded from PRIORITY_ROOT) calls infx.srt_slurm.synthetic_acceptance.selected_recipes, which lazily imports srtctl.core.config.generate_override_configs; that package is only ever present via $MEASURED_ROOT/utils/srt-slurm/src (e2e-tests.yml:335-336), which is the PR's own branch/submodule pointer. A PR that points its utils/srt-slurm submodule gitlink at a malicious commit gets that code executed inside the trusted PRIORITY_ROOT python process when get-jobs runs (same for profile.yml:111-116). python -P only disables automatic sys.path prepending, not PYTHONPATH, so this import is not blocked. The output of this step feeds CONFIG_JSON/job outputs that downstream benchmark-dispatch jobs (which do use secrets.INFERENCEX_OFFICIAL_RO_HF_TOKEN, secrets.MODAL_TOKEN_ID/SECRET) consume, so compromising this step can also poison what those privileged jobs run.
Verification: e2e-tests.yml:335 runs git -C "$MEASURED_ROOT" submodule update --init utils/srt-slurm (PR-controlled submodule), and :336 sets PYTHONPATH="$PRIORITY_ROOT:$MEASURED_ROOT/utils/srt-slurm/src", so srtctl resolves only from the measured tree.
Description
Follow-up to #3567; the two PRs belong together. #3567 aligns the two srt-slurm recipes whose container had drifted from the master-config
image. This PR adds the check that would have caught the drift: each sweep now binds every planned srt-slurm point to the recipe variant it would launch, and stops before dispatching any job if a point doesn't bind.Problem. An srt-slurm point names its container in two places: the master config's
imageand the recipe'smodel.container. An image bump has to edit both, as #3446 did, but nothing compares them before benchmark jobs start.infx/srt_slurm/single_node.pyrejects a mismatch (Single-node SRT image: recipe/matrix ...) only after the benchmark job has started on its cluster runner, prepared an srt-slurm checkout and installed srtctl. [AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446's first sweep failed this way. So would every point of the two keys fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 fixes, because feat(srt): run AgentX on native srt-slurm #3428 ported their recipes from the legacy scripts as they were before the image bumps in config(dsv41flash): re-sweep MI355X on the first ROCm 10.0 nightly carrying vllm#58510 / 在首个包含 vllm#58510 的 ROCm 10.0 nightly 上重新扫描 MI355X #3420 and [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334.#3567 first added a pytest scan for this. Review found four problems with it:
run-sweep, which is a separate workflow.EVAL_CONFIG_FILE.AGENTS.mdtest rules.#3567 dropped the scan, and this PR replaces it.
Changes
infx/srt_slurm/preflight.py(python -m infx.srt_slurm.preflight) reads the matrix on stdin. If every srt-slurm point binds, it writes the matrix unchanged to stdout; otherwise it prints each problem once with the points it affects and exits 1.benchmark-tmpl.ymlexports for the point and calls the runtime's ownselect_recipe. Exactly one variant must match on engine, model, image, precision, parallelism, GPU count, speculative decoding, concurrency and KV offloading. The check runs per point, so a recipe shared by keys with different images is checked against each key's image.CONFIG_FILEandEVAL_CONFIG_FILEwith srtctl's own override expansion. Each variant'smodel.containermust be the point's image (in either registry spelling) or a container alias thatconfigs/runners.yamldefines for the runner's cluster. A declaredidentity.container.imagemust name the same image. TileRT points are skipped because they pair their own containers.run-sweep.yml:setupinitializes the srt-slurm submodule and runs the validator afterbenchmark_schema --plan, beforeci_priority. If the validator fails,setupfails, socanary-selectand the benchmark jobs, which all require a successfulsetup, never start.e2e-tests.yml(whichtrusted-external-sweep.ymlandclaude.ymlalso dispatch) andprofile.yml: the trusted tooling's validator checks the measured tree's recipes, runner config and srt-slurm submodule. The step is skipped when either tree predates the module, so measuring an older revision is unaffected.infx/tests/srt_slurm/test_preflight.pyuses small temporary recipes and runner inventories, per theAGENTS.mdtest rules. It covers:EVAL_CONFIG_FILEmismatch;A sweep that selects the B200 key without #3567's fix stops at
setupwith:Validation
test_preflight.py: 14 passed. Each of eight deliberate regressions in the validator fails at least one test:EVAL_CONFIG_FILE, single-node points, or the alias lookup for bare cluster ids;identity.container.image, or rejecting the pyxis spelling;ci.ymlcommands, run locally: Lint is clean, and Tests has 2,349 passed. Twotest_slurm_clicases fail only on the test host because it has a realsacct; they fail the same way on cleanmain.dsv4-fp4-b200-sglang-agentic-hicache-mtp(12 points) anddsv41flash-fp4-mi355x-vllm-agentic-dspark(16 points) are fixed by fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567.glm5.2-fp8-mi325x-sglang-agentic-mtp(7 points) still drifts onmain.run-sweepsetupreplayed in a throwaway worktree with the workflow's exact commands underbash -e:ci_priority.Notes for reviewers
setup; it would fail at runtime anyway.glm5.2-fp8-mi325x-sglang-agentic-mtpstill drifts onmain, so a sweep that selects it fails atsetupuntil its recipe is fixed. Only the points a sweep plans are checked, so drift in other keys doesn't block unrelated PRs.single_node_environmentmirrors howbenchmark-tmpl.ymlexports matrix fields to the job. A change to that mapping needs the same change here.run-sweeponly triggers on PRs that editperf-changelog.yaml, so this PR's own checks don't run the newsetupstep. The replay above stands in for that.perf-changelog.yamlentry, because no benchmark config or recipe changes.@SemiAnalysisAI/core.AI model disclosure
Related Issue
No issue. Follow-up to #3567. Related: #3428, #3446.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
改动说明
本 PR 是 #3567 的后续,两者是一个整体。#3567 修正了两个 container 与主配置
image不一致的 srt-slurm 配方;本 PR 加上本可以发现这种不一致的检查:每次 sweep 都会把每个计划中的 srt-slurm 点绑定到它实际会启动的配方 variant,任何一个点绑定失败,就在派发任何任务之前停止。问题: srt-slurm 点在两个地方声明 container:主配置的
image和配方的model.container。升级镜像必须同时修改两处(如 #3446),但在 benchmark 任务开始之前,没有任何检查比较这两处。infx/srt_slurm/single_node.py要等 benchmark 任务已在集群 runner 上启动、准备好 srt-slurm checkout 并安装 srtctl 之后,才会拒绝不一致的点(Single-node SRT image: recipe/matrix ...)。[AMD] Update GLM-5.2 MI355X image to 20260924 daily / 更新 GLM-5.2 MI355X 镜像至 20260924 daily #3446 的第一次 sweep 就是这样失败的。fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 修复的两个 key 的每个点也会这样失败,因为 feat(srt): run AgentX on native srt-slurm #3428 移植这些配方时,依据的是 config(dsv41flash): re-sweep MI355X on the first ROCm 10.0 nightly carrying vllm#58510 / 在首个包含 vllm#58510 的 ROCm 10.0 nightly 上重新扫描 MI355X #3420 和 [Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 #3334 升级镜像之前的旧脚本。#3567 最初为此加了一个 pytest 扫描。Review 指出它有四个问题:
run-sweepworkflow。EVAL_CONFIG_FILE。AGENTS.md的测试规则。#3567 已移除该扫描,由本 PR 取代。
改动:
infx/srt_slurm/preflight.py(python -m infx.srt_slurm.preflight)从 stdin 读入 matrix。所有 srt-slurm 点都能绑定时,原样输出到 stdout;否则每个问题只打印一次并列出受影响的点,然后以退出码 1 结束。benchmark-tmpl.yml为该点导出的环境变量,调用运行时自己的select_recipe。必须在引擎、模型、镜像、精度、并行配置、GPU 数、投机解码、并发和 KV offloading 上恰好匹配一个 variant。由于逐点检查,被多个不同镜像的 key 共享的配方会分别对照每个 key 的镜像。CONFIG_FILE和EVAL_CONFIG_FILE。每个 variant 的model.container必须是该点的镜像(两种 registry 写法均可),或configs/runners.yaml为该 runner 所在 cluster 配置的 container alias。声明了identity.container.image时,它必须是同一个镜像。TileRT 点使用自己配对的 container,因此跳过。run-sweep.yml:setup先初始化 srt-slurm submodule,在benchmark_schema --plan之后、ci_priority之前运行 validator。validator 失败时setup失败,canary-select和所有要求setup成功的 benchmark 任务都不会启动。e2e-tests.yml(trusted-external-sweep.yml和claude.yml也通过它派发)和profile.yml:用可信工具树里的 validator 检查被测代码树的配方、runner 配置和 srt-slurm submodule。任一棵树还没有这个模块时跳过该步骤,因此测量旧版本不受影响。infx/tests/srt_slurm/test_preflight.py按AGENTS.md的测试规则,只使用临时目录中的小型配方和 runner 配置,覆盖:EVAL_CONFIG_FILE不一致;上方的示例输出,是一个选中 B200 key、但没有 #3567 修复的 sweep 在
setup停止时打印的内容。验证:
test_preflight.py:14 个通过。对 validator 故意做的 8 种破坏,每一种都会让至少一个测试失败:EVAL_CONFIG_FILE、单节点点,或裸 cluster id 的 alias 查找;identity.container.image,或不接受 pyxis 写法;ci.yml的命令:Lint 干净,Tests 2,349 个通过。两个test_slurm_cli用例只在测试机上失败,因为它装有真实的sacct;在干净的main上同样失败。dsv4-fp4-b200-sglang-agentic-hicache-mtp(12 个点)和dsv41flash-fp4-mi355x-vllm-agentic-dspark(16 个点)由 fix(srt): align B200 and MI355X recipe images with master configs / fix(srt):对齐 B200 与 MI355X 配方镜像与主配置 #3567 修复。glm5.2-fp8-mi325x-sglang-agentic-mtp(7 个点)在main上仍不一致。bash -e重放run-sweep的setup:ci_priority。审阅注意事项:
setup失败;这些 sweep 在运行时本来也会失败。glm5.2-fp8-mi325x-sglang-agentic-mtp在main上仍不一致,选中它的 sweep 在其配方修复之前会在setup失败。validator 只检查 sweep 计划中的点,其他 key 的不一致不会挡住无关的 PR。single_node_environment复刻了benchmark-tmpl.yml把 matrix 字段导出给任务的方式;那边的映射改动时,这里也要同步修改。run-sweep只在修改perf-changelog.yaml的 PR 上触发,因此本 PR 自己的检查不会运行新的setup步骤,由上面的重放代替。perf-changelog.yaml记录,因为本 PR 不改任何 benchmark 配置或配方。@SemiAnalysisAI/core。AI 模型使用说明
关联 issue
无。#3567 的后续;相关 PR:#3428、#3446。
改动类型
新功能。