refactor: migrate fixed-sequence recipes to SRT-Slurm / 将定长配方迁移至 SRT-Slurm - #3352
Conversation
开始原生单节点 SRT-Slurm 迁移,添加 H200 SGLang 8k1k 并行候选配方并复用现有基准客户端;生产路由保持不变。
|
Thanks for the contribution!
中文感谢你的贡献!
|
为单节点迁移的 changelog 条目补充草稿 PR 链接。
将原生 H200 SRT 试点接入端到端工作流,显式校验配方、保留结果和功耗产物,并保持生产路由不变。
覆盖原生单节点作业的提交失败、Slurm 失败、产物保留及定向取消行为。
删除 SRT 准备阶段未使用的 AIPerf 排空参数要求,使固定序列工作流能进入实际提交;新增缺省参数回归覆盖。
将配方指定镜像直接交给原生 SRT/Pyxis,移除试点对旧 squash 缓存就绪状态的依赖,保留模型预检查。
在试点提交前执行原生 SRT 二进制准备,并覆盖准备失败时不提交作业的行为。
将 H200 DeepSeek-R1 MTP 与 Qwen3.5 EP8 配方迁移到并行原生 SRT 试点,保留聊天模板及随并发变化的图捕获,并使用 4/16/64 逐点验证回归。
移除独立迁移指南及入口链接,迁移范围与验证证据保留在 PR 中。
合并最新 main,保留上游与迁移分支的全部变更及性能日志条目。
将 H100、H200、B200 和 B300 的 21 个活跃 SGLang 定长配方及全部 181 个配置点切换到原生 SRT-Slurm。共享提交、eval 和产物处理,并保留逐点拓扑与推测解码设置。
将八个定长 TRT 配方迁移至原生 SRT,保留 63 个测试点的引擎参数、客户端和 eval token 预算。
合并 main 的五个新提交,保留现有单节点迁移及全部上游更改;按追加顺序逐字节保留性能变更日志。
将剩余 AMD SGLang/ATOM 与 RTX 定长配方切换到原生 SRT 配置,保留测试点及服务参数;固定直接 ATOM 服务草稿依赖并扩展行为验证。
合并最新 main,保留上游 AgentX 调优及全部原生 SRT 迁移记录。
为 RTX 原生 Docker 路径显式传递 eval 模型名称,并验证实际启动器的命令、失败传播及容器清理。
按原生 SRT 语义逐项传递容器选项,增加映射行为回归测试并输出提交失败原因。
复用作业检出目录存放 AMD 原生运行时临时文件,由现有 runner 清理流程回收,移除额外 scratch-root 配置。
修复 AMD 原生启动命令的尾部换行,在 accounting 不可用时通过 Slurm 控制器验证终态,并等待启动失败的日志进程退出。
移除 Docker 迁移支持和 50 个已被原生配方替代的定长 Bash 实现。Slurm 定长作业必须提供 SRT 配方,保留 AgentX、多节点和显式采集器的独立执行路径。
归档仅支持 Docker 的 RTX 定长配置及脚本,移除无调用方的 runner 路由;保留历史 changelog 条目,但不为已归档的精确配置键生成作业。
修复容器虚拟环境中的客户端依赖安装,移除不兼容的用户目录安装参数。
保留节点独占预留,同时将原生单节点服务和客户端步骤限制为配方指定的 GPU 数量,避免采集空闲设备功耗。
|
Claude finished @cquil11's task in 9m 35s —— View job Review of PR #3352
Summary: No blocking issues found. One 🟡 warning posted inline: the native single-node path only reuses an existing
Limitations: the |
| # Reuse only a valid cache for this exact image. Missing caches are imported | ||
| # by native Pyxis inside the same benchmark allocation. | ||
| SRT_CONTAINER="$IMAGE" | ||
| if [[ -n "${SRT_SQUASH_FILE:-}" && -r "$SRT_SQUASH_FILE" ]] && unsquashfs -s "$SRT_SQUASH_FILE" >/dev/null 2>&1; then | ||
| SRT_CONTAINER="$SRT_SQUASH_FILE" | ||
| fi |
There was a problem hiding this comment.
🟡 WARNING: The native single-node path never populates the squash cache. It only reuses an existing valid SRT_SQUASH_FILE; otherwise it hands the raw docker://-style reference to Pyxis, and launch_b200-nb.sh unsets SRT_SQUASH_FILE so it always does this.
Why it matters: Unlike the retained enroot import -o $SQUASH_FILE docker://$IMAGE blocks in every legacy launcher, nothing on this path ever writes the .sqsh, so a fresh image tag (i.e. every image bump PR) is pulled from the registry by Pyxis independently for each srun step in the allocation (worker, frontend, benchmark client, eval) on every run, inside the GPU allocation. The unsquashfs -s probe also runs on the runner/login host, so on pools where the squash dir is node-local (e.g. B200_SQUASH_DIR on Nscale) it silently falls back to the pull path even when a cache exists on the compute node. This also removes the reproducibility guarantee the enroot-import pattern gives (the exact same imported image used across steps and runs).
Fix: Either import once per job before submission, reusing the existing per-pool import helpers, e.g.
if [[ -n "${SRT_SQUASH_FILE:-}" ]] && ! unsquashfs -s "$SRT_SQUASH_FILE" >/dev/null 2>&1; then
srun -N 1 -A "$SLURM_ACCOUNT" -p "$SLURM_PARTITION" --time="$ENROOT_IMPORT_TIME_LIMIT" \
bash -c "enroot import -o '$SRT_SQUASH_FILE' 'docker://${IMAGE/nvcr.io\//nvcr.io#}'"
fior document here (and in the launchers that unset SRT_SQUASH_FILE) why per-step Pyxis pulls are acceptable for these pools.
There was a problem hiding this comment.
this is gonna be deprecated in later cleanup anyway so I don't think this matters?
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it's a large (151-file) migration of single-node benchmark execution to a new YAML-recipe-driven SRT-Slurm path, touching CI workflows, benchmark scripts, runner launchers, master configs, and a submodule pin, a human look would still be worthwhile.
What was reviewed: the new srt_fixed_sequence.sh/srt_eval.sh launcher scripts for AGENTS.md bash conventions (env-var validation via check_env_vars, no silent fallback defaults); a sample recipe YAML (dsr1/atom/mi355x-fp4-mtp/8k1k.yaml) for structure and the synthetic-acceptance/served-model-name handling described in the PR; and the overall diff shape (deleted per-model bash scripts, new recipe tree, connector module, runner/launcher and master-config updates).
Extended reasoning...
This PR is a large, multi-system migration (151 files, ~6.6k/~5.1k line diff) spanning CI workflow config, dozens of new/removed benchmark scripts and YAML recipes, runner launcher scripts, master configs, a connector module, and an upstream submodule pin bump — none of which is auth/crypto-sensitive but all of which affects benchmark correctness and CI behavior broadly. The bug hunter ran to a dry streak with no findings, and spot checks of the new launcher script and a sample recipe followed the repo's documented bash/config conventions. There is also an inline reviewer comment thread from a third party (cquil11) on one recipe file whose resolution status cannot be independently confirmed from available metadata, and no prior review from this session exists on this PR, so approval is not appropriate given the size and open thread.
Move the shared SRT submodule from v2.22.1 (3cbc5dd) to upstream release v2.23.2 (8dace5f). Upstream delta is three additive changes: vllm-router multi-node hybrid-DP rank offsets, SGLang leader IP honoring the cluster network interface, and nsys/DSight coverage for SGLang and Dynamo. No recipe or connector changes required; utils/ and runners/ tests pass.
Replace the MI300X container preamble with a host setup hook that fails on MEC firmware older than 177 instead of exporting HSA_NO_SCRATCH_RECLAIM in every container.
Post-eval steps dropped the recipe's srun_options, so ATOM evals on MI355X ran in a read-only container. Apply NVIDIA/srt-slurm#504 to the job clone until it merges.
…e port Drop the pilot, cutover and eval-fix changelog entries, the planner support for changelog keys retired later in the PR that only those entries needed, and the doc note about historical smokes.
…mage to v0.5.20-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 (#3334) * feat(dsv4): update B200 SGLang AgentX HiCache MTP image to v0.5.20-cu130 Update dsv4-fp4-b200-sglang-agentic-hicache-mtp from lmsysorg/sglang:v0.5.19-cu130 to lmsysorg/sglang:v0.5.20-cu130 (sglang v0.5.20, build commit 94602c9c). SGLang v0.5.20 removed the deprecated --cuda-graph-max-bs alias (sgl-project/sglang#38375), so the recipe passes the same value through --cuda-graph-max-bs-decode. Model, topology, DSpark settings, HiCache settings and all points are unchanged. 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的镜像从 lmsysorg/sglang:v0.5.19-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130(sglang v0.5.20,构建提交 94602c9c)。SGLang v0.5.20 移除了已弃用的 --cuda-graph-max-bs 别名(sgl-project/sglang#38375),因此配方改用 --cuda-graph-max-bs-decode 传入同一数值。模型、拓扑、DSpark 设置、HiCache 设置及所有测试点保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * chore(changelog): record the B200 DeepSeek-V4-Pro SGLang AgentX image update Append the perf-changelog entry for dsv4-fp4-b200-sglang-agentic-hicache-mtp moving to lmsysorg/sglang:v0.5.20-cu130 (PR #3334). 为 dsv4-fp4-b200-sglang-agentic-hicache-mtp 更新至 lmsysorg/sglang:v0.5.20-cu130 追加 perf-changelog 条目(PR #3334)。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * fix(ci): verify only newly submitted signoffs (#3340) 仅对新提交的签核自动运行验证,编辑清单后需手动重试,并修正文档及验证器提示。 * perf(amd): enable ATOM DSpark K6 with RCCL DEP / ATOM 全并发启用 DSpark K6 与 RCCL DEP (#2912) * perf(amd): switch DSV4 AgentX to native RCCL DEP 将 DeepSeek-V4-Pro MI355X ATOM AgentX 的 c48 及以上测试切换到本地验证过的原生 RCCL DEP 配置,并保持低并发 TP 测试不变。 同步固定的 post-merge ATOM 镜像、EP8 元数据、关闭 TBO/EPLB、真实 MTP 接受率以及本地验证过的路由和 AIPerf 参数。 * docs(perf): link DSV4 RCCL DEP pull request 将 DeepSeek-V4-Pro RCCL DEP 性能变更记录中的占位链接替换为实际的 InferenceX PR 链接。 * fix(amd): trim DSV4 AgentX DEP overrides 精简 DeepSeek-V4-Pro AgentX RCCL DEP 配置,移除与 CLI 或公共默认值重复的环境变量,并补齐本地验证使用的 3600 秒 warmup grace。 * fix(amd): restore DSV4 AgentX scheduling controls Restore the request-equivalent weight, prefill delayer, and decode interval requested for the AgentX run. Remove the newly introduced terminal MTP overrides while keeping the rest of the cleanup unchanged. * fix(amd): remove redundant AgentX timeout overrides Keep the 3600-second agentic warmup allowance, but rely on the server keep-alive setting and AIPerf default benchmark grace period. * fix(amd): keep DSV4 state checkpoints at 8K Remove the DEP-only 32K override so both TP and DEP retain the original 8192-token state checkpoint interval. * perf(amd): capture dense DSV4 DEP decode graphs 为 DSV4-Pro 的 DPA/DEP 路径补齐小 batch CUDA graph,避免非二次幂 decode batch 向上填充造成的 attention、MoE 和 RCCL 无效计算。保留 TP 路径默认行为,并为高并发配置保留大 batch graph。 * perf(amd): include upstream DSV4 DEP fixes in AgentX Pin the September 12 ATOM nightly containing the RCCL EP sentinel fix and AITER stage2 tuning fix. Record the server's installed sources and bundled tuning rows in the AgentX artifacts while retaining dense decode graphs, FP8 KV, and the current scheduler settings. 固定到包含 RCCL EP sentinel 与 AITER stage2 调优修复的 9 月 12 日 ATOM nightly;保存服务端实际源码版本及镜像内 tuning 行,保留 dense decode graph、FP8 KV 和现行调度设置。镜像包含 AITER#4159,但本配方不启用其 BF16 FlyDSL paged SWA 分支。 * fix(ci): provide exact PR context to priority classifier Supply immutable base/head revisions and the permitted diff command, document read-only tool use, and allow 16 turns for structured classification. 为优先级分类器提供确切的 Base/Head 版本及允许执行的 Diff 命令,明确只读工具用法,并将结构化分类的轮数上限调整为 16。 * perf(amd): enable DSpark K6 across the ATOM AgentX sweep 将 PR #2912 的十个 ATOM AgentX 性能点切换为固定版本 0813 checkpoint、 DSpark K6/q7 和 golden AL 3.77;C256 全量 GSM8K 保留真实验证。 修复 draft_model 路由与共享缓存挂载,统一服务与客户端 tokenizer, 增加 DSpark head、66 个分片及 tokenizer 检查并记录运行身份。 保留既有镜像优化、原生 RCCL DEP 和 dense FULL graph 配置。 * fix(amd): preserve inference mode in DSpark TP MoE callbacks 修复 DSpark q7 触发 AITER TP 通信融合 MoE 补齐路径时的 inference tensor 原地写入错误。仅对固定源码应用 InferenceMode callback 补丁,保留通信融合与 graph,保存补丁哈希,并覆盖原始失败、补齐、对齐和重复应用场景。 * fix(amd): use upstream ATOM inference-mode fix 将 MI355X ATOM AgentX 配方切换到包含 ROCm/ATOM#2233 的 pr2233-4f3a808 镜像,并移除运行时 AITER 源码补丁、对应测试和补丁清单记录。 * perf(amd): use BF16 KV through concurrency 16 将 MI355X ATOM AgentX 的 C1、C2、C4、C8 和 C16 任务切换为 BF16 KV cache;C48 及以上 DEP 任务继续使用 FP8 KV cache。 * perf(amd): pin DSv4 ATOM AgentX to official nightly_202609161445 Restore run-sweep.yml to the PR base and replace the temporary rocm/atom-dev:pr2233-4f3a808 image with official nightly_202609161445, which includes the merged ROCm/ATOM#2233 inference-mode fix. 将 run-sweep.yml 还原到 PR base,并把临时镜像 pr2233-4f3a808 换成 包含已合入 ATOM#2233 修复的官方 nightly_202609161445。 * fix: restore missing config-keys in perf-changelog merge The main merge left the first PR #2912 changelog entry without its - config-keys: header, so process_changelog could not parse the added YAML. 修复合并后第一条 2912 changelog 条目缺少 - config-keys: 导致的 YAML 解析失败。 * docs: drop unrelated CI procedure changes 将两份 CI 流程文档恢复为 main 版本,使其不再出现在本 PR 的变更中。 * docs: align CI procedures with PR base 将中英文 CI 流程文档恢复为 PR 当前基础版本,确保它们不再出现在本 PR 的变更中。 * docs(perf): correct DSv4-Pro ATOM AgentX changelog image to nightly_202609161445 Align the perf-changelog entry with the recipe's actual image (rocm/atom-dev:nightly_202609161445 in configs/amd-master.yaml), replacing the stale nightly_202609071454 reference and its now-mismatched digest and build-commit details. Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs(perf): align dsv4-fp4-mi355x-atom-agentic-mtp changelog with PR 2912 Rewrite the dsv4-fp4-mi355x-atom-agentic-mtp entry to match the final recipe: EAGLE MTP -> DSpark (draft_model) on the DeepSeek-V4-Pro-0813 checkpoint, NUM_SPEC_TOKENS 3 -> 6 (K6), SPEC_DECODE_AL 2.49 -> 3.77, and the RCCL DEP band on TP8/DPA8/EP8. Correct the state checkpoint interval to STATE_CHECKPOINT_INTERVAL_TOKENS=8192 (was mis-stated as 32768) and note eval verifies real DSpark drafts, not MTP. Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs(perf): align dsv4-fp4-mi355x-atom-agentic-mtp changelog with PR 2912 Rewrite the dsv4-fp4-mi355x-atom-agentic-mtp entry to match the final recipe: EAGLE MTP -> DSpark (draft_model) on the DeepSeek-V4-Pro-0813 checkpoint, NUM_SPEC_TOKENS 3 -> 6 (K6), SPEC_DECODE_AL 2.49 -> 3.77, and the RCCL DEP band on TP8/DPA8/EP8. Correct the state checkpoint interval to STATE_CHECKPOINT_INTERVAL_TOKENS=8192 (was mis-stated as 32768) and note eval verifies real DSpark drafts, not MTP. Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: seungrokj <[email protected]> Co-authored-by: Chun Fang <[email protected]> Co-authored-by: seungrokj <[email protected]> Co-authored-by: Claude Opus 4.6 <[email protected]> * docs: define draft precision as serve-the-draft-as-it-ships (#3348) Rewrite the Draft-model precision rule in CONTRIBUTING.md, its Chinese mirror, the checklist item, and verifier Check 13 so the baseline is the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling. Forbid only submission-side changes that lower draft precision below that default. Add the Qwen FP8 embedded MTP head and the SGLang DSpark wo_a FP8-to-BF16 load conversion as worked examples of what is compliant. 将 CONTRIBUTING.md 中的 Draft 模型精度规则、其中文镜像、审阅清单条目以及 验证器检查 13 重写为:基线是随所服务 checkpoint 一同发布的 draft,保持其 存储精度,并采用锁定上游镜像的默认加载处理。仅禁止提交方将 draft 精度降到 该默认值以下的改动。新增 Qwen FP8 内嵌 MTP head 与 SGLang DSpark wo_a FP8 转 BF16 加载转换两个合规示例。 * Fix MI300X AMD partition and match its eight-node runner pool (#3350) Submit to compute-0 and retire mi300x-amd_08 from both inventory groups. Benchmark commands and resource settings are unchanged. * Consolidate deprecated configs and remove retired runtime support (#3349) * chore: consolidate deprecated configs by vendor * style: separate deprecated configs with blank lines * chore: remove retired model routing and stale agent guidance * chore: archive retired server registry entries and fix deprecation guidance * chore: archive remaining retired disaggregated registry variants * chore: remove retired DSV4 fixed-sequence helper branches * docs: correct remaining GLM AgentX policy comment * Add B300 Qwen3.5 397B FP8 disaggregated AgentX recipes (#3268) * feat: add B300 Qwen3.5 FP8 disaggregated AgentX starter * chore: link B300 starter performance changelog * perf: colocate B300 Qwen3.5 prefill and decode workers * fix: match B300 Mamba cache flag to pinned SGLang * perf: use latest SGLang nightly for B300 Qwen3.5 FP8 * fix: use Dynamo compatibility with latest SGLang * fix: use intra-node NVLink for colocated B300 PD * perf: select measured B300 FP8 HiCache PD candidate * fix: propagate failed B300 Slurm runs after collecting artifacts * Prepare expanded measured B300 FP8 qualification curve * perf(b300): qualify measured concurrency 48 PD point * perf(b300): preserve existing aggregate publication scopes * style(b300): normalize final config newline * fix: require validated B300 Qwen3.5 AgentX power telemetry * fix: bind B300 power windows to matrix concurrency * perf: qualify B300 FP8 write-through PD at concurrency 96 * style: format B300 power contract fixture * fix: declare the B300 AgentX formal power window contract * docs: consolidate the B300 FP8 performance changelog * Avoid unverified host-memory pressure for B300 H298 recipes * fix: allow complete B300 AgentX runs on eligible nodes * perf: retain B300 FP8 candidates supported by canonical qualification * feat(models): acquire verified original BF16 MTP shards atomically * test(models): reject corrupted provenance and unexpected asset files * feat(b300): stage original BF16 MTP assets for Qwen FP8 recipes * feat(b300): prepare C48 disaggregation with 64 decode slots * perf(b300): add screened TP2 EP2 write-through candidate * perf(b300): select measured C24 decode capacity * perf: expand B300 FP8 disaggregated AgentX frontier * fix: allow full B300 TP2 sweep startup and drain time * perf: select B300 1P2D throughput frontier point * perf: select B300 1P2D C32 frontier point * perf: add B300 1P2D C64 frontier point * perf: refine B300 FP8 disaggregation frontier candidates * Refine B300 FP8 frontier with decode cache and DP attention * perf: refine B300 FP8 disaggregated frontier * perf(b300): qualify six FP8 disaggregated frontier candidates * perf(b300): qualify C4 latency endpoint and drop regressed C40 * perf(b300): select basic DCGM counters for Qwen AgentX * perf(b300): add higher-capacity C48 FP8 disaggregation candidate * perf: expand B300 FP8 sweep with screened C12 and C44 points * perf: add screened B300 FP8 C56 ReplaySSM capacity point * test: remove B300 SRT status test * test: remove MTP acquisition test * fix(b300): use the official FP8 embedded MTP head * fix(b300): disable inherited AgentX power collection * fix(b300): narrow node exclusion and simplify recipe metadata * [Klaud Cold] Update qwen3.5-fp8-h200-sglang-agentic-hicache-mtp SGLang image to nightly-dev-cu13-20260916-c9a8fba9 / 将 qwen3.5-fp8-h200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 nightly-dev-cu13-20260916-c9a8fba9 (#3200) * perf: update qwen3.5-fp8-h200-sglang-agentic-hicache-mtp SGLang image to nightly-dev-cu13-20260916-c9a8fba9 Move the H200 Qwen3.5 FP8 AgentX HiCache MTP recipe from lmsysorg/sglang:nightly-dev-cu13-20260914-4358a161 to lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9 (sgl-project/sglang@c9a8fba9, index digest sha256:509c2742b9441fe9182ae3dd7dac8840ba48e03d2d44b7b24e80d428f117fd85). The 09-16 nightly is the first cu13 build carrying sgl-project/sglang#39516, the HiCache startup fallback for the pinned sglang-kernel 0.4.7 wheel that broke every registered host pool on the 09-15 build. Recipe script, MTP settings, HiCache arms and the concurrency grid are unchanged. 将 H200 Qwen3.5 FP8 AgentX HiCache MTP 配方的 SGLang 镜像从 lmsysorg/sglang:nightly-dev-cu13-20260914-4358a161 更新为 lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9(sgl-project/sglang@c9a8fba9)。 09-16 nightly 是首个包含 sgl-project/sglang#39516 的 cu13 构建,该修复为固定的 sglang-kernel 0.4.7 wheel 提供 HiCache 启动回退,09-15 构建中所有已注册主机池因此无法启动。 配方脚本、MTP 设置、HiCache 分支及并发网格保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * perf: add changelog entry for the qwen3.5-fp8-h200-sglang-agentic-hicache-mtp image bump Append the perf-changelog entry for moving qwen3.5-fp8-h200-sglang-agentic-hicache-mtp to lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9 after the targeted smoke run (35135979333) passed the c2 throughput point and the c24 gsm8k eval. 为 qwen3.5-fp8-h200-sglang-agentic-hicache-mtp 更新至 lmsysorg/sglang:nightly-dev-cu13-20260916-c9a8fba9 追加 perf-changelog 条目; 定向冒烟运行(35135979333)已通过 c2 吞吐点与 c24 gsm8k 评测。 Co-Authored-By: Claude Fable 5.1 <[email protected]> --------- Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <[email protected]> Co-authored-by: adibarra <[email protected]> * [Klaud Cold] Update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 / 将 minimaxm3-fp8-mi300x-vllm-agentic-mtp 的 vLLM ROCm 镜像更新至 v0.29.0 (#3063) * feat: update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 Move the MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 recipe from the unstable commit nightly vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to the v0.29.0 release image vllm/vllm-openai-rocm:v0.29.0 (digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1). Recipe, search space and evals are unchanged. 将 MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 配方的镜像从不稳定的提交 nightly vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 更新为 v0.29.0 发布镜像 vllm/vllm-openai-rocm:v0.29.0 (digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1)。 配方、搜索空间与评测保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * docs: add perf-changelog entry for the minimaxm3-fp8-mi300x-vllm-agentic-mtp v0.29.0 image bump Record the vLLM ROCm image move from nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to the v0.29.0 release image for PR #3063 after the smoke run passed. 为 PR #3063 记录 vLLM ROCm 镜像从 nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 更新至 v0.29.0 发布镜像的 perf-changelog 条目(冒烟运行已通过)。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * docs: shorten PR 3063 performance changelog 缩短 PR 3063 的性能变更日志条目。 * docs: keep PR 3063 changelog concise 保持 PR 3063 性能变更日志简洁。 --------- Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <[email protected]> Co-authored-by: adibarra <[email protected]> * [Klaud Cold] Update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 / 将 minimaxm3-fp8-mi300x-vllm-agentic-mtp 的 vLLM ROCm 镜像更新至 v0.29.0 (#3063) * feat: update minimaxm3-fp8-mi300x-vllm-agentic-mtp vLLM ROCm image to v0.29.0 Move the MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 recipe from the unstable commit nightly vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to the v0.29.0 release image vllm/vllm-openai-rocm:v0.29.0 (digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1). Recipe, search space and evals are unchanged. 将 MI300X MiniMax-M3 MXFP8 AgentX EAGLE3 配方的镜像从不稳定的提交 nightly vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 更新为 v0.29.0 发布镜像 vllm/vllm-openai-rocm:v0.29.0 (digest sha256:e5e47f6aaab675c252c381f0dac237b31b10d87bb74d092b07fb4065efd7f5a1)。 配方、搜索空间与评测保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * docs: add perf-changelog entry for the minimaxm3-fp8-mi300x-vllm-agentic-mtp v0.29.0 image bump Record the vLLM ROCm image move from nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to the v0.29.0 release image for PR #3063 after the smoke run passed. 为 PR #3063 记录 vLLM ROCm 镜像从 nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 更新至 v0.29.0 发布镜像的 perf-changelog 条目(冒烟运行已通过)。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * docs: shorten PR 3063 performance changelog 缩短 PR 3063 的性能变更日志条目。 * docs: keep PR 3063 changelog concise 保持 PR 3063 性能变更日志简洁。 --------- Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <[email protected]> Co-authored-by: adibarra <[email protected]> * [Klaud Cold] Update glm5.2-fp8-mi325x-sglang-agentic-mtp SGLang ROCm image to v0.5.19-rocm720-mi30x / 将 glm5.2-fp8-mi325x-sglang-agentic-mtp 的 SGLang ROCm 镜像更新至 v0.5.19-rocm720-mi30x (#2980) * chore: bump glm5.2-fp8-mi325x-sglang-agentic-mtp SGLang image to v0.5.19-rocm720-mi30x Update the GLM-5.2 FP8 MI325X SGLang AgentX MTP family from lmsysorg/sglang:v0.5.16-rocm720-mi30x to the v0.5.19 release image lmsysorg/sglang:v0.5.19-rocm720-mi30x (Docker Hub digest sha256:a0ffcdd013af6d79c9b7c4348c42b14a79a48f6c43dd3ca0235ec3802da45f75). Model, precision, topology, speculative settings, workloads and recipe script are unchanged. 将 GLM-5.2 FP8 MI325X SGLang AgentX MTP 配方的镜像从 lmsysorg/sglang:v0.5.16-rocm720-mi30x 更新为 v0.5.19 发布镜像 lmsysorg/sglang:v0.5.19-rocm720-mi30x。模型、精度、拓扑、投机解码参数、 工作负载与配方脚本保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * chore: add perf-changelog entry for glm5.2-fp8-mi325x-sglang-agentic-mtp image bump Append the changelog entry for updating the GLM-5.2 FP8 MI325X SGLang AgentX MTP image to lmsysorg/sglang:v0.5.19-rocm720-mi30x (PR #2980). 为 glm5.2-fp8-mi325x-sglang-agentic-mtp 镜像更新至 lmsysorg/sglang:v0.5.19-rocm720-mi30x 追加性能变更日志条目(PR #2980)。 Co-Authored-By: Claude Fable 5.1 <[email protected]> * docs: shorten PR 2980 performance changelog 缩短 PR 2980 的性能变更日志条目。 --------- Co-authored-by: Klaud-Cold <[email protected]> Co-authored-by: Claude Fable 5.1 <[email protected]> Co-authored-by: adibarra <[email protected]> * [Klaud Cold] Add MI325X TP4 and TP2 DeepSeek-V4.1-Flash vLLM AgentX arms with Engram host offload / 新增 Engram 主机卸载的 MI325X TP4 与 TP2 DeepSeek-V4.1-Flash vLLM AgentX 臂 (#3337) * [Klaud Cold] Add MI325X TP4 and TP2 DeepSeek-V4.1-Flash vLLM AgentX arms with Engram host offload / 新增 Engram 主机卸载的 MI325X TP4 与 TP2 DeepSeek-V4.1-Flash vLLM AgentX 臂 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Repin dsv41flash-mi325x-tp2-tp4-engram-offload to the ROCm 10.0 nightly channel / 将镜像重新固定到 ROCm 10.0 nightly 渠道 Match SemiAnalysisAI/InferenceX#3326, which moved MI355X to nightly-rocm100-3df4ae15 on the ROCm 10.0 channel. Same vLLM commit, different ROCm runtime. 与 SemiAnalysisAI/InferenceX#3326 保持一致:相同的 vLLM commit,ROCm 10.0 运行时。 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * chore: refresh PR #3337 for sweep reuse [skip-sweep] Sync with origin/main after the green sweep run 35637516947; the reuse gate authorizes that run on this head. 在绿色 sweep 运行 35637516947 之后与 origin/main 同步;reuse gate 在此 head 上授权该运行。 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 5 (1M context) <[email protected]> * chore(ci): run the priority classifier on Opus 5 at low effort [skip-sweep] (#3360) * chore: bump srt-slurm to v2.18.0 and stream benchmark output (#3367) Move the utils/srt-slurm submodule from 984180e (#448) to v2.18.0 (2ac4eb1), 28 commits ahead, which adds benchmark.stream_output (NVIDIA/srt-slurm#483). Enable it for every srt-slurm recipe through setup_srt_slurm so benchmark.out is mirrored into the Actions job log while the client runs. The TileRT fork pin predates the option and is left unchanged. * Add DSV4 B200 Dynamo+SGLang AgentX configs / 添加 DSV4 B200 Dynamo+SGLang AgentX 配置 (#3257) * feat(dsv4): add B200 Dynamo SGLang AgentX configs * docs(perf): link DSV4 B200 AgentX PR * fix(dsv4): expand c256 decode KV headroom --------- Co-authored-by: hshrivastava-droid <[email protected]> Co-authored-by: Rohit Nagraj <[email protected]> * [Klaud Cold] Add MI300X TP4 DeepSeek-V4.1-Flash vLLM AgentX arm with Engram host offload / 新增 Engram 主机卸载的 MI300X TP4 DeepSeek-V4.1-Flash vLLM AgentX 臂 (#3336) * [Klaud Cold] Add MI300X TP4 and TP2 DeepSeek-V4.1-Flash vLLM AgentX arms with Engram host offload / 新增 Engram 主机卸载的 MI300X TP4 与 TP2 DeepSeek-V4.1-Flash vLLM AgentX 臂 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Repin dsv41flash-mi300x-tp2-tp4-engram-offload to the ROCm 10.0 nightly channel / 将镜像重新固定到 ROCm 10.0 nightly 渠道 Match SemiAnalysisAI/InferenceX#3326, which moved MI355X to nightly-rocm100-3df4ae15 on the ROCm 10.0 channel. Same vLLM commit, different ROCm runtime. 与 SemiAnalysisAI/InferenceX#3326 保持一致:相同的 vLLM commit,ROCm 10.0 运行时。 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * Drop the MI300X TP2 arm: measured infeasible on a 192 GB card / 移除 MI300X TP2 臂:在 192 GB 卡上实测不可行 Run 35671005506 reported "Available KV cache memory: -13.51 GiB" at TP2 concurrency 1 and the engine refused to start, even with the indexer buffer already halved to 4096 batched tokens. TP8 and TP4 are unaffected; TP4 passed at concurrency 1 and 32 in the same run. 运行 35671005506 在 TP2 并发 1 下报告 "Available KV cache memory: -13.51 GiB" 并拒绝启动,此时 indexer 缓冲区已减半至 4096。TP8 与 TP4 不受影响。 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * chore: refresh PR #3336 for sweep reuse [skip-sweep] Sync with origin/main after the green sweep run 35678619120; the reuse gate authorizes that run on this head. 在绿色 sweep 运行 35678619120 之后与 origin/main 同步;reuse gate 在此 head 上授权该运行。 Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 5 (1M context) <[email protected]> * fix: warn when sign-off lacks a reuse command (#3368) 中文:签署审核缺少授权复用命令时改为警告,同步裁定发布逻辑、双语文档和回归测试。 * [PowerX] retain healthy window measurements in audit sidecars / 在审计文件中保留健康窗口测量 (#3357) * fix: retain healthy power windows in audit sidecars 保留健康窗口的独立测量和逐点失败证据,保持包完整性检查、required-power 及发布判定不变。 * docs: clarify retained power diagnostics 精简中英文窗口保留说明,明确审计顶层逐 GPU 诊断字段不构成发布许可。 * [B200][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance (#3346) * perf(b200): tune DeepSeek V4.1 host Engram layout on latest SGLang * chore(b200): link performance candidate to PR 3346 * perf(b200): add STP and retain native DSpark FP8 projections * fix(b200): import pinned images and handle strided draft inputs * perf(b200): compare GPU-resident Engram tables for STP * fix(b200): use shipped default DSpark implementation * perf(b200): service STP decodes between long prefills * perf(b200): probe TP2 DSpark with bounded prefill memory * perf(b200): schedule DSpark decode between long prefills * perf(b200): compare GPU-resident Engram for TP4 DSpark * fix(b200): release completed MXFP4 loader temporaries * fix(b200): mark loader patch utility executable * perf(b200): test more DSpark SWA prefix tails * fix(b200): avoid TP2 loader allocation fragmentation * perf(b200): compare DSpark prefill service at interval four * perf(b200): compare host DSpark with larger SWA retention * perf(b200): qualify TP2 DSpark across concurrency * perf(b200): test more TP2 SWA prefix retention * docs(b200): document TP2 loader memory patch waiver * perf(b200): select DSpark curve with concurrency-scaled SWA capacity * perf(b200): qualify stock TP2 loader with expandable allocator * Test stock B200 loader on latest SGLang nightly AI-assisted by gpt-6-astra with high reasoning. * Document latest-nightly stock B200 loader qualification * refactor: remove redundant B200 recipe startup checks * [B300][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance (#3342) * feat(b300): add DeepSeek V4.1 SGLang nightly AgentX recipe * fix(b300): record recipe pull request provenance * feat(b300): add non-speculative SGLang AgentX arm * docs: consolidate performance changelog into one PR entry * fix(b300): preserve native DSpark FP8 draft projections * fix(b300): normalize strided native FP8 projection inputs * fix: restore default SGLang DSpark precision * perf(b300): expand DSpark parallelism and interleave decode * Launch B300 SGLang AgentX through an owned Slurm batch job * Match B300 batch PMI environment to non-MPI container step * Use supported PMIx container steps for B300 AgentX batch jobs * perf(b300): scope final AgentX sweep to DSpark candidates * Test B300 low-concurrency prefix retention and checkpoint prefetch * Scale B300 TP4 prefix retention within measured KV budget * Pin B300 to the September 22 SGLang nightly * Test larger B300 TP2 KV budget for high concurrency * Increase high-concurrency TP2 cache memory on B300 * Screen higher prefill duty for B300 TP2 DSpark * perf(amd): tune MiniMax-M3 ATOM AgentX memory, graphs, indexer CP, and offload / 调优 MiniMax-M3 ATOM AgentX 显存、图捕获、Indexer CP 与卸载 (#3189) * perf(minimaxm3-atom): raise util, capture the real batch sizes, add indexer CP and a TP2 offload curve Six changes to the MiniMax-M3 MXFP4 ATOM agentic recipe, all measured on 8xMI355X against the same cc-traces replay: - gpu-memory-utilization 0.90 -> 0.95. - Declare cudagraph capture sizes. The stock list pads a running batch of 9 to 16 and 17 to 32; this workload keeps only 22-41% of the offered concurrency in decode at any instant, so those are the batches that actually replay. - Enable ATOM_M3_INDEXER_CP on TP4 past CONC=24. Each rank scores all four index heads over 1/TP of the blocks instead of one head over all of them. - Restore EAGLE3 on the CONC=40/48 bands. The forward step is 37-41 ms with or without it, so dropping it handed back the whole acceptance-length multiplier. - Route CONC=20/25/30 to the chunk-256 + ATOM_SLRU offload tier and derive LMCACHE_LOOKUP_SERVER_WORKER_IDS from TP instead of hardcoding four ranks. - Spread TP2 offload ranks across NUMA nodes; pinning 256 GB/rank takes 45 min with both ranks on node 0 and 21 s with one per node. Matrix: drop TP2 c5, add the TP2 c20/25/30 offload curve. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * refactor(minimaxm3-atom): drop the dead per-band knobs The per-CONC case set MAX_NUM_SEQS, MAX_NUM_BATCHED_TOKENS and GPU_MEM_UTIL in all seven bands and two lines later overwrote all three unconditionally: 21 of its 49 assignments could never be read. What actually survived it was NUM_SPEC_TOKENS, SPEC_DECODE_AL and STATE_OFFLOAD_CPU_GIB, and after restoring EAGLE3 on 40/48 only CONC=56 differs on the first two. Collapsed to defaults plus three overrides; STATE_OFFLOAD_CPU_GIB keeps its exact per-band values. ATOM_ENABLE_REPLAYSSM is removed rather than collapsed: it is only read by gdn_attn, which gates the gated-delta-net path, and MiniMax-M3 is pure attention with no SSM layers, so the value never reached anything. Verified by replaying every point in the amd-master matrix (19 combinations of TP, CONC and offload) through a stubbed benchmark_lib and diffing the resulting server argv plus every exported env var: byte-identical apart from the removed variable. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * refactor(minimaxm3-atom): one table per concurrency, and drop the tier M3 cannot use The hybrid offload tier is unreachable for this model. ATOM's M3 connector is PAGE-only -- the docstring says MiniMax-M3 has no recurrent state and no SLOT sidecar, and it raises if slot_regions is non-empty -- so OFFLOAD_STATE and the staging knobs behind STATE_OFFLOAD_CPU_GIB never applied. Its LMCACHE_CHUNK_SIZE is justified by "block-size(128) x dcp(8)", but this model's search space sets no dcp-size at all. Removed, along with the comment claiming M3 carries a per-request recurrent state. An unlisted concurrency now fails instead of silently taking a tier that does not fit. CONC=56 is removed for the same reason: nothing in the matrix runs it here (the 56 in this file's neighbourhood belongs to Kimi K3, which has its own script), and it was the only band that disabled the draft model. With it gone NUM_SPEC_TOKENS is constant, so the guard around SPEC_ARGS goes too. The remaining per-concurrency facts now live in one table. Previously the CP gate was a numeric comparison and the offload tier an enumeration, so adding a concurrency above 24 would silently enable indexer CP on a point nobody measured, while forgetting the tier failed loudly. Both are declared per band now: the loud failure is kept, the silent one becomes a missing speedup. Verified the same way as the previous cleanup: all 19 matrix points replayed through a stubbed benchmark_lib, server argv and every exported env var byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * perf(minimaxm3-atom): spread TP4 offload ranks across NUMA nodes, move to the rocm7 nightly Pinning 256 GB/rank with all four TP4 ranks on NUMA node 0 takes 27m08s (12:13:10 "NUMA mapping" -> 12:40:18 "Created backend: LocalCPUBackend"); the same pinning with TP2 split one rank per node takes 21 s. TP4 now selects 0,1,4,5 for the same reason TP2 selects 0,4. Only applied when the caller has not already pinned a device set, and only on the offload path -- a run without the CPU tier pins no host memory. The image moves to the rocm7 line of ATOM's nightly release (232f62e3), which carries the M3 dense-split work. Checked against the measured baseline rather than assumed: triton 3.7.0 and torch 2.10.0+rocm7.2.4.git3d3aa833 are identical to the image every number in this PR was measured on, which matters because the rocm10.0 line of the same build carries triton 3.8.0 and that version regresses this model's Gluon kernels. LMCache in the new image is 0.5.5rc3, not the 0.4.5 the matrix declared, so the two MiniMax-M3 offload rows are corrected; the other models keep their own versions. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * docs(perf): record MiniMax-M3 AgentX tuning sweep 记录 MiniMax-M3 ATOM AgentX 调优、矩阵变化、测试结果与已知 CPU DRAM 元数据限制,以便上游 PR 启动完整 AgentX sweep。 * docs(perf): link MiniMax-M3 tuning pull request 将 MiniMax-M3 ATOM AgentX 性能 changelog 的占位链接替换为上游 PR #3189。 * fix(ci): provide exact PR context to priority classifier Supply immutable base/head revisions and the permitted diff command, document read-only tool use, and allow 16 turns for structured classification. 为优先级分类器提供确切的 Base/Head 版本及允许执行的 Diff 命令,明确只读工具用法,并将结构化分类的轮数上限调整为 16。 * ci(agentx): run MiniMax-M3 throughput only 禁用本次 changelog 的自动 vendor eval,只保留 19 个 AgentX 吞吐任务,避免为每个并发额外展开评测。 * perf(minimaxm3-atom): extend indexer CP down to CONC=15 on TP4 The gate was CONC > 24, chosen from a kernel microbenchmark over uniform contexts. Three same-commit A/B pairs (1800 s per arm, TP4 + EAGLE3) put the real crossover below 15: CONC 15 interactivity p90 +3.7% ISL-normalised throughput +0.67% CONC 24 +12.3% +0.24% CONC 28 +16.7% +2.7% Monotone and never negative. Throughput is neutral everywhere (only CONC=28 clears the 0.6% noise floor) while interactivity rises with the running batch, which is what CP actually shortens. Prefix hit rates are identical within each pair (96.7/96.7, 97.5/97.5, 97.2/97.3), so the gain is not a caching artifact. CONC=20 is interpolated between the two measured points, not run on its own. Nothing below 15 is measured and the microbenchmark is negative at batch ~2, so the band stops there. CONC=20 appears in both search-space rows, so it is split out of the offload enumeration: TP4 gets CP without a tier, and the tp_size == 4 gate keeps CP off the TP2 offload curve. 25 and 30 are TP2-only and stay without a CP declaration, which would never take effect but would misreport the point. Verified with the 19-point argv+env replay: the only difference against the previous revision is ATOM_M3_INDEXER_CP="1" on tp4_c15, tp4_c20 and tp4_c24. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> * docs(perf): record expanded MiniMax-M3 indexer CP band 记录 TP4 C15/C20/C24 新增 indexer-only CP 的 A/B 证据,并说明 TP2 C20 仍由 TP-size gate 保持关闭。 * ci: restore default PR priority classifier Revert the 16-turn exact-diff prompt so this recipe PR does not change shared sweep CI. 还原默认的优先级分类器配置,避免本配方 PR 修改共享 sweep workflow。 Co-authored-by: Cursor <[email protected]> * perf(minimaxm3-atom): switch to official nightly_202609170640_minimax_m3 image Replace the developer-named ATOM nightly so this submission uses an official tag. 将 ATOM 镜像替换为不含开发者用户名的官方 tag nightly_202609170640_minimax_m3。 Co-authored-by: Cursor <[email protected]> * switch to 0917 nightly atom image * perf(minimaxm3): enable agentic eval by removing no-evals opt-out Drop `no-evals: true` from the minimaxm3-fp4-mi355x-atom-agentic-mtp changelog entry so plan generation no longer skips eval rows; the model's automatic minimax-vendor (minimax_m3_full) agentic eval then runs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs(perf): correct MiniMax-M3 ATOM AgentX changelog image to nightly_202609171455 Align the perf-changelog entry with the recipe's actual image (rocm/atom-dev:nightly_202609171455 in configs/amd-master.yaml), replacing the stale nightly_202609170640_minimax_m3 reference in both the English and Chinese descriptions. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(minimaxm3): size LMCache CPU pool per-rank from TOTAL_CPU_DRAM_GB/TP The cpu256 offload tier hardcoded LMCACHE_MAX_LOCAL_CPU_SIZE=256 per rank, independent of the matrix generator's TOTAL_CPU_DRAM_GB budget. Derive it as TOTAL_CPU_DRAM_GB / TP so the per-rank pool respects the declared budget and is TP-independent (node_DRAM * dram-utilization / 8). Bump the atom arm's dram-utilization 0.20 -> 0.687 so the offload arms land ~256 GB/rank. Co-Authored-By: Claude Opus 4.6 <[email protected]> * Update perf-changelog.yaml * docs(minimaxm3): expand ATOM AgentX changelog and record EAGLE3-GQA draft online-quant finding Update the minimaxm3-fp4-mi355x-atom-agentic-mtp changelog entry to reflect every change in SemiAnalysisAI/InferenceX#3189 and to record the EAGLE3-GQA draft online-quant verification: - Raise ATOM GPU memory utilization 0.90 -> 0.95 and declare a fine-grained CUDA graph ladder (capture sizes 1-64) sized for the 22-41%-of-CONC running batch, trimmed by ModelRunner to min(2*CONC, 8192). TP4 C24 A/B: throughput +9.2%, interactivity +50.7%, unchanged estimated graph memory. - Enable ATOM_M3_INDEXER_CP for TP4 C15/C20/C24/C28/C32/C40/C48; the TP-size gate (tp_size == sparse_num_index_heads == 4) keeps it off the TP2 C20 offload point. Restore EAGLE3 K3 (golden AL 2.78) at C40/C48. - Collapse per-band knobs into one case; add the TP2 LMCache DRAM-offload curve at C20/C25/C30 on the chunk-256 ATOM_SLRU tier; size LMCACHE_MAX_LOCAL_CPU_SIZE as TOTAL_CPU_DRAM_GB/TP; cover every rank in LMCACHE_LOOKUP_SERVER_WORKER_IDS. Remove TP2 C5 and the unreachable hybrid/state-offload and C56 branches. - NUMA-spread offload GPUs (0,4 for TP2; 0,1,4,5 for TP4), cutting TP2 host-memory pinning from 45 min to 21 s. Move to rocm/atom-dev:nightly_202609171455 and LMCache 0.5.5rc3+rocm7.2.4 on both offload arms. - Set dram-utilization to 0.687 so the recorded budget matches the real ~256 GB per-rank pinning (512 GB total TP2, 1,024 GB TP4), independent of TP, replacing the earlier 0.20 metadata that under-described the allocation. - Verified against ATOM commit b6e2a39433d73b586c366ca90e661f098eb463e3 that --online_quant_config ptpc_fp8 does not change the Inferact/MiniMax-M3-EAGLE3-GQA draft precision: eagle3_llama.py builds every draft linear (qkv_proj, o_proj, gate_up_proj, down_proj, fc) with no quant_config, so should_stream_online_quant short-circuits to False and OnlineQuantStreamer.maybe_create finds no candidates. The draft runs at native bf16; only the target MiniMax-M3 weights are online- quantized to ptpc_fp8, and exclude_layer applies to the target alone. Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Wang Yiting <[email protected]> Co-authored-by: Claude Opus 5 (1M context) <[email protected]> Co-authored-by: Cursor <[email protected]> Co-authored-by: billishyahao <[email protected]> Co-authored-by: seungrokj <[email protected]> Co-authored-by: seungrokj <[email protected]> * feat(agentx): retune GLM5.2 FP4 MI355X ATOM recipe on _0921 docker (#3359) * feat(agentx): retune GLM5.2 FP4 MI355X ATOM recipe on _0921 docker Signed-off-by: zhuyuhua-v <[email protected]> * docs(glm5.2): note MTP block stays BF16 under --online_quant_config (PR #3359) Changelog for the GLM-5.2 FP4 MI355X ATOM AgentX recipe retune: - Track the updated GLM-5.2 standalone recipe (ROCm/ATOM PR 2345) on the MI355X ATOM AgentX config. - Bump the image from rocm/atom-dev:nightly_202609151450 to rocm/atom-dev:nightly_202609211553, which carries ROCm/ATOM PR 2345. - Drop the LMCache DRAM offload from the small-concurrency TP4 arm: conc [2, 4, 8, 10] moves from kv-offloading dram with the lmcache 0.4.5 backend to kv-offloading none, so those points serve off the native GPU prefix cache alone and no longer start an LMCache CPU tier. The TP4+DCP4 arm at conc [16, 24, 32, 40, 48] keeps LMCache at 256 GiB per rank (dram-utilization 0.171 unchanged), and the TP8 arm stays GPU-resident as before. - Add --index_cache_dtype fp4 and --block-size 64 to the ATOM server launch on every point. Neither flag was set before, so both took the ATOM default. - Drop ATOM_USE_FLYDSL_GATHER_KV_B_PROJ=0, which pinned gather_kv_b_proj to the Triton path on the DCP prefill-context and MTP verify paths; the updated recipe no longer sets it. - The MTP block data precision is not touched by --online_quant_config: the exclude_layer list adds "model.layers.78.*" so layer 78 (the GLM-5.2 MTP head) is skipped by online quantization and stays in native BF16, as it ships unquantized in amd/GLM-5.2-MXFP4. The expert excludes only reach layers 0-77, so without this exclude the whole MTP block would be online-quantized to ptpc_fp8 while the target's experts stay MXFP4. Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Signed-off-by: zhuyuhua-v <[email protected]> Co-authored-by: seungrokj <[email protected]> Co-authored-by: seungrokj <[email protected]> Co-authored-by: Claude Opus 4.6 <[email protected]> * config(dsv41flash): sweep MI355X at TP=2 and TP=4 through concurrency 128 / MI355X DSv4.1-Flash 在 TP=2 与 TP=4 下扫描至并发 128 (#3326) * config(dsv41flash): sweep MI355X at TP=2 and TP=4 through concurrency 128 Add a TP=2 search space next to TP=4 for the MI355X DSv4.1-Flash AgentX entry, and extend both rows to concurrency 128. TP=2 became feasible once the Engram tables moved to host memory: at 47.2 GiB of device memory per rank they previously forced four GPUs per server, and two GPUs per server doubles the servers per node. The image moves to the ROCm 10 nightly and carries a placeholder tag until that build is published. The digest is pinned before the sweep runs, which is why the PR opens as a draft. 中文:为 MI355X DSv4.1-Flash AgentX 配置在 TP=4 之外新增 TP=2 搜索空间,并将两者 的并发扩展到 128。Engram 表已驻留主机内存,不再占用每 rank 47.2 GiB 设备内存 (此前这迫使每个服务器使用四张 GPU),因此 TP=2 可行,且每节点可容纳的服务器数 量翻倍。镜像切换到 ROCm 10 nightly,在该构建发布前暂用占位 tag;digest 固定后再 运行 sweep,因此本 PR 以 draft 形式提交。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): record the sweep PR link in the changelog 中文:在 changelog 中记录本次 sweep 的 PR 链接。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): let vLLM pick max_num_seqs for the MI355X AgentX recipe A fixed max_num_seqs of 128 caps in-flight sequences at the outer concurrency once the sweep reaches c128, leaving no headroom for AgentX subagent fan-out. Use the MI355X API-server default of 1024 instead, and derive the CUDA graph ceiling from max(2 * CONC, 128): 1024 through c64, 2048 at c128. Supersedes PR #3111, which carried this recipe change and was closed in favour of this sweep. 中文:扫描到 c128 时,固定的 max_num_seqs 128 会把在途序列数限制在外层并发上, 使 AgentX 子代理扇出没有余量。改用 MI355X API server 默认值 1024,并由 max(2 * CONC, 128) 推导 CUDA graph 上限:c64 及以下为 1024,c128 为 2048。 本次提交取代已关闭的 PR #3111(该 PR 原本包含此 recipe 改动)。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): keep the placeholder on the existing ROCm nightly channel Per #3215, the MI355X DSv4.1-Flash entry stays on the existing upstream ROCm nightly channel rather than moving to nightly-rocm100. Drop the rocm10 marker from the placeholder tag; the pin is filled in with the nightly this sweep actually runs on before dispatch. 中文:依据 #3215,MI355X DSv4.1-Flash 条目继续使用现有的上游 ROCm nightly 渠道, 不切换到 nightly-rocm100。占位 tag 去掉 rocm10 标记;派发前再填入本次 sweep 实际运行的 nightly。 Co-authored-by: Cursor <[email protected]> * Revert "config(dsv41flash): keep the placeholder on the existing ROCm nightly channel" This reverts commit 039194e50d92a80c847d770729e106237e2b36c4. The MI355X DSv4.1-Flash AgentX sweep runs on the ROCm 10 nightly after all, so restore the nightly-rocm100 placeholder. The pin is filled in with the published nightly before dispatch. 中文:还原提交 039194e50d92a80c847d770729e106237e2b36c4。 MI355X DSv4.1-Flash AgentX sweep 最终仍在 ROCm 10 nightly 上运行,因此恢复 nightly-rocm100 占位 tag。派发前再填入已发布的 nightly。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): pin the MI355X AgentX sweep to the ROCm 10 nightly Replace the placeholder with vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984, published 2026-09-21T05:50:03Z with digest sha256:eccb72b74b8c04ce7406d9212e200b129f58e24be642795a857f755a94fdb3a1. The tag is active and linux/amd64, and its manifest resolves from registry-1.docker.io at that digest. This sweep qualifies the new image; the previous green run on nightly-eed1f3d0 is not evidence for it. 将占位 tag 替换为 vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984, 该镜像于 2026-09-21T05:50:03Z 发布,digest 为 sha256:eccb72b74b8c04ce7406d9212e200b129f58e24be642795a857f755a94fdb3a1。 该 tag 状态为 active、架构为 linux/amd64,其 manifest 可从 registry-1.docker.io 按该 digest 解析。 本次 sweep 用于验证新镜像;此前在 nightly-eed1f3d0 上的绿色运行不能作为其证据。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): offload Engram to host memory at TP=2 on MI355X vllm-project/vllm#57491 widened the two ROCm is_cuda() gates to is_cuda_alike(), so the pinned nightly resolves an Engram config on gfx950 and offloads the tables to pinned host memory by default. Set cpu_offload explicitly per TP instead of taking that default. TP=2 needs the offload: the tables cost 47.2 GiB per rank at TP=4, so 94.4 GiB at TP=2, which does not fit beside half of the 511 GB checkpoint on a 288 GiB card. TP=4 keeps them resident so it stays comparable with the validated concurrency 1-32 run. vllm-project/vllm#57491 将 ROCm 上的两处 is_cuda() 判断放宽为 is_cuda_alike(),因此所固定的 nightly 在 gfx950 上会解析 Engram 配置, 并默认将表下放到锁页主机内存。这里按 TP 显式设置 cpu_offload,而不是 沿用该默认值。 TP=2 需要该下放:表在 TP=4 时每 rank 占用 47.2 GiB,TP=2 时为 94.4 GiB, 无法与 511 GB 检查点的一半同时放入 288 GiB 的单卡。TP=4 保持表常驻 GPU, 以便与已验证的并发 1-32 运行保持可比。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): disable SWA bounded replay on MI355X Every TP=2 and TP=4 point of run 35567570539 died with HSA_STATUS_ERROR_MEMORY_FAULT, taking the worker and then the engine core down and aborting AgentX warmup. The fault follows the first launch of _pad_replayed_slots_kernel in all 16 server logs, which is the first prefix hit carrying a replay start. vllm-project/vllm#56227 added SWA bounded replay, default on, after the eed1f3d0 pin and before nightly-rocm100-3df4ae153. It pads the replayed tokens' slots in the prefix-cacheable groups, but the window clamp it relies on landed in the FlashInfer and FlashMLA kernels; the ROCm sparse SWA path only gained the replay_start kwarg, and consumes it on the decode path while the fault is in prefill. Pass --no-swa-bounded-replay until ROCm clamps too. Prefix caching itself stays on. 运行 35567570539 的所有 TP=2 与 TP=4 数据点都以 HSA_STATUS_ERROR_MEMORY_FAULT 崩溃,先后带崩 worker 与 engine core, 并中止 AgentX warmup。16 份 server 日志中,该故障均紧随 _pad_replayed_slots_kernel 的首次启动,即首个带 replay start 的前缀命中。 vllm-project/vllm#56227 在 eed1f3d0 与 nightly-rocm100-3df4ae153 之间 引入了默认开启的 SWA bounded replay。它会填充被重放 token 在可前缀缓存 分组中的 slot,但其依赖的窗口钳制只落在 FlashInfer 与 FlashMLA 内核中; ROCm 稀疏 SWA 路径仅新增了 replay_start 参数,且只在 decode 路径消费它, 而故障发生在 prefill。在 ROCm 同样实现钳制之前传入 --no-swa-bounded-replay。前缀缓存本身保持开启。 Co-authored-by: Cursor <[email protected]> * config(dsv41flash): fix the MI355X KV split at high concurrency Concurrency 64 collapsed on both arms of run 35574132719 because the KV pool was too small to hold the AgentX working set. TP=2 fell to a 17.6% prefix cache hit rate, 187 s TTFT and 150 tok/s, against 94.8%, 1.3 s and 957 tok/s at c32. Two settings were spending device memory that the KV pool needed. The Engram tables stayed resident at TP=4. vllm-project/vllm#57491 widened the two ROCm is_cuda() gates to is_cuda_alike(), so gfx950 can now offload them to pinned host memory as every NVIDIA arm has since #2963. Measured here at TP=4 with 16384 batched tokens, resident tables leave 37.96 GiB of KV cache and 14.15x maximum concurrency at 1M context, against 84.54 GiB and 31.52x offloaded. The sparse-attention indexer and its companion per-rank buffers scale at roughly 4.4 MiB per batched token, so the upstream 16384 was the larger cost. Size it per TP instead, 4096 at TP=2 and 8192 at TP=4, and cap max_num_seqs at the shape graph capture already covers rather than the MI355X API-server default of 1024. TP=2 then keeps 79.34 GiB (39.44x) and TP=4 keeps 121.03 GiB (54.15x), against the B300 arm's 132.48 GiB (60.17x). B300 runs 8192 at TP=4 and the Blackwell TP=2 arms run 4096 (#3320, #3321). Sweep concurrency 64 and 128 to confirm the high end first. Concurrency 1-32 is backfilled once those land, since every point now runs the new memory split. Co-authored-by: Cursor Agent <[email protected]> Signed-off-by: Fangzhou Ai <[email protected]> * config(dsv41flash): change the MI355X memory split only where KV ran short The previous commit applied one memory split per TP across every concurrency. Concurrency 1-32 never ran short of KV, so it paid for changes it did not need: offloaded Engram lookups go to pinned host memory over UVA, and a smaller prefill chunk costs TTFT. Scope each change to the points that were actually starved, which also leaves concurrency 1-32 and TP=4 c64 byte-identical to the settings run 35574132719 already measured on this image and flag set, so those points are combined with this sweep rather than re-run. Dividing measured KV tokens by concurrency gives the per-request budget and separates the failures cleanly. Every point that held had 232K or more; TP=2 c64 collapsed at 122K. TP=4 c1-c64 resident, 16384 14.83M tokens 232K and up TP=4 c128 offload, 8192 49.63M 388K TP=2 c1-c32 offload, 16384 7.84M 245K and up TP=2 c64 offload, 8192 24.68M 386K TP=2 c128 offload, 4096 33.34M 260K Probing the configurations directly also turned up two things. TP=2 c128 at 4096 with the API-server default of 1024 sequences faults during profiling with HSA_STATUS_ERROR_EXCEPTION at M=1024, N=129280, K=256, reproduced on two GPU pairs: DSpark verifies 1+5 tokens per sequence, so a decode batch of max_num_seqs needs six times that many token slots and 4096 leaves four. Cap max_num_seqs at the graph-capture shape wherever the chunk falls below that bound, which is TP=2 c128 alone. And the reason c128 never completed before is that at 16384 with capture at 2048, TP=2 c128 held 7.72 GiB, 3.02M tokens, 23.6K per request. Co-authored-by: Cursor Agent <[email protected]> Signed-off-by: Fangzhou Ai <[email protected]> * config(dsv41flash): measure the full MI355X concurrency range on the new pin * docs(dsv41flash): retire the MI355X draft status and repoint GPU validation The MI355X section still described a draft recipe at TP4 concurrency 1-32 and cited run 34710937012 on nightly-eed1f3d0 as the pinned image. Both predate this branch: the arm now sweeps TP4 and TP2 at concurrency 1-128 on nightly-rocm100-3df4ae15. Lead with the current pin, keep the earlier run only as the superseded-pin note explaining why its points do not carry onto this image, and add the merged upstream recipe #1006 beside #968. MODELS.md and MODELS_zh.md carried a second DeepSeek-V4.1-Flash row holding the MI355X arm as pending GPU validation. Drop it; the active row already covers the model. --------- Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]> Co-authored-by: functionstackx <[email protected]> Co-authored-by: Chun Fang <[email protected]> * [AMD] Update GLM-5.2 MI355X image, HiCache capacity, and sweep / 更新 GLM-5.2 MI355X 镜像、HiCache 容量与 sweep (#3329) * perf: update GLM-5.2 ROCm image and HiCache path 将 GLM-5.2 ROCm 镜像更新到 20260920,切换 HiCache kernel/page_first,并调整并发矩阵。 * perf: split GLM-5.2 low and high concurrency points TP8/EP1 仅保留并发 1、2、4;TP4/EP4 HiCache 仅运行并发 12、14、16,并补充实际 PR 链接。 * Keep existing GLM-5.2 HiCache settings * Use default GLM-5.2 HiCache backend and layout * Add missing GLM-5.2 concurrency points * perf: increase GLM-5.2 HiCache capacity 将 GLM-5.2 TP4/EP4 HiCache 调整为每 rank 180 GB,并将 SA DRAM 利用率提高到 0.85。 * fix changelog --------- Co-authored-by: Chun Fang <[email protected]> * fix(review): bind verifier verdicts to signoffs (#3371) 让清单编辑触发验证,并按签署资源隔离验证评论。 * [TileRT] GLM-5.3 MI355X AgentX parallel KV inject experiments, reverted pending tile-ai/TileRT#66 / GLM-5.3 MI355X AgentX 并行 KV 注入实验(已回退,待 tile-ai/TileRT#66) (#3376) * [GB300][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance (#3344) * perf(gb300): add nightly SGLang V4.1 Flash DSpark candidates * chore(gb300): link performance changelog to PR 3344 * perf(gb300): add non-speculative SGLang AgentX candidates * chore(gb300): consolidate performance changelog entry * perf(gb300): resolve local SGLang checkpoint for Engram cache advice * perf(gb300): retain native FP8 DSpark draft projections * fix(gb300): make native draft FP8 inputs contiguous * perf(gb300): test GPU-resident Engram tables on TP4 * Use shipped nightly DSpark precision on GB300 * Interleave GB300 STP decode with long prefills * perf(gb300): prioritize DSpark across the AgentX curve * perf(gb300): test larger SWA prefix cache for DSpark * Test larger SWA cache on GB300 TP2 DSpark * Select DSpark and scale SWA retention for the GB300 sweep * perf(gb300): select qualified prefill capacity on latest nightly * chore(gb300): remove documentation changes from performance PR * [H200][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance (#3341) * perf(h200): qualify DeepSeek V4.1 Flash on latest SGLang nightly * chore(h200): link performance qualification PR * perf(h200): enable huge pages for Engram host shards * perf(h200): add native non-speculative SGLang throughput arm * chore(h200): consolidate performance changelog entry * fix(sglang): retain native FP8 DSpark attention projections * style(sglang): sort native projection patch imports * chore(sglang): store context-free native projection patch * fix(h200): make native draft FP8 inputs contiguous * perf(h200): test GPU-resident Engram for non-speculative serving * perf(h200): interleave STP decode with long prefix prefill * fix: restore default SGLang DSpark precision * perf(h200): add TP4 host Engram throughput comparison * perf(h200): prioritize TP4 DSpark and prevent prefill starvation * perf(h200): tune block32 kernels and extend DSpark concurrency * perf(h200): compare native Marlin MoE with stock DSpark * perf(h200): retain DSpark after non-speculative comparisons * perf(h200): retain more SWA prefixes within the KV budget * perf(h200): select measured DSpark MoE backends by TP * perf: consolidate H200 DSpark topology and launch tuning * perf: qualify H200 DSpark with exact EP1 baseline topologies * perf: select complete H200 EP1 qualification grid * fix(h200): allow full high-concurrency AgentX warmup * [CI] select Python 3.12 for result processing / 固定结果处理 Python 3.12 (#3397) * fix: select Python 3.12 for result processing 中文:在分配 GPU 前准备独立的结果处理 Python,并在固定序列和 AgentX 路径中显式使用;配置缺失时仍保留失败产物。 * fix: use managed Python for H200 AgentX power results 在 H200 AgentX DCGM 结果处理路径中校验并使用受管 Python 解释器,保留失败产物;补充回归测试和中英文恢复说明。 * docs(review): explicitly ban FP8 NextN MoE flag (#3403) 明确禁止启用 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE,并同步中英文贡献指南及审阅清单。 * [TileRT] Bump GLM-5.3 FP8 MI355X AgentX to tilert 0.1.6.post2 / GLM-5.3 FP8 MI355X AgentX 升级至 tilert 0.1.6.post2 (#3389) * [TileRT] Bump GLM-5.3 FP8 MI355X AgentX to tilert 0.1.6.post2 * Set changelog pr-link * [MI355X][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance Add the native SGLang DeepSeek V4.1 Flash DSpark recipe for MI355X, with radix caching enabled and supported long-context memory limits. Validated by official sweep 35943469224: six canonical performance cells and six full GSM8K evaluations. Reuse this sweep for production publication. * [GB200][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance Qualify the native SGLang DeepSeek V4.1 Flash DSpark recipe on GB200 with radix caching and full context. Official sweep 35876607563 passed sixteen canonical performance cells and sixteen full GSM8K evaluations across TP2/EP1 and TP4/EP1 at C1 through C128. Reuse the qualified sweep for production publication. * [H200][SGLang][AgentX] Qualify DSpark decoder SWA bounded replay (#3392) Enable decoder SWA bounded replay and pin the qualified SGLang nightly for H200 DSpark. Reuse full sweep 35893977246: 16 canonical performance cells and 16 complete GSM8K evaluations passed; final staging verified. * [H100][SGLang][AgentX] Improve DeepSeek V4.1 Flash performance (#3345) Restore the first fully qualified H100 SGLang recipe and pinned September 22 image, including TP8/EP8 and DP8 coverage. Reuse successful full sweep 35690159495. Preserve native context, prefix caching and shipped DSpark precision. * refactor: migrate fixed-sequence recipes to SRT-Slurm / 将定长配方迁移至 SRT-Slurm (#3352) * refactor: start native single-node SRT-Slurm migration 开始原生单节点 SRT-Slurm 迁移,添加 H200 SGLang 8k1k 并行候选配方并复用现有基准客户端;生产路由保持不变。 * docs: link single-node migration changelog to draft PR 为单节点迁移的 changelog 条目补充草稿 PR 链接。 * feat: wire native H200 SRT pilot into end-to-end workflow 将原生 H200 SRT 试点接入端到端工作流,显式校验配方、保留结果和功耗产物,并保持生产路由不变。 * test: cover native single-node job failures and artifacts 覆盖原生单节点作业的提交失败、Slurm 失败、产物保留及定向取消行为。 * fix: drop unused AIPerf inputs from SRT setup 删除 SRT 准备阶段未使用的 AIPerf 排空参数要求,使固定序列工作流能进入实际提交;新增缺省参数回归覆盖。 * fix: let native SRT stage the pilot container image 将配方指定镜像直接交给原生 SRT/Pyxis,移除试点对旧 squash 缓存就绪状态的依赖,保留模型预检查。 * fix: bootstrap native SRT binaries before pilot submission 在试点提交前执行原生 SRT 二进制准备,并覆盖准备失败时不提交作业的行为。 * feat: port H200 MTP and Qwen recipes to native SRT 将 H200 DeepSeek-R1 MTP 与 Qwen3.5 EP8 配方迁移到并行原生 SRT 试点,保留聊天模板及随并发变化的图捕获,并使用 4/16/64 逐点验证回归。 * docs: remove migration guide 移除独立迁移指南及入口链接,迁移范围与验证证据保留在 PR 中。 * feat: migrate fixed-sequence SGLang recipes to native SRT 将 H100、H200、B200 和 B300 的 21 个活跃 SGLang 定长配方及全部 181 个配置点切换到原生 SRT-Slurm。共享提交、eval 和产物处理,并保留逐点拓扑与推测解码设置。 * feat: migrate fixed-sequence TRT recipes to native SRT 将八个定长 TRT 配方迁移至原生 SRT,保留 63 个测试点的引擎参数、客户端和 eval token 预算。 * feat: convert remaining fixed-sequence recipes to native SRT 将剩余 AMD SGLang/ATOM 与 RTX 定长配方切换到原生 SRT 配置,保留测试点及服务参数;固定直接 ATOM 服务草稿依赖并扩展行为验证。 * fix: forward native Docker eval model identity 为 RTX 原生 Docker 路径显式传递 eval 模型名称,并验证实际启动器的命令、失败传播及容器清理。 * fix: apply AMD container options as native leaf overrides 按原生 SRT 语义逐项传递容器选项,增加映射行为回归测试并输出提交失败原因。 * refactor: reuse native job workspace for AMD scratch files 复用作业检出目录存放 AMD 原生运行时临时文件,由现有 runner 清理流程回收,移除额外 scratch-root 配置。 * fix: complete AMD native launch and terminal status handling 修复 AMD 原生启动命令的尾部换行,在 accounting 不可用时通过 Slurm 控制器验证终态,并等待启动失败的日志进程退出。 * refactor: require native recipes for the Slurm fixed-sequence cutover 移除 Docker 迁移支持和 50 个已被原生配方替代的定长 Bash 实现。Slurm 定长作业必须提供 SRT 配方,保留 AgentX、多节点和显式采集器的独立执行路径。 * refactor: make single-node fixed-sequence coverage SRT-only 归档仅支持 Docker 的 RTX 定长配置及脚本,移除无调用方的 runner 路由;保留历史 changelog 条目,但不为已归档的精确配置键生成作业。 * fix: allow client dependencies in container virtual environments 修复容器虚拟环境中的客户端依赖安装,移除不兼容的用户目录安装参数。 * fix: limit native single-node steps to serving GPUs 保留节点独占预留,同时将原生单节点服务和客户端步骤限制为配方指定的 GPU 数量,避免采集空闲设备功耗。 * fix: restore repository workdir for native container steps 原生容器步骤从已有仓库挂载目录启动,避免 Python 动态模块在根目录下导入失败,并保留调用方运行时覆盖。 * fix: stream AMD power samples without input buffering 逐行写出 AMD 功耗 CSV,避免 awk 输入缓冲在采样器关闭时丢失末尾数据;新增真实流式读取回归测试。 * fix: match registered H200 runner labels 将 H200 DGXC runner 清单对齐到实际注册的两位数字标签,恢复精确 smoke 选择和 sweep 调度。 * chore: pin srt-slurm submodule to main [skip ci] * refactor(srt): return single-node migration to NVIDIA upstream 将单节点迁移切回 NVIDIA 上游 srt-slurm。ATOM 使用原生 AToMesh 及单独固定的官方路由器镜像;TRT-LLM 使用原生模型名称转发。保留 worker 镜像、服务参数、并发范围与拓扑,不增加 Bash 分支。 * chore: bump srt-slurm submodule to v2.23.2 Move the shared SRT submodule from v2.22.1 (3cbc5dd) to upstream release v2.23.2 (8dace5f). Upstream delta is three additive changes: vllm-router multi-node hybrid-DP rank offsets, SGLang leader IP honoring the cluster network interface, and nsys/DSight coverage for SGLang and Dynamo. No recipe or connector changes required; utils/ and runners/ tests pass. * refactor(srt): check MI300X firmware in a setup hook Replace the MI300X container preamble with a host setup hook that fails on MEC firmware older than 177 instead of exporting HSA_NO_SCRATCH_RECLAIM in every container. * fix(srt): apply pending upstream srt-slurm patches at setup Post-eval steps dropped the recipe's srun_options, so ATOM evals on MI355X ran in a read-on…
Converts all 50 active Slurm single-node fixed-sequence configs to native SRT recipes: 35 SGLang, eight TRT-LLM, and seven ATOM, covering all 407 existing points. This includes DeepSeek-R1 and Qwen3.5 on H100/H200/B200/B300 and MI300X/MI325X/MI355X. Models, images, concurrency ranges, topology, and serving settings are preserved. Every converted entry exists in current main's active master configs; the deprecation cleanup in #3349 is included. The two Docker-only RTX PRO 6000 fixed-sequence configs are now retired into the existing NVIDIA archive, with original settings and scripts preserved.
Fixed-sequence jobs in the eleven Slurm pools now require an SRT recipe. There is no fallback to the former single-node Bash implementation; the 50 replaced scripts are removed. Active single-node fixed-sequence coverage is SRT-only: the two RTX Docker configs and their scripts are archived, and their unused runner mappings, launcher, and runtime settings are removed. There is no active fixed-sequence Bash/Docker fallback. Changelog planning preserves entries naming archived configs without scheduling them; unknown keys and retired append-only selections still fail. AgentX, existing multi-node paths, and explicitly selected SPEED-Bench collectors retain their current execution paths. AgentX migration will be a separate PR.
Recipes own serving and workload settings; pool profiles own scheduling, mounts, model locations, and image caches. The connector selects exactly one native variant, validates it against the matrix, and applies runtime inputs. Submission, allocation state, cancellation, result/power staging, and native provenance are checked explicitly. Native eval dispatch uses the existing InferenceX eval implementation and propagates eval or staging failures.
Concurrency-dependent settings use existing upstream SRT
zip_override_*support. Lists within a group pair by index and overridebase; they are not a Cartesian product. For example, DeepSeek-R1 H200 TRT MTP uses MTP3 without DP attention at c4–32, then MTP1 with DP attention at c64–256, with corresponding batch/token budgets. Qwen3.5 H200 SGLang varies CUDA graph batch size with concurrency. InferenceX selects the indexed variant matching the job's concurrency and topology.The SRT submodule now uses NVIDIA upstream main
3cbc5dd256af2bfd2fed09b724628c3f5456c85f, including merged AMD, MoRI, ATOM/AToMesh, and native TRT-LLM served-model-name support. This migration no longer depends on fork #27. The seven ATOM recipes select nativefrontend.type: atomeshwith a single aggregate worker. Their historical worker images remain unchanged; the router alone uses officialrocm/atom-dev:nightly_202609161445, pinned by digestsha256:f00a9dc588a380038fdad54cd57a956043cb316ddda54ebb1b62a24f8caa343bthroughfrontend.container_image. The cached May worker image does not include AToMesh; the official router image tag/digest and cached executable were verified. TRT recipes use nativeengine.served_model_nameforwarding, removing the now-duplicateextra_argsflag. No worker image, model argument, concurrency range, topology, or benchmark setting was tuned, and no Bash branches were added. The pre-existing TileRT legacy checkout remains outside this migration.Upstream-switch validation at
66781f6c0: 152 focused CPU tests passed (122 InferenceX integration checks and 30 upstream ATOM/TRT/SGLang checks). Two stubbed-launcher tests required an unsandboxed rerun because macOS denied/dev/stderr; both passed there. All 50 configs / 407 points passed native schema, topology, single-point submission and worker-command generation for throughput and eval-only modes (814 checks), including 73 AToMesh variants and 63 TRT variants with exactly one native served-name argument. The before/after recipe comparison confirms port-only changes. Changelog matrix generation and historical byte preservation passed. Fresh GPU validation of this upstream pin and AToMesh request path is pending: all ten MI355X compute nodes were allocated at the capacity check, so no new job was queued. The hardware measurements below remain evidence for their explicitly recorded earlier pins, not for this change.Validation:
Local CPU suite at
8f3f0d551after retiring the remaining Docker configs and fixing archived changelog selection: 1,991 passed, 1 skipped, and 187 subtests passed. Ruff, Bash syntax, changelog validation, and historical changelog byte preservation passed.All 407 points passed native schema, topology, submission selection, and expanded legacy/native command comparisons. Eval-only generation also passed for all 407 points. These checks do not establish hardware qualification.
Qwen3.5 H200 c4, runtime
a88a27a7cdf04d9f7b9e87492e6abdeadd38fc2c: 40/40 requests, 432.42 output tokens/s, about -0.2% versus the matching published baseline.Qwen3.5 H200 c16, same runtime: 160/160 requests, 1002.34 output tokens/s, about +0.6% versus the matching published baseline. Both runs use SGLang
v0.5.14-cu130, FP8, TP8/EP8, 8k1k, and pass required eight-GPU power validation.Qwen3.5 H200 c64, runtime
2763e8d002efd1786d7c52e30c73410c0d0803c2: 640/640 requests, 1741.50 output tokens/s, about -1.5% versus the matching published baseline. Required eight-GPU power validation passed (maximum gap 1.183 s); Slurm completed with0:0, and result collection succeeded.DeepSeek-R1 H200 c4, runtime
348f773eb6fba29cdb4216d41240cff122a4522e: 40/40 requests, 344.78 output tokens/s, about -0.04% versus the matching published baseline; required power validation passed.DeepSeek-R1 TRT MTP H200 c4, runtime
b4724b08bb1488105ff6fba32a46ad37470d6417, SRT2ac4eb1367dd2a78f597a72ca91afe4211d76b38: FP8, TRT-LLM1.3.0rc14, TP8/EP8, MTP3, 8k1k, concurrency 4. Native E2E passed: 40/40 measured requests, 416.27 output tokens/s, correct served model name, eight-GPU power validation, result collection, and allocation cleanup. This exercises the existing recipeextra_argspath without [NVIDIA] Add GB200 DSR1 FP4 TRT #26. A sequential legacy control at main7e257acec661842d9b7f5651f92b4c29ea162fd7completed 40/40 requests at 414.86 output tokens/s with identical input/output token totals: native throughput was +0.34%. The control workflow failed result discovery; its raw JSON was later included in the power-audit artifact. Offline validation of the unmodified telemetry in the verified cluster UTC timezone passed for all eight GPUs. That recovery does not change the failed workflow conclusion. This is one point on different H200 nodes, not full performance or accuracy qualification. Both allocations are released.DeepSeek-R1 TRT MTP H200 c64, runtime
b4724b08, SRT2ac4eb1: FP8, TRT-LLM1.3.0rc14, TP8/EP8 with DP attention, MTP1, 8k1k, 640/640 measured requests, 1622.04 output tokens/s. E2E, eight-GPU power validation, result collection, and cleanup passed. This verifies the higher-concurrency recipe variant; no fresh c64 control has been completed.Qwen3.5 H200 eval integration, runtime
a67435c0f8d2cb2e2d0c61c829ca4197a1368e64: SGLangv0.5.14-cu130, FP8, TP8/EP8, concurrency 4, GSM8K first 16 samples. All requests completed; strict and flexible exact match were 0.9375 (15/16). The workflow score gate failed against Qwen's 0.94 threshold. Native eval receipt was0, metadata/artifacts staged, and Slurm completed0:0. This proves the exercised execution/artifact path, not accuracy qualification; the limited-sample failure remains visible and the threshold is unchanged.MI300X Qwen3.5 c4 startup check, runtime
571fa51b3, SRTc29ef7c: FP8, SGLangv0.5.12-rocm720-mi30x, TP8/EP1, 8k1k. The allocation failed before server startup because the cluster preamble ended with a newline before the native&&join; zero measured requests. The allocation was released. Commit0bd3bf5f7fixes both affected AMD preambles and adds controller-based terminal verification for pools without working accounting; native shell execution and negative status cases passed locally. The hardware retry at8e279188ereached server readiness but failed before benchmark requests becausepip --useris incompatible with the ROCm image virtual environment. Commitc430090a8removes the forced user-site install. The AMD c4 hardware retry at6f60a8bf2, SRTc29ef7c, passed E2E: 40/40 measured requests, 238.58 output tokens/s versus 241.17 in the matching 2026-05-17 InferenceX API baseline (-1.07%). Mean TTFT was 1.0813 s versus 0.9888 s (+9.35%); mean TPOT was 15.227 ms versus 15.143 ms (+0.56%). Required power passed on exactly eight GPUs (maximum sample gap 2 s); artifacts were collected and Slurm completed0:0, with the allocation released. This historical comparison does not establish latency parity or accuracy qualification.DeepSeek-R1 B200 SGLang MTP c1, runtime
c430090a8, SRTc29ef7c: FP4, SGLangv0.5.16-cu130, TP4/EP1, 8k1k. 10/10 requests completed, 297.44 output tokens/s versus 296.69 in the matching InferenceX API baseline dated 2026-08-06 (+0.25%); mean TTFT -3.59%, mean TPOT -0.02%. Slurm completed0:0and released the allocation. The workflow failed required power validation because the exclusive eight-GPU allocation exposed eight devices to the client instead of four; this is measured throughput evidence, not a green smoke. Commit6f60a8bf2limits native server/client steps to the validated serving GPU count while retaining node exclusivity. The verification run at6f60a8bf2passed E2E: 10/10 requests, 297.74 output tokens/s (+0.36% against the same API baseline), mean TTFT -3.79%, mean TPOT -0.11%. Power validation observed exactly four GPUs (maximum sample gap 1.071 s), artifacts were collected, and Slurm completed0:0with its allocation released. The five-recipe, 16-point representative batch is paced one allocation at a time. H200 c4 has now passed, completing initial low-concurrency coverage for all five recipes. The B200 Qwen TRT-MTP c4 run completed all three TP2/TP4/TP8 variants serially with required power validation. MI355X ATOM c4 now passes required power and cleanup after the sampler fix, as detailed below. The H200 c4 smoke passed at77961ed42with required power and no evals. MI300X c16 has since passed, as detailed below. H200 c64 has also passed, including its DP attention/MTP1 settings, as detailed below. B200 c32 has also passed with exactly four power-valid GPUs, as detailed below. H200 c256 and MI300X c64 have also passed, completing both recipes’ selected low/medium/high points. B200 SGLang c256 has now passed as well, completing its selected c1/c32/c256 points. B200 Qwen TRT c128 has also passed. MI355X ATOM c32 and c256 have now passed, completing the selected batch at 16/16. The final MI355X ATOM c256 point has passed. All 16 selected points across five recipes are complete: 11,290/11,290 measured requests, required power validation and cleanup passed, and all task-owned allocations are released. The batch always ran at most one task-owned GPU allocation. These no-eval smokes do not qualify all 50 converted configs or establish full performance or accuracy parity. The Qwen TP8 c4 historical delta remains unresolved: current exact CLI selection expands all three TP variants and rejects their shared experiment name as ambiguous, so no three-variant rerun was dispatched for a single follow-up point.Qwen3.5 B200 TRT-MTP c4, TP2, runtime
6f60a8bf2, SRTc29ef7c: FP4, TRT-LLM1.3.0rc18, TP2/EP1, DP attention off, 8k1k. The benchmark job passed with 40/40 measured requests and 770.39 output tokens/s versus 769.23 in the matching 2026-06-24 API baseline (+0.15%). Mean TTFT +2.80%; mean TPOT -0.68%. Required power passed on exactly two GPUs (maximum sample gap 1.040 s); Slurm16906completed0:0and released its allocation. TP4 and TP8 also passed as recorded below; the full three-variant workflow completed successfully. No accuracy eval was requested.Qwen3.5 B200 TRT-MTP c4, TP8, runtime
6f60a8bf2, SRTc29ef7c: FP4, TRT-LLM1.3.0rc18, TP8/EP8, DP attention off, 8k1k. The benchmark job passed with 40/40 measured requests and 1042.85 output tokens/s versus 1081.62 in the matching 2026-06-24 API baseline (-3.58%). Mean TTFT +3.55%; mean TPOT +2.17%. Required power passed on eight GPUs (maximum sample gap 1.001 s); Slurm16907completed0:0and released its allocation before TP4 started. This throughput decrease is flagged for a fresh comparison; a single historical-baseline smoke cannot identify its cause or establish performance parity.Qwen3.5 B200 TRT-MTP c4, TP4, runtime
6f60a8bf2, SRTc29ef7c: FP4, TRT-LLM1.3.0rc18, TP4/EP4, DP attention off, 8k1k. 40/40 requests passed, 942.20 output tokens/s versus 917.70 in the matching 2026-06-24 API baseline (+2.67%); mean TTFT -6.64%, mean TPOT -1.68%. Required power passed on exactly four GPUs (maximum sample gap 1.001 s); Slurm16908completed0:0and released its allocation. All three c4 TP variants and result collection finished successfully.MI355X DSR1 ATOM-MTP c4 startup failure, runtime
b3a5a7f39, SRTc29ef7c: FP4, ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511, TP8/EP1, 8k1k. Python 3.12 raisedIndexErrorinimportlib.cache_from_sourcewhile importing a PyTorch generated module from/withPYTHONPYCACHEPREFIXenabled. Zero benchmark requests; no throughput or accuracy result. Slurm45548failed1:0and released its allocation. The same import-path failure was reproduced locally. Commit161aa243arestores the legacy container working-directory behavior through nativesrun_options.container-workdir=/infmax-workspace, using the existing repository mount and preserving caller overrides. All 69 focused CPU tests, native command generation, Ruff and changelog validation passed. The retry at838607d8acompleted 40/40 requests, confirming the startup fix; its separate power-validation failure is recorded below.MI355X DSR1 ATOM-MTP c4 measurement, runtime
838607d8a, SRTc29ef7c: FP4, TP8/EP1, DP attention off, 8k1k. 40/40 requests completed, 611.90 output tokens/s versus 605.22 in the matching 2026-05-20 API baseline (+1.10%); mean TTFT -2.56%, mean TPOT -0.35%. Slurm45549completed0:0and released its allocation. The workflow failed required power validation: all eight GPUs were observed, but CSV samples stopped 3.68–4.68 s before the benchmark ended (benchmark_window_not_bracketed). This is measured throughput, not a green smoke or accuracy qualification. The existing CSV filter was reproduced buffering live pipe input withmawkdespitefflush(). Commitdd955705creplaces it with a shared Bash line reader, preserving validation and benchmark settings. All 237 focused CPU tests, a Linux streaming check, Ruff, shell syntax and append-only changelog validation passed. The c4 verification run atdd955705cpassed E2E: 40/40 requests, 610.71 output tokens/s (+0.91%) against the same May 20 API baseline; mean TTFT -0.36%, mean TPOT -1.39%. Required power passed on all eight GPUs with a maximum sample gap of 2 s. Slurm45550completed0:0and released its allocation. No accuracy eval was requested. That MI355X allocation and the subsequent H200 c4 allocation are released.H200 exact selection exposed stale single-digit entries in
configs/runners.yaml:_7matched no registered_07runner. Commit77961ed42aligns both H200 labels with all 14 live registered two-digit names. Eight runner-filter CPU tests, exact one-point generation and append-only changelog validation passed. Benchmark settings are unchanged.H200 DSR1 TRT-MTP c4, previous fork pin, runtime
77961ed42, SRTc29ef7c: FP8, TRT-LLM1.3.0rc14, TP8/EP8, DP attention off, MTP3, 8k1k. E2E passed, 40/40 requests, 413.62 output tokens/s versus 413.01 in the matching 2026-05-18 API baseline (+0.15%); mean TTFT -1.97%, mean TPOT -0.36%. Required power passed on all eight GPUs (maximum sample gap 1.164 s); source revisions and artifacts were verified. Slurm88762completed0:0and released its allocation. No accuracy eval was requested.MI300X Qwen3.5 SGLang c16, runtime
77961ed42, SRTc29ef7c: FP8, SGLangv0.5.12-rocm720-mi30x, TP8/EP1, DP attention off, no speculation, 8k1k. E2E passed, 160/160 measured requests, 650.52 output tokens/s versus 652.14 in the matching 2026-05-17 API baseline (-0.25%); mean TTFT +0.89%, mean TPOT +0.21%. Required power passed on all eight GPUs (maximum sample gap 2 s). Source revisions and artifacts were verified. The native terminal-state gate and workflow cleanup passed; Slurm allocation1159was independently confirmed absent from the live queue. No accuracy eval; the historical comparison does not establish full performance parity.H200 DSR1 TRT-MTP c64, previous fork pin, runtime
77961ed42, SRTc29ef7c: FP8, TRT-LLM1.3.0rc14, TP8/EP8, DP attention enabled, MTP1, 8k1k. E2E passed, 640/640 measured requests, 1629.27 output tokens/s versus 1628.59 in the matching 2026-05-18 API baseline (+0.04%); mean TTFT -3.64%, mean TPOT +0.29%. Required power passed on all eight GPUs (maximum sample gap 1.154 s). The emitted TRT config confirms the concurrency-dependent settings. Source revisions, artifacts and cleanup were verified; Slurm88777completed0:0and released its allocation. No accuracy eval; the historical comparison does not establish full performance parity.B200 DSR1 SGLang-MTP c32, runtime
77961ed42, SRTc29ef7c: FP4, SGLangv0.5.16-cu130, TP4/EP1, DP attention off, 8k1k. E2E passed, 320/320 measured requests, 1699.06 output tokens/s versus 1717.36 in the matching 2026-08-06 API baseline (-1.07%); mean TTFT -2.17%, mean TPOT +1.20%. Required power passed on exactly four GPUs (maximum sample gap 1.072 s); native submission usessrun_options.gpus-per-node=4. Source revisions, artifacts and cleanup were verified; Slurm16910completed0:0and released its allocation. No accuracy eval; the historical comparison does not establish full performance parity.H200 DSR1 TRT-MTP c256, runtime
77961ed42, SRTc29ef7c: FP8, TRT-LLM1.3.0rc14, TP8/EP8, DP attention enabled, MTP1, 8k1k. E2E passed, 2560/2560 measured requests, 2639.48 output tokens/s versus 2636.47 in the matching 2026-05-18 API baseline (+0.11%); mean TTFT -1.69%, mean TPOT -0.02%. Required power passed on all eight GPUs (maximum sample gap 1.169 s); emitted TRT settings includemax_batch_size: 32. Source revisions, artifacts and cleanup were verified; Slurm88778completed0:0and released its allocation. The selected H200 c4/c64/c256 points have all passed; no accuracy eval or full performance qualification is claimed.MI300X Qwen3.5 SGLang c64, runtime
77961ed42, SRTc29ef7c: FP8, SGLangv0.5.12-rocm720-mi30x, TP8/EP1, DP attention off, no speculation, 8k1k. E2E passed, 640/640 measured requests, 1385.23 output tokens/s versus 1387.16 in the matching 2026-05-17 API baseline (-0.14%); mean TTFT +4.29% (2.0803 s versus 1.9947 s), mean TPOT -0.09%. Required power passed on all eight GPUs (maximum sample gap 2 s). Source revisions, artifacts and cleanup were verified; Slurm1160completed0:0and released its allocation. MI300X selected c4/c16/c64 are complete. No accuracy eval; the historical comparison does not establish full performance parity.B200 DSR1 SGLang-MTP c256, runtime
77961ed42, SRTc29ef7c: FP4, SGLangv0.5.16-cu130, TP4/EP4, DP attention enabled, 8k1k. E2E passed, 2560/2560 measured requests, 3843.97 output tokens/s versus 3904.54 in the matching 2026-08-06 API baseline (-1.55%); mean TTFT +1.89%, mean TPOT +1.68%. Server visibility and required power both covered exactly four GPUs (maximum sample gap 1.003 s). Source revisions, artifacts and cleanup were verified; Slurm16911completed0:0and released its allocation. Selected B200 SGLang c1/c32/c256 are complete, including the EP4/DP-attention transition. No accuracy eval; historical comparisons do not establish full performance parity.B200 Qwen3.5 TRT-MTP c128, runtime
77961ed42, SRTc29ef7c: FP4, TRT-LLM1.3.0rc18, TP8/EP8, DP attention enabled, 8k1k. E2E passed, 1280/1280 measured requests, 8977.02 output tokens/s versus 8801.95 in the matching 2026-06-24 API baseline (+1.99%); mean TTFT -4.04%, mean TPOT +2.34%. Required power passed on all eight GPUs (maximum sample gap 1.140 s). Source revisions, artifacts and cleanup were verified; Slurm16912completed0:0and released its allocation. This covers the selected DP-attention transition; it does not resolve the separate TP8 c4 historical throughput decrease or qualify accuracy.MI355X DSR1 ATOM-MTP c32, runtime
77961ed42, SRTc29ef7c: FP4, ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511, TP8/EP1, DP attention off, 8k1k. E2E passed, 320/320 measured requests, 2002.39 output tokens/s versus 1978.91 in the matching 2026-05-20 API baseline (+1.19%); mean TTFT +1.28%, mean TPOT -1.14%. Required power passed on all eight GPUs (maximum sample gap 2 s). Source revisions, artifacts and cleanup were verified; Slurm45584completed0:0onmia1-p01-g14and released its allocation. No accuracy eval; the historical comparison does not establish full performance parity.MI355X DSR1 ATOM-MTP c256, runtime
77961ed42, SRTc29ef7c: FP4, ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511, TP8/EP1, DP attention off, 8k1k. E2E passed, 2560/2560 measured requests, 3492.79 output tokens/s versus 3471.59 in the matching 2026-05-20 API baseline (+0.61%); mean TTFT +5.88% (2.8468 s), mean TPOT -0.91% (69.409 ms). Required power passed on all eight GPUs (maximum sample gap 2 s). Source revisions, artifacts and cleanup were verified; Slurm45587completed0:0onmia1-p01-g14and released its allocation. This completes the 16-point representative batch. No accuracy eval; the historical comparison does not establish latency or full performance parity.Keep draft. The completed representative batch passed 16/16 fixed-sequence points across five recipes on the previous fork pin
c29ef7c, with 11,290/11,290 measured requests, required power and cleanup; six additional NVIDIA points used older pins. None of those measurements qualify the new upstream/AtoMesh path. The B200 Qwen TRT TP8 c4 historical throughput delta remains -3.58% versus the June 24 API baseline. Broader recipe and pool coverage, accuracy evaluation, fresh performance comparison and cancellation under load remain incomplete. Future GPU checks stay bounded to one task-owned allocation at a time. This replaces closed #3351.中文
将全部 50 个活跃的 Slurm 单节点定长配置切换为原生 SRT 配方:35 个 SGLang、8 个 TRT-LLM、7 个 ATOM,覆盖全部 407 个现有测试点。包含 H100/H200/B200/B300 和 MI300X/MI325X/MI355X 上的 DeepSeek-R1 与 Qwen3.5。模型、镜像、并发范围、拓扑和服务设置保持一致。每个迁移条目均存在于当前 main 的活跃 master 配置中;已包含 #3349 的弃用清理。两个仅使用 Docker 的 RTX PRO 6000 定长配置现已退役至现有 NVIDIA 归档,原始设置与脚本均予保留。
十一个 Slurm 池的定长作业现在必须提供 SRT 配方,不再回退到旧版单节点 Bash 实现;50 个已替代脚本已删除。活跃的单节点定长覆盖仅使用 SRT:两个 RTX Docker 配置及其脚本已归档,不再使用的 runner 映射、启动器和运行时设置已移除,不保留活跃的定长 Bash/Docker 回退。Changelog 规划保留引用归档配置的历史条目,但不生成对应作业;未知键名及
append-only选择已退役配置仍会报错。AgentX、现有多节点路径及显式选择的 SPEED-Bench 采集器继续使用当前执行路径。AgentX 迁移将单独提交 PR。配方负责服务与工作负载设置;集群配置负责调度、挂载、模型位置及镜像缓存。连接器精确选择一个原生变体,与矩阵核对后再注入运行时参数。提交、分配状态、取消、结果与功耗产物及原生来源信息均有显式检查。原生 eval 调度复用现有 InferenceX 实现,并将评测或产物准备失败传递给作业。
随并发变化的设置使用 SRT 上游已有的
zip_override_*功能。同一组内的列表按索引配对并覆盖base,不会生成笛卡尔积。例如,DeepSeek-R1 H200 TRT MTP 在 c4–32 使用 MTP3、不启用 DP attention,在 c64–256 改用 MTP1 和 DP attention,并相应调整 batch/token 预算。Qwen3.5 H200 SGLang 则随并发调整 CUDA graph batch size。InferenceX 根据作业的并发与拓扑选择对应索引的变体。SRT 子模块现使用 NVIDIA 上游 main
3cbc5dd256af2bfd2fed09b724628c3f5456c85f,包含已合入的 AMD、MoRI、ATOM/AToMesh 及原生 TRT-LLM 服务模型名称支持。本迁移不再依赖分叉 #27。七个 ATOM 配方选择原生frontend.type: atomesh,保留一个聚合 worker。历史 worker 镜像保持不变;仅路由器通过frontend.container_image使用官方rocm/atom-dev:nightly_202609161445,并固定摘要sha256:f00a9dc588a380038fdad54cd57a956043cb316ddda54ebb1b62a24f8caa343b。已确认缓存的五月 worker 镜像不包含 AToMesh,并验证官方路由器镜像的 tag、摘要及缓存中的可执行文件。TRT 配方使用原生engine.served_model_name转发,移除现已重复的extra_args参数。未调整 worker 镜像、模型参数、并发范围、拓扑或基准设置,也未新增 Bash 分支。原有 TileRT 旧版独立检出路径不属于本次迁移。上游切换在
66781f6c0上的验证:152 项定向 CPU 测试通过,其中 122 项为 InferenceX 集成检查,30 项为上游 ATOM/TRT/SGLang 检查。两个使用模拟 launcher 的测试因 macOS 沙箱禁止访问/dev/stderr而在沙箱外重跑,均通过。全部 50 个配置 / 407 个测试点通过吞吐及 eval-only 模式的原生 schema、拓扑、单点提交与 worker 命令生成检查,共 814 项检查,涵盖 73 个 AToMesh 变体以及仅传入一次原生服务模型名称参数的 63 个 TRT 变体。配方前后对比确认仅做移植适配。Changelog 矩阵生成及历史字节保留检查通过。此上游固定版本和 AToMesh 请求路径的新实机验证仍待完成:检查时 MI355X 十个计算节点全部占用,因此未提交新排队作业。下方实机测量仅证明其明确记录的较早版本,不能作为本次切换的验证。验证:
8f3f0d551退役剩余 Docker 配置并修复归档 Changelog 选择后的本地 CPU 测试:1,991 项通过、1 项跳过,另有 187 个子测试通过。Ruff、Bash 语法、changelog 校验及历史字节保留检查通过。全部 407 个测试点通过原生 schema、拓扑、提交变体选择和展开后的旧版/原生命令对比;407 个 eval-only 配置生成检查也通过。这些检查不代表硬件验收完成。
Qwen3.5 H200 c4,运行时
a88a27a7cdf04d9f7b9e87492e6abdeadd38fc2c:40/40 请求完成,输出吞吐 432.42 tokens/s,相对匹配的已发布基线约 -0.2%。Qwen3.5 H200 c16,相同运行时:160/160 请求完成,输出吞吐 1002.34 tokens/s,相对匹配的已发布基线约 +0.6%。两次运行均使用 SGLang
v0.5.14-cu130、FP8、TP8/EP8、8k1k,八张 GPU 的强制功耗校验通过。Qwen3.5 H200 c64,运行时
2763e8d002efd1786d7c52e30c73410c0d0803c2:640/640 请求完成,输出吞吐 1741.50 tokens/s,相对匹配的已发布基线约 -1.5%。八张 GPU 的强制功耗校验通过(最大采样间隔 1.183 s);Slurm 以0:0完成,结果汇总成功。DeepSeek-R1 H200 c4,运行时
348f773eb6fba29cdb4216d41240cff122a4522e:40/40 请求完成,输出吞吐 344.78 tokens/s,相对匹配的已发布基线约 -0.04%;强制功耗校验通过。DeepSeek-R1 TRT MTP H200 c4,运行时
b4724b08bb1488105ff6fba32a46ad37470d6417、SRT2ac4eb1367dd2a78f597a72ca91afe4211d76b38:FP8、TRT-LLM1.3.0rc14、TP8/EP8、MTP3、8k1k、并发 4。原生 E2E 通过:40/40 个计量请求完成,输出吞吐 416.27 tokens/s,服务模型名称正确,八张 GPU 的功耗校验、结果汇总和资源清理均通过。这验证了现有配方extra_args路径,无需 [NVIDIA] Add GB200 DSR1 FP4 TRT #26。随后串行运行的旧版对照使用 main7e257acec661842d9b7f5651f92b4c29ea162fd7,40/40 请求完成,输出吞吐 414.86 tokens/s;输入和输出 token 总数完全一致,原生吞吐高 0.34%。对照工作流因未找到结果文件而失败,但原始 JSON 随后包含在上传的功耗审计产物中。按已核实的集群 UTC 时区离线校验未经修改的遥测,八张 GPU 均通过。恢复产物不会改变工作流失败的结论。这只是不同 H200 节点上的一个测试点,不代表完整性能或准确性验收。两个资源分配均已释放。DeepSeek-R1 TRT MTP H200 c64,运行时
b4724b08、SRT2ac4eb1:FP8、TRT-LLM1.3.0rc14、TP8/EP8、DP attention、MTP1、8k1k,640/640 个计量请求完成,输出吞吐 1622.04 tokens/s。E2E、八张 GPU 的功耗校验、结果收集及清理均通过。验证了高并发配方变体;尚未完成新的 c64 对照运行。Qwen3.5 H200 eval 集成检查,运行时
a67435c0f8d2cb2e2d0c61c829ca4197a1368e64:SGLangv0.5.14-cu130、FP8、TP8/EP8、并发 4,GSM8K 前 16 个样本。全部请求完成,strict/flexible exact match 均为 0.9375(15/16)。工作流分数门槛未通过,Qwen 要求 0.94。原生 eval 退出记录为0,元数据及产物已准备,Slurm 以0:0完成。这只验证实际执行和产物路径,不代表准确性验收;保留小样本失败记录,阈值未变。MI300X Qwen3.5 c4 启动检查,运行时
571fa51b3、SRTc29ef7c:FP8、SGLangv0.5.12-rocm720-mi30x、TP8/EP1、8k1k。集群前置命令尾部的换行与原生&&拼接冲突,资源分配在服务启动前失败;计量请求数为零。资源已释放。0bd3bf5f7修复了两个受影响的 AMD 前置命令,并在 accounting 不可用的集群上通过控制器核验终态;原生 shell 实际执行和失败状态用例已在本地通过。基于8e279188e的实机重试已到达服务就绪阶段,但pip --user与 ROCm 镜像的虚拟环境不兼容,因而在发送基准请求前失败。c430090a8移除了强制用户目录安装参数。基于6f60a8bf2、SRTc29ef7c的 AMD c4 实机重试 E2E 通过:40/40 个计量请求完成,输出吞吐 238.58 tokens/s;匹配的 2026-05-17 InferenceX API 基线为 241.17 tokens/s,变化 -1.07%。平均 TTFT 为 1.0813 s,基线为 0.9888 s(+9.35%);平均 TPOT 为 15.227 ms,基线为 15.143 ms(+0.56%)。八张 GPU 的强制功耗校验通过,最大采样间隔 2 s;产物已收集,Slurm 以0:0完成并释放资源。历史基线比较不能证明延迟完全一致,也不代表准确性验收。DeepSeek-R1 B200 SGLang MTP c1,运行时
c430090a8、SRTc29ef7c:FP4、SGLangv0.5.16-cu130、TP4/EP1、8k1k。10/10 请求完成,输出吞吐 297.44 tokens/s;匹配的 InferenceX API 基线日期为 2026-08-06,吞吐为 296.69 tokens/s,变化 +0.25%;平均 TTFT -3.59%,平均 TPOT -0.02%。Slurm 以0:0完成并释放资源。节点独占分配向客户端暴露了八张 GPU,而预期为四张,因此工作流未通过必需的功耗校验;这是已测吞吐证据,不是通过的 smoke。6f60a8bf2在保留节点独占的同时,将原生服务和客户端步骤限制为已验证的推理 GPU 数量。基于6f60a8bf2的验证运行 E2E 通过:10/10 请求完成,输出吞吐 297.74 tokens/s(相同 API 基线相比 +0.36%),平均 TTFT -3.79%,平均 TPOT -0.11%。功耗校验恰好观测到四张 GPU,最大采样间隔 1.071 s;产物已收集,Slurm 以0:0完成并释放资源。五个配方的 16 个代表点每次仅运行一个分配。H200 c4 现已通过,五个配方均已完成首轮低并发验证。B200 Qwen TRT-MTP c4 运行已串行完成 TP2/TP4/TP8,强制功耗校验均通过。修复采样器后,MI355X ATOM c4 已通过必需的功耗校验和资源清理,详情见下文。基于77961ed42的 H200 c4 smoke 已通过必需功耗校验,本次未运行 eval。MI300X c16 现已通过,详情见下文。H200 c64 也已通过,包括 DP attention/MTP1 配置,详情见下文。B200 c32 也已通过,功耗校验恰好覆盖四张 GPU,详情见下文。H200 c256 和 MI300X c64 均已通过,这两个配方选定的低、中、高并发测试点已全部完成。B200 SGLang c256 现也已通过,其选定的 c1/c32/c256 均已完成。B200 Qwen TRT c128 现也已通过。MI355X ATOM c32 和 c256 均已通过,选定批次已全部完成,达到 16/16。最后一个 MI355X ATOM c256 测试点已通过。五个配方的全部 16 个选定测试点均已完成:11,290/11,290 个计量请求完成,必需功耗校验及清理通过,本任务的全部资源分配均已释放。整个批次始终最多仅有一个 GPU 资源分配。这些未运行 eval 的 smoke 不代表全部 50 个已转换配置通过验收,也不能证明完整性能或准确性一致。Qwen TP8 c4 相对历史基线的差异仍未查明:当前 CLI 会展开三个 TP 变体,且因它们共享实验名称而拒绝精确名称选择,因此未为单个复测点重跑三个变体。Qwen3.5 B200 TRT-MTP c4、TP2,运行时
6f60a8bf2、SRTc29ef7c:FP4、TRT-LLM1.3.0rc18、TP2/EP1、不启用 DP attention、8k1k。基准作业通过,40/40 个计量请求完成,输出吞吐 770.39 tokens/s;匹配的 2026-06-24 API 基线为 769.23 tokens/s,变化 +0.15%。平均 TTFT +2.80%,平均 TPOT -0.68%。功耗校验恰好观测到两张 GPU,最大采样间隔 1.040 s;Slurm16906以0:0完成并释放资源。TP4 和 TP8 也已通过,结果如下;三个变体的完整工作流已成功结束。本次未运行准确性评测。Qwen3.5 B200 TRT-MTP c4、TP8,运行时
6f60a8bf2、SRTc29ef7c:FP4、TRT-LLM1.3.0rc18、TP8/EP8、不启用 DP attention、8k1k。基准作业通过,40/40 个计量请求完成,输出吞吐 1042.85 tokens/s;匹配的 2026-06-24 API 基线为 1081.62 tokens/s,变化 -3.58%。平均 TTFT +3.55%,平均 TPOT +2.17%。八张 GPU 的强制功耗校验通过,最大采样间隔 1.001 s;Slurm16907以0:0完成,并在 TP4 启动前释放资源。这一下降已列为后续新鲜对照的关注项;单次历史基线 smoke 不能确定原因或证明性能一致。Qwen3.5 B200 TRT-MTP c4、TP4,运行时
6f60a8bf2、SRTc29ef7c:FP4、TRT-LLM1.3.0rc18、TP4/EP4、不启用 DP attention、8k1k。40/40 请求通过,输出吞吐 942.20 tokens/s;匹配的 2026-06-24 API 基线为 917.70 tokens/s,变化 +2.67%;平均 TTFT -6.64%,平均 TPOT -1.68%。四张 GPU 的强制功耗校验通过,最大采样间隔 1.001 s;Slurm16908以0:0完成并释放资源。三个 c4 TP 变体及结果收集均已成功结束。MI355X DSR1 ATOM-MTP c4 启动失败,运行时
b3a5a7f39、SRTc29ef7c:FP4、ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511、TP8/EP1、8k1k。启用PYTHONPYCACHEPREFIX且从/启动时,Python 3.12 在导入 PyTorch 动态生成模块的importlib.cache_from_source路径抛出IndexError。基准请求数为零,无吞吐或准确性结果。 Slurm45548以1:0失败并释放资源。已在本地复现相同导入路径错误。161aa243a通过原生srun_options.container-workdir=/infmax-workspace恢复旧版容器工作目录行为,使用现有仓库挂载并保留调用方覆盖。69 项定向 CPU 测试、原生命令生成、Ruff 和 changelog 校验均通过。基于838607d8a的重试完成了 40/40 个请求,确认启动问题已修复;另一个功耗校验失败记录如下。MI355X DSR1 ATOM-MTP c4 测量,运行时
838607d8a、SRTc29ef7c:FP4、TP8/EP1、不启用 DP attention、8k1k。40/40 请求完成,输出吞吐 611.90 tokens/s;匹配的 2026-05-20 API 基线为 605.22 tokens/s,变化 +1.10%;平均 TTFT -2.56%,平均 TPOT -0.35%。Slurm45549以0:0完成并释放资源。工作流未通过必需的功耗校验:八张 GPU 均有数据,但 CSV 样本在基准结束前 3.68–4.68 s 停止(benchmark_window_not_bracketed)。这是吞吐测量结果,不是通过的 smoke 或准确性验收。已复现现有 CSV 过滤器在mawk下即使调用fflush()仍缓冲管道输入的问题。dd955705c将其替换为共享 Bash 逐行读取器,保留功耗校验和基准设置。237 项定向 CPU 测试、Linux 流式检查、Ruff、shell 语法及追加式 changelog 校验均通过。基于dd955705c的 c4 验证运行 E2E 通过:40/40 请求完成,输出吞吐 610.71 tokens/s(+0.91%),对比相同的 5 月 20 日 API 基线;平均 TTFT -0.36%,平均 TPOT -1.39%。八张 GPU 的必需功耗校验通过,最大采样间隔为 2 s。Slurm45550以0:0完成并释放资源。本次未运行准确性评测,该 MI355X 分配及后续 H200 c4 分配均已释放。H200 精确选择暴露了
configs/runners.yaml中过期的单数字条目:_7无法匹配实际注册的_07runner。77961ed42将两个 H200 标签清单与 14 个在线注册的两位数字名称对齐。8 项 runner 过滤 CPU 测试、精确单点生成及追加式 changelog 校验均通过,基准设置未变。H200 DSR1 TRT-MTP c4,此前分叉固定版本,运行时
77961ed42、SRTc29ef7c:FP8、TRT-LLM1.3.0rc14、TP8/EP8、不启用 DP attention、MTP3、8k1k。E2E 通过,40/40 请求完成,输出吞吐 413.62 tokens/s;匹配的 2026-05-18 API 基线为 413.01 tokens/s,变化 +0.15%;平均 TTFT -1.97%,平均 TPOT -0.36%。八张 GPU 的必需功耗校验通过,最大采样间隔 1.164 s;源版本和产物均已核实。Slurm88762以0:0完成并释放资源,本次未运行准确性评测。MI300X Qwen3.5 SGLang c16,运行时
77961ed42、SRTc29ef7c:FP8、SGLangv0.5.12-rocm720-mi30x、TP8/EP1、不启用 DP attention、无推测解码、8k1k。E2E 通过,160/160 个计量请求完成,输出吞吐 650.52 tokens/s;匹配的 2026-05-17 API 基线为 652.14 tokens/s,变化 -0.25%;平均 TTFT +0.89%,平均 TPOT +0.21%。八张 GPU 的必需功耗校验通过,最大采样间隔 2 s。源版本及产物已核实;原生终态校验与工作流清理均通过,另行检查实时队列确认 Slurm 分配1159已释放。本次未运行准确性评测,历史基线比较不代表完整性能验收。H200 DSR1 TRT-MTP c64,此前分叉固定版本,运行时
77961ed42、SRTc29ef7c:FP8、TRT-LLM1.3.0rc14、TP8/EP8、启用 DP attention、MTP1、8k1k。E2E 通过,640/640 个计量请求完成,输出吞吐 1629.27 tokens/s;匹配的 2026-05-18 API 基线为 1628.59 tokens/s,变化 +0.04%;平均 TTFT -3.64%,平均 TPOT +0.29%。八张 GPU 的必需功耗校验通过,最大采样间隔 1.154 s。实际生成的 TRT 配置确认了随并发变化的设置。源版本、产物与清理均已核实;Slurm88777以0:0完成并释放资源。本次未运行准确性评测,历史基线比较不代表完整性能验收。B200 DSR1 SGLang-MTP c32,运行时
77961ed42、SRTc29ef7c:FP4、SGLangv0.5.16-cu130、TP4/EP1、不启用 DP attention、8k1k。E2E 通过,320/320 个计量请求完成,输出吞吐 1699.06 tokens/s;匹配的 2026-08-06 API 基线为 1717.36 tokens/s,变化 -1.07%;平均 TTFT -2.17%,平均 TPOT +1.20%。必需功耗校验恰好覆盖四张 GPU,最大采样间隔 1.072 s;原生提交使用srun_options.gpus-per-node=4。源版本、产物与清理均已核实;Slurm16910以0:0完成并释放资源。本次未运行准确性评测,历史基线比较不代表完整性能验收。H200 DSR1 TRT-MTP c256,运行时
77961ed42、SRTc29ef7c:FP8、TRT-LLM1.3.0rc14、TP8/EP8、启用 DP attention、MTP1、8k1k。E2E 通过,2560/2560 个计量请求完成,输出吞吐 2639.48 tokens/s;匹配的 2026-05-18 API 基线为 2636.47 tokens/s,变化 +0.11%;平均 TTFT -1.69%,平均 TPOT -0.02%。八张 GPU 的必需功耗校验通过,最大采样间隔 1.169 s;实际 TRT 设置包含max_batch_size: 32。源版本、产物与清理均已核实;Slurm88778以0:0完成并释放资源。选定的 H200 c4/c64/c256 均已通过;本次未运行准确性评测,也不代表完整性能验收。MI300X Qwen3.5 SGLang c64,运行时
77961ed42、SRTc29ef7c:FP8、SGLangv0.5.12-rocm720-mi30x、TP8/EP1、不启用 DP attention、无推测解码、8k1k。E2E 通过,640/640 个计量请求完成,输出吞吐 1385.23 tokens/s;匹配的 2026-05-17 API 基线为 1387.16 tokens/s,变化 -0.14%;平均 TTFT +4.29%(2.0803 s 对比 1.9947 s),平均 TPOT -0.09%。八张 GPU 的必需功耗校验通过,最大采样间隔 2 s。源版本、产物与清理均已核实;Slurm1160以0:0完成并释放资源。选定的 MI300X c4/c16/c64 已全部通过。本次未运行准确性评测,历史基线比较不代表完整性能验收。B200 DSR1 SGLang-MTP c256,运行时
77961ed42、SRTc29ef7c:FP4、SGLangv0.5.16-cu130、TP4/EP4、启用 DP attention、8k1k。E2E 通过,2560/2560 个计量请求完成,输出吞吐 3843.97 tokens/s;匹配的 2026-08-06 API 基线为 3904.54 tokens/s,变化 -1.55%;平均 TTFT +1.89%,平均 TPOT +1.68%。服务可见设备与必需功耗校验均恰好覆盖四张 GPU,最大采样间隔 1.003 s。源版本、产物与清理均已核实;Slurm16911以0:0完成并释放资源。选定的 B200 SGLang c1/c32/c256 已全部完成,覆盖 EP4/DP attention 配置切换。本次未运行准确性评测,历史基线比较不代表完整性能验收。B200 Qwen3.5 TRT-MTP c128,运行时
77961ed42、SRTc29ef7c:FP4、TRT-LLM1.3.0rc18、TP8/EP8、启用 DP attention、8k1k。E2E 通过,1280/1280 个计量请求完成,输出吞吐 8977.02 tokens/s;匹配的 2026-06-24 API 基线为 8801.95 tokens/s,变化 +1.99%;平均 TTFT -4.04%,平均 TPOT +2.34%。八张 GPU 的必需功耗校验通过,最大采样间隔 1.140 s。源版本、产物与清理均已核实;Slurm16912以0:0完成并释放资源。本次覆盖选定的 DP attention 配置切换,但不能解释 TP8 c4 相对历史基线的吞吐下降,也不代表准确性验收。MI355X DSR1 ATOM-MTP c32,运行时
77961ed42、SRTc29ef7c:FP4、ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511、TP8/EP1、不启用 DP attention、8k1k。E2E 通过,320/320 个计量请求完成,输出吞吐 2002.39 tokens/s;匹配的 2026-05-20 API 基线为 1978.91 tokens/s,变化 +1.19%;平均 TTFT +1.28%,平均 TPOT -1.14%。八张 GPU 的必需功耗校验通过,最大采样间隔 2 s。源版本、产物与清理均已核实;Slurm45584在mia1-p01-g14上以0:0完成并释放资源。本次未运行准确性评测,历史基线比较不代表完整性能验收。MI355X DSR1 ATOM-MTP c256,运行时
77961ed42、SRTc29ef7c:FP4、ATOMrocm7.2.3_ubuntu24.04_py3.12_pytorch_release_2.10.0_atom20260511、TP8/EP1、不启用 DP attention、8k1k。E2E 通过,2560/2560 个计量请求完成,输出吞吐 3492.79 tokens/s;匹配的 2026-05-20 API 基线为 3471.59 tokens/s,变化 +0.61%;平均 TTFT +5.88%(2.8468 s),平均 TPOT -0.91%(69.409 ms)。八张 GPU 的必需功耗校验通过,最大采样间隔 2 s。源版本、产物与清理均已核实;Slurm45587在mia1-p01-g14上以0:0完成并释放资源。至此 16 个代表性测试点全部完成。本次未运行准确性评测,历史基线比较不能证明延迟或完整性能一致。保持草稿。此前分叉固定版本
c29ef7c上的代表性批次已通过五个配方的 16/16 个固定序列长度测试点,11,290/11,290 个计量请求完成,必需功耗及清理检查通过;另六个 NVIDIA 测试点使用更早版本。这些测量均不能验收新的上游/AtoMesh 路径。B200 Qwen TRT TP8 c4 相对 6 月 24 日 API 历史基线的吞吐差异仍为 -3.58%。更广的配方与集群池覆盖、准确性评测、新鲜性能对照以及负载下取消验证仍未完成。后续 GPU 检查继续限制为同一时间至多一个本任务分配。本 PR 替代已关闭的 #3351。