Skip to content

[Klaud Cold] Update qwen3.5-fp4-b200-sglang-mtp SGLang image to v0.5.20-cu130 / 将 qwen3.5-fp4-b200-sglang-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 - #3411

Closed
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-fbf1925cac848db7-ea7bf153004ad90b
Closed

Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-fbf1925cac848db7-ea7bf153004ad90b

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Goal: Update SGLang image from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.20-cu130.
Baseline: 2026-09-01 · lmsysorg/sglang:v0.5.14-cu130
Mean latency · Sources: API 1, API 2

Point Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
8k/1k c4 TP2 EP1 8eb5f1 3,918.19 435.75 367.36 4.11
8k/1k c4 TP2 EP1 cf6161 N/A N/A N/A N/A
8k/1k c4 TP4 EP1 381171 2,327.98 258.9 369.08 3.39
8k/1k c4 TP4 EP1 8ee4bc N/A N/A N/A N/A
8k/1k c8 TP2 EP1 7e1f79 5,435.3 610.85 644.43 5.67
8k/1k c8 TP2 EP1 92b789 N/A N/A N/A N/A
8k/1k c16 TP2 EP1 753299 N/A N/A N/A N/A
8k/1k c16 TP2 EP1 a1e139 7,332.27 812.23 914.61 8.57
8k/1k c16 TP2 EP2 3b18fb 7,645.5 846.93 792.16 8.33
8k/1k c16 TP2 EP2 5d2eb9 N/A N/A N/A N/A
8k/1k c32 TP2 EP1 0f0493 5,393.35 602.95 1,111.01 25.04
8k/1k c32 TP2 EP1 704873 N/A N/A N/A N/A
8k/1k c32 TP2 EP2 43060d N/A N/A N/A N/A
8k/1k c32 TP2 EP2 6a22e8 10,257.18 1,146.7 1,145.02 12.38
8k/1k c64 TP2 EP1 496e07 N/A N/A N/A N/A
8k/1k c64 TP2 EP1 5a1a9a 12,729.04 1,411.77 1,951.9 20.16
8k/1k c64 TP2 EP2 095c9a N/A N/A N/A N/A
8k/1k c64 TP2 EP2 94c355 13,668.76 1,515.99 1,808.23 18.77

Note: 8k/1k c4 TP2 EP1 cf6161, 8k/1k c4 TP4 EP1 8ee4bc, 8k/1k c8 TP2 EP1 92b789, 8k/1k c16 TP2 EP1 753299, 8k/1k c16 TP2 EP2 5d2eb9, 8k/1k c32 TP2 EP1 704873, 8k/1k c32 TP2 EP2 43060d, 8k/1k c64 TP2 EP1 496e07, 8k/1k c64 TP2 EP2 095c9a: unavailable.

Eval: N/A

中文

**目标:**将 SGLang 镜像从 lmsysorg/sglang:v0.5.14-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130。
**基线:**2026-09-01 · lmsysorg/sglang:v0.5.14-cu130
平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。

Update the qwen3.5-fp4-b200-sglang-mtp master image and its srt-slurm recipe
container from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.20-cu130
(digest sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f,
sglang commit 94602c9c2b7cbdb8efd5c52802dac6a1c180089e). Rename the recipe's
cuda-graph-max-bs setting to cuda-graph-max-bs-decode because v0.5.20 removed
the deprecated alias (sgl-project/sglang#38375); the alias already stored into
cuda_graph_max_bs_decode on v0.5.14, so the captured decode batch sizes are
unchanged. Model, topology, EAGLE/MTP settings, workload and all points are
unchanged.

将 qwen3.5-fp4-b200-sglang-mtp 的主镜像及其 srt-slurm 配方容器从
lmsysorg/sglang:v0.5.14-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130。由于
v0.5.20 移除了已弃用的 cuda-graph-max-bs 别名(sgl-project/sglang#38375),
将配方中的该设置重命名为 cuda-graph-max-bs-decode;其取值与语义不变。模型、
拓扑、EAGLE/MTP 设置、负载与全部测点均保持不变。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@Klaud-Cold

Klaud-Cold commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator Author

Initial attempt · Passed · Run 36043528375 / attempt 1 · 2026-09-24 19:14 UTC
lmsysorg/sglang:v0.5.20-cu130 · a8b6c33148e9 · Mean latency
Change: Bump lmsysorg/sglang:v0.5.14-cu130 (commit 49e384ce, FlashInfer 0.6.12) to v0.5.20-cu130 (tag commit 94602c9c equals the image build label, digest sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f, sgl-kernel 0.4.7, FlashInfer 0.6.18) and rename the recipe's cuda-graph-max-bs to cuda-graph-max-bs-decode because sgl-project/sglang#38375 removed that deprecated alias, whose dest was already cuda_graph_max_bs_decode on v0.5.14; all other flags, backend choices and the BF16 MTP-head handling for this modelopt_mixed checkpoint are unchanged per sgl-project/sglang@v0.5.14...v0.5.20.

Point Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
8k/1k c4 TP4 EP1 8ee4bc 328.23 (N/A) 220.47 (N/A) 2.75 (N/A)
8k/1k c4 TP2 EP1 cf6161 495.47 (N/A) 293.52 (N/A) 3.62 (N/A)
8k/1k c16 TP2 EP2 5d2eb9 976.62 (N/A) 679.89 (N/A) 7.24 (N/A)

Note: All rows: request errors unavailable.
Note: All rows: Δ N/A: point unavailable.

Eval Score ↑ Samples
gsm8k/em_strict · c32 96.97% (N/A) N/A/1,319 (old/new)
gsm8k/em_strict · c32 96.66% (N/A) N/A/1,319 (old/new)

Note: All rows: Δ N/A: no matched eval baseline.

Next: Finish with outcome failed and release the family: the frozen 2026-09-01 roster contains nine pre-srt-slurm identities (c4 TP4 EP1 381171, c4 TP2 EP1 8eb5f1, c8 TP2 EP1 7e1f79, c16 TP2 EP1 a1e139, c32 TP2 EP1 0f0493, c64 TP2 EP1 5a1a9a, c16 TP2 EP2 3b18fb, c32 TP2 EP2 6a22e8, c64 TP2 EP2 94c355) that check-final cannot find in the current family after #3352, so no final sweep is dispatched.

中文

初次尝试 · 已通过 · Run 36043528375 / attempt 1 · 2026-09-24 19:14 UTC
lmsysorg/sglang:v0.5.20-cu130 · a8b6c33148e9 · 平均延迟
**变更:**将 lmsysorg/sglang:v0.5.14-cu130(提交 49e384ce)更新为 v0.5.20-cu130(标签提交 94602c9c 与镜像构建标签一致,摘要 sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f,sgl-kernel 0.4.7,FlashInfer 0.6.18),并因 sgl-project/sglang#38375 移除了已弃用别名而将配方中的 cuda-graph-max-bs 重命名为 cuda-graph-max-bs-decode(v0.5.14 上该别名的目标字段已是 cuda_graph_max_bs_decode,取值不变);其余参数、后端选项以及该 modelopt_mixed 检查点的 BF16 MTP 头处理在两个标签间均未变化。实测数值及异常说明见上表。
**下一步:**以 failed 结果收尾并释放该配置族:冻结的 2026-09-01 基线包含九个 srt-slurm 迁移前的测点标识(c4 TP4 EP1 381171、c4 TP2 EP1 8eb5f1、c8 TP2 EP1 7e1f79、c16 TP2 EP1 a1e139、c32 TP2 EP1 0f0493、c64 TP2 EP1 5a1a9a、c16 TP2 EP2 3b18fb、c32 TP2 EP2 6a22e8、c64 TP2 EP2 94c355),在 #3352 之后 check-final 无法在当前配置族中找到它们,因此不触发最终 sweep。

….20-cu130 bump

Append the perf-changelog entry for updating the qwen3.5-fp4-b200-sglang-mtp
SGLang image from v0.5.14-cu130 to v0.5.20-cu130 with the cuda-graph-max-bs to
cuda-graph-max-bs-decode rename, after the targeted smoke run 36043528375
passed its three throughput points and both GSM8K evals.

为 qwen3.5-fp4-b200-sglang-mtp 追加 perf-changelog 条目:SGLang 镜像从
v0.5.14-cu130 更新为 v0.5.20-cu130,并将 cuda-graph-max-bs 重命名为
cuda-graph-max-bs-decode;定向 smoke 运行 36043528375 的三个吞吐测点与两个
GSM8K 评测均已通过。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Cleanup pending. Stop and confirm owned runs before closing.

中文

failed · 清理待完成。先停止并确认自有运行结束,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 24, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-fbf1925cac848db7-ea7bf153004ad90b branch September 24, 2026 19:15
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Repairs: 0 · Runs: 36043528375, 36046391915
All owned runs ended. PR closed; branch deleted for retry.

中文

failed · 修复次数:0 · 运行:36043528375, 36046391915
所有自有运行均已结束。PR 已关闭;分支已删除,可重新尝试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant