[Klaud Cold] Update qwen3.5-fp4-b200-sglang-mtp SGLang image to v0.5.20-cu130 / 将 qwen3.5-fp4-b200-sglang-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 - #3411
Conversation
Update the qwen3.5-fp4-b200-sglang-mtp master image and its srt-slurm recipe container from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.20-cu130 (digest sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f, sglang commit 94602c9c2b7cbdb8efd5c52802dac6a1c180089e). Rename the recipe's cuda-graph-max-bs setting to cuda-graph-max-bs-decode because v0.5.20 removed the deprecated alias (sgl-project/sglang#38375); the alias already stored into cuda_graph_max_bs_decode on v0.5.14, so the captured decode batch sizes are unchanged. Model, topology, EAGLE/MTP settings, workload and all points are unchanged. 将 qwen3.5-fp4-b200-sglang-mtp 的主镜像及其 srt-slurm 配方容器从 lmsysorg/sglang:v0.5.14-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130。由于 v0.5.20 移除了已弃用的 cuda-graph-max-bs 别名(sgl-project/sglang#38375), 将配方中的该设置重命名为 cuda-graph-max-bs-decode;其取值与语义不变。模型、 拓扑、EAGLE/MTP 设置、负载与全部测点均保持不变。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
|
Initial attempt · Passed · Run 36043528375 / attempt 1 · 2026-09-24 19:14 UTC
Note: All rows: request errors unavailable.
Note: All rows: Δ N/A: no matched eval baseline. Next: Finish with outcome failed and release the family: the frozen 2026-09-01 roster contains nine pre-srt-slurm identities (c4 TP4 EP1 381171, c4 TP2 EP1 8eb5f1, c8 TP2 EP1 7e1f79, c16 TP2 EP1 a1e139, c32 TP2 EP1 0f0493, c64 TP2 EP1 5a1a9a, c16 TP2 EP2 3b18fb, c32 TP2 EP2 6a22e8, c64 TP2 EP2 94c355) that check-final cannot find in the current family after #3352, so no final sweep is dispatched. 中文初次尝试 · 已通过 · Run 36043528375 / attempt 1 · 2026-09-24 19:14 UTC |
….20-cu130 bump Append the perf-changelog entry for updating the qwen3.5-fp4-b200-sglang-mtp SGLang image from v0.5.14-cu130 to v0.5.20-cu130 with the cuda-graph-max-bs to cuda-graph-max-bs-decode rename, after the targeted smoke run 36043528375 passed its three throughput points and both GSM8K evals. 为 qwen3.5-fp4-b200-sglang-mtp 追加 perf-changelog 条目:SGLang 镜像从 v0.5.14-cu130 更新为 v0.5.20-cu130,并将 cuda-graph-max-bs 重命名为 cuda-graph-max-bs-decode;定向 smoke 运行 36043528375 的三个吞吐测点与两个 GSM8K 评测均已通过。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
|
failed · Cleanup pending. Stop and confirm owned runs before closing. 中文failed · 清理待完成。先停止并确认自有运行结束,再关闭 PR。 |
|
failed · Repairs: 0 · Runs: 36043528375, 36046391915 中文failed · 修复次数:0 · 运行:36043528375, 36046391915 |
Goal: Update SGLang image from
lmsysorg/sglang:v0.5.14-cu130tolmsysorg/sglang:v0.5.20-cu130.Baseline: 2026-09-01 ·
lmsysorg/sglang:v0.5.14-cu130Mean latency · Sources: API 1, API 2
Note: 8k/1k c4 TP2 EP1 cf6161, 8k/1k c4 TP4 EP1 8ee4bc, 8k/1k c8 TP2 EP1 92b789, 8k/1k c16 TP2 EP1 753299, 8k/1k c16 TP2 EP2 5d2eb9, 8k/1k c32 TP2 EP1 704873, 8k/1k c32 TP2 EP2 43060d, 8k/1k c64 TP2 EP1 496e07, 8k/1k c64 TP2 EP2 095c9a: unavailable.
Eval: N/A
中文
**目标:**将 SGLang 镜像从
lmsysorg/sglang:v0.5.14-cu130更新为lmsysorg/sglang:v0.5.20-cu130。**基线:**2026-09-01 ·
lmsysorg/sglang:v0.5.14-cu130平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。