Skip to content

[Klaud Cold] Update qwen3.5-fp4-b200-sglang SGLang image to v0.5.20-cu130 / 将 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.20-cu130 - #3412

Closed
Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-1fa5c7d01fd1a2fd-ed0569e9ff1aef5d
Closed

Klaud-Cold wants to merge 1 commit into
mainfrom
klaud/auto-1fa5c7d01fd1a2fd-ed0569e9ff1aef5d

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Goal: Update SGLang image from lmsysorg/sglang:v0.5.19-cu130 to lmsysorg/sglang:v0.5.20-cu130.
Baseline: 2026-09-10 · lmsysorg/sglang:v0.5.19-cu130
Mean latency · Sources: API 1, API 2

Point Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
8k/1k c4 TP2 EP1 4f6a03 2,830.4 314.86 253.64 5.94
8k/1k c4 TP2 EP1 fc5d95 N/A N/A N/A N/A
8k/1k c4 TP4 EP1 ac33ee 1,705.01 189.67 179.37 4.96
8k/1k c4 TP4 EP1 ee8a0f N/A N/A N/A N/A
8k/1k c8 TP2 EP1 67aca8 N/A N/A N/A N/A
8k/1k c8 TP2 EP1 b85636 4,431.64 498.19 400.3 7.47
8k/1k c16 TP2 EP1 2ac0c0 6,442.34 713.85 508.12 10.36
8k/1k c16 TP2 EP1 db75ed N/A N/A N/A N/A
8k/1k c32 TP2 EP1 94d9ff N/A N/A N/A N/A
8k/1k c32 TP2 EP1 a2240f 8,753.16 978.85 784.47 15.14
8k/1k c64 TP2 EP1 73e4d1 11,509.37 1,276.86 1,269.76 23.28
8k/1k c64 TP2 EP1 fce6c1 N/A N/A N/A N/A
8k/1k c128 TP2 EP1 8f48a1 N/A N/A N/A N/A
8k/1k c128 TP2 EP1 d2b2e2 14,443.07 1,599.79 2,088.39 37.18

Note: 8k/1k c4 TP2 EP1 fc5d95, 8k/1k c4 TP4 EP1 ee8a0f, 8k/1k c8 TP2 EP1 67aca8, 8k/1k c16 TP2 EP1 db75ed, 8k/1k c32 TP2 EP1 94d9ff, 8k/1k c64 TP2 EP1 fce6c1, 8k/1k c128 TP2 EP1 8f48a1: unavailable.

Eval: N/A

中文

**目标:**将 SGLang 镜像从 lmsysorg/sglang:v0.5.19-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130。
**基线:**2026-09-10 · lmsysorg/sglang:v0.5.19-cu130
平均延迟 · 来源: API 1, API 2;数值及异常说明见上表。

Move qwen3.5-fp4-b200-sglang and its 8k1k SRT recipe from
lmsysorg/sglang:v0.5.19-cu130 to lmsysorg/sglang:v0.5.20-cu130 and rename
the removed --cuda-graph-max-bs alias to --cuda-graph-max-bs-decode.

将 qwen3.5-fp4-b200-sglang 及其 8k1k SRT 配方的镜像从
lmsysorg/sglang:v0.5.19-cu130 更新为 lmsysorg/sglang:v0.5.20-cu130,
并将已移除的 --cuda-graph-max-bs 别名改为 --cuda-graph-max-bs-decode。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Initial 0/5 · deferred (no dispatch) · frozen-baseline coverage cannot be satisfied
lmsysorg/sglang:v0.5.20-cu130 · b47a5d0d6c6277ee49e86f40a9784b2a749437f6 · 8k/1k · TP4/EP1 + TP2/EP1 · Mean latency
Change: Move the master image and the recipe model.container from v0.5.19-cu130 to v0.5.20-cu130 and rename cuda-graph-max-bs to cuda-graph-max-bs-decode, because sgl-project/sglang#38375 removed the deprecated alias between the two tags (v0.5.19 alias, v0.5.20 field).

Upstream evidence. v0.5.20 (2026-09-18, tag commit 94602c9c2b7cbdb8efd5c52802dac6a1c180089e) is the newest release; the v0.5.20-cu130 index digest is sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f (amd64 sha256:b27fce60bc5494c118c4910702812bcfa8cee67abcdd1ff8b0902f21647552f4, image label ai.sglang.build.commit = the tag commit, CUDA 13.0.3, FlashInfer 0.6.18, sglang-kernel 0.4.6.post1 → 0.4.7, sgl-deep-gemm 0.1.7 → 0.2.0, torch 2.13.0 and transformers 5.12.1 unchanged). Every other recipe flag and choice (trtllm_mha, flashinfer_trtllm, modelopt_fp4, fp8_e4m3, enable-symm-mem, mamba-ssm-dtype, scheduler-recv-interval, tokenizer-worker-num) still exists in the v0.5.20 field declarations. The digest is recorded rather than written into the config because the B200 Nscale native single-node path passes the raw image to Pyxis on a squash-cache miss and this pool's Enroot does not accept Docker's @sha256: form (launcher); the sibling B200 refresh #3334 uses the same plain-tag spelling. The pre-existing runners/srt-slurm/patches/504-post-eval-srun-options.patch patches the srt-slurm orchestrator only, not the engine.

Blocker. The frozen baseline (2026-09-10, producer run 34452191302, head baed6629) holds 14 points: 7 with published results from the pre-SRT recipe and 7 unpublished points for the current SRT recipe. Point identities differ only in the srt-recipe field that #3352 (merged 2026-09-24) added to every fixed-sequence point, so no matrix generated from this head can contain the 7 published keys. check-final and finish require every frozen key, so final validation would fail regardless of benchmark results. Affected points (all 8k/1k): c4 TP2 EP1 4f6a03, c4 TP4 EP1 ac33ee, c8 TP2 EP1 b85636, c16 TP2 EP1 2ac0c0, c32 TP2 EP1 a2240f, c64 TP2 EP1 73e4d1, c128 TP2 EP1 d2b2e2. No GPU work was dispatched; the same condition applies to every fixed-sequence family migrated by #3352 until either the baseline point identity normalizes srt-recipe or a post-migration result is published for the old image. Published evals for this family exist only for 2026-04-08, so eval deltas would have been N/A.

Next: Finish with failed so the family returns to the pool once the roster identity or published baseline is updated.

中文

初始 0/5 · 已推迟(未调度) · 冻结基线覆盖无法满足
lmsysorg/sglang:v0.5.20-cu130 · b47a5d0d6c6277ee49e86f40a9784b2a749437f6 · 8k/1k · TP4/EP1 + TP2/EP1 · 平均延迟
**变更:**将主配置镜像和配方 model.container 从 v0.5.19-cu130 更新为 v0.5.20-cu130,并将 cuda-graph-max-bs 改为 cuda-graph-max-bs-decode,因为 sgl-project/sglang#38375 在两个标签之间移除了该弃用别名。上游证据、镜像摘要及未写入配置的原因见上文。
**阻塞:**冻结基线包含 7 个来自迁移前配方的已发布点,其标识与当前 SRT 配方仅相差 #3352 新增的 srt-recipe 字段,因此本分支生成的任何矩阵都无法包含这些键,check-final 与 finish 必然失败。受影响的点见上文列表;未调度任何 GPU 任务。
**下一步:**以 failed 结束,待基线点标识或已发布基线更新后让该系列回到候选池。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Cleanup pending. Stop and confirm owned runs before closing.

中文

failed · 清理待完成。先停止并确认自有运行结束,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 24, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-1fa5c7d01fd1a2fd-ed0569e9ff1aef5d branch September 24, 2026 18:48
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

failed · Repairs: 0 · Runs: —
All owned runs ended. PR closed; branch deleted for retry.

中文

failed · 修复次数:0 · 运行:—
所有自有运行均已结束。PR 已关闭;分支已删除,可重新尝试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant