Skip to content

[Klaud Cold] Retire GLM-5.1 and its B200 TileRT config / [Klaud Cold] 退役 GLM-5.1 及其 B200 TileRT 配置 - #3638

Merged
functionstackx merged 2 commits into
mainfrom
klaud/retire-glm51
Oct 1, 2026
Merged

functionstackx merged 2 commits into
mainfrom
klaud/retire-glm51

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Description

Retires GLM-5.1 completely. glm5.1-fp8-b200-tilert was the last GLM-5.1 config and the last Single-turn 1k1k config, so this also ends 1k1k coverage.

Removed

  • glm5.1-fp8-b200-tilert from nvidia-master.yaml (1k1k + 8k1k, B200 Nscale, vLLM prefill + TileRT 0.1.5 decode)
  • benchmarks/multi_node/srt-slurm-recipes/glm5.1/ (both recipes) and srt-slurm-recipes/configs/tilert-b200-setup.sh
  • B200 Nscale launcher rows used only by this config:
    • policy.NATIVE_SRT_LANES GLM-5.1 TileRT match
    • power.py GLM-5.1 fixed-sequence DCGM rule
    • models.OVERRIDES["b200-nscale"] (GLM-5.1-FP8@shared)
    • TILERT_ENV (its only row was b200-nscale's UCX settings); runtime_env now just layers the cluster env and the given settings
    • the /tilert_weights lane mount
  • runners.yaml b200-nscale: GLM-5.1-FP8 / GLM-5.1-FP8@shared entries and the now-unused shared-models / tilert-weights volumes
  • glm5.1 gsm8k eval threshold
  • Docs: the "TileRT fixed-sequence recipes" section (EN/ZH), the "use 1k1k only for GLM-5.1" notes in configuration-procedures / testing (EN/ZH), the stale tilert-weights volume name in CONFIGS.md, and the GLM-5.1 exception in the claude.yml agent prompts and .claude/commands/add-model-hardware.md
  • MODELS.md / MODELS_zh.md: the 1k1k scenario row and the GLM-5 / GLM-5.1 support-matrix row now record the retirement, plus a dated retirement note

Kept

  • Generic TileRT support (driver, synthetic acceptance, tilert framework handling), still used by glm5.3-fp8-mi355x-tilert-agentic. The TileRT driver test moved from the GLM-5.1 B200 lane to the GLM-5.3 MI355X lane instead of being deleted.
  • srt_fixed_sequence.sh, shared with the DSR1 B200 recipes
  • Historical text: README news items, the dated 2026-09-21 parity audit, and perf-changelog.yaml
  • operatorx/testlists/*.json GLM-5.1 kernel shapes (OperatorX microbenchmarks, not InferenceX configs)

No perf-changelog entry, matching earlier retirement PRs (#3618, #3595, #3467). No benchmark is affected, so no sweep label.

Validation

  • infx.matrix.generate full-sweep over both master configs: 1715 rows, 0 glm5.1, 0 ISL-1024
  • pytest infx/tests (serial, srt-slurm submodule initialized): all pass apart from results/test_collect_eval_results.py, which fails to import because a results extra isn't installed locally. It fails the same way on main.

AI model disclosure

  • Model/version: claude-opus-5-5[1m] (Claude Opus 5.5, 1M context), via Claude Code
  • Role: found every GLM-5.1 reference, made the deletions and test/doc updates, ran local validation, and drafted this PR

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries (N/A: retirement only)
中文

改动说明

彻底退役 GLM-5.1。glm5.1-fp8-b200-tilert 是最后一个 GLM-5.1 配置,也是最后一个单轮 1k1k 配置,因此本 PR 同时结束了 1k1k 覆盖。

删除

  • nvidia-master.yaml 中的 glm5.1-fp8-b200-tilert(1k1k + 8k1k,B200 Nscale,vLLM prefill + TileRT 0.1.5 decode)
  • benchmarks/multi_node/srt-slurm-recipes/glm5.1/(两个配方)及 srt-slurm-recipes/configs/tilert-b200-setup.sh
  • 仅供该配置使用的 B200 Nscale 启动器条目:native lane 匹配、DCGM 功耗规则、GLM-5.1-FP8@shared checkpoint 覆盖、TILERT_ENV(唯一一行是 b200-nscale 的 UCX 设置;runtime_env 现在只叠加集群 env 与传入设置)、/tilert_weights 挂载
  • runners.yaml b200-nscale 中的 GLM-5.1 checkpoint 条目及不再使用的 shared-models / tilert-weights 卷
  • glm5.1 的 gsm8k eval 阈值
  • 文档:「TileRT 固定序列长度配方」章节(中英)、「1k1k 仅用于 GLM-5.1」的说明、CONFIGS.md 中过时的 tilert-weights 卷名,以及 claude.yml 智能体提示和 .claude/commands/add-model-hardware.md 中的 GLM-5.1 例外
  • MODELS.md / MODELS_zh.md:1k1k 场景行与 GLM-5 / GLM-5.1 支持矩阵行改为记录退役,并新增带日期的退役说明

保留

  • 通用 TileRT 支持,仍由 glm5.3-fp8-mi355x-tilert-agentic 使用;TileRT driver 测试从 GLM-5.1 B200 lane 移至 GLM-5.3 MI355X lane,而非删除
  • 与 DSR1 B200 配方共用的 srt_fixed_sequence.sh
  • 历史文本:README 新闻条目、2026-09-21 一致性核查、perf-changelog.yaml
  • operatorx/testlists/*.json 中的 GLM-5.1 kernel 形状(OperatorX 微基准,不属于 InferenceX 配置)

与此前的退役 PR(#3618、#3595、#3467)一致,不新增 perf-changelog 条目;不影响任何基准测试,因此不加 sweep 标签。

验证

  • 对两个主配置运行 infx.matrix.generate full-sweep:1715 行,glm5.1 0 行,ISL-1024 0 行
  • pytest infx/tests(串行,已初始化 srt-slurm 子模块):除 results/test_collect_eval_results.py 外全部通过;该文件因本地未安装 results 可选依赖而无法 import,在 main 上同样失败

AI 模型使用说明

  • 模型/版本:claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用
  • 工作内容:查找全部 GLM-5.1 引用,完成删除及测试/文档更新,运行本地验证,起草本 PR

🤖 Generated with Claude Code

Remove glm5.1-fp8-b200-tilert (the last GLM-5.1 and last 1k1k config),
its srt-slurm recipes, the B200 TileRT setup script, the b200-nscale
launcher rows (native lane, power rule, checkpoint override, TILERT_ENV,
/tilert_weights mount), the GLM-5.1 runner entries and eval threshold,
and the docs/agent guidance that carved out the GLM-5.1 exception.
Generic TileRT support stays for GLM-5.3 MI355X; its driver test moves
to that lane.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward and thorough retirement change — correctly scoped and consistent across code, config, tests, and docs.

What was reviewed:

  • Launch policy/driver removals (policy.py NATIVE_SRT_LANES/TILERT_ENV, lanes.py mount, models.py override, power.py rule) — confirmed each was only reachable via the deleted glm5.1 b200-nscale TileRT config, and generic TileRT support (used by glm5.3-fp8-mi355x-tilert-agentic) is untouched.
  • runtime_env simplification after dropping TILERT_ENV — verified other TileRT callers (drivers/srt/run.py, config.py, __init__.py) don't depend on the removed cluster-specific UCX env.
  • Config/runner removals (nvidia-master.yaml, runners.yaml, thresholds.yaml) and deleted recipe files — grepped the repo for dangling references (shared-models, tilert-weights, tilert-b200-setup, active glm5.1 configs) and found none outside historical perf-changelog.yaml entries, which the PR intentionally leaves untouched.
  • Test updates in test_launch_policy.py, test_srt_driver.py, test_srt_policy.py — confirmed the TileRT driver test was moved (not deleted) to the GLM-5.3 MI355X lane, consistent with the PR description.
Extended reasoning...

The PR removes the last GLM-5.1 config and all code/config/docs paths exclusively serving it, across launch policy (policy.py, lanes.py, models.py, power.py), config files (nvidia-master.yaml, runners.yaml, thresholds.yaml), deleted recipe files, and docs/tests. No security-sensitive surface (auth, crypto, injection) is touched — this is CI/benchmark orchestration config and test code. I traced each removed code branch to confirm it was only reachable by the deleted glm5.1 config and verified via grep that no dangling references to removed volumes/scripts/configs remain outside intentionally-preserved historical changelog entries; tests were moved rather than deleted to cover the retained generic TileRT path (glm5.3-fp8-mi355x-tilert-agentic).

This review covers commit 0099e53, which is no longer the latest commit on this pull request; later commits are not covered by it.

@functionstackx
functionstackx merged commit 49460fc into main Oct 1, 2026
5 checks passed
@functionstackx
functionstackx deleted the klaud/retire-glm51 branch October 1, 2026 18:58
adibarra added a commit that referenced this pull request Oct 1, 2026
Ports #3646 (AIPerf harness v1.0.6): the agentx scenario now owns the
replay flags it locks, so replay_argv and REPLAY_ENV drop them. Drops the
multi-node sweep's CLIENT_BACKEND, BENCHMARK_SERVED_MODEL_NAME and
NUM_PROMPTS overrides, whose only recipes #3638 retired.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant