Skip to content

[GLM-5.2 B200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 - #3400

Open
edwingao28 wants to merge 10 commits into
mainfrom
fix/glm52-nvidia-draft-quantization-off
Open

edwingao28 wants to merge 10 commits into
mainfrom
fix/glm52-nvidia-draft-quantization-off

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Set SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0 in the GLM-5.2 B200 AgentX recipes so the NextN/MTP draft MoE keeps its shipped precision, per the draft-precision rule. Images, topology and workload are unchanged.

Testing: Full sweep (attempt 4) green: 5/5 AgentX points; GSM8K strict 0.966 (C48) and 0.965 (C64).

中文

在 GLM-5.2 B200 AgentX 配方中设置 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0,按 draft 精度规则让 NextN/MTP draft MoE 保持原始发布精度。镜像、拓扑和工作负载不变。

测试: 完整 sweep(第 4 次尝试)全绿:5/5 个 AgentX 点;GSM8K strict 0.966(C48)和 0.965(C64)。

AI model disclosure

  • Model/version: Claude Opus 5.5 (claude-opus-5-5).
  • Role: Implementation and drafting.

Related Issue

Related to #3228. / 关联 #3228。

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@edwingao28
edwingao28 force-pushed the fix/glm52-nvidia-draft-quantization-off branch from 0c215c1 to 6629a45 Compare September 23, 2026 21:44
@edwingao28 edwingao28 added full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) and removed full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) labels Sep 23, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-nvidia-draft-quantization-off branch from 6629a45 to ead3d05 Compare September 23, 2026 21:58
@edwingao28 edwingao28 changed the title [GLM-5.2] Disable NVIDIA draft MoE quantization / 关闭 NVIDIA draft MoE 量化 [GLM-5.2] Disable B200 draft MoE quantization / 关闭 B200 draft MoE 量化 Sep 23, 2026
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 23, 2026
@edwingao28 edwingao28 removed the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2] Disable B200 draft MoE quantization / 关闭 B200 draft MoE 量化 [GLM-5.2 B200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 B200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 [GLM-5.2 B200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 Sep 26, 2026
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 26, 2026
@edwingao28
edwingao28 marked this pull request as ready for review September 27, 2026 07:46
@edwingao28
edwingao28 requested a review from a team September 27, 2026 07:46
Set SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0 so the NextN/MTP draft keeps its shipped precision.

设置 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0,使 NextN/MTP draft 保持原始发布精度。
@edwingao28
edwingao28 force-pushed the fix/glm52-nvidia-draft-quantization-off branch from b3155e8 to 5348346 Compare September 27, 2026 07:55
@edwingao28 edwingao28 changed the title [GLM-5.2 B200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 [GLM-5.2 B200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 Sep 27, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and did not find any bugs beyond the two candidates already ruled out. Beyond those, I also checked that prepare_srt_power's override forwarding is consistent end-to-end: apply_srt_recipe passes its trailing arguments straight into plan_commands, so the SRTCTL_RECIPE_ARGS built by prepare_srt_power (and inspected via --inspect-power) are the same arguments later submitted — no override-ordering mismatch. I also checked the new install_srt_slurm/prepare_srt_power helpers and the USES_GLM52_RECOVERY branch in launch_b200-nscale-slurm.sh against the repo's bash conventions (no set -u, check_env_vars used for required inputs, no unconditional-default overrides replacing caller-supplied CONFIG_FILE/SRTCTL_RECIPE_ARGS) and found no violations. Because this touches Slurm launcher branching and a new native power-telemetry validation path, a human look is still worthwhile.

Extended reasoning...

The diff adds a new PowerX telemetry-validation path (inspect_power in infx/srt_slurm/synthetic_acceptance.py), new bash helpers (install_srt_slurm, prepare_srt_power, srt_slurm_version) in runners/slurm_utils.sh, and a new USES_GLM52_RECOVERY branch in runners/launch_b200-nscale-slurm.sh, plus recipe/config changes (agg/disagg-variants.yaml, power-counters.csv) and a vendored patch file. No auth/crypto surface; the sensitive surface is override/precedence correctness for a Slurm job launcher and schema validation via SrtConfig. I independently verified override-argument consistency between inspect_power and plan_commands and checked the new bash code against the repo's mandatory bash conventions (no nounset, check_env_vars usage, no silent defaults), finding nothing beyond the two already-ruled-out candidates, which is why I deferred with a short note rather than approving outright given the change's size and launcher-critical nature.

This review covers commit b3155e8, which is no longer the latest commit on this pull request; later commits are not covered by it.

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

Copy link
Copy Markdown
Collaborator Author

Merging origin/main in (the repo's merge-from-main flow, not a rebase, so run 36304573738 stays pinnable) as soon as its last job, 1P×4D c48, finishes — pushing now would cancel that sweep. The merge is already prepared locally; only conflict is the perf-changelog tail after #3525.


Generated by Claude Code

Copy link
Copy Markdown
Collaborator Author

/use 36304573738


Generated by Claude Code

@edwingao28

Copy link
Copy Markdown
Collaborator Author

/use 36304573738

将 main 合并进本分支以跟上 #3525 等重构;仅解决 perf-changelog.yaml 尾部冲突,
配方改动与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU
再次合并 main 以保持可合并状态;仅解决 perf-changelog.yaml 尾部冲突,
配方与基准矩阵不变,保留 5348346 以复用已 /use 的 sweep 36304573738。

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01FvrigUBZgG9bcFtWHPccAU

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants