Skip to content

[GLM-5.2 GB300] Preserve draft precision and extend scheduler watchdog / [GLM-5.2 GB300] 保留 draft 精度并延长调度 watchdog - #3402

Draft
edwingao28 wants to merge 1 commit into
mainfrom
fix/glm52-gb300-draft-quantization-off
Draft

edwingao28 wants to merge 1 commit into
mainfrom
fix/glm52-gb300-draft-quantization-off

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Keep the draft at shipped precision and retain the 1800-second disaggregated watchdog. Images, topology and workload remain unchanged.

Testing: Changelog and native matrix pass; Full sweep started; results pending.

Blocker: Prior HiCache/NCCL failures remain unresolved.

中文

使 draft 保持原始发布精度,保留分离式 1800 秒 watchdog;镜像、拓扑和工作负载不变。

测试: changelog 及原生矩阵验证通过;完整 sweep 已启动,结果待验证。

阻塞: 此前 HiCache/NCCL 失败仍待确认恢复。

AI 模型: Claude Opus 5.5 (claude-opus-5-5) 完成初始实施与起草;GPT-6(具体版本不可确认)完成恢复、验证和委派复核。

AI model disclosure

  • Claude Opus 5.5 (claude-opus-5-5): initial implementation/drafting.
  • GPT-6, exact version unavailable: recovery, validation and delegated review.

Related Issue

Related to #3228. / 关联 #3228。

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from 0079a3a to 66a5f68 Compare September 23, 2026 21:59
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 23, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 edwingao28 removed the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2] Disable GB300 draft MoE quantization / 关闭 GB300 draft MoE 量化 [GLM-5.2 GB300] require PowerX and isolate HiCache load waits / 启用功耗并隔离 HiCache 加载等待 Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 GB300] require PowerX and isolate HiCache load waits / 启用功耗并隔离 HiCache 加载等待 [GLM-5.2 GB300] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 Sep 26, 2026
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 26, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from 581f91f to d8b13ce Compare September 27, 2026 08:07
@edwingao28 edwingao28 changed the title [GLM-5.2 GB300] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 [GLM-5.2 GB300] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 Sep 27, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from d8b13ce to 945377b Compare September 27, 2026 18:56
@edwingao28 edwingao28 changed the title [GLM-5.2 GB300] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 [GLM-5.2 GB300] disable draft MoE FP8 conversion and scope HiCache fix / 关闭 draft MoE FP8 转换并限定 HiCache 修复 Sep 27, 2026
@edwingao28 edwingao28 removed the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 27, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from 945377b to 9bced94 Compare September 27, 2026 21:55
@edwingao28 edwingao28 added full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures and removed engine-patch labels Sep 27, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from 9bced94 to 3cf29c6 Compare September 27, 2026 23:58
@edwingao28 edwingao28 changed the title [GLM-5.2 GB300] disable draft MoE FP8 conversion and scope HiCache fix / 关闭 draft MoE FP8 转换并限定 HiCache 修复 [GLM-5.2 GB300] disable draft MoE FP8 conversion and raise scheduler watchdog / 关闭 draft MoE FP8 转换并提高调度 watchdog Sep 27, 2026
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@cursor
cursor Bot force-pushed the fix/glm52-gb300-draft-quantization-off branch 2 times, most recently from 4d4d437 to 1903073 Compare September 28, 2026 18:38
Set SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0 so the NextN/MTP draft keeps its shipped
precision, and use the 1800 s scheduler watchdog of the GLM-5.2 B200/GB200
recipes for disaggregated prefill and decode.

设置 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0,使 NextN/MTP draft 保持原始发布精度;
分离式 prefill 和 decode 采用 GLM-5.2 B200/GB200 配方的 1800 秒调度 watchdog。

Co-authored-by: Wenyao Gao <[email protected]>
@edwingao28
edwingao28 force-pushed the fix/glm52-gb300-draft-quantization-off branch from 1903073 to fc03e75 Compare September 29, 2026 07:49
@edwingao28 edwingao28 changed the title [GLM-5.2 GB300] disable draft MoE FP8 conversion and raise scheduler watchdog / 关闭 draft MoE FP8 转换并提高调度 watchdog [GLM-5.2 GB300] Preserve draft precision and extend scheduler watchdog / [GLM-5.2 GB300] 保留 draft 精度并延长调度 watchdog Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants