[GLM-5.2 GB300] Preserve draft precision and extend scheduler watchdog / [GLM-5.2 GB300] 保留 draft 精度并延长调度 watchdog - #3402
Draft
edwingao28 wants to merge 1 commit into
Draft
edwingao28 wants to merge 1 commit into
edwingao28 wants to merge 1 commit into
Conversation
Contributor
|
Thanks for the contribution!
中文感谢你的贡献!
|
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 23, 2026 21:59
0079a3a to
66a5f68
Compare
Contributor
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36538942770 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36538942770 |
7 of 10 tasks
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 27, 2026 08:07
581f91f to
d8b13ce
Compare
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 27, 2026 18:56
d8b13ce to
945377b
Compare
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 27, 2026 21:55
945377b to
9bced94
Compare
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 27, 2026 23:58
9bced94 to
3cf29c6
Compare
Collaborator
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
cursor
Bot
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
2 times, most recently
from
September 28, 2026 18:38
4d4d437 to
1903073
Compare
Set SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0 so the NextN/MTP draft keeps its shipped precision, and use the 1800 s scheduler watchdog of the GLM-5.2 B200/GB200 recipes for disaggregated prefill and decode. 设置 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE=0,使 NextN/MTP draft 保持原始发布精度; 分离式 prefill 和 decode 采用 GLM-5.2 B200/GB200 配方的 1800 秒调度 watchdog。 Co-authored-by: Wenyao Gao <[email protected]>
edwingao28
force-pushed
the
fix/glm52-gb300-draft-quantization-off
branch
from
September 29, 2026 07:49
1903073 to
fc03e75
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Keep the draft at shipped precision and retain the 1800-second disaggregated watchdog. Images, topology and workload remain unchanged.
Testing: Changelog and native matrix pass; Full sweep started; results pending.
Blocker: Prior HiCache/NCCL failures remain unresolved.
中文
使 draft 保持原始发布精度,保留分离式 1800 秒 watchdog;镜像、拓扑和工作负载不变。
测试: changelog 及原生矩阵验证通过;完整 sweep 已启动,结果待验证。
阻塞: 此前 HiCache/NCCL 失败仍待确认恢复。
AI 模型: Claude Opus 5.5 (
claude-opus-5-5) 完成初始实施与起草;GPT-6(具体版本不可确认)完成恢复、验证和委派复核。AI model disclosure
claude-opus-5-5): initial implementation/drafting.Related Issue
Related to #3228. / 关联 #3228。
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.