Skip to content

[GLM-5.2 GB200] Preserve draft precision and disable background UCX progress / [GLM-5.2 GB200] 保留 draft 精度并关闭 UCX 后台进展 - #3401

Draft
edwingao28 wants to merge 1 commit into
mainfrom
fix/glm52-gb200-draft-quantization-off
Draft

edwingao28 wants to merge 1 commit into
mainfrom
fix/glm52-gb200-draft-quantization-off

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Keep the draft at shipped precision. Disable background UCX progress only for GB200 disaggregated prefill; retain images, workload and golden acceptance.

Testing: Changelog, native matrix and setup selection pass; Full sweep started; results pending.

Blocker: Engine-patch waiver requires reviewer approval.

中文

使 draft 保持原始发布精度,仅关闭 GB200 分离式 prefill 的 UCX 后台进展;保留镜像、工作负载及 golden acceptance。

测试: changelog、原生矩阵及 setup 选择验证通过;完整 sweep 已启动,结果待验证。

阻塞: 补丁例外仍需审阅者批准。

AI 模型: Claude Opus 5.5 (claude-opus-5-5) 完成初始实施与起草;GPT-6(具体版本不可确认)完成恢复、验证和委派复核。

AI model disclosure

  • Claude Opus 5.5 (claude-opus-5-5): initial implementation/drafting.
  • GPT-6, exact version unavailable: recovery, validation and delegated review.

Related Issue

Related to #3228. / 关联 #3228。

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from 91efe8b to e908cc4 Compare September 23, 2026 21:59
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 edwingao28 changed the title [GLM-5.2] Disable GB200 draft MoE quantization / 关闭 GB200 draft MoE 量化 [GLM-5.2 GB200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 [GLM-5.2 GB200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 Sep 26, 2026
@edwingao28 edwingao28 added full-sweep-fail-fast engine-patch Modifies inference engine or serving-stack code; apply patchwork CI priority and removed full-sweep-fail-fast engine-patch Modifies inference engine or serving-stack code; apply patchwork CI priority labels Sep 26, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from 68d3ce2 to 274c69d Compare September 27, 2026 07:55
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 [GLM-5.2 GB200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 Sep 27, 2026
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@cursor
cursor Bot force-pushed the fix/glm52-gb200-draft-quantization-off branch 2 times, most recently from 6113b30 to fca8de5 Compare September 28, 2026 18:38
…refill

Keep the draft at shipped precision and disable only the two background UCX progress controls for disaggregated prefill. Preserve images, workload, strict synchronization and golden acceptance. Request the scoped engine-patch waiver.

保留 GLM-5.2 GB200 draft 的原始精度,仅关闭分离式 prefill 的两项 UCX 后台进展控制;保留镜像、工作负载、严格同步及 golden acceptance,并申请对应补丁例外。
@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from fca8de5 to 9fbe9bf Compare September 29, 2026 07:48
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 [GLM-5.2 GB200] Preserve draft precision and disable background UCX progress / [GLM-5.2 GB200] 保留 draft 精度并关闭 UCX 后台进展 Sep 29, 2026
@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants