Add Qwen3.8-Flash-Next MXFP4 AgentX on MI355X / 新增 MI355X Qwen3.8-Flash-Next MXFP4 AgentX 配方 - #3129
functionstackx wants to merge 4 commits into
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
将新增配方的性能日志关联到 PR #3129。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=35016198611 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=35016198611 |
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs, but because this adds a new bring-up recipe with GPU validation explicitly still pending, a human look would still be worthwhile.
What was reviewed: the new benchmark script (env/AITER flags, cuda-graph sizing, eval vs. agentic-replay flow), the new qwen3.8next-fp4-mi355x-sglang-agentic entry in configs/amd-master.yaml and the matching append-only perf-changelog.yaml entry, and the new model+framework-specific /ix mount override in runners/launch_mi355x-amds.sh. Also checked that the new script omits --reasoning-parser, unlike every other Qwen3.8-Flash-Next SGLang recipe in the repo — worth a second look for eval-scoring impact.
Extended reasoning...
Overview
This PR is a single-commit, four-file addition of a new AgentX benchmark recipe (qwen3.8next-fp4-mi355x-sglang-agentic): a new benchmark script under benchmarks/single_node/agentic/, a new scenario block in configs/amd-master.yaml, an append-only entry in perf-changelog.yaml, and a narrow conditional branch in runners/launch_mi355x-amds.sh that reroutes the repo bind mount from /workspace to /ix for this exact model+framework pair so the bring-up image's bundled SGLang/AITER packages aren't shadowed. No engine code, auth, or data-handling paths are touched.
Security risks
None identified. The image is pinned by digest, the checkpoint is a public HF repo, and the runner change is an if gate scoped to one literal MODEL/FRAMEWORK pair — it cannot affect any other recipe's mount behavior. No secrets, network exposure, or privilege changes are involved.
Level of scrutiny
This is config/recipe-pattern work, mechanically similar to many existing sibling recipes, which argues for a lighter touch. But two things push against an outright approval: (1) the PR's own checklist states GPU validation and full-sweep/eval evidence are still pending — this recipe has not yet been run on hardware; (2) the new script is the only Qwen3.8-Flash-Next SGLang recipe in the tree that does not pass --reasoning-parser to sglang.launch_server, while every sibling (qwen3.8next_fp8_h100_mtp.sh, qwen3.8next_fp8_h200_sglang_mtp.sh, qwen3.8next_fp4_b200_sglang_mtp.sh, qwen3.8next_fp4_b300_sglang_mtp.sh) passes --reasoning-parser auto. This was already surfaced and investigated this run; I independently confirmed the sibling-consistency pattern but could not confirm from the code alone whether the omission is benign for this checkpoint (e.g., if this MXFP4 variant doesn't emit <think> tags by default) or would silently skew GSM8K/lm-eval answer extraction. That residual uncertainty, combined with pending GPU validation, is enough that I don't have high confidence a human doesn't need to look.
Other factors
The runner mount-override pattern mirrors an existing documented /ix + INFMAX_CONTAINER_WORKSPACE convention used for other AMD agentic recipes, so it's not novel risk. The perf-changelog entry is appended at the tail with an intentional PR_NUMBER placeholder, consistent with the documented pre-merge convention. No tests are added or modified, and the change doesn't touch any CODEOWNER-restricted logic beyond the standard config/runner files already covered by existing patterns.
This review covers commit 37349d1, which is no longer the latest commit on this pull request; later commits are not covered by it.
新增 MI355X 上 Qwen3.8-Flash-Next Quark MXFP4 AgentX 配方,使用固定的上游 SGLang 镜像及完整并发矩阵。
将新增配方的性能日志关联到 PR #3129。
仅为新增 Qwen3.8-Flash-Next MXFP4 SGLang 配方转换 Enroot digest URI,并增加冷缓存导入和挂载路由行为测试。
…N MTP Agentic recipes run with speculative decoding only (MODELS.md), and every other qwen3.8next SGLang AgentX arm is an -agentic-mtp key. Rename the key and script to the MTP form, add the cookbook NEXTN card (3 steps, topk 1, 4 draft tokens), pin throughput acceptance to the golden AL 2.32 while EVAL_ONLY keeps real verification, drop --mem-fraction-static to 0.85 for draft-head headroom, and use --cuda-graph-max-bs-decode like the other MI355X SGLang MTP arms. Launcher routing tests follow the new script name. 智能体配方按 MODELS.md 仅运行投机解码,且其他 qwen3.8next SGLang AgentX 配方均为 -agentic-mtp。本次将配置键与脚本改为 MTP 形式,加入 cookbook 的 NEXTN 参数 (3 步、topk 1、4 草稿 token),吞吐固定黄金 AL 2.32、评测保留真实验证, --mem-fraction-static 降至 0.85,并使用 --cuda-graph-max-bs-decode。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
c94bdad to
63449de
Compare
|
InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Description
Add
qwen3.8next-fp4-mi355x-sglang-agentic-mtpusing the public AMD Quark MXFP4 checkpoint.cluster:mi355x-amds, TP8/EP1, concurrency 1/4/8/12/16, no KV offloading.match-expectedandreal-draft-token; evals do not enable simulated acceptance.lmsysorg/sglang-rocm:qwen38flashnextto digestsha256:f2e3928cd5be1d7bf9bfb783d57e9f5b35baa8e09d4248e5fc44da43c85c35e4./ixso the repository does not hide SGLang and AITER under the image's/workspace.The new checkpoint keeps PLE in BF16. The separate FP8 recipe in #2754 hits an assertion when loading quantized PLE scales (SGLang #36616); this PR does not modify that recipe or suppress the assertion.
Validation
Local checks on
63449de4aab8defd5faa1718ebd3eba10edca92e:git diff --checkpassed.utils/matrix_logic/andrunners/test_slurm_utils.pypassed, including launcher routing and cold-cache Enroot import behavior.--all-evalsselected all five MTP concurrency points for lm-eval.GPU validation is blocked. In sweep 35016198611, concurrency-1 eval, the model reached warmup, then all eight ranks reported
HSA_STATUS_ERROR_EXCEPTION/hipErrorLaunchFailure. The scheduler aborted at 02:10 UTC on September 16; surviving HTTP workers continued returning failed health checks until the 500-minute CI timeout. The remaining sweep was cancelled to release capacity. No eval score or throughput result passed.SGLang #36601 documents remaining shared-expert loader and ROCm QSA fixes for this checkpoint. The validated path in cookbook PR #36919 uses TP8/EP8, page size 64, explicit AITER attention/MoE, and exact unmerged source over an immutable official nightly. This PR currently has neither that source override nor the revised topology. An engine-source override requires explicit authorization; no assertion has been disabled.
A local opt-in readiness fix detects confirmed scheduler/fatal GPU errors even when the HTTP wrapper survives. It passed 402 matrix, launcher, server-watch, and readiness tests, but is not yet pushed to avoid automatically launching another sweep against the known-failing image. The PR retains
full-sweep-fail-fastandall-evals; a complete sweep and evals must pass on the final commit.Checklist
/reuse-sweep-runafter the final green sweep.The FP8 cookbook is not MXFP4 validation. No merge is requested as part of this bring-up.
中文
新增
qwen3.8next-fp4-mi355x-sglang-agentic-mtp,使用公开的amd/Qwen3.8-Flash-Next-Quark-MXFP4权重,在cluster:mi355x-amds上运行 SGLang,配置为 TP8/EP1,并发度 1/4/8/12/16,不启用 KV offloading。启动参数参考上方链接的 MI355X FP8 balanced 配方,并将权重替换为 Quark MXFP4。当前版本启用原生 NEXTN MTP,参数为 3 steps、top-k 1、4 draft tokens。吞吐测试采用仓库中的黄金接受长度 2.32,使用
match-expected和real-draft-token;评测不启用模拟接受率。上游 cookbook 尚未验证这个 MXFP4 组合。镜像通过 digest 固定;仓库挂载到
/ix,避免覆盖镜像/workspace下的 SGLang 和 AITER。冷缓存导入将镜像引用转换为 Enroot 支持的 digest URI,保留原始镜像与缓存标识,转换仅影响此模型、框架和镜像仓库。保留标准 256k AgentX 轨迹和一小时性能测试,不修改推理引擎、不缩短测试、不跳过正确性断言,也不降低评测阈值。新权重的 PLE 保持 BF16;本 PR 不修改 #2754 的 FP8 配方,也不屏蔽其 scale 断言。
在提交
63449de4aab8defd5faa1718ebd3eba10edca92e上,Bash 语法、diff 检查、矩阵与 launcher 的 392 项测试均通过,包括冷缓存 Enroot 导入行为。--all-evals为全部五个 MTP 并发点选择 lm-eval。这些本地检查不证明 SGLang MXFP4/NEXTN 的 GPU 运行兼容性。GPU sweep 35016198611 的并发度 1 评测 在 warmup 阶段出现
HSA_STATUS_ERROR_EXCEPTION/hipErrorLaunchFailure。scheduler 于 9 月 16 日 02:10 UTC 中止,但残留 HTTP worker 持续返回健康检查失败,直到 500 分钟 CI 超时。其余 sweep 已取消以释放资源,没有通过的评测分数或吞吐结果。SGLang #36601 包含该权重所需的 shared-expert loader 和 ROCm QSA 修复。cookbook PR #36919 的验证路径采用 TP8/EP8、page size 64、显式 AITER attention/MoE,并在固定 digest 的官方 nightly 上挂载精确的未合并源码。本 PR 尚未采用该源码覆盖或新拓扑;覆盖引擎源码需明确授权,不会屏蔽断言。
本地已添加可选的就绪检查修复,在 HTTP wrapper 仍存活时也能检测 scheduler 中止或致命 GPU 错误。402 项相关测试通过,暂未推送,以免对已知失败的镜像再次自动启动 sweep。PR 保留
full-sweep-fail-fast和all-evals;合并前仍需最终提交的完整绿色 sweep、对应上游 cookbook、CODEOWNER 审查和授权维护者的/reuse-sweep-run。本次任务不执行合并。