Skip to content

[AMD][MI355X] DSv4.1-Flash vLLM: enable sparse-MLA Gluon kernel / DSv4.1-Flash vLLM:启用 sparse-MLA Gluon kernel - #3555

Closed
Fangzhou-Ai wants to merge 10 commits into
mainfrom
amd/dsv41flash-mi355x-vllm-attention-config-indexer
Closed

Fangzhou-Ai wants to merge 10 commits into
mainfrom
amd/dsv41flash-mi355x-vllm-attention-config-indexer

Conversation

@Fangzhou-Ai

@Fangzhou-Ai Fangzhou-Ai commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Add VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True to
dsv41flash-fp4-mi355x-vllm-agentic-dspark to force
vllm-project/vllm#53492's
Gluon sparse-MLA kernel path (matches what was actually set in the local
validation runs referenced below).

Unpins image/model.container to vllm/vllm-openai-rocm:nightly-rocm100-TBD
(only the trailing hash needs updating later): the current pin predates this
recipe's cherry-picked AMD vLLM PRs and is unrelated to them. A maintainer
should fill in the hash once these land upstream and are released in a
nightly image:

Together these bring better topk/mHC/MLA kernels with better
compute/communication stream overlap for DSv4.1-Flash on MI355X.

This PR previously also added --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}' and a
--block-size 128 workaround to exercise
vllm-project/vllm#58671's
ROCm paged MXFP4 sparse-logits indexer kernel. Both are dropped: #58671 isn't
merging in time, and the --block-size 128 pin's entire justification (vLLM's
block-size auto-resolution otherwise picking 64 across this model's 4+
attention backends) was specifically about the 128 that
DeepseekV4ROCMAiterMLASparseBackend.get_preferred_block_size requests once
the fp4 indexer is active — moot without it. A live A/B test (TP2, conc=16,
matched 900s profiling window, #58208 reverted upstream via
vllm-project/vllm#59125 so
the dense fallback path doesn't crash) measured the dense fp8 indexer path
that #58671 would otherwise replace: interactivity p50/p90 -14.7%/-14.9%,
e2e latency p50/p90 +9.5%/+10.8% worse, throughput/GPU unchanged (+0.5%,
within noise), TTFT unchanged. So #58671 is a meaningful interactivity/tail-
latency win, not just a config toggle, but it's still not ready to merge, so
this recipe leaves it out for now.

This PR does not touch --max-num-batched-tokens: the existing TP2 c64 (8192)
/ c128 (4096) values in this recipe already avoid the KV-headroom regression
that a separate, since-closed draft (#3451) flagged as unvalidated.

Duplicate-check: searched open PRs for dsv41flash mi355x, attention-config,
indexer_kv_dtype, mxfp4 indexer, VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA
and block-size 128 on this repo; no open PR adds this env var or touches
this recipe's image pin.

AI model disclosure

  • Model/version: Claude Sonnet 5 (Cursor)
  • Role: Investigated ATOM's official AgentX benchmark results for this model,
    reviewed the unmerged AMD vLLM cherry-pick PRs, ran a KV-headroom
    root-cause investigation and validation sweep for a related TP2 c64
    regression, root-caused a pre-existing #58208 Triton bug hit only by the
    dense fp8 indexer fallback path at AgentX-scale context, fixed the DCO on
    and helped land its revert (vllm-project/vllm#59125), ran a live A/B test
    measuring #58671's actual perf contribution once that confound was
    removed, and revised this PR to drop #58671-specific changes once it was
    clear it wouldn't merge in time. The submitting human reviewed the
    recipe/config changes and is responsible for validating the env var on the
    current image once one is available.

Related Issue

Fixes #

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally (draft: pending a nightly image with the three remaining cherry-picks; the env var was validated against a stale pre-migration copy of this recipe)
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR
中文

为 dsv41flash-fp4-mi355x-vllm-agentic-dspark 添加
VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True,以强制启用
vllm-project/vllm#53492 的
Gluon sparse-MLA kernel 路径(与下方本地验证运行实际设置的配置保持一致)。

将 image/model.container 取消固定为
vllm/vllm-openai-rocm:nightly-rocm100-TBD(之后只需更新末尾的哈希值):当前
固定的镜像早于本配方所 cherry-pick 的 AMD vLLM PR,与其无关。待以下改动合并
至上游并随 nightly 镜像发布后,维护者应填入实际哈希值:

这些改动共同为 MI355X 上的 DSv4.1-Flash 带来更优的 topk/mHC/MLA kernel,以及
更好的计算/通信流重叠模式。

本 PR 此前还添加了
--attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'
与 --block-size 128 的绕过配置,用于启用
vllm-project/vllm#58671 的
ROCm 分页 MXFP4 稀疏 logits indexer kernel。二者均已移除:#58671 无法及时合并;
而 --block-size 128 的全部依据(vLLM 的 block-size 自动推断在该模型 4 个以上
attention backend 同时生效时会退化为选出 64)本身就是针对 fp4 indexer 生效后
DeepseekV4ROCMAiterMLASparseBackend.get_preferred_block_size 所请求的 128 —
在不启用该 indexer 的情况下已不适用。一次真实的 A/B 测试(TP2、并发 16、匹配
900 秒 profiling 窗口,并已通过
vllm-project/vllm#59125 在
上游回滚 #58208 以避免稠密回退路径崩溃)测得:#58671 本应替换的稠密 fp8
indexer 路径,其 interactivity p50/p90 降低 14.7%/14.9%,e2e 延迟 p50/p90
增加 9.5%/10.8%(更差),throughput/GPU 基本不变(+0.5%,在噪声范围内),
TTFT 不变。也就是说 #58671 确实带来了可观的 interactivity/尾部延迟提升,
而不只是一个配置开关,但目前仍未准备好合并,因此本配方暂不启用它。

本 PR 未改动 --max-num-batched-tokens:该配方现有的 TP2 c64(8192)/
c128(4096)取值已经规避了另一个已关闭的草稿 PR(#3451)标记为"尚未验证"的
KV 显存余量回退问题。

重复性检查:已在本仓库搜索 dsv41flash mi355x、attention-config、
indexer_kv_dtype、mxfp4 indexer、VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA
与 block-size 128 相关的未关闭 PR,未发现任何未关闭 PR 添加该环境变量或
改动本配方的镜像固定值。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

Fangzhou-Ai added a commit that referenced this pull request Sep 28, 2026
…log.yaml 的 pr-link 更新为实际 PR #3555

Co-authored-by: Cursor Agent (Claude Sonnet 5) <[email protected]>
Signed-off-by: Fangzhou Ai <[email protected]>
@Fangzhou-Ai
Fangzhou-Ai force-pushed the amd/dsv41flash-mi355x-vllm-attention-config-indexer branch from 14bb21b to d37d6ab Compare September 29, 2026 04:02
@Fangzhou-Ai Fangzhou-Ai changed the title [AMD][MI355X] DSv4.1-Flash vLLM: enable ROCm paged MXFP4 sparse indexer / DSv4.1-Flash vLLM:启用 ROCm 分页 MXFP4 稀疏 indexer [AMD][MI355X] DSv4.1-Flash vLLM: enable sparse-MLA Gluon kernel / DSv4.1-Flash vLLM:启用 sparse-MLA Gluon kernel Sep 29, 2026
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA
env var and the image unpin for #58208/#58655, which are unrelated. Drop the
--attention-config flag and its --block-size 128 workaround, both of which
existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer.

Signed-off-by: Fangzhou Ai <[email protected]>
Co-authored-by: Cursor <[email protected]>
@Fangzhou-Ai
Fangzhou-Ai force-pushed the amd/dsv41flash-mi355x-vllm-attention-config-indexer branch from d37d6ab to 3e3ad54 Compare September 29, 2026 04:16
Fangzhou-Ai added a commit that referenced this pull request Sep 29, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <[email protected]>
Co-authored-by: Cursor <[email protected]>
vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89
is the first published ROCm nightly build whose underlying vLLM commit
(36768d1bfd39094681cdbc8cb37d4b31c0729c89) includes all three PRs this
recipe was waiting on: vllm-project/vllm#58655, #53492 and #58208.

Signed-off-by: Fangzhou Ai <[email protected]>
Co-authored-by: Cursor <[email protected]>
@Fangzhou-Ai
Fangzhou-Ai marked this pull request as ready for review September 29, 2026 05:52
@Fangzhou-Ai
Fangzhou-Ai requested a review from a team September 29, 2026 05:52
@Fangzhou-Ai Fangzhou-Ai added the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Sep 29, 2026
@claude

claude Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @Fangzhou-Ai's task in 31s —— View job


Review of PR #3555

I found one blocking issue. The recipe and master config are consistent: both are pinned to vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89, and the new VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 'True' env var is added to the recipe (agentic.yaml). The changelog entry is not at the end of the file:

The PR body is also stale. It still describes unpinning to nightly-rocm100-TBD, and the "tested locally" box refers to waiting for a nightly image. Commit d305501 already pins that image, so the body should be updated before sign-off.

Comment thread inferencex-e2e/perf-changelog.yaml Outdated
Comment on lines +9007 to +9012
- config-keys:
- dsv41flash-fp4-mi355x-vllm-agentic-dspark
description:
- "Force VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True for vllm-project/vllm#53492's Gluon sparse-MLA kernel. Unpin image to TBD pending #58208 (faster DSA candidate topk selection), #58655 (fused mHC Triton kernel) and #53492 (sparse MLA Gluon kernel) landing in a nightly release: better topk/mHC/MLA kernels with better compute/communication stream overlap. Does not include vllm-project/vllm#58671's MXFP4 sparse-logits indexer: it isn't merging in time, and reverting a dependency it exposed (#58208's DSA candidate-block topk, vllm-project/vllm#59125) showed the fp8 dense indexer path #58671 would otherwise replace costs roughly 15% interactivity and 10% tail e2e latency at TP2 c16, with no measurable throughput/GPU change."
- "设置 VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True 以启用 vllm-project/vllm#53492 的 Gluon sparse-MLA kernel。将镜像取消固定为 TBD,待 #58208(更快的 DSA 候选块 topk 选择)、#58655(融合 mHC Triton kernel)与 #53492(sparse MLA Gluon kernel)合并至上游并随 nightly 镜像发布:更优的 topk/mHC/MLA kernel,以及更好的计算/通信流重叠模式。不包含 vllm-project/vllm#58671 的 MXFP4 稀疏 logits indexer:该 PR 未能及时合并;回滚它所暴露的一个依赖问题(#58208 的 DSA 候选块 topk,vllm-project/vllm#59125)后测得,#58671 本应替换的 fp8 稠密 indexer 路径在 TP2 c16 下约有 15% 的 interactivity 与 10% 的尾部 e2e 延迟损失,而 throughput/GPU 无明显变化。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3555

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 BLOCKING: This entry is in the middle of the file. The main merge added five newer entries after it (#3459, #3430, #3426, #3421, …), so it now sits between #3477 and #3459 and is no longer at the end.

Why it matters: perf-changelog.yaml is read chronologically, oldest at the top and newest at the bottom. New entries have to be appended to the END of the file, and an entry in the middle breaks that ordering.

Fix: Move this entry to the very bottom of the file, after the #3421 entry. While you're there, fix the description. It still says "Unpin image to TBD pending #58208 … landing in a nightly release", but commit d305501 pinned both the recipe and the master config to nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89. Say that it was repinned to this image. Most neighbouring entries also include scenario-type: [agentic-coding], so consider adding it for consistency.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked whether the new VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 'True' env var's string casing (it differs from the sibling AITER flags in the same block, which use '1') could cause it to be silently ignored by vLLM's env-var boolean parsing — but this 'True'/'False' string pattern for boolean env vars is already used elsewhere in this repo's recipes (e.g. SGLANG_MOONCAKE_CUSTOM_MEM_POOL: 'True'), so it's not a distinct concern.

Extended reasoning...

The diff re-pins the DSv4.1-Flash MI355X vLLM ROCm nightly image in agentic.yaml and amd-master.yaml, adds one new AITER env var, and appends a bilingual perf-changelog.yaml entry — pure config/YAML with no auth, injection, or engine-patching surface. A confirmed finding (the changelog text saying the image is left "unpinned to TBD" while the actual diff pins a concrete new nightly tag) is being posted as an inline comment, which alone warrants deferring to a human rather than approving. I additionally checked the new env var's string-boolean formatting for a parsing risk and ruled it out via a precedent elsewhere in the repo.

Comment thread inferencex-e2e/perf-changelog.yaml Outdated
- config-keys:
- dsv41flash-fp4-mi355x-vllm-agentic-dspark
description:
- "Force VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True for vllm-project/vllm#53492's Gluon sparse-MLA kernel. Unpin image to TBD pending #58208 (faster DSA candidate topk selection), #58655 (fused mHC Triton kernel) and #53492 (sparse MLA Gluon kernel) landing in a nightly release: better topk/mHC/MLA kernels with better compute/communication stream overlap. Does not include vllm-project/vllm#58671's MXFP4 sparse-logits indexer: it isn't merging in time, and reverting a dependency it exposed (#58208's DSA candidate-block topk, vllm-project/vllm#59125) showed the fp8 dense indexer path #58671 would otherwise replace costs roughly 15% interactivity and 10% tail e2e latency at TP2 c16, with no measurable throughput/GPU change."

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Maintainers reading this changelog entry will believe the DSv4.1-Flash image is left unpinned pending upstream landing, when the actual config change pins a concrete nightly tag now. perf-changelog.yaml:9010 says "Unpin image to TBD pending #58208 ... #58655 ... and #53492 ... landing in a nightly release", but agentic.yaml:12 and amd-master.yaml:1392 pin container/image to the concrete tag nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89, and the recipe comment at agentic.yaml:9-11 asserts this nightly already "includes" #58655, #53492 and #58208. …

Why this was flagged

…Fix: make the changelog text match the applied config — either state the image is pinned to the new tag now (and confirm/record which of #58208/#58655/#53492 that tag actually contains), or actually leave image/container unpinned as TBD if those PRs have not landed yet; the two changed files must agree on whether the fix is already shipped or still pending.

perf-changelog.yaml is append-only and is the historical record maintainers consult to understand why a config changed; this entry (PR #3555) says the image is being left as TBD pending three PRs landing in a future nightly release. But the same diff sets a concrete container/image tag in agentic.yaml:12 and amd-master.yaml:1392, and the comments there (agentic.yaml:9-11, amd-master.yaml:1390-1391) claim that tag already includes #58655/#53492/#58208. A future engineer investigating a perf regression or trying to know whether these PRs already ship will get contradictory answers depending on which file they read, and cannot tell if the pinned nightly genuinely contains the claimed fixes.

Verification: Severity: nit. The inconsistency is real and present in the diff. perf-changelog.yaml:9010 (English) and its Chinese twin state "Unpin image to TBD pending #58208 ... #58655 ... and #53492 ... landing in a nightly release" — i.e. the entry claims the image is being left unpinned/TBD. But the same diff does the opposite: agentic.yaml:12 sets `container:… | nit. The inconsistency is real and…

@github-actions

Copy link
Copy Markdown
Contributor

Fangzhou-Ai added a commit to Fangzhou-Ai/recipes that referenced this pull request Sep 30, 2026
…parse indexer on MI355X

Add to the MI355X TP override:
- VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=1 for vllm-project/vllm#53492's
  gfx950-only Gluon sparse-MLA kernel.
- --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'
  for vllm-project/vllm#58671's ROCm paged MXFP4 sparse indexer, replacing
  the dense fp8 indexer path.
- --block-size 128, which the fp4 indexer and SparseMLA backend prefer once
  active (auto-resolution otherwise falls back to 64).

Both vLLM PRs are merged and gfx950-scoped. Mirrors
SemiAnalysisAI/InferenceX#3555 and
SemiAnalysisAI/InferenceX#3571.

Signed-off-by: fai <[email protected]>
Co-authored-by: Cursor <[email protected]>
@Fangzhou-Ai

Copy link
Copy Markdown
Collaborator Author

recipe PR merged vllm-project/recipes#1049 @seungrokj @chunfangamd

@Fangzhou-Ai Fangzhou-Ai removed the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Sep 30, 2026
@chunfangamd

Copy link
Copy Markdown
Collaborator

/stage-results 36528242520

@github-actions

Copy link
Copy Markdown
Contributor

@chunfangamd /stage-results requires a completed run from a PR using one of: full-sweep-enabled, non-canary-full-sweep-enabled, full-sweep-fail-fast.

@chunfangamd

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 36528242520

@chunfangamd chunfangamd left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Validation / reuse: run 36528242520 (attempt 3) on in-PR commit 6b2a44e: 16/16 agentic points green (TP2 + TP4, c1–c128). /reuse-sweep-run 36528242520 is posted. Recipe/config values are unchanged since 6b2a44e; later commits are main merges and changelog/comment edits.
  • Evals: agentic eval, TP4 c128: GSM8K strict 0.9742 / flexible 0.9735 (threshold 0.90). Same image, real DSpark block rejection.
  • Recipe: vllm-project/recipes#1049 (MERGED). The MI355X single_node_tp entry matches TP2 (TP4 documented), --moe-backend aiter, DSpark k=5 with adaptive verification off, --no-swa-bounded-replay, Engram offload and VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA. The recipe also enables the vllm-project/vllm#58671 MXFP4 indexer and --block-size 128. #58671 landed after the pinned build, so that part is deferred to #3571.
  • Image: upstream vllm/vllm-openai-rocm:nightly-rocm100-36768d1b (server reports 0.30.1rc1.dev312+g36768d1bf). It includes vllm-project/vllm#53492 and #58655. #58208 was reverted by vllm-project/vllm#59125 before this build.
  • Draft precision: the draft is the embedded DSpark head mtp.0-2 of deepseek-ai/DeepSeek-V4.1-Flash (HF dba1be0a). Shipped (safetensors headers): MXFP4 routed experts, FP8-E4M3 attention/shared-expert/main_proj with UE8M0 32×32 scales, BF16 norms and heads. Effective: the same, through the pinned image's default loading (expert_dtype resolved to 'fp4', MXFP8 linear, DSpark draft model loaded). There are no draft quantization, dtype or KV overrides; KV is the default fp8_ds_mla, shared with the target. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is unset (vLLM). The Gluon sparse-MLA path computes in BF16 and does not quantize Q.
  • Golden AL: all 16 throughput jobs run rejection_sample_method: synthetic with synthetic_acceptance_length: 3.51, matching dsv41flash_dspark.yaml (thinking_on, k=5).
  • Patches / append-only: neither is used, so no waiver is needed.
  • Pareto (P90 tput/GPU vs E2EL, one curve, results_bmk): 16 points measured, 0 invalid, 7 on the frontier (TP4 c8/c16/c32, TP2 c8/c16/c32/c64).

Signed: @chunfangamd

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

❌❌❌ REJECTED ❌❌❌

@chunfangamd: blocking on Check 3. The linked merged recipe's MI355X serve command sets the MXFP4 sparse indexer, and this PR's launch command does not. As a result, the benchmarked configuration is not reproducible from the published recipe.

❌ Check 3 (Recipe linked, merged, complete): FAIL — a major kernel/KV-dtype arg contradicts the merged recipe. vllm-project/recipes#1049 (MERGED 2026-09-30T00:14Z) sets --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}' (plus --block-size 128) in the single_node_tp.hardware_overrides.mi355x entry. inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/mi355x-fp4-mtp/agentic.yaml omits it and runs the default dense FP8 indexer. The PR body's own A/B test puts that difference at roughly 10% on P90 E2EL. The recipe's "MI355X benchmark server" section does not match either: it uses a TP4 command on nightly-eed1f3d0 without VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA. The pinned nightly-rocm100-36768d1b image also predates vllm-project/vllm#58671, so the recipe's MI355X command cannot run on it. To fix this, either publish an upstream recipe revision that documents the non-MXFP4-indexer path for this build, or land the indexer args together with an image that contains #58671 (for example, via #3571). The other major args match: TP2 (TP4 documented), --moe-backend aiter, DSpark k=5 with adaptive verification off, --no-swa-bounded-replay, Engram offload, and the sparse-MLA env var. Informational only: the per-point max-num-batched-tokens, max-cudagraph-capture-size and concurrency values are InferenceX tuning.

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @chunfangamd is listed directly for inferencex-e2e/configs/amd-master.yaml; the other changed paths fall only under the catch-all.

✅ Check 1 (Passing sweep on in-PR commit): PASS — commit 6b2a44e, which is still in the PR, has 16/16 agentic / jobs (TP2 and TP4, c1–c128) plus agentic eval / at success in run 36528242520 (attempt 3). The fixed-sequence jobs were skipped because the config is agentic-only. Between 6b2a44e and the pinned 0b77984, the only change to this recipe and config entry is comments.

✅ Check 2 (Evals pass): PASS — GSM8K em_strict 0.9742 and em_flexible 0.9735 (n=1319) in the eval_results_all artifact, on the same image vllm/vllm-openai-rocm:nightly-rocm100-36768d1b…, with real block rejection (job).

✅ Check 4 (Reuse command): PASS — /reuse-sweep-run 36528242520 was posted by chunfangamd (COLLABORATOR) on a line of its own.

✅ Check 5 (Latest checklist template): PASS — all 17 items in the current PR_REVIEW_CHECKLIST.md, including the recipe sub-item and the Pareto item, are present and checked.

✅ Check 6 (Upstream images / engine-first): PASS — framework: vllm with the upstream image vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89.

✅ Check 7 (No deprecated models/scenarios): PASS — MODELS.md lists dsv41flash agentic coding as active on 2026-09-30.

✅ Check 8 (No architecture hacks): PASS — no --hf-overrides, model-config edits or layer/expert trimming; the env var only selects the sparse-MLA kernel.

✅ Check 9 (Spec decode uses chat templates): PASS — the replay client runs with --endpoint-type chat against /v1/chat/completions (per benchmark.log).

✅ Check 10 (No engine patches): PASS — the diff has no patches, heredoc source edits, monkey-patching or engine wheel installs.

✅ Check 11 (Agentic golden AL): PASS — all throughput points render rejection_sample_method: synthetic with synthetic_acceptance_length: 3.51, which matches dsv41flash_dspark.yaml at thinking_on with k=5 (3.51).

➖ Check 12 (Append-only): N/A — the new changelog entry does not set append-only: true.

✅ Check 13 (Draft runs as shipped): PASS — the draft is the checkpoint's embedded DSpark head (method='dspark', model='deepseek-ai/DeepSeek-V4.1-Flash'), loaded with the image defaults (expert_dtype resolved to 'fp4', DSpark draft model loaded, kv_cache_dtype=auto). There are no draft quantization, dtype or KV overrides. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE does not apply (vLLM) and is unset.

✅ Check 14 (Pareto coverage): PASS — one curve (dsv41flash agentic, MI355X vLLM FP4, P90 E2EL vs total tput/GPU, run 36528242520 attempt 3, same image): 16 measured points, 0 invalid, 7 on the frontier per pareto_coverage (results_bmk).

Assessed commit: 0b779843f285935674a0ebdb1f4254c28b0829b5.

Fangzhou-Ai added a commit that referenced this pull request Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <[email protected]>
Co-authored-by: Cursor <[email protected]>
Fangzhou-Ai added a commit that referenced this pull request Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <[email protected]>
Co-authored-by: Cursor <[email protected]>
@Fangzhou-Ai Fangzhou-Ai closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants