[AMD][MI355X] DSv4.1-Flash vLLM: enable sparse-MLA Gluon kernel / DSv4.1-Flash vLLM:启用 sparse-MLA Gluon kernel - #3555
Fangzhou-Ai wants to merge 10 commits into
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
…log.yaml 的 pr-link 更新为实际 PR #3555 Co-authored-by: Cursor Agent (Claude Sonnet 5) <[email protected]> Signed-off-by: Fangzhou Ai <[email protected]>
14bb21b to
d37d6ab
Compare
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA env var and the image unpin for #58208/#58655, which are unrelated. Drop the --attention-config flag and its --block-size 128 workaround, both of which existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer. Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]>
d37d6ab to
3e3ad54
Compare
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]>
vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89 is the first published ROCm nightly build whose underlying vLLM commit (36768d1bfd39094681cdbc8cb37d4b31c0729c89) includes all three PRs this recipe was waiting on: vllm-project/vllm#58655, #53492 and #58208. Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]>
|
Claude finished @Fangzhou-Ai's task in 31s —— View job Review of PR #3555I found one blocking issue. The recipe and master config are consistent: both are pinned to
The PR body is also stale. It still describes unpinning to |
| - config-keys: | ||
| - dsv41flash-fp4-mi355x-vllm-agentic-dspark | ||
| description: | ||
| - "Force VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True for vllm-project/vllm#53492's Gluon sparse-MLA kernel. Unpin image to TBD pending #58208 (faster DSA candidate topk selection), #58655 (fused mHC Triton kernel) and #53492 (sparse MLA Gluon kernel) landing in a nightly release: better topk/mHC/MLA kernels with better compute/communication stream overlap. Does not include vllm-project/vllm#58671's MXFP4 sparse-logits indexer: it isn't merging in time, and reverting a dependency it exposed (#58208's DSA candidate-block topk, vllm-project/vllm#59125) showed the fp8 dense indexer path #58671 would otherwise replace costs roughly 15% interactivity and 10% tail e2e latency at TP2 c16, with no measurable throughput/GPU change." | ||
| - "设置 VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True 以启用 vllm-project/vllm#53492 的 Gluon sparse-MLA kernel。将镜像取消固定为 TBD,待 #58208(更快的 DSA 候选块 topk 选择)、#58655(融合 mHC Triton kernel)与 #53492(sparse MLA Gluon kernel)合并至上游并随 nightly 镜像发布:更优的 topk/mHC/MLA kernel,以及更好的计算/通信流重叠模式。不包含 vllm-project/vllm#58671 的 MXFP4 稀疏 logits indexer:该 PR 未能及时合并;回滚它所暴露的一个依赖问题(#58208 的 DSA 候选块 topk,vllm-project/vllm#59125)后测得,#58671 本应替换的 fp8 稠密 indexer 路径在 TP2 c16 下约有 15% 的 interactivity 与 10% 的尾部 e2e 延迟损失,而 throughput/GPU 无明显变化。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3555 |
There was a problem hiding this comment.
🔴 BLOCKING: This entry is in the middle of the file. The main merge added five newer entries after it (#3459, #3430, #3426, #3421, …), so it now sits between #3477 and #3459 and is no longer at the end.
Why it matters: perf-changelog.yaml is read chronologically, oldest at the top and newest at the bottom. New entries have to be appended to the END of the file, and an entry in the middle breaks that ordering.
Fix: Move this entry to the very bottom of the file, after the #3421 entry. While you're there, fix the description. It still says "Unpin image to TBD pending #58208 … landing in a nightly release", but commit d305501 pinned both the recipe and the master config to nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89. Say that it was repinned to this image. Most neighbouring entries also include scenario-type: [agentic-coding], so consider adding it for consistency.
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked whether the new VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 'True' env var's string casing (it differs from the sibling AITER flags in the same block, which use '1') could cause it to be silently ignored by vLLM's env-var boolean parsing — but this 'True'/'False' string pattern for boolean env vars is already used elsewhere in this repo's recipes (e.g. SGLANG_MOONCAKE_CUSTOM_MEM_POOL: 'True'), so it's not a distinct concern.
Extended reasoning...
The diff re-pins the DSv4.1-Flash MI355X vLLM ROCm nightly image in agentic.yaml and amd-master.yaml, adds one new AITER env var, and appends a bilingual perf-changelog.yaml entry — pure config/YAML with no auth, injection, or engine-patching surface. A confirmed finding (the changelog text saying the image is left "unpinned to TBD" while the actual diff pins a concrete new nightly tag) is being posted as an inline comment, which alone warrants deferring to a human rather than approving. I additionally checked the new env var's string-boolean formatting for a parsing risk and ruled it out via a precedent elsewhere in the repo.
| - config-keys: | ||
| - dsv41flash-fp4-mi355x-vllm-agentic-dspark | ||
| description: | ||
| - "Force VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True for vllm-project/vllm#53492's Gluon sparse-MLA kernel. Unpin image to TBD pending #58208 (faster DSA candidate topk selection), #58655 (fused mHC Triton kernel) and #53492 (sparse MLA Gluon kernel) landing in a nightly release: better topk/mHC/MLA kernels with better compute/communication stream overlap. Does not include vllm-project/vllm#58671's MXFP4 sparse-logits indexer: it isn't merging in time, and reverting a dependency it exposed (#58208's DSA candidate-block topk, vllm-project/vllm#59125) showed the fp8 dense indexer path #58671 would otherwise replace costs roughly 15% interactivity and 10% tail e2e latency at TP2 c16, with no measurable throughput/GPU change." |
There was a problem hiding this comment.
🟡 (optional) Maintainers reading this changelog entry will believe the DSv4.1-Flash image is left unpinned pending upstream landing, when the actual config change pins a concrete nightly tag now. perf-changelog.yaml:9010 says "Unpin image to TBD pending #58208 ... #58655 ... and #53492 ... landing in a nightly release", but agentic.yaml:12 and amd-master.yaml:1392 pin container/image to the concrete tag nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89, and the recipe comment at agentic.yaml:9-11 asserts this nightly already "includes" #58655, #53492 and #58208. …
Why this was flagged
…Fix: make the changelog text match the applied config — either state the image is pinned to the new tag now (and confirm/record which of #58208/#58655/#53492 that tag actually contains), or actually leave image/container unpinned as TBD if those PRs have not landed yet; the two changed files must agree on whether the fix is already shipped or still pending.
perf-changelog.yaml is append-only and is the historical record maintainers consult to understand why a config changed; this entry (PR #3555) says the image is being left as TBD pending three PRs landing in a future nightly release. But the same diff sets a concrete container/image tag in agentic.yaml:12 and amd-master.yaml:1392, and the comments there (agentic.yaml:9-11, amd-master.yaml:1390-1391) claim that tag already includes #58655/#53492/#58208. A future engineer investigating a perf regression or trying to know whether these PRs already ship will get contradictory answers depending on which file they read, and cannot tell if the pinned nightly genuinely contains the claimed fixes.
Verification: Severity: nit. The inconsistency is real and present in the diff. perf-changelog.yaml:9010 (English) and its Chinese twin state "Unpin image to TBD pending #58208 ... #58655 ... and #53492 ... landing in a nightly release" — i.e. the entry claims the image is being left unpinned/TBD. But the same diff does the opposite: agentic.yaml:12 sets `container:… | nit. The inconsistency is real and…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36528242520 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36528242520 |
…parse indexer on MI355X Add to the MI355X TP override: - VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=1 for vllm-project/vllm#53492's gfx950-only Gluon sparse-MLA kernel. - --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}' for vllm-project/vllm#58671's ROCm paged MXFP4 sparse indexer, replacing the dense fp8 indexer path. - --block-size 128, which the fp4 indexer and SparseMLA backend prefer once active (auto-resolution otherwise falls back to 64). Both vLLM PRs are merged and gfx950-scoped. Mirrors SemiAnalysisAI/InferenceX#3555 and SemiAnalysisAI/InferenceX#3571. Signed-off-by: fai <[email protected]> Co-authored-by: Cursor <[email protected]>
|
recipe PR merged vllm-project/recipes#1049 @seungrokj @chunfangamd |
|
/stage-results 36528242520 |
|
@chunfangamd |
|
/reuse-sweep-run 36528242520 |
…m#58208 vllm-project/vllm#58208 was reverted by vllm-project/vllm#59125 before the pinned build, so the image only carries #58655 and #53492. Comment-only change. Co-authored-by: Cursor <[email protected]>
There was a problem hiding this comment.
As a PR reviewer and CODEOWNER, I have reviewed this and have:
- Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
- Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
- Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
- Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
- Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
- Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison. - For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
- Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
- Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
- If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
- If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
- Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
- I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
- Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/
<PR_NUMBER>.md— named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section. - If this PR uses
append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it. - If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
- Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.
Additional detail section:
- Validation / reuse: run 36528242520 (attempt 3) on in-PR commit
6b2a44e: 16/16 agentic points green (TP2 + TP4, c1–c128)./reuse-sweep-run 36528242520is posted. Recipe/config values are unchanged since6b2a44e; later commits are main merges and changelog/comment edits. - Evals: agentic eval, TP4 c128: GSM8K strict 0.9742 / flexible 0.9735 (threshold 0.90). Same image, real DSpark
blockrejection. - Recipe: vllm-project/recipes#1049 (MERGED). The MI355X
single_node_tpentry matches TP2 (TP4 documented),--moe-backend aiter, DSpark k=5 with adaptive verification off,--no-swa-bounded-replay, Engram offload andVLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA. The recipe also enables the vllm-project/vllm#58671 MXFP4 indexer and--block-size 128. #58671 landed after the pinned build, so that part is deferred to #3571. - Image: upstream
vllm/vllm-openai-rocm:nightly-rocm100-36768d1b(server reports0.30.1rc1.dev312+g36768d1bf). It includes vllm-project/vllm#53492 and #58655. #58208 was reverted by vllm-project/vllm#59125 before this build. - Draft precision: the draft is the embedded DSpark head
mtp.0-2ofdeepseek-ai/DeepSeek-V4.1-Flash(HFdba1be0a). Shipped (safetensors headers): MXFP4 routed experts, FP8-E4M3 attention/shared-expert/main_projwith UE8M0 32×32 scales, BF16 norms and heads. Effective: the same, through the pinned image's default loading (expert_dtype resolved to 'fp4', MXFP8 linear,DSpark draft model loaded). There are no draft quantization, dtype or KV overrides; KV is the defaultfp8_ds_mla, shared with the target.SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis unset (vLLM). The Gluon sparse-MLA path computes in BF16 and does not quantize Q. - Golden AL: all 16 throughput jobs run
rejection_sample_method: syntheticwithsynthetic_acceptance_length: 3.51, matchingdsv41flash_dspark.yaml(thinking_on, k=5). - Patches / append-only: neither is used, so no waiver is needed.
- Pareto (P90 tput/GPU vs E2EL, one curve, results_bmk): 16 points measured, 0 invalid, 7 on the frontier (TP4 c8/c16/c32, TP2 c8/c16/c32/c64).
Signed: @chunfangamd
❌❌❌ REJECTED ❌❌❌@chunfangamd: blocking on Check 3. The linked merged recipe's MI355X serve command sets the MXFP4 sparse indexer, and this PR's launch command does not. As a result, the benchmarked configuration is not reproducible from the published recipe. ❌ Check 3 (Recipe linked, merged, complete): FAIL — a major kernel/KV-dtype arg contradicts the merged recipe. vllm-project/recipes#1049 (MERGED 2026-09-30T00:14Z) sets Passed and not applicable checks✅ Check 0 (CODEOWNER): PASS — ✅ Check 1 (Passing sweep on in-PR commit): PASS — commit ✅ Check 2 (Evals pass): PASS — GSM8K em_strict 0.9742 and em_flexible 0.9735 (n=1319) in the ✅ Check 4 (Reuse command): PASS — ✅ Check 5 (Latest checklist template): PASS — all 17 items in the current ✅ Check 6 (Upstream images / engine-first): PASS — ✅ Check 7 (No deprecated models/scenarios): PASS — ✅ Check 8 (No architecture hacks): PASS — no ✅ Check 9 (Spec decode uses chat templates): PASS — the replay client runs with ✅ Check 10 (No engine patches): PASS — the diff has no patches, heredoc source edits, monkey-patching or engine wheel installs. ✅ Check 11 (Agentic golden AL): PASS — all throughput points render ➖ Check 12 (Append-only): N/A — the new changelog entry does not set ✅ Check 13 (Draft runs as shipped): PASS — the draft is the checkpoint's embedded DSpark head ( ✅ Check 14 (Pareto coverage): PASS — one curve (dsv41flash agentic, MI355X vLLM FP4, P90 E2EL vs total tput/GPU, run 36528242520 attempt 3, same image): 16 measured points, 0 invalid, 7 on the frontier per Assessed commit: |
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]>
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <[email protected]> Co-authored-by: Cursor <[email protected]>
Description
Add
VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=Truetodsv41flash-fp4-mi355x-vllm-agentic-dsparkto forcevllm-project/vllm#53492's
Gluon sparse-MLA kernel path (matches what was actually set in the local
validation runs referenced below).
Unpins
image/model.containertovllm/vllm-openai-rocm:nightly-rocm100-TBD(only the trailing hash needs updating later): the current pin predates this
recipe's cherry-picked AMD vLLM PRs and is unrelated to them. A maintainer
should fill in the hash once these land upstream and are released in a
nightly image:
Together these bring better topk/mHC/MLA kernels with better
compute/communication stream overlap for DSv4.1-Flash on MI355X.
This PR previously also added
--attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'and a--block-size 128workaround to exercisevllm-project/vllm#58671's
ROCm paged MXFP4 sparse-logits indexer kernel. Both are dropped: #58671 isn't
merging in time, and the
--block-size 128pin's entire justification (vLLM'sblock-size auto-resolution otherwise picking 64 across this model's 4+
attention backends) was specifically about the 128 that
DeepseekV4ROCMAiterMLASparseBackend.get_preferred_block_sizerequests oncethe fp4 indexer is active — moot without it. A live A/B test (TP2, conc=16,
matched 900s profiling window,
#58208reverted upstream viavllm-project/vllm#59125 so
the dense fallback path doesn't crash) measured the dense fp8 indexer path
that
#58671would otherwise replace: interactivity p50/p90 -14.7%/-14.9%,e2e latency p50/p90 +9.5%/+10.8% worse, throughput/GPU unchanged (+0.5%,
within noise), TTFT unchanged. So
#58671is a meaningful interactivity/tail-latency win, not just a config toggle, but it's still not ready to merge, so
this recipe leaves it out for now.
This PR does not touch
--max-num-batched-tokens: the existing TP2 c64 (8192)/ c128 (4096) values in this recipe already avoid the KV-headroom regression
that a separate, since-closed draft (#3451) flagged as unvalidated.
Duplicate-check: searched open PRs for
dsv41flash mi355x,attention-config,indexer_kv_dtype,mxfp4 indexer,VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLAand
block-size 128on this repo; no open PR adds this env var or touchesthis recipe's image pin.
AI model disclosure
reviewed the unmerged AMD vLLM cherry-pick PRs, ran a KV-headroom
root-cause investigation and validation sweep for a related TP2 c64
regression, root-caused a pre-existing
#58208Triton bug hit only by thedense fp8 indexer fallback path at AgentX-scale context, fixed the DCO on
and helped land its revert (
vllm-project/vllm#59125), ran a live A/B testmeasuring
#58671's actual perf contribution once that confound wasremoved, and revised this PR to drop
#58671-specific changes once it wasclear it wouldn't merge in time. The submitting human reviewed the
recipe/config changes and is responsible for validating the env var on the
current image once one is available.
Related Issue
Fixes #
Type of Change
Checklist
perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR中文
为
dsv41flash-fp4-mi355x-vllm-agentic-dspark添加VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True,以强制启用vllm-project/vllm#53492 的
Gluon sparse-MLA kernel 路径(与下方本地验证运行实际设置的配置保持一致)。
将
image/model.container取消固定为vllm/vllm-openai-rocm:nightly-rocm100-TBD(之后只需更新末尾的哈希值):当前固定的镜像早于本配方所 cherry-pick 的 AMD vLLM PR,与其无关。待以下改动合并
至上游并随 nightly 镜像发布后,维护者应填入实际哈希值:
这些改动共同为 MI355X 上的 DSv4.1-Flash 带来更优的 topk/mHC/MLA kernel,以及
更好的计算/通信流重叠模式。
本 PR 此前还添加了
--attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'与
--block-size 128的绕过配置,用于启用vllm-project/vllm#58671 的
ROCm 分页 MXFP4 稀疏 logits indexer kernel。二者均已移除:#58671 无法及时合并;
而
--block-size 128的全部依据(vLLM 的 block-size 自动推断在该模型 4 个以上attention backend 同时生效时会退化为选出 64)本身就是针对 fp4 indexer 生效后
DeepseekV4ROCMAiterMLASparseBackend.get_preferred_block_size所请求的 128 —在不启用该 indexer 的情况下已不适用。一次真实的 A/B 测试(TP2、并发 16、匹配
900 秒 profiling 窗口,并已通过
vllm-project/vllm#59125 在
上游回滚
#58208以避免稠密回退路径崩溃)测得:#58671本应替换的稠密 fp8indexer 路径,其 interactivity p50/p90 降低 14.7%/14.9%,e2e 延迟 p50/p90
增加 9.5%/10.8%(更差),throughput/GPU 基本不变(+0.5%,在噪声范围内),
TTFT 不变。也就是说
#58671确实带来了可观的 interactivity/尾部延迟提升,而不只是一个配置开关,但目前仍未准备好合并,因此本配方暂不启用它。
本 PR 未改动
--max-num-batched-tokens:该配方现有的 TP2 c64(8192)/c128(4096)取值已经规避了另一个已关闭的草稿 PR(#3451)标记为"尚未验证"的
KV 显存余量回退问题。
重复性检查:已在本仓库搜索
dsv41flash mi355x、attention-config、indexer_kv_dtype、mxfp4 indexer、VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA与
block-size 128相关的未关闭 PR,未发现任何未关闭 PR 添加该环境变量或改动本配方的镜像固定值。