Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,9 @@ base:
name: dsv41flash-fp4-mi355x-vllm-agentic
model:
path: hf:deepseek-ai/DeepSeek-V4.1-Flash
container: vllm/vllm-openai-rocm:nightly-rocm100-7f1a5398e9610d96c473931a26c0e12bbe0d0423
# Repinned to a nightly build that includes vllm-project/vllm#58655
# (mHC fused Triton seams) and #53492 (sparse MLA Gluon kernel).
container: vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89
precision: fp4
resources:
gpu_type: mi355x
Expand Down Expand Up @@ -54,6 +56,7 @@ base:
VLLM_ROCM_USE_AITER: '1'
VLLM_ROCM_USE_AITER_MOE: '1'
AITER_TRITON_LOG_LEVEL: ERROR
VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA: 'True'
# DeepseekV41ForCausalLM is not torch-compiled upstream; breakable
# graphs keep FULL_AND_PIECEWISE capture working.
VLLM_USE_BREAKABLE_CUDAGRAPH: '1'
Expand Down
5 changes: 3 additions & 2 deletions inferencex-e2e/configs/amd-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1391,8 +1391,9 @@ dsv4-fp4-mi355x-sglang-agentic-mtp:
# Two GPUs per server doubles the servers per node and is the layout that
# decides whether DSv4.1-Flash is throughput- or interactivity-bound here.
dsv41flash-fp4-mi355x-vllm-agentic-dspark:
# ROCm 10.0 nightly channel, shared with kimik3-fp4-mi355x-vllm-agentic-mtp.
image: vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098
# Repinned to a nightly build that includes vllm-project/vllm#58655 and
# #53492 (see agentic.yaml).
image: vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89
model: deepseek-ai/DeepSeek-V4.1-Flash
model-prefix: dsv41flash
runner: cluster:mi355x-amds
Expand Down
7 changes: 7 additions & 0 deletions inferencex-e2e/perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -9120,3 +9120,10 @@
- "Restore the DCP8 LMCache bands on the native srt-slurm recipe: concurrency 14 and 16 (DSpark 3, ReplaySSM) and 48, 56 and 72 (no draft) run ATOM's in-process lmcache_offload connector through roles.agg.args.extra-kv-connectors (srt-slurm patch 507), with 128 GB/rank up to 48 and 192 GB/rank at 56 and 72. Concurrency 1 and 4 stay GPU-resident."
- "No change to the Inferact/Kimi-K3-DSpark draft's precision: online_quant_config still excludes every draft linear (layers.*, context_proj), so its weights and activations stay BF16, and it keeps the target's FP8 KV cache (kv_cache_dtype fp8). FlyDSL FP8 prefill attention applies only to the target, since the draft runs its block pass as decode attention."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3407

- config-keys:
- dsv41flash-fp4-mi355x-vllm-agentic-dspark
description:
- "Force VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True to enable vllm-project/vllm#53492's gfx950 Gluon sparse-MLA kernel. Repin the recipe and master config to vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89, which includes #53492 and vllm-project/vllm#58655 (fused mHC Triton kernel). vllm-project/vllm#58208 (DSA candidate topk) was reverted by vllm-project/vllm#59125 before this build, and vllm-project/vllm#58671's MXFP4 sparse indexer landed after it; the indexer is left to #3571."
- "设置 VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True,启用 vllm-project/vllm#53492 的 gfx950 Gluon sparse-MLA kernel。将配方与主配置固定到 vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89,该镜像包含 #53492 与 vllm-project/vllm#58655(融合 mHC Triton kernel)。vllm-project/vllm#58208(DSA 候选块 topk)在该构建之前已被 vllm-project/vllm#59125 回滚;vllm-project/vllm#58671 的 MXFP4 稀疏 indexer 在该构建之后才合入,留待 #3571。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3555
Loading