From 234bddbc22eb2b288641c9369a675f0397ce498f Mon Sep 17 00:00:00 2001 From: Honglie Yi Date: Wed, 30 Sep 2026 07:37:03 +0000 Subject: [PATCH 1/3] perf(amd): bump DSV4 ATOM AgentX image to nightly_202609291501 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Image-only update of dsv4-fp4-mi355x-atom-agentic-mtp on its native srt-slurm recipe; points, DSpark K6, golden AL and env are unchanged. 将 dsv4-fp4-mi355x-atom-agentic-mtp 的原生 srt-slurm 配方镜像更新为 nightly_202609291501;测试点、DSpark K6、golden AL 和环境变量均不变。 Co-Authored-By: Claude Opus 5.5 (1M context) --- .../dsv4/atom/mi355x-fp4-mtp/agentic.yaml | 2 +- inferencex-e2e/configs/amd-master.yaml | 2 +- inferencex-e2e/docs/configuration-procedures.md | 8 ++++++-- inferencex-e2e/docs/configuration-procedures_zh.md | 7 +++++-- inferencex-e2e/perf-changelog.yaml | 10 ++++++++++ 5 files changed, 23 insertions(+), 6 deletions(-) diff --git a/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml b/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml index 49bdd63b51..592a91827b 100644 --- a/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml +++ b/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml @@ -6,7 +6,7 @@ base: name: dsv4-fp4-mi355x-atom-agentic model: path: hf:deepseek-ai/DeepSeek-V4-Pro-0813 - container: rocm/atom-dev:nightly_202609161445 + container: rocm/atom-dev:nightly_202609291501 precision: fp4 resources: gpu_type: mi355x diff --git a/inferencex-e2e/configs/amd-master.yaml b/inferencex-e2e/configs/amd-master.yaml index 8feef1b86f..516e2f46ac 100644 --- a/inferencex-e2e/configs/amd-master.yaml +++ b/inferencex-e2e/configs/amd-master.yaml @@ -924,7 +924,7 @@ dsv4-fp4-mi355x-vllm-agentic-mtp: # Both throughput bands use thinking_on golden AL 3.77; eval uses real # acceptance. Keep the historical config key for changelog compatibility. dsv4-fp4-mi355x-atom-agentic-mtp: - image: rocm/atom-dev:nightly_202609161445 + image: rocm/atom-dev:nightly_202609291501 model: deepseek-ai/DeepSeek-V4-Pro-0813 model-prefix: dsv4 runner: cluster:mi355x-amds diff --git a/inferencex-e2e/docs/configuration-procedures.md b/inferencex-e2e/docs/configuration-procedures.md index d0ed2d739b..0403d4eac9 100644 --- a/inferencex-e2e/docs/configuration-procedures.md +++ b/inferencex-e2e/docs/configuration-procedures.md @@ -269,7 +269,7 @@ All ten AgentX throughput points use DSpark K6 (target verification length 7) and the committed golden AL 3.77. C1/2/4/8/16 use TP8/EP1; C48/64/96/128/256 use TP8/DPA8/EP8 with native RCCL. Each point runs for 3600 seconds. The C256 full GSM8K eval omits forced acceptance. Keep the -pinned `rocm/atom-dev:nightly_202609161445` image and GPU-only KV. C1 through C16 use +pinned `rocm/atom-dev:nightly_202609291501` image and GPU-only KV. C1 through C16 use BF16 KV, while C48 and above retain FP8 KV; all points use the FP4 index cache, 8192-token checkpoints and DEP dense FULL graph ladder. Fixed q7 graphs are captured in each new server; confirm target and DSpark draft capture in @@ -282,8 +282,12 @@ record model/source identity and requested settings. Successful startup, graph capture and requests require runtime log evidence. The pinned image is the official ATOM nightly -`rocm/atom-dev:nightly_202609161445`, which includes the merged +`rocm/atom-dev:nightly_202609291501` (ATOM `0.1.7.dev46+g74fd942b0`, ROCm 7.2.4), +which includes the merged [ROCm/ATOM#2233](https://github.com/ROCm/ATOM/pull/2233) inference-mode fix. +From this image ATOM stores the checkpoint's `ue8m0` FP8 block scales as E8M0 +on gfx950 by default ([ROCm/ATOM#2419](https://github.com/ROCm/ATOM/pull/2419)); +the powers-of-two scales are represented exactly. The recipe does not patch AITER source at runtime; TP communication fusion, DSpark K6 and graph capture use the implementation shipped in the image. diff --git a/inferencex-e2e/docs/configuration-procedures_zh.md b/inferencex-e2e/docs/configuration-procedures_zh.md index a659d68c46..76d604935d 100644 --- a/inferencex-e2e/docs/configuration-procedures_zh.md +++ b/inferencex-e2e/docs/configuration-procedures_zh.md @@ -248,7 +248,7 @@ DSpark Markov/confidence head、全部 66 个分片的 header 与 payload 边界 全部十个 AgentX 性能点使用 DSpark K6(target 验证长度为 7)和已提交的 golden AL 3.77。C1/2/4/8/16 使用 TP8/EP1;C48/64/96/128/256 使用 TP8/DPA8/EP8 原生 RCCL。每个性能点运行 3600 秒。C256 全量 GSM8K 不传强制 -接受率参数。保留固定的 `rocm/atom-dev:nightly_202609161445` 镜像和 GPU KV;C1 至 C16 +接受率参数。保留固定的 `rocm/atom-dev:nightly_202609291501` 镜像和 GPU KV;C1 至 C16 使用 BF16 KV,C48 及以上继续使用 FP8 KV,所有任务均使用 FP4 index cache、 8192-token checkpoint 和 DEP dense FULL graph 阶梯。每个新服务进程重新捕获固定 q7 图;必须从 `server.log` 确认 target 和 DSpark draft capture 完成。confidence @@ -259,8 +259,11 @@ schedule 和 ragged verification 保持关闭。 `runtime_manifest.json` 和 `server_command.txt` 保存模型/源码身份及请求的配置。 成功启动、graph capture 和请求执行仍需运行时日志证明。 -固定镜像为官方 ATOM nightly `rocm/atom-dev:nightly_202609161445`,已包含已合入的 +固定镜像为官方 ATOM nightly `rocm/atom-dev:nightly_202609291501`(ATOM `0.1.7.dev46+g74fd942b0`, +ROCm 7.2.4),已包含已合入的 [ROCm/ATOM#2233](https://github.com/ROCm/ATOM/pull/2233) inference-mode 修复。 +自该镜像起,ATOM 在 gfx950 上默认以 E8M0 存储检查点的 `ue8m0` FP8 block scale +([ROCm/ATOM#2419](https://github.com/ROCm/ATOM/pull/2419)),2 的幂次 scale 可被精确表示。 配方不再在运行时修改 AITER 源码;TP 通信融合、DSpark K6 和 graph capture 直接使用镜像内实现。 diff --git a/inferencex-e2e/perf-changelog.yaml b/inferencex-e2e/perf-changelog.yaml index 2b184929a0..1ac72db004 100644 --- a/inferencex-e2e/perf-changelog.yaml +++ b/inferencex-e2e/perf-changelog.yaml @@ -9127,3 +9127,13 @@ - "Add Kimi-K3 MXFP4 vLLM agentic-coding on MI355X (TP8, DSpark): dcp1 c1/c4 GPU-resident and c8-c14 with SimpleCPUOffload DRAM offload, plus a new dcp8 c44/c48/c70 throughput band (mtp synthetic acceptance, no draft); vLLM ROCm image vllm/vllm-openai-rocm:nightly-rocm100-36768d1bfd39094681cdbc8cb37d4b31c0729c89" - "The Inferact/Kimi-K3-DSpark draft (K3DSparkModel, model_type k3_dspark, 5 layers, hidden 7168, torch_dtype bfloat16, no quantization_config) keeps every layer at its pristine dtype. vLLM builds all its modules with quant_config from get_draft_quant_config() (vllm/models/kimi_k3/nvidia/dspark_mla.py), which returns None because K3DSparkModel is explicitly excluded from the DeepSeek-V4 branch that would set draft.quantization = target.quantization (vllm/config/speculative.py:1431); so context_proj (ReplicatedLinear), context_kv_proj (MergedColumnParallelLinear) and each decoder layer's MLA q/kv projections and dense KimiMLP (gate/up/down) load unquantized in BF16 -- no ptpc_fp8, no mxfp4, no INT4 weights. The draft has no FusedMoE/block_sparse_moe, so VLLM_ROCM_USE_AITER_MOE_SITUV2 (A8W4) never touches it, and its KV cache stays fp8 (kv_cache_dtype). This PR only changes draft depth (num_speculative_tokens 4->7) and enables INT4 custom quick all-reduce (VLLM_ROCM_QUICK_REDUCE_QUANTIZATION=INT4, vllm/distributed/device_communicators/quick_all_reduce.py); because the draft shares the target's TP8 group (vllm/v1/spec_decode/draft_model.py), that INT4 applies to its tensor-parallel reductions too -- a collective-reduction transport precision, not any draft weight, activation, or KV dtype." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3561 + +- config-keys: + - dsv4-fp4-mi355x-atom-agentic-mtp + scenario-type: + - agentic-coding + description: + - "Update the ATOM image from rocm/atom-dev:nightly_202609161445 (ATOM 0.1.6rc1.dev477+g709c25be0) to rocm/atom-dev:nightly_202609291501 (digest sha256:e0d7134efff68c251acdc2612a9b6c238bc4da5d613832b2bd1251f34bd93355; ATOM 0.1.7.dev46+g74fd942b0, ROCm 7.2.4). Image-only change on the existing native srt-slurm recipe: all ten points (TP8/EP1 at concurrency 1, 2, 4, 8, 16; TP8/DPA8/EP8 native RCCL at 48, 64, 96, 128, 256), DSpark K6 with the thinking_on golden AL 3.77 for throughput and real acceptance for eval, KV and index-cache dtypes, graph capture and all environment settings are unchanged." + - "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid." + - "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only." + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX From a65d950cb1f2b6858b8bbbe02fbc1c2251abf6a3 Mon Sep 17 00:00:00 2001 From: Honglie Yi Date: Wed, 30 Sep 2026 07:37:43 +0000 Subject: [PATCH 2/3] chore: set perf-changelog pr-link to #3605 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 将 perf-changelog 的 pr-link 设为 #3605。 Co-Authored-By: Claude Opus 5.5 (1M context) --- inferencex-e2e/perf-changelog.yaml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/inferencex-e2e/perf-changelog.yaml b/inferencex-e2e/perf-changelog.yaml index 1ac72db004..3ada1feafe 100644 --- a/inferencex-e2e/perf-changelog.yaml +++ b/inferencex-e2e/perf-changelog.yaml @@ -9136,4 +9136,4 @@ - "Update the ATOM image from rocm/atom-dev:nightly_202609161445 (ATOM 0.1.6rc1.dev477+g709c25be0) to rocm/atom-dev:nightly_202609291501 (digest sha256:e0d7134efff68c251acdc2612a9b6c238bc4da5d613832b2bd1251f34bd93355; ATOM 0.1.7.dev46+g74fd942b0, ROCm 7.2.4). Image-only change on the existing native srt-slurm recipe: all ten points (TP8/EP1 at concurrency 1, 2, 4, 8, 16; TP8/DPA8/EP8 native RCCL at 48, 64, 96, 128, 256), DSpark K6 with the thinking_on golden AL 3.77 for throughput and real acceptance for eval, KV and index-cache dtypes, graph capture and all environment settings are unchanged." - "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid." - "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only." - pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3605 From ba3a288251e7c07bab2a31362bd6e176d79bb4e2 Mon Sep 17 00:00:00 2001 From: seungrokj Date: Wed, 30 Sep 2026 16:04:49 -0700 Subject: [PATCH 3/3] perf(amd): bump DSV4 ATOM AgentX image to nightly_202609291501 Update the ATOM image from rocm/atom-dev:nightly_202609161445 (ATOM 0.1.6rc1.dev477+g709c25be0) to rocm/atom-dev:nightly_202609291501 (digest sha256:e0d7134efff68c251acdc2612a9b6c238bc4da5d613832b2bd1251f34bd93355; ATOM 0.1.7.dev46+g74fd942b0, ROCm 7.2.4). Image-only change on the existing native srt-slurm recipe: all ten points (TP8/EP1 at concurrency 1, 2, 4, 8, 16; TP8/DPA8/EP8 native RCCL at 48, 64, 96, 128, 256), DSpark K6 with the thinking_on golden AL 3.77 for throughput and real acceptance for eval, KV and index-cache dtypes, graph capture and all environment settings are unchanged. Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid. No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only. Co-Authored-By: Claude Opus 5.5