diff --git a/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml b/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml index 49bdd63b51..592a91827b 100644 --- a/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml +++ b/inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml @@ -6,7 +6,7 @@ base: name: dsv4-fp4-mi355x-atom-agentic model: path: hf:deepseek-ai/DeepSeek-V4-Pro-0813 - container: rocm/atom-dev:nightly_202609161445 + container: rocm/atom-dev:nightly_202609291501 precision: fp4 resources: gpu_type: mi355x diff --git a/inferencex-e2e/configs/amd-master.yaml b/inferencex-e2e/configs/amd-master.yaml index 27d23c28a5..a852be37b6 100644 --- a/inferencex-e2e/configs/amd-master.yaml +++ b/inferencex-e2e/configs/amd-master.yaml @@ -946,7 +946,7 @@ dsv4-fp4-mi355x-vllm-agentic-mtp: # Both throughput bands use thinking_on golden AL 3.77; eval uses real # acceptance. Keep the historical config key for changelog compatibility. dsv4-fp4-mi355x-atom-agentic-mtp: - image: rocm/atom-dev:nightly_202609161445 + image: rocm/atom-dev:nightly_202609291501 model: deepseek-ai/DeepSeek-V4-Pro-0813 model-prefix: dsv4 runner: cluster:mi355x-amds diff --git a/inferencex-e2e/docs/configuration-procedures.md b/inferencex-e2e/docs/configuration-procedures.md index 3a9a647039..d2df756784 100644 --- a/inferencex-e2e/docs/configuration-procedures.md +++ b/inferencex-e2e/docs/configuration-procedures.md @@ -275,7 +275,7 @@ All ten AgentX throughput points use DSpark K6 (target verification length 7) and the committed golden AL 3.77. C1/2/4/8/16 use TP8/EP1; C48/64/96/128/256 use TP8/DPA8/EP8 with native RCCL. Each point runs for 3600 seconds. The C256 full GSM8K eval omits forced acceptance. Keep the -pinned `rocm/atom-dev:nightly_202609161445` image and GPU-only KV. C1 through C16 use +pinned `rocm/atom-dev:nightly_202609291501` image and GPU-only KV. C1 through C16 use BF16 KV, while C48 and above retain FP8 KV; all points use the FP4 index cache, 8192-token checkpoints and DEP dense FULL graph ladder. Fixed q7 graphs are captured in each new server; confirm target and DSpark draft capture in @@ -288,8 +288,12 @@ record model/source identity and requested settings. Successful startup, graph capture and requests require runtime log evidence. The pinned image is the official ATOM nightly -`rocm/atom-dev:nightly_202609161445`, which includes the merged +`rocm/atom-dev:nightly_202609291501` (ATOM `0.1.7.dev46+g74fd942b0`, ROCm 7.2.4), +which includes the merged [ROCm/ATOM#2233](https://github.com/ROCm/ATOM/pull/2233) inference-mode fix. +From this image ATOM stores the checkpoint's `ue8m0` FP8 block scales as E8M0 +on gfx950 by default ([ROCm/ATOM#2419](https://github.com/ROCm/ATOM/pull/2419)); +the powers-of-two scales are represented exactly. The recipe does not patch AITER source at runtime; TP communication fusion, DSpark K6 and graph capture use the implementation shipped in the image. diff --git a/inferencex-e2e/docs/configuration-procedures_zh.md b/inferencex-e2e/docs/configuration-procedures_zh.md index 399f0f52e7..0d023deb1c 100644 --- a/inferencex-e2e/docs/configuration-procedures_zh.md +++ b/inferencex-e2e/docs/configuration-procedures_zh.md @@ -255,7 +255,7 @@ DSpark Markov/confidence head、全部 66 个分片的 header 与 payload 边界 全部十个 AgentX 性能点使用 DSpark K6(target 验证长度为 7)和已提交的 golden AL 3.77。C1/2/4/8/16 使用 TP8/EP1;C48/64/96/128/256 使用 TP8/DPA8/EP8 原生 RCCL。每个性能点运行 3600 秒。C256 全量 GSM8K 不传强制 -接受率参数。保留固定的 `rocm/atom-dev:nightly_202609161445` 镜像和 GPU KV;C1 至 C16 +接受率参数。保留固定的 `rocm/atom-dev:nightly_202609291501` 镜像和 GPU KV;C1 至 C16 使用 BF16 KV,C48 及以上继续使用 FP8 KV,所有任务均使用 FP4 index cache、 8192-token checkpoint 和 DEP dense FULL graph 阶梯。每个新服务进程重新捕获固定 q7 图;必须从 `server.log` 确认 target 和 DSpark draft capture 完成。confidence @@ -266,8 +266,11 @@ schedule 和 ragged verification 保持关闭。 `runtime_manifest.json` 和 `server_command.txt` 保存模型/源码身份及请求的配置。 成功启动、graph capture 和请求执行仍需运行时日志证明。 -固定镜像为官方 ATOM nightly `rocm/atom-dev:nightly_202609161445`,已包含已合入的 +固定镜像为官方 ATOM nightly `rocm/atom-dev:nightly_202609291501`(ATOM `0.1.7.dev46+g74fd942b0`, +ROCm 7.2.4),已包含已合入的 [ROCm/ATOM#2233](https://github.com/ROCm/ATOM/pull/2233) inference-mode 修复。 +自该镜像起,ATOM 在 gfx950 上默认以 E8M0 存储检查点的 `ue8m0` FP8 block scale +([ROCm/ATOM#2419](https://github.com/ROCm/ATOM/pull/2419)),2 的幂次 scale 可被精确表示。 配方不再在运行时修改 AITER 源码;TP 通信融合、DSpark K6 和 graph capture 直接使用镜像内实现。 diff --git a/inferencex-e2e/perf-changelog.yaml b/inferencex-e2e/perf-changelog.yaml index b415cf56a5..9b9c48a0ef 100644 --- a/inferencex-e2e/perf-changelog.yaml +++ b/inferencex-e2e/perf-changelog.yaml @@ -9158,3 +9158,13 @@ - "Port the DeepSeek-V4-Pro-0813 MI355X ATOM 1P1D AgentX config from the legacy amd_utils path to a native srt-slurm recipe (benchmarks/multi_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4/agentx/disagg-lmcache-dspark.yaml), one override variant per point. Server flags, env, AToMesh routing policies and decode CUDA-graph capture sizes match the legacy server_atom.sh for every tier: TP8 at concurrency 1 and 16, DP attention with prefill TBO at 64 and 128, and DP attention with ATOM's in-process LMCache CPU offload (lmcache_offload, 187 GB per prefill rank) at 256. Golden acceptance 3.01 for DSpark with three draft tokens is now injected by the srt-slurm path, the same value the legacy models_atom.yaml hardcoded." - "The LMCache tier uses srt-slurm's extra-kv-connectors (NVIDIA/srt-slurm#507, in v2.36.0) to add lmcache_offload next to the generated Mooncake connector." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3543 + +- config-keys: + - dsv4-fp4-mi355x-atom-agentic-mtp + scenario-type: + - agentic-coding + description: + - "Update the ATOM image from rocm/atom-dev:nightly_202609161445 (ATOM 0.1.6rc1.dev477+g709c25be0) to rocm/atom-dev:nightly_202609291501 (digest sha256:e0d7134efff68c251acdc2612a9b6c238bc4da5d613832b2bd1251f34bd93355; ATOM 0.1.7.dev46+g74fd942b0, ROCm 7.2.4). Image-only change on the existing native srt-slurm recipe: all ten points (TP8/EP1 at concurrency 1, 2, 4, 8, 16; TP8/DPA8/EP8 native RCCL at 48, 64, 96, 128, 256), DSpark K6 with the thinking_on golden AL 3.77 for throughput and real acceptance for eval, KV and index-cache dtypes, graph capture and all environment settings are unchanged." + - "Relevant ATOM changes in the range: V4 decode reuses the sparse prefill ASM (ROCm/ATOM#2271), greedy sampler picks via aiter.topk_select (ROCm/ATOM#2244), and FP8 block scales declared scale_fmt ue8m0 are now stored as E8M0 on gfx950 by default (ROCm/ATOM#2419; previously FP32 unless ATOM_FP8_BLOCKSCALE_USE_E8M0_SCALE=1). The upstream DeepSeek-V4 recipes are unchanged across the range, and every recipe flag and choice (all2all-backend rccl, dp-load-balance least_tokens, moe-backend standard) remains valid." + - "No data-type or precision change to the DeepSeek-V4-Pro-0813 DSpark draft: no online quantization is configured, so it keeps its checkpoint precision; the E8M0 scale storage represents the checkpoint's power-of-two block scales exactly. kv-cache-dtype and index-cache-dtype touch cache storage only." + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3605