Conversation
…cipe Track recipes/Agentic-Kimi-K3.md as retuned in ROCm/ATOM#2382: enable FlyDSL FP8 prefill attention on every band, and hold a ready prefill for four decode passes from concurrency 16 up. The published concurrency set and every other launch argument are unchanged. 将 MI355X Kimi-K3 FP4 ATOM AgentX 提交切换到 0924 镜像,并跟随 ROCm/ATOM#2382 重调后的 recipe:全部并发开启 FlyDSL FP8 prefill attention;并发 16 及以上时让就绪的 prefill 等待 4 个 decode 轮次。 已发布的并发点集合与其余启动参数保持不变。 Co-Authored-By: Claude Opus 5 <[email protected]>
|
Thanks for the contribution!
中文感谢你的贡献!
|
将 perf-changelog 条目的 pr-link 指向 PR 3407。 Co-Authored-By: Claude Opus 5 <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36211860108 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36211860108 |
…1613 Co-Authored-By: Claude Opus 4.6 <[email protected]>
|
InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
…x-0924 Port the ATOM image bump and FlyDSL FP8 prefill attention onto the native srt-slurm recipe; the legacy script is deleted on main.
Carry srt-slurm patch 507 so ATOM aggregate workers accept extra-kv-connectors, and restore the DCP8 bands from the legacy config and ROCm/ATOM recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 with DSpark 3 and ReplaySSM, 48, 56 and 72 without a draft, all on the in-process lmcache_offload connector (128 GB/rank, 192 GB/rank at 56 and 72). The PrefillDelayer applies from concurrency 16 up. The single-node adapter now reads ATOM's decode-context-parallel-size for the DCP_SIZE check, as it does for vLLM.
Track the checked-in ATOM recipe
recipes/Agentic-Kimi-K3.mdas retuned in ROCm/ATOM#2382, on imagekimi_k3_agentic_0924.The published concurrency set [1, 4, 14, 16, 48, 56, 72] is unchanged; no points are added or dropped.
ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1on every band. ATOM defaults it to0, so the recipe's prefill attention path was not reached before this change.ATOM_PREFILL_DECODE_INTERVAL=4andATOM_PREFILL_DELAYER_MAX_QUEUE_MS=5000hold a ready prefill for four decode passes instead of interleaving it into every step. This is a threshold, not a band: concurrency 1, 4 and 14 run without it.max-num-seqs,max-num-batched-tokens,gpu-memory-utilization, the CUDA-graph ladder,dcp-size, draft depth, synthetic acceptance, ReplaySSM placement,AITER_REUSE_IDENTICAL_COMM_GROUPSand LMCache sizing are unchanged from #3207.Re-created on an in-repo
amd/branch so sweep dispatch and labels (AMD,agentx,full-sweep-enabled) apply.AI model disclosure
Prepared with Claude Code using
claude-opus-5(recipe reconciliation, edits, changelog entry). No other model contributed.Port to native srt-slurm and LMCache
Merged
mainin; single-node AgentX now runs on the native recipeinferencex-e2e/benchmarks/single_node/srt-slurm-recipes/kimik3/atom/mi355x-fp4-mtp/agentic.yaml, and the legacy script is gone.rocm/atom-dev:nightly_202609251613),ATOM_USE_FLYDSL_FP8_PREFILL_ATTN=1, and the PrefillDelayer from concurrency 16 up.recipes/Agentic-Kimi-K3.md: concurrency 14 and 16 (DSpark 3, ReplaySSM) and 48, 56 and 72 (no draft) on ATOM's in-processlmcache_offloadconnector, 128 GB/rank up to 48 and 192 GB/rank at 56 and 72, chunk size 1024,PYTHONHASHSEED=0. Concurrency 1 and 4 are unchanged.roles.agg.args.extra-kv-connectorsfrom srt-slurm patch507-lmcache-server-atom-sglang.patch(feat: add LMCache support for ATOM and SGLang srt-slurm#32).decode-context-parallel-sizefor thedcp-sizecheck, as it already does for vLLM.