Skip to content

[Klaud Cold] [TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / [Klaud Cold] [TileRT] GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill - #3623

Open
functionstackx wants to merge 7 commits into
mainfrom
klaud/glm5.3-fp8-mi355x-tilert-post3
Open

functionstackx wants to merge 7 commits into
mainfrom
klaud/glm5.3-fp8-mi355x-tilert-post3

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Upstream port of #3617 (fork PR by @CrimsonDump, #3617) so the sweep can run from an in-repo branch. Commits are cherry-picked unchanged; origin/main was merged in (perf-changelog conflict with #3605 resolved by appending this entry after it) and the changelog pr-link now points at this PR.

本 PR 是 #3617(@CrimsonDump 的 fork PR)的上游移植,以便从仓库内分支运行 sweep。提交保持不变;已合入 origin/main(与 #3605 的 perf-changelog 冲突通过把本条目追加在其后解决),changelog 的 pr-link 改为指向本 PR。

Description

Now that #3552 has merged, this re-targets main (it replaces #3563, which GitHub closed automatically when #3552's branch was deleted). Updates the glm5.3-fp8-mi355x-tilert-agentic recipe (AgentX, concurrency 1, 3600 s) in two ways. Nothing else in the recipe changes: decode image, 1M context, bf16 KV, layer-sharded on-GPU PD buffers, GLM5_AR_N=2.

1. tilert 0.1.6.post2 → 0.1.6.post3 (PyPI, 2026-09-28, sha256:d6fbf0a55be1fbde…), on both TileRT ranks; router metadata follows. The engine .so files are byte-identical to post2; the change is in tilert/pd_vllm (4 files, +622 / −12). Three prefill→decode latency features are now on by default:

Feature Switch (post3 default) What it does
Multi-sender staging TILERT_PD_SENDERS=8 (was 1) Every prefill TP rank extracts and RDMA-sends the layers it owns, instead of rank 0 sending all 79
Per-layer pipelined send TILERT_PD_PIPELINE=1 (was 0) A layer's KV is extracted and written as soon as that layer is computed, overlapping the transfer with the rest of the forward
Router-side incremental tokenization TILERT_ROUTER_TOKENIZE=1 (was 0) The router renders the chat template and tokenizes through a per-segment cache (only the new turn is tokenized), then sends vLLM token ids on /v1/completions. A start-up self-check against vLLM /tokenize, and a cross-check of vLLM's own --chat-template / --default-chat-template-kwargs, keep it off on any mismatch; tools, logprobs, structured output and non-text content always take vLLM's chat path

2. Prefill image → vLLM 0.28 + ATOM (ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built on rocm/atom-dev:vllm-v0.28.0-nightly_20260923: vLLM 0.28.1.dev0, ATOM 0.1.7.dev17, ROCm 7.2.4). The only additions on top of the base image are the dependencies the TileRT connector needs and WORKDIR /app (with WORKDIR /, any PYTHONPYCACHEPREFIX makes the ATOM plugin fail at import). In the recipe's prefill role, enforce-eager is dropped and args add CUDA graphs (compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"}), async-scheduling, load-format: fastsafetensors, enable-prefix-caching and max-num-batched-tokens: 16384; env adds the two AITER settings of the in-tree GLM-5.2 ATOM MI355X agentic recipe (benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml), which uses the same prefix caching and chunk size.

Evidence (local 2×8 MI350X, same recipe, concurrency 1, 3600 s, unfiltered AgentX corpus, 1M context)

The 3600 s run used the pre-release build that post3 was cut from, with the three features switched on by environment variables; post3 is that code with the defaults flipped (verified: identical AST, and a clean-environment import reads 8 / 1 / 1). The local runs went through the bash launcher that #3552 replaces, with the same vLLM and TileRT settings as this recipe; the srt-slurm path itself is exercised by the sweep.

post2, official MI355X run (#3389) this PR, local MI350X
submission_valid / coverage true / 100% true / 100% (TTFT and ITL)
requests / errors 271 / 0 294 / 0
TTFT p50 / p90 / p99 (ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 309k / 491k 333k / 531k

Paired by request (same ISL and OSL, n = 271): TTFT p90 4348 → 1721 ms, per-request TTFT ratio median 0.344. Step by step on the same MI350X pair (each step changes one thing):

Step TTFT p90, paired Per-request ratio
hardware only: MI355X → MI350X, post2 on both +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM (+ flags above) −41.7% 0.810
multi-sender + pipelined send −32.5% 0.595
router-side tokenization −17.3% 0.743

Decode is untouched: against post2 on the same MI350X pair, same-batch intvty p50/p90 moves +0.21% / +0.16%. (Against the MI355X run it is ~7% lower, which is the MI350X/MI355X clock difference; the official run on MI355X is the number that counts.)

Accuracy, GSM8K (lm-eval, 1319 questions, 5-shot, real MTP) on the same stack: strict / flexible 0.9742 / 0.9742 and 0.9757 / 0.9757 on the two pre-release builds (post1 on vLLM 0.24: 0.9765 / 0.9757). In the second run the router cross-checked all 1319 prompts against vLLM /tokenize: 0 mismatches.

Checklist notes

AI model disclosure

  • Model/version: claude-opus-5-5[1m] (Claude Opus 5.5, 1M context), via Claude Code
  • Role: implemented the post3 pd_vllm changes, ran the local validation, and drafted this PR

Related Issue

N/A

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR
中文

改动说明

#3552 合入后改为针对 main(替代 #3563——#3552 的分支被删除时 GitHub 自动关闭了它)。更新 glm5.3-fp8-mi355x-tilert-agentic 配方(AgentX,并发 1,3600 秒),共两处。其余不变:decode 镜像、1M 上下文、bf16 KV、按层分片的 GPU PD 缓冲、GLM5_AR_N=2。

1. tilert 0.1.6.post2 → 0.1.6.post3(PyPI,2026-09-28,sha256:d6fbf0a55be1fbde…),两侧 TileRT rank 同步,router 元数据随之更新。引擎 .so 与 post2 逐字节相同,改动全在 tilert/pd_vllm(4 个文件,+622 / −12)。三项 prefill→decode 时延优化改为默认开启:

功能 开关(post3 默认) 作用
多发送端暂存 TILERT_PD_SENDERS=8(原为 1) prefill 的每个 TP rank 抽取并 RDMA 发送自己负责的层,而不是由 rank 0 发全部 79 层
逐层流水线发送 TILERT_PD_PIPELINE=1(原为 0) 每层算完即抽取并写出该层 KV,传输与后续层的前向重叠
router 侧增量分词 TILERT_ROUTER_TOKENIZE=1(原为 0) router 渲染 chat 模板,经按段缓存分词(只对新一轮分词),再把 token id 经 /v1/completions 交给 vLLM。启动时与 vLLM /tokenize 自检,并核对 vLLM 自己的 --chat-template / --default-chat-template-kwargs,任何不一致即保持关闭;带 tools、logprobs、结构化输出或非纯文本内容的请求一律走 vLLM 的 chat 路径

2. prefill 镜像改为 vLLM 0.28 + ATOM(ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于 rocm/atom-dev:vllm-v0.28.0-nightly_20260923:vLLM 0.28.1.dev0、ATOM 0.1.7.dev17、ROCm 7.2.4)。在基底镜像之上只加了 TileRT connector 所需的依赖和 WORKDIR /app(若为 WORKDIR /,只要设了 PYTHONPYCACHEPREFIX,ATOM 插件在 import 时就会失败)。配方的 prefill 角色去掉 enforce-eager,args 加上 CUDA graph(compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"})、async-scheduling、load-format: fastsafetensors、enable-prefix-caching 与 max-num-batched-tokens: 16384;env 加上在树 GLM-5.2 ATOM MI355X agentic 配方(benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml)的两项 AITER 设置,该配方也用同样的 prefix caching 与 chunk 大小。

证据(本地 2×8 MI350X,同一配方,并发 1,3600 秒,未过滤 AgentX 语料,1M 上下文)

3600 秒那一轮用的是切出 post3 的预发布版本,三项功能用环境变量打开;post3 就是这份代码把默认值翻转(已验证:AST 相同,干净环境 import 读到 8 / 1 / 1)。本地轮次走的是 #3552 所替换的 bash 启动流程,vLLM 与 TileRT 设置与本配方相同;srt-slurm 流程本身由 sweep 验证。

post2,MI355X 官方轮次(#3389) 本 PR,本地 MI350X
submission_valid / 覆盖率 true / 100% true / 100%(TTFT 与 ITL)
请求 / 错误 271 / 0 294 / 0
TTFT p50 / p90 / p99(ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 30.9 万 / 49.1 万 33.3 万 / 53.1 万

按请求配对(ISL 与 OSL 均相同,n = 271):TTFT p90 4348 → 1721 ms,每请求 TTFT 比值中位 0.344。同一对 MI350X 上逐步拆解(每一步只改一处):

步骤 配对 TTFT p90 每请求比值
仅换硬件:MI355X → MI350X,两边都是 post2 +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM(含上述参数) −41.7% 0.810
多发送端 + 流水线发送 −32.5% 0.595
router 侧分词 −17.3% 0.743

decode 未改:与同一对 MI350X 上的 post2 相比,同批 intvty p50/p90 变化 +0.21% / +0.16%。(与 MI355X 那轮相比约低 7%,即 MI350X 与 MI355X 的频率差;以 MI355X 上的官方轮次为准。)

精度,GSM8K(lm-eval,1319 题,5-shot,真实 MTP),同一套配置:两个预发布版本 strict / flexible 分别为 0.9742 / 0.9742 与 0.9757 / 0.9757(vLLM 0.24 上的 post1:0.9765 / 0.9757)。第二轮中 router 对全部 1319 条提示与 vLLM /tokenize 逐条对照:0 条不一致。

清单说明

AI 模型使用说明

  • 模型/版本:claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用
  • 工作内容:实现 post3 的 pd_vllm 改动、完成本地验证、起草本 PR

关联 issue

无

改动类型

配置变更

🤖 Generated with Claude Code

CrimsonDump and others added 3 commits October 1, 2026 04:42
…TOM prefill

Stacked on the declarative srt-slurm TileRT recipe (#3552). Bump tilert
0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill
role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1
(the master config's PREFILL_IMAGE, which the launcher stages, follows).
post3 turns on multi-sender staging, per-layer pipelined KV send and
router-side incremental chat tokenization by default. The prefill role drops
enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling,
fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2
ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry.

基于声明式 srt-slurm TileRT 配方(#3552)。tilert 由 0.1.6.post2 升级到
0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1(master config 中供启动器预导
镜像的 PREFILL_IMAGE 同步更新)。post3 默认开启多发送端暂存、逐层流水线发送 KV
与 router 侧增量对话分词。prefill 角色去掉 enforce-eager,开启 CUDA graph
(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、prefix caching、
16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方的 AITER 设置。
追加 perf-changelog 条目。

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
将 tilert post3 条目的 pr-link 指向 #3617。

Co-Authored-By: Claude Opus 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
…5x-tilert-post3

# Conflicts:
#	inferencex-e2e/perf-changelog.yaml
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@functionstackx functionstackx added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Oct 1, 2026
Comment thread inferencex-e2e/perf-changelog.yaml Outdated
Comment on lines +9171 to +9179

- config-keys:
- glm5.3-fp8-mi355x-tilert-agentic
scenario-type:
- agentic-coding
description:
- "Bump tilert 0.1.6.post2 -> 0.1.6.post3 (PyPI, 2026-09-28) on both TileRT ranks; router metadata follows. post3 turns on three prefill-to-decode latency features by default: multi-sender staging (TILERT_PD_SENDERS=8, each TP rank extracts and sends its own layers), per-layer pipelined KV send (TILERT_PD_PIPELINE=1, a layer is extracted and RDMA-written as soon as it is computed), and router-side incremental chat tokenization (TILERT_ROUTER_TOKENIZE=1, the router tokenizes only the new turn through a per-segment cache and hands vLLM token ids; a start-up self-check against vLLM /tokenize keeps it off on any mismatch). Move the prefill rank to a vLLM 0.28 + ATOM image (ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built on rocm/atom-dev:vllm-v0.28.0-nightly_20260923) with CUDA graphs, async scheduling, prefix caching and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Decode image, 1M context, bf16 KV and layer-sharded PD buffers unchanged. TileRT reports AgentX TTFT p90 4.35 s -> 1.60 s at concurrency 1 (local 2x8 MI350X against the post2 MI355X run)."
- "两侧 TileRT rank 上 tilert 由 0.1.6.post2 升级到 0.1.6.post3(PyPI,2026-09-28),router 元数据随之更新。post3 默认打开三项 prefill 到 decode 的时延优化:多发送端暂存(TILERT_PD_SENDERS=8,每个 TP rank 抽取并发送自己负责的层)、逐层流水线发送 KV(TILERT_PD_PIPELINE=1,每层算完即抽取并 RDMA 写出)、router 侧增量对话分词(TILERT_ROUTER_TOKENIZE=1,router 经按段缓存只对新一轮分词,把 token id 交给 vLLM;启动自检与 vLLM /tokenize 不一致即保持关闭)。prefill rank 改用 vLLM 0.28 + ATOM 镜像(ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于 rocm/atom-dev:vllm-v0.28.0-nightly_20260923),开启 CUDA graph、异步调度、prefix caching,并沿用 GLM-5.2 ATOM MI355X agentic 配方的 AITER 设置。decode 镜像、1M 上下文、bf16 KV 与按层分片的 PD 缓冲不变。TileRT 报告并发 1 下 AgentX TTFT p90 由 4.35 秒降至 1.60 秒(本地 2x8 MI350X 对比 post2 的 MI355X 官方轮次)。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3617

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The newly appended changelog entry's pr-link still points at the fork PR #3617, not at this PR, contradicting the PR's own description ("the changelog pr-link now points at this PR") and leaving the merged record misattributed. validate_added_pr_link() in infx/workflows/validate_perf_changelog.py requires an appended entry's pr-link to equal pull/{pr_number} of the PR actually being merged, or be an allowed placeholder; pull/3617 is neither once this PR's real number differs. canonicalize_appended_links() then raises ChangelogValidationError instead of rewriting it, since it only rewrites placeholders, not stale literal links. Fix: before merge, set the new entry's pr-link (line 9179) to this PR's own canonical pull/ URL or to an allowed placeholder, not the source fork PR's link.

Why this was flagged

The appended perf-changelog.yaml entry at lines 9171-9179 carries pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3617, the fork PR by @ CrimsonDump being ported, not the number of the PR this diff belongs to (the PR text itself says it is a port of #3617 into a new, different PR). validate_added_pr_link (infx/workflows/validate_perf_changelog.py:139-150) requires link == expected (built from the real pr_number) or a value in PR_LINK_PLACEHOLDERS; pull/3617 matches neither. canonicalize_appended_links (infx/workflows/prepare_perf_changelog_merge.py:79-114) only auto-corrects placeholders, so it raises ChangelogValidationError for a stale literal link, blocking the automated merge-prep/reuse path and contradicting the PR's own claim that the link was already fixed.

Verification: The appended changelog block (inferencex-e2e/perf-changelog.yaml:9172-9179) ends at line 9179 with pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3617 — the fork PR being ported, not this PR, contradicting the PR description. validate_added_pr_link (infx/workflows/validate_perf_changelog.py:139-150) requires the link equal pull/{pr_number} or be a placeholder; pull/3617 is neither, so it raises ChangelogValidationError.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1 sets
MOONCAKE_DISABLE_HIP_DMABUF=1, so Mooncake registers the 12.54 GB VRAM
staging buffer with plain ibv_reg_mr, which fails with EINVAL (-202) on
all eight ranks and kills prefill startup. Setting it to "0" restores the
dmabuf path the post2 run registered through.
With only gpu-memory-utilization=0.85, vLLM sized the KV cache to
121.89 GiB before TileRT allocated its 12.54 GiB per-rank staging buffer,
and ATOM skips CUDA graph memory profiling on ROCm (5.3 GiB), so the
rejection-sampler warmup hit HIP OOM on every rank. 100 GiB still holds
~1.14M tokens (93.9 KB/token), above the 1M max-model-len.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants