Skip to content

[TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / [TileRT] GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill - #3563

Open
CrimsonDump wants to merge 1 commit into
SemiAnalysisAI:feat/tilert-mi355x-agentxfrom
CrimsonDump:feat/glm5.3-fp8-mi355x-tilert-post3
Open

CrimsonDump wants to merge 1 commit into
SemiAnalysisAI:feat/tilert-mi355x-agentxfrom
CrimsonDump:feat/glm5.3-fp8-mi355x-tilert-post3

Conversation

@CrimsonDump

@CrimsonDump CrimsonDump commented Sep 29, 2026 •

Copy link
Copy Markdown

Description

Stacked on #3552 (the declarative srt-slurm TileRT recipe); merge after it. Updates the glm5.3-fp8-mi355x-tilert-agentic recipe (AgentX, concurrency 1, 3600 s) in two ways. Nothing else in the recipe changes: decode image, 1M context, bf16 KV, layer-sharded on-GPU PD buffers, GLM5_AR_N=2.

1. tilert 0.1.6.post2 → 0.1.6.post3 (PyPI, 2026-09-28, sha256:d6fbf0a55be1fbde…), on both TileRT ranks; router metadata follows. The engine .so files are byte-identical to post2; the change is in tilert/pd_vllm (4 files, +622 / −12). Three prefill→decode latency features are now on by default:

Feature Switch (post3 default) What it does
Multi-sender staging TILERT_PD_SENDERS=8 (was 1) Every prefill TP rank extracts and RDMA-sends the layers it owns, instead of rank 0 sending all 79
Per-layer pipelined send TILERT_PD_PIPELINE=1 (was 0) A layer's KV is extracted and written as soon as that layer is computed, overlapping the transfer with the rest of the forward
Router-side incremental tokenization TILERT_ROUTER_TOKENIZE=1 (was 0) The router renders the chat template and tokenizes through a per-segment cache (only the new turn is tokenized), then sends vLLM token ids on /v1/completions. A start-up self-check against vLLM /tokenize, and a cross-check of vLLM's own --chat-template / --default-chat-template-kwargs, keep it off on any mismatch; tools, logprobs, structured output and non-text content always take vLLM's chat path

2. Prefill image → vLLM 0.28 + ATOM (ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built on rocm/atom-dev:vllm-v0.28.0-nightly_20260923: vLLM 0.28.1.dev0, ATOM 0.1.7.dev17, ROCm 7.2.4). The only additions on top of the base image are the dependencies the TileRT connector needs and WORKDIR /app (with WORKDIR /, any PYTHONPYCACHEPREFIX makes the ATOM plugin fail at import). In the recipe's prefill role, enforce-eager is dropped and args add CUDA graphs (compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"}), async-scheduling, load-format: fastsafetensors, enable-prefix-caching and max-num-batched-tokens: 16384; env adds the two AITER settings of the in-tree GLM-5.2 ATOM MI355X agentic recipe (benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml), which uses the same prefix caching and chunk size.

Evidence (local 2×8 MI350X, same recipe, concurrency 1, 3600 s, unfiltered AgentX corpus, 1M context)

The 3600 s run used the pre-release build that post3 was cut from, with the three features switched on by environment variables; post3 is that code with the defaults flipped (verified: identical AST, and a clean-environment import reads 8 / 1 / 1). The local runs went through the bash launcher that #3552 replaces, with the same vLLM and TileRT settings as this recipe; the srt-slurm path itself is exercised by the sweep.

post2, official MI355X run (#3389) this PR, local MI350X
submission_valid / coverage true / 100% true / 100% (TTFT and ITL)
requests / errors 271 / 0 294 / 0
TTFT p50 / p90 / p99 (ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 309k / 491k 333k / 531k

Paired by request (same ISL and OSL, n = 271): TTFT p90 4348 → 1721 ms, per-request TTFT ratio median 0.344. Step by step on the same MI350X pair (each step changes one thing):

Step TTFT p90, paired Per-request ratio
hardware only: MI355X → MI350X, post2 on both +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM (+ flags above) −41.7% 0.810
multi-sender + pipelined send −32.5% 0.595
router-side tokenization −17.3% 0.743

Decode is untouched: against post2 on the same MI350X pair, same-batch intvty p50/p90 moves +0.21% / +0.16%. (Against the MI355X run it is ~7% lower, which is the MI350X/MI355X clock difference; the official run on MI355X is the number that counts.)

Accuracy, GSM8K (lm-eval, 1319 questions, 5-shot, real MTP) on the same stack: strict / flexible 0.9742 / 0.9742 and 0.9757 / 0.9757 on the two pre-release builds (post1 on vLLM 0.24: 0.9765 / 0.9757). In the second run the router cross-checked all 1319 prompts against vLLM /tokenize: 0 mismatches.

Checklist notes

AI model disclosure

  • Model/version: claude-opus-5-5[1m] (Claude Opus 5.5, 1M context), via Claude Code
  • Role: implemented the post3 pd_vllm changes, ran the local validation, and drafted this PR

Related Issue

N/A

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR
中文

改动说明

基于 #3552(声明式 srt-slurm TileRT 配方),需在其之后合入。更新 glm5.3-fp8-mi355x-tilert-agentic 配方(AgentX,并发 1,3600 秒),共两处。其余不变:decode 镜像、1M 上下文、bf16 KV、按层分片的 GPU PD 缓冲、GLM5_AR_N=2。

1. tilert 0.1.6.post2 → 0.1.6.post3(PyPI,2026-09-28,sha256:d6fbf0a55be1fbde…),两侧 TileRT rank 同步,router 元数据随之更新。引擎 .so 与 post2 逐字节相同,改动全在 tilert/pd_vllm(4 个文件,+622 / −12)。三项 prefill→decode 时延优化改为默认开启:

功能 开关(post3 默认) 作用
多发送端暂存 TILERT_PD_SENDERS=8(原为 1) prefill 的每个 TP rank 抽取并 RDMA 发送自己负责的层,而不是由 rank 0 发全部 79 层
逐层流水线发送 TILERT_PD_PIPELINE=1(原为 0) 每层算完即抽取并写出该层 KV,传输与后续层的前向重叠
router 侧增量分词 TILERT_ROUTER_TOKENIZE=1(原为 0) router 渲染 chat 模板,经按段缓存分词(只对新一轮分词),再把 token id 经 /v1/completions 交给 vLLM。启动时与 vLLM /tokenize 自检,并核对 vLLM 自己的 --chat-template / --default-chat-template-kwargs,任何不一致即保持关闭;带 tools、logprobs、结构化输出或非纯文本内容的请求一律走 vLLM 的 chat 路径

2. prefill 镜像改为 vLLM 0.28 + ATOM(ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于 rocm/atom-dev:vllm-v0.28.0-nightly_20260923:vLLM 0.28.1.dev0、ATOM 0.1.7.dev17、ROCm 7.2.4)。在基底镜像之上只加了 TileRT connector 所需的依赖和 WORKDIR /app(若为 WORKDIR /,只要设了 PYTHONPYCACHEPREFIX,ATOM 插件在 import 时就会失败)。配方的 prefill 角色去掉 enforce-eager,args 加上 CUDA graph(compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"})、async-scheduling、load-format: fastsafetensors、enable-prefix-caching 与 max-num-batched-tokens: 16384;env 加上在树 GLM-5.2 ATOM MI355X agentic 配方(benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml)的两项 AITER 设置,该配方也用同样的 prefix caching 与 chunk 大小。

证据(本地 2×8 MI350X,同一配方,并发 1,3600 秒,未过滤 AgentX 语料,1M 上下文)

3600 秒那一轮用的是切出 post3 的预发布版本,三项功能用环境变量打开;post3 就是这份代码把默认值翻转(已验证:AST 相同,干净环境 import 读到 8 / 1 / 1)。本地轮次走的是 #3552 所替换的 bash 启动流程,vLLM 与 TileRT 设置与本配方相同;srt-slurm 流程本身由 sweep 验证。

post2,MI355X 官方轮次(#3389) 本 PR,本地 MI350X
submission_valid / 覆盖率 true / 100% true / 100%(TTFT 与 ITL)
请求 / 错误 271 / 0 294 / 0
TTFT p50 / p90 / p99(ms) 2344 / 4347 / 13149 854 / 1601 / 7899
ISL p50 / p90 30.9 万 / 49.1 万 33.3 万 / 53.1 万

按请求配对(ISL 与 OSL 均相同,n = 271):TTFT p90 4348 → 1721 ms,每请求 TTFT 比值中位 0.344。同一对 MI350X 上逐步拆解(每一步只改一处):

步骤 配对 TTFT p90 每请求比值
仅换硬件:MI355X → MI350X,两边都是 post2 +9.6% 0.956
prefill vLLM 0.24 → vLLM 0.28 + ATOM(含上述参数) −41.7% 0.810
多发送端 + 流水线发送 −32.5% 0.595
router 侧分词 −17.3% 0.743

decode 未改:与同一对 MI350X 上的 post2 相比,同批 intvty p50/p90 变化 +0.21% / +0.16%。(与 MI355X 那轮相比约低 7%,即 MI350X 与 MI355X 的频率差;以 MI355X 上的官方轮次为准。)

精度,GSM8K(lm-eval,1319 题,5-shot,真实 MTP),同一套配置:两个预发布版本 strict / flexible 分别为 0.9742 / 0.9742 与 0.9757 / 0.9757(vLLM 0.24 上的 post1:0.9765 / 0.9757)。第二轮中 router 对全部 1319 条提示与 vLLM /tokenize 逐条对照:0 条不一致。

清单说明

AI 模型使用说明

  • 模型/版本:claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用
  • 工作内容:实现 post3 的 pd_vllm 改动、完成本地验证、起草本 PR

关联 issue

无

改动类型

配置变更

🤖 Generated with Claude Code

https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

CrimsonDump added a commit to CrimsonDump/InferenceX that referenced this pull request Sep 29, 2026
将 tilert post3 条目的 pr-link 指向 SemiAnalysisAI#3563。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hi @CrimsonDump can u stack ur PR on top of #3552 , we moving tileRT to an declaractive yaml instead of bash scripts

@cquil11 @Oseltamivir

…TOM prefill

Stacked on the declarative srt-slurm TileRT recipe (SemiAnalysisAI#3552). Bump tilert
0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill
role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1.
post3 turns on multi-sender staging, per-layer pipelined KV send and
router-side incremental chat tokenization by default. The prefill role drops
enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling,
fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2
ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry.

基于声明式 srt-slurm TileRT 配方(SemiAnalysisAI#3552)。tilert 由 0.1.6.post2 升级到
0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、
逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager,
开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、
prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方
的 AITER 设置。追加 perf-changelog 条目。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
@CrimsonDump
CrimsonDump force-pushed the feat/glm5.3-fp8-mi355x-tilert-post3 branch from e3ace0b to 8aab900 Compare September 29, 2026 02:40
@CrimsonDump
CrimsonDump changed the base branch from main to feat/tilert-mi355x-agentx September 29, 2026 02:40
@CrimsonDump

CrimsonDump commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

hi @CrimsonDump can u stack ur PR on top of #3552 , we moving tileRT to an declaractive yaml instead of bash scripts

@cquil11 @Oseltamivir

sure, please check again @functionstackx

@functionstackx

Copy link
Copy Markdown
Collaborator

thanks @CrimsonDump created an upstream branch PR here and started sweeping it #3565

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants