[TileRT] GLM-5.3 FP8 MI355X AgentX: tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill / [TileRT] GLM-5.3 FP8 MI355X AgentX:tilert 0.1.6.post3 + vLLM 0.28/ATOM prefill - #3563
Open
CrimsonDump wants to merge 1 commit into
Open
CrimsonDump wants to merge 1 commit into
CrimsonDump wants to merge 1 commit into
Conversation
CrimsonDump
requested review from
1am9trash,
billishyahao,
chunfangamd,
seungrokj and
yctseng0211
as code owners
September 29, 2026 02:27
CrimsonDump
added a commit
to CrimsonDump/InferenceX
that referenced
this pull request
Sep 29, 2026
将 tilert post3 条目的 pr-link 指向 SemiAnalysisAI#3563。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
functionstackx
left a comment
Collaborator
There was a problem hiding this comment.
hi @CrimsonDump can u stack ur PR on top of #3552 , we moving tileRT to an declaractive yaml instead of bash scripts
…TOM prefill Stacked on the declarative srt-slurm TileRT recipe (SemiAnalysisAI#3552). Bump tilert 0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1. post3 turns on multi-sender staging, per-layer pipelined KV send and router-side incremental chat tokenization by default. The prefill role drops enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling, fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2 ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry. 基于声明式 srt-slurm TileRT 配方(SemiAnalysisAI#3552)。tilert 由 0.1.6.post2 升级到 0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像 ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、 逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager, 开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、 prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方 的 AITER 设置。追加 perf-changelog 条目。 Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL
CrimsonDump
force-pushed
the
feat/glm5.3-fp8-mi355x-tilert-post3
branch
from
September 29, 2026 02:40
e3ace0b to
8aab900
Compare
Author
sure, please check again @functionstackx |
Collaborator
|
thanks @CrimsonDump created an upstream branch PR here and started sweeping it #3565 |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Stacked on #3552 (the declarative srt-slurm TileRT recipe); merge after it. Updates the
glm5.3-fp8-mi355x-tilert-agenticrecipe (AgentX, concurrency 1, 3600 s) in two ways. Nothing else in the recipe changes: decode image, 1M context, bf16 KV, layer-sharded on-GPU PD buffers,GLM5_AR_N=2.1. tilert
0.1.6.post2→0.1.6.post3(PyPI, 2026-09-28,sha256:d6fbf0a55be1fbde…), on both TileRT ranks; router metadata follows. The engine.sofiles are byte-identical to post2; the change is intilert/pd_vllm(4 files, +622 / −12). Three prefill→decode latency features are now on by default:TILERT_PD_SENDERS=8(was 1)TILERT_PD_PIPELINE=1(was 0)TILERT_ROUTER_TOKENIZE=1(was 0)/v1/completions. A start-up self-check against vLLM/tokenize, and a cross-check of vLLM's own--chat-template/--default-chat-template-kwargs, keep it off on any mismatch; tools, logprobs, structured output and non-text content always take vLLM's chat path2. Prefill image → vLLM 0.28 + ATOM (
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1, built onrocm/atom-dev:vllm-v0.28.0-nightly_20260923: vLLM0.28.1.dev0, ATOM0.1.7.dev17, ROCm 7.2.4). The only additions on top of the base image are the dependencies the TileRT connector needs andWORKDIR /app(withWORKDIR /, anyPYTHONPYCACHEPREFIXmakes the ATOM plugin fail at import). In the recipe's prefill role,enforce-eageris dropped andargsadd CUDA graphs (compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"}),async-scheduling,load-format: fastsafetensors,enable-prefix-cachingandmax-num-batched-tokens: 16384;envadds the two AITER settings of the in-tree GLM-5.2 ATOM MI355X agentic recipe (benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml), which uses the same prefix caching and chunk size.Evidence (local 2×8 MI350X, same recipe, concurrency 1, 3600 s, unfiltered AgentX corpus, 1M context)
The 3600 s run used the pre-release build that post3 was cut from, with the three features switched on by environment variables; post3 is that code with the defaults flipped (verified: identical AST, and a clean-environment import reads 8 / 1 / 1). The local runs went through the bash launcher that #3552 replaces, with the same vLLM and TileRT settings as this recipe; the srt-slurm path itself is exercised by the sweep.
submission_valid/ coveragePaired by request (same ISL and OSL, n = 271): TTFT p90 4348 → 1721 ms, per-request TTFT ratio median 0.344. Step by step on the same MI350X pair (each step changes one thing):
Decode is untouched: against post2 on the same MI350X pair, same-batch intvty p50/p90 moves +0.21% / +0.16%. (Against the MI355X run it is ~7% lower, which is the MI350X/MI355X clock difference; the official run on MI355X is the number that counts.)
Accuracy, GSM8K (lm-eval, 1319 questions, 5-shot, real MTP) on the same stack: strict / flexible 0.9742 / 0.9742 and 0.9757 / 0.9757 on the two pre-release builds (post1 on vLLM 0.24: 0.9765 / 0.9757). In the second run the router cross-checked all 1319 prompts against vLLM
/tokenize: 0 mismatches.Checklist notes
rocm/atom-devfamily the in-tree ATOM recipes use.AI model disclosure
claude-opus-5-5[1m](Claude Opus 5.5, 1M context), via Claude Codepd_vllmchanges, ran the local validation, and drafted this PRRelated Issue
N/A
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR中文
改动说明
基于 #3552(声明式 srt-slurm TileRT 配方),需在其之后合入。更新
glm5.3-fp8-mi355x-tilert-agentic配方(AgentX,并发 1,3600 秒),共两处。其余不变:decode 镜像、1M 上下文、bf16 KV、按层分片的 GPU PD 缓冲、GLM5_AR_N=2。1. tilert
0.1.6.post2→0.1.6.post3(PyPI,2026-09-28,sha256:d6fbf0a55be1fbde…),两侧 TileRT rank 同步,router 元数据随之更新。引擎.so与 post2 逐字节相同,改动全在tilert/pd_vllm(4 个文件,+622 / −12)。三项 prefill→decode 时延优化改为默认开启:TILERT_PD_SENDERS=8(原为 1)TILERT_PD_PIPELINE=1(原为 0)TILERT_ROUTER_TOKENIZE=1(原为 0)/v1/completions交给 vLLM。启动时与 vLLM/tokenize自检,并核对 vLLM 自己的--chat-template/--default-chat-template-kwargs,任何不一致即保持关闭;带 tools、logprobs、结构化输出或非纯文本内容的请求一律走 vLLM 的 chat 路径2. prefill 镜像改为 vLLM 0.28 + ATOM(
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1,基于rocm/atom-dev:vllm-v0.28.0-nightly_20260923:vLLM0.28.1.dev0、ATOM0.1.7.dev17、ROCm 7.2.4)。在基底镜像之上只加了 TileRT connector 所需的依赖和WORKDIR /app(若为WORKDIR /,只要设了PYTHONPYCACHEPREFIX,ATOM 插件在 import 时就会失败)。配方的 prefill 角色去掉enforce-eager,args加上 CUDA graph(compilation-config: {"cudagraph_mode":"FULL_AND_PIECEWISE"})、async-scheduling、load-format: fastsafetensors、enable-prefix-caching与max-num-batched-tokens: 16384;env加上在树 GLM-5.2 ATOM MI355X agentic 配方(benchmarks/single_node/srt-slurm-recipes/glm5.2/atom/mi355x-fp4-mtp/agentic.yaml)的两项 AITER 设置,该配方也用同样的 prefix caching 与 chunk 大小。证据(本地 2×8 MI350X,同一配方,并发 1,3600 秒,未过滤 AgentX 语料,1M 上下文)
3600 秒那一轮用的是切出 post3 的预发布版本,三项功能用环境变量打开;post3 就是这份代码把默认值翻转(已验证:AST 相同,干净环境 import 读到 8 / 1 / 1)。本地轮次走的是 #3552 所替换的 bash 启动流程,vLLM 与 TileRT 设置与本配方相同;srt-slurm 流程本身由 sweep 验证。
submission_valid/ 覆盖率按请求配对(ISL 与 OSL 均相同,n = 271):TTFT p90 4348 → 1721 ms,每请求 TTFT 比值中位 0.344。同一对 MI350X 上逐步拆解(每一步只改一处):
decode 未改:与同一对 MI350X 上的 post2 相比,同批 intvty p50/p90 变化 +0.21% / +0.16%。(与 MI355X 那轮相比约低 7%,即 MI350X 与 MI355X 的频率差;以 MI355X 上的官方轮次为准。)
精度,GSM8K(lm-eval,1319 题,5-shot,真实 MTP),同一套配置:两个预发布版本 strict / flexible 分别为 0.9742 / 0.9742 与 0.9757 / 0.9757(vLLM 0.24 上的 post1:0.9765 / 0.9757)。第二轮中 router 对全部 1319 条提示与 vLLM
/tokenize逐条对照:0 条不一致。清单说明
rocm/atom-dev同一系列。AI 模型使用说明
claude-opus-5-5[1m](Claude Opus 5.5,1M 上下文),经 Claude Code 使用pd_vllm改动、完成本地验证、起草本 PR关联 issue
无
改动类型
配置变更
🤖 Generated with Claude Code
https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL