Skip to content

Port GLM-5.3 MI355X TileRT AgentX to srt-slurm - #3552

Draft
cquil11 wants to merge 7 commits into
feat/tilert-role-engine-portsfrom
feat/tilert-mi355x-agentx
Draft

cquil11 wants to merge 7 commits into
feat/tilert-role-engine-portsfrom
feat/tilert-mi355x-agentx

Conversation

@cquil11

@cquil11 cquil11 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Based on #3551. Ports glm5.3-fp8-mi355x-tilert-agentic to native vLLM prefill, TileRT decode, and the TileRT router.

Preserves TP8 1P1D, images, 1M context, BF16 KV, MTP, golden acceptance, concurrency 1, and local-checkpoint tokenization. Removes the old AMD TileRT server and launch path. The launcher mounts prepared checkpoints; the setup script only installs dependencies.

Validation: full AgentX config sweep passed: 3,600-second profile, 271 successful requests, zero errors, and 11 successful warmup requests. Required vLLM server metrics passed validation. This run does not validate AgentX power collection.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@cquil11
cquil11 force-pushed the feat/tilert-mi355x-agentx branch from 811fdc6 to 34217c7 Compare September 29, 2026 01:39
@functionstackx
functionstackx added this pull request to stack #3564 September 29, 2026 02:30
CrimsonDump added a commit to CrimsonDump/InferenceX that referenced this pull request Sep 29, 2026
…TOM prefill

Stacked on the declarative srt-slurm TileRT recipe (SemiAnalysisAI#3552). Bump tilert
0.1.6.post2 -> 0.1.6.post3 (router metadata follows) and move the prefill
role to the vLLM 0.28 + ATOM image ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1.
post3 turns on multi-sender staging, per-layer pipelined KV send and
router-side incremental chat tokenization by default. The prefill role drops
enforce-eager and adds CUDA graphs (FULL_AND_PIECEWISE), async scheduling,
fastsafetensors loading, prefix caching, a 16384-token chunk and the GLM-5.2
ATOM MI355X agentic recipe's AITER settings. Append the perf-changelog entry.

基于声明式 srt-slurm TileRT 配方(SemiAnalysisAI#3552)。tilert 由 0.1.6.post2 升级到
0.1.6.post3(router 元数据随之更新),prefill 角色改用 vLLM 0.28 + ATOM 镜像
ghcr.io/tile-ai/tilert-rocm-prefill:0.1.6.post1。post3 默认开启多发送端暂存、
逐层流水线发送 KV 与 router 侧增量对话分词。prefill 角色去掉 enforce-eager,
开启 CUDA graph(FULL_AND_PIECEWISE)、异步调度、fastsafetensors 加载、
prefix caching、16384 token 的 chunk,并沿用 GLM-5.2 ATOM MI355X agentic 配方
的 AITER 设置。追加 perf-changelog 条目。

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016SS3MCfU8mhe9buef6pNBL

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant