Skip to content

[DNM][Experimental] Add MI355X TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 MI355X TP2 V4 Flash SGLang 8k1k MTP 扫描 - #3518

Open
Oseltamivir wants to merge 3 commits into
mainfrom
benchmark/dsv4flash-mi355x-tp2-sglang-eagle-8k1k
Open

Oseltamivir wants to merge 3 commits into
mainfrom
benchmark/dsv4flash-mi355x-tp2-sglang-eagle-8k1k

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Validation

  • GPU sweep and evals: pending (full-sweep-fail-fast).

摘要

验证

  • GPU 扫描与评测:待 CI(full-sweep-fail-fast)。

Mirror the B300/B200 DeepSeek-V4-Flash fixed-sequence recipes on MI355X with
SGLang's MI355X Flash FP4 low-latency flags at TP2 and EAGLE 2/1/3.
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment thread perf-changelog.yaml Outdated
description:
- "Add DeepSeek-V4-Flash MI355X TP2 SGLang 8k1k serving at concurrency 1, 2, 4, 8, 16, 32, 64, 128 on lmsysorg/sglang-rocm v0.5.20-rocm720-mi35x-20260926, with bundled MTP via EAGLE (2 steps, top-k 1, 3 draft tokens), real verification, GPU-resident weights/KV, and the DeepSeek-V4 chat encoder. Flags follow SGLang's MI355X Flash FP4 low-latency recipe at TP2."
- "新增 DeepSeek-V4-Flash MI355X TP2 SGLang 8k1k serving,并发为 1、2、4、8、16、32、64、128,镜像为 lmsysorg/sglang-rocm v0.5.20-rocm720-mi35x-20260926;通过 EAGLE 使用原生 MTP(2 steps、top-k 1、3 draft tokens),采用真实验证、GPU 常驻权重/KV 和 DeepSeek-V4 chat 编码器。参数沿用 SGLang MI355X Flash FP4 低延迟配方,改为 TP2。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) The new perf-changelog entry ships with a placeholder pr-link (.../pull/TBD) instead of the real PR number, so once merged this permanent, append-only record can never be corrected. AGENTS.md requires perf-changelog.yaml to be append-only and byte-sensitive (only append at the tail, never edit existing bytes), so if this placeholder lands it is stuck wrong forever, unlike every other entry in the file which cites a real PR number (e.g. pull/3503 just above it). Fix: replace pull/TBD with the actual PR number before merge; if the number truly isn't known yet, the entry should be finalized in a follow-up commit before merge rather than merged with a fake link.

Why this was flagged

Trigger: this PR is merged with perf-changelog.yaml:8997 still containing pr-link https://github.com/SemiAnalysisAI/InferenceX/pull/TBD. The file is documented as append-only and byte-sensitive (AGENTS.md), so no later commit can edit this line in place without violating that invariant. Every other entry in perf-changelog.yaml uses a real PR URL (e.g. pull/3503); this is the only occurrence of pull/TBD in the file (grep count 1). Consequence: anyone auditing performance history via this changelog gets a dead/wrong link for this recipe permanently, unlike the base branch where every entry resolves to its actual PR.

Verification: nit. The placeholder is real and present: perf-changelog.yaml:8997 reads pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD, while every other entry cites a real PR number (e.g. line 8988/8976 .../pull/3503, .../pull/3493). Grep confirms pull/TBD occurs exactly once in the file. AGENTS.md line 123 confirms the convention the candidate cites: "Every change that can affect…

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Set cuda-graph-backend-{decode,prefill}=disabled, which also skips target
verify and draft capture, and drop the per-concurrency decode graph batch.
@functionstackx functionstackx changed the title Add MI355X TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 MI355X TP2 V4 Flash SGLang 8k1k MTP 扫描 [DNM][Experimental] Add MI355X TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 MI355X TP2 V4 Flash SGLang 8k1k MTP 扫描 Sep 27, 2026
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants