[DNM][Experimental] Add MI355X TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 MI355X TP2 V4 Flash SGLang 8k1k MTP 扫描 - #3518
Oseltamivir wants to merge 3 commits into
Conversation
Mirror the B300/B200 DeepSeek-V4-Flash fixed-sequence recipes on MI355X with SGLang's MI355X Flash FP4 low-latency flags at TP2 and EAGLE 2/1/3.
|
Thanks for the contribution!
中文感谢你的贡献!
|
| description: | ||
| - "Add DeepSeek-V4-Flash MI355X TP2 SGLang 8k1k serving at concurrency 1, 2, 4, 8, 16, 32, 64, 128 on lmsysorg/sglang-rocm v0.5.20-rocm720-mi35x-20260926, with bundled MTP via EAGLE (2 steps, top-k 1, 3 draft tokens), real verification, GPU-resident weights/KV, and the DeepSeek-V4 chat encoder. Flags follow SGLang's MI355X Flash FP4 low-latency recipe at TP2." | ||
| - "新增 DeepSeek-V4-Flash MI355X TP2 SGLang 8k1k serving,并发为 1、2、4、8、16、32、64、128,镜像为 lmsysorg/sglang-rocm v0.5.20-rocm720-mi35x-20260926;通过 EAGLE 使用原生 MTP(2 steps、top-k 1、3 draft tokens),采用真实验证、GPU 常驻权重/KV 和 DeepSeek-V4 chat 编码器。参数沿用 SGLang MI355X Flash FP4 低延迟配方,改为 TP2。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD |
There was a problem hiding this comment.
🟡 (optional) The new perf-changelog entry ships with a placeholder pr-link (.../pull/TBD) instead of the real PR number, so once merged this permanent, append-only record can never be corrected. AGENTS.md requires perf-changelog.yaml to be append-only and byte-sensitive (only append at the tail, never edit existing bytes), so if this placeholder lands it is stuck wrong forever, unlike every other entry in the file which cites a real PR number (e.g. pull/3503 just above it). Fix: replace pull/TBD with the actual PR number before merge; if the number truly isn't known yet, the entry should be finalized in a follow-up commit before merge rather than merged with a fake link.
Why this was flagged
Trigger: this PR is merged with perf-changelog.yaml:8997 still containing pr-link https://github.com/SemiAnalysisAI/InferenceX/pull/TBD. The file is documented as append-only and byte-sensitive (AGENTS.md), so no later commit can edit this line in place without violating that invariant. Every other entry in perf-changelog.yaml uses a real PR URL (e.g. pull/3503); this is the only occurrence of pull/TBD in the file (grep count 1). Consequence: anyone auditing performance history via this changelog gets a dead/wrong link for this recipe permanently, unlike the base branch where every entry resolves to its actual PR.
Verification: nit. The placeholder is real and present: perf-changelog.yaml:8997 reads pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/TBD, while every other entry cites a real PR number (e.g. line 8988/8976 .../pull/3503, .../pull/3493). Grep confirms pull/TBD occurs exactly once in the file. AGENTS.md line 123 confirms the convention the candidate cites: "Every change that can affect…
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36328557374 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36328557374 |
Set cuda-graph-backend-{decode,prefill}=disabled, which also skips target
verify and draft capture, and drop the per-concurrency decode graph batch.
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
Summary
dsv4flash-fp4-mi355x-sglang-mtp: DeepSeek-V4-Flash on MI355X at TP2, 8k1k fixed sequence, concurrency 1, 2, 4, 8, 16, 32, 64, 128. Companion to B300 #3516 and B200 #3517.EAGLEwith 2 steps, top-k 1, and 3 draft tokens, matching the other two PRs.cuda-graph-backend-decode: disabled,cuda-graph-backend-prefill: disabled), so decode, prefill, and EAGLE target-verify/draft all run eagerly; startup reports zero capture time for every graph phase.lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260926, the newest ROCm 7.2.0 MI35x build.dsv4attention, page 256, fp8 KV, shared-expert fusion, mem-fraction 0.90, SWA ratio 0.15, chunked prefill 16384), reduced from TP8 to TP2. Radix cache is disabled for the fixed-sequence workload; admission and the decode graph batch are pinned to each concurrency.--dsv4fixed-sequence client passthrough and test as [DNM][Experimental] Add B300 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B300 TP2 V4 Flash SGLang 8k1k MTP 扫描 #3516/[DNM][Experimental] Add B200 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B200 TP2 V4 Flash SGLang 8k1k MTP 扫描 #3517 (identical content), so this PR stands alone.Validation
full-sweep-fail-fast).摘要
dsv4flash-fp4-mi355x-sglang-mtp:DeepSeek-V4-Flash 在 MI355X 上以 TP2 运行 8k1k 固定序列,并发为 1、2、4、8、16、32、64、128,与 B300 #3516 和 B200 #3517 配套。EAGLE使用原生 MTP(2 steps、top-k 1、3 draft tokens),与另外两个 PR 一致。cuda-graph-backend-decode: disabled、cuda-graph-backend-prefill: disabled),decode、prefill 及 EAGLE 验证/draft 均以 eager 模式运行;启动日志中各 graph 阶段捕获耗时均为 0。lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260926。Serving 参数与 ROCm 环境变量沿用 SGLang MI355X Flash FP4 低延迟配方,由 TP8 改为 TP2;固定序列负载关闭 radix cache。--dsv4客户端参数透传及测试,可独立合并。验证
full-sweep-fail-fast)。