[DNM][Experimental] Add B300 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B300 TP2 V4 Flash SGLang 8k1k MTP 扫描 - #3516
functionstackx wants to merge 7 commits into
Conversation
新增 B300 TP2 DeepSeek V4 Flash SGLang 8k1k serving 扫描,使用原生 EAGLE 三步 MTP。
|
Thanks for the contribution!
中文感谢你的贡献!
|
将新增 V4 Flash 扫描记录关联到 PR #3516。
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36328552743 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36328552743 |
移除仅适用于 AgentX 的 KV_OFFLOADING 元数据,修复固定序列 CI 配方选择失败。
提交前将 V4 Flash 下载到 B300 runner 的 home 可写目录,保留其他模型路径,并测试下载失败时停止提交。
There was a problem hiding this comment.
Looks good — a self-contained new benchmark recipe/sweep entry with a minimal, backward-compatible change to the shared fixed-sequence client script. Reviewed the new dsv4flash-fp4-b300-sglang-mtp recipe/sweep entry, the srt_fixed_sequence.sh allowlist addition for --dsv4, and confirmed --dsv4 is already handled downstream by benchmark_lib.sh's run_benchmark_serving (this PR only forwards an existing flag, adds no new client behavior). The model-prefix startswith-matching concern raised during triage is pre-existing sweep-filter behavior shared by other keys (e.g. dsv41flash-*), not something this PR introduces.
Extended reasoning...
Change adds a new single-node SGLang B300 TP2 8k1k MTP recipe/sweep entry for DeepSeek-V4-Flash, a one-line allowlist widening in srt_fixed_sequence.sh to forward an already-existing --dsv4 client flag, plus docs, changelog, and a new subprocess-based test. No auth, crypto, or data-exposure surface is touched; the only candidate issue (prefix-filter overlap with dsv4) is a pre-existing sweep-filter pattern also present for other keys, not new. Small, mechanical, and well-tested, so a human need not additionally review before merge.
This review covers commit 4bde51e, which is no longer the latest commit on this pull request; later commits are not covered by it.
主机下载显式使用 home 下的 HF 和 Xet 缓存,避免继承容器路径导致权限错误。
The shared Lustre copy is read through mmap page faults at ~10 MB/s, so TP2 weight loading outlasted the 1800 s health timeout. DeepSeek-V4-Flash is now staged under /scratch/models on every B300 node; drop the host-side download path and list it in STAGED_MODELS. Switch MTP to 2 steps, top-k 1, 3 draft tokens.
Set cuda-graph-backend-{decode,prefill}=disabled, which also skips target
verify and draft capture, and drop the per-concurrency decode graph batch.
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Description
Add
dsv4flash-fp4-b300-sglang-mtpfor the exactdeepseek-ai/DeepSeek-V4-Flashcheckpoint: B300 TP2, 8192 input / 1024 output tokens, concurrency 1, 2, 4, 8, 16, 32, 64, 128. This is separate from V4 Pro and V4.1 Flash.EAGLE, with 2 steps, top-k 1, and 3 draft tokens. The requester explicitly clarified this choice rather than the distinct EAGLE3 algorithm.cuda-graph-backend-decode: disabled,cuda-graph-backend-prefill: disabled), so decode, prefill, and EAGLE target-verify/draft all run eagerly; startup reports zero capture time for every graph phase.nightly-dev-cu13-20260925-8ca82118to its amd64 digest. Keep the checkpoint's shipped draft head and upstream precision handling; explicitly disableSGLANG_NVFP4_CKPT_FP8_NEXTN_MOE.--dsv4client option throughsrt_fixed_sequence.sh, using the shared DeepSeek-V4 chat encoder. This checkpoint has no HF chat template.Upstream recipe reference: DeepSeek-V4 B300 configuration. Its original Flash recipe uses bundled EAGLE 3/1/4 at TP4; this PR tests TP2 and an 8k1k fixed-sequence workload. TP2 runtime performance and accuracy remain unverified until CI passes.
Checkpoint metadata: DeepSeek-V4-Flash config. The target has FP4 experts, FP8 quantization metadata, and one bundled NextN layer. This PR does not override draft weights, draft quantization, or draft KV dtype.
Validation
ruff check infx,ruff format --check infx, Bash syntax, andgit diff --checkpass.KV_OFFLOADING: none, while fixed-sequence CI supplies an empty value. Removed that optional metadata; serving and GPU residency settings are unchanged. Revalidated all eight points using the actual fixed-sequence CI environment./data/models/DeepSeek-V4-Flashdid not exist yet. An interim host-sidehf downloadinto the runner home (Lustre) let the following canary start, but the server never reachedLoad weight end: reproduced on dsxe-sa-b300-prd0-gpu-00, py-spy showed every MoE loader thread blocked incuMemcpyHtoDAsyncwhile mmap page faults read the Lustre copy at ~10 MB/s (~4 h for 160 GB), past the 1800 s health timeout./scratch/models, 46 shards / 159,617,149,040 bytes verified on all 17 up nodes). The download path is removed and the checkpoint is listed inSTAGED_MODELS. On-node repro from NVMe with a 192-CPU exclusive allocation:Load weight endat 339 s (incl. 244 s MHC JIT prewarm), server ready in ~10 min, smoke completion correct.AI model disclosure
GPT 6 Astra Fast prepared the configuration, client change, test, PR, and CI monitoring. No delegated agents.
中文
新增
dsv4flash-fp4-b300-sglang-mtp:使用原始deepseek-ai/DeepSeek-V4-Flashcheckpoint,在 B300 上运行 TP2、8192 输入 / 1024 输出 Token 的 serving 扫描,并发为 1、2、4、8、16、32、64、128。此配置与 V4 Pro、V4.1 Flash 分开。用户已明确选择通过
EAGLE使用原生 MTP(2 steps、top-k 1、3 draft tokens),不是独立的 EAGLE3 算法。镜像固定为官方 CUDA 13 nightlynightly-dev-cu13-20260925-8ca82118的 amd64 digest。保留 checkpoint 自带的 draft head 及上游默认精度处理,显式关闭SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE;不修改引擎、不使用外部 draft、不卸载至 CPU、不模拟接受长度。cuda-graph-backend-decode: disabled、cuda-graph-backend-prefill: disabled),decode、prefill 及 EAGLE 验证/draft 均以 eager 模式运行;启动日志中各 graph 阶段捕获耗时均为 0。共享固定序列客户端新增对既有
--dsv4参数的透传,使用 DeepSeek-V4 chat 编码器,因为 checkpoint 没有 HF chat template。权重与 KV 保留在 GPU。扫描仅覆盖新增配置,默认在 c64/c128 运行 eval,不请求合并。上游 B300 配方使用 TP4 和 EAGLE 3/1/4;本 PR 验证 TP2 及 8k1k 固定序列场景,GPU 性能和准确率仍待 CI 验证。未覆盖 draft 权重、量化或 KV dtype。
本地验证:矩阵生成得到八个点;SRT schema 和逐点选择通过;设置
RUNNER_NAME=local-validation后 111 项针对性测试通过,包括通过真实 B300 启动器验证下载路径传递及下载失败时停止提交;Ruff、Bash 语法与 diff 检查通过。首次测试因既有 H100 fixture 缺少RUNNER_NAME失败,没有改动 H100 代码或测试。AI 模型使用说明:GPT 6 Astra Fast 负责配置、客户端修改、测试、PR 和 CI 监控,未委派其他 agent。
首次 canary 在服务启动前失败:配方使用了仅适用于 AgentX 的
KV_OFFLOADING: none,但固定序列 CI 传入空值。现已移除此可选元数据,并按实际 CI 环境重新验证八个点;Serving 设置和 GPU 常驻策略不变,修正后提交的 GPU 扫描与 eval 仍待完成。第二次 canary 已通过配方选择,但因
/data/models/DeepSeek-V4-Flash尚不存在而失败。临时的主机端hf download(下载至 runner home,Lustre)使后续 canary 得以启动,但服务始终未到达Load weight end:在 gpu-00 上复现,py-spy 显示所有 MoE 加载线程阻塞在cuMemcpyHtoDAsync,mmap 缺页读取 Lustre 副本仅约 10 MB/s,超过 1800 s 健康检查超时。现 DeepSeek-V4-Flash 已预置于全部 17 个可用节点的本地 NVMe(/scratch/models),移除下载路径并加入STAGED_MODELS;节点复现 339 s 完成权重加载,约 10 分钟服务就绪。