Skip to content

[DNM][Experimental] Add B300 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B300 TP2 V4 Flash SGLang 8k1k MTP 扫描 - #3516

Open
functionstackx wants to merge 7 commits into
mainfrom
benchmark/dsv4flash-b300-tp2-sglang-eagle-8k1k
Open

functionstackx wants to merge 7 commits into
mainfrom
benchmark/dsv4flash-b300-tp2-sglang-eagle-8k1k

Conversation

@functionstackx

@functionstackx functionstackx commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Add dsv4flash-fp4-b300-sglang-mtp for the exact deepseek-ai/DeepSeek-V4-Flash checkpoint: B300 TP2, 8192 input / 1024 output tokens, concurrency 1, 2, 4, 8, 16, 32, 64, 128. This is separate from V4 Pro and V4.1 Flash.

  • Use bundled MTP through EAGLE, with 2 steps, top-k 1, and 3 draft tokens. The requester explicitly clarified this choice rather than the distinct EAGLE3 algorithm.
  • CUDA/HIP graphs are disabled (cuda-graph-backend-decode: disabled, cuda-graph-backend-prefill: disabled), so decode, prefill, and EAGLE target-verify/draft all run eagerly; startup reports zero capture time for every graph phase.
  • Pin the official SGLang CUDA 13 nightly nightly-dev-cu13-20260925-8ca82118 to its amd64 digest. Keep the checkpoint's shipped draft head and upstream precision handling; explicitly disable SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE.
  • Keep weights and KV on GPU, with real speculative verification. No engine patches, external draft, CPU offloading, or simulated acceptance.
  • Forward the existing --dsv4 client option through srt_fixed_sequence.sh, using the shared DeepSeek-V4 chat encoder. This checkpoint has no HF chat template.
  • Scope the sweep to the new fixed-sequence key. Default eval selection marks c64/c128. No merge requested.

Upstream recipe reference: DeepSeek-V4 B300 configuration. Its original Flash recipe uses bundled EAGLE 3/1/4 at TP4; this PR tests TP2 and an 8k1k fixed-sequence workload. TP2 runtime performance and accuracy remain unverified until CI passes.

Checkpoint metadata: DeepSeek-V4-Flash config. The target has FP4 experts, FP8 quantization metadata, and one bundled NextN layer. This PR does not override draft weights, draft quantization, or draft KV dtype.

Validation

  • Matrix generation: eight TP2 8k1k points, c1–128; c64 and c128 marked for eval.
  • Upstream SRT schema validation: all variants pass.
  • Native recipe selection: exactly one matching variant per point, two serving GPUs, EAGLE enabled, no simulated acceptance.
  • Focused SRT tests: unchanged from main after dropping the download-path cases; CI runs the suite.
  • ruff check infx, ruff format --check infx, Bash syntax, and git diff --check pass.
  • Initial canary failed before server launch because the recipe supplied AgentX-only KV_OFFLOADING: none, while fixed-sequence CI supplies an empty value. Removed that optional metadata; serving and GPU residency settings are unchanged. Revalidated all eight points using the actual fixed-sequence CI environment.
  • The next canary passed recipe selection but failed because /data/models/DeepSeek-V4-Flash did not exist yet. An interim host-side hf download into the runner home (Lustre) let the following canary start, but the server never reached Load weight end: reproduced on dsxe-sa-b300-prd0-gpu-00, py-spy showed every MoE loader thread blocked in cuMemcpyHtoDAsync while mmap page faults read the Lustre copy at ~10 MB/s (~4 h for 160 GB), past the 1800 s health timeout.
  • DeepSeek-V4-Flash is now staged on node-local NVMe (/scratch/models, 46 shards / 159,617,149,040 bytes verified on all 17 up nodes). The download path is removed and the checkpoint is listed in STAGED_MODELS. On-node repro from NVMe with a 192-CPU exclusive allocation: Load weight end at 339 s (incl. 244 s MHC JIT prewarm), server ready in ~10 min, smoke completion correct.
  • GPU sweep and evals: pending on the corrected head (EAGLE 2/1/3, NVMe weights).

AI model disclosure

GPT 6 Astra Fast prepared the configuration, client change, test, PR, and CI monitoring. No delegated agents.

中文

新增 dsv4flash-fp4-b300-sglang-mtp:使用原始 deepseek-ai/DeepSeek-V4-Flash checkpoint,在 B300 上运行 TP2、8192 输入 / 1024 输出 Token 的 serving 扫描,并发为 1、2、4、8、16、32、64、128。此配置与 V4 Pro、V4.1 Flash 分开。

用户已明确选择通过 EAGLE 使用原生 MTP(2 steps、top-k 1、3 draft tokens),不是独立的 EAGLE3 算法。镜像固定为官方 CUDA 13 nightly nightly-dev-cu13-20260925-8ca82118 的 amd64 digest。保留 checkpoint 自带的 draft head 及上游默认精度处理,显式关闭 SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE;不修改引擎、不使用外部 draft、不卸载至 CPU、不模拟接受长度。

  • 关闭 CUDA/HIP graph(cuda-graph-backend-decode: disabled、cuda-graph-backend-prefill: disabled),decode、prefill 及 EAGLE 验证/draft 均以 eager 模式运行;启动日志中各 graph 阶段捕获耗时均为 0。

共享固定序列客户端新增对既有 --dsv4 参数的透传,使用 DeepSeek-V4 chat 编码器,因为 checkpoint 没有 HF chat template。权重与 KV 保留在 GPU。扫描仅覆盖新增配置,默认在 c64/c128 运行 eval,不请求合并。

上游 B300 配方使用 TP4 和 EAGLE 3/1/4;本 PR 验证 TP2 及 8k1k 固定序列场景,GPU 性能和准确率仍待 CI 验证。未覆盖 draft 权重、量化或 KV dtype。

本地验证:矩阵生成得到八个点;SRT schema 和逐点选择通过;设置 RUNNER_NAME=local-validation 后 111 项针对性测试通过,包括通过真实 B300 启动器验证下载路径传递及下载失败时停止提交;Ruff、Bash 语法与 diff 检查通过。首次测试因既有 H100 fixture 缺少 RUNNER_NAME 失败,没有改动 H100 代码或测试。

AI 模型使用说明:GPT 6 Astra Fast 负责配置、客户端修改、测试、PR 和 CI 监控,未委派其他 agent。

首次 canary 在服务启动前失败:配方使用了仅适用于 AgentX 的 KV_OFFLOADING: none,但固定序列 CI 传入空值。现已移除此可选元数据,并按实际 CI 环境重新验证八个点;Serving 设置和 GPU 常驻策略不变,修正后提交的 GPU 扫描与 eval 仍待完成。

第二次 canary 已通过配方选择,但因 /data/models/DeepSeek-V4-Flash 尚不存在而失败。临时的主机端 hf download(下载至 runner home,Lustre)使后续 canary 得以启动,但服务始终未到达 Load weight end:在 gpu-00 上复现,py-spy 显示所有 MoE 加载线程阻塞在 cuMemcpyHtoDAsync,mmap 缺页读取 Lustre 副本仅约 10 MB/s,超过 1800 s 健康检查超时。现 DeepSeek-V4-Flash 已预置于全部 17 个可用节点的本地 NVMe(/scratch/models),移除下载路径并加入 STAGED_MODELS;节点复现 339 s 完成权重加载,约 10 分钟服务就绪。

新增 B300 TP2 DeepSeek V4 Flash SGLang 8k1k serving 扫描,使用原生 EAGLE 三步 MTP。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

将新增 V4 Flash 扫描记录关联到 PR #3516。
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

移除仅适用于 AgentX 的 KV_OFFLOADING 元数据,修复固定序列 CI 配方选择失败。
提交前将 V4 Flash 下载到 B300 runner 的 home 可写目录,保留其他模型路径,并测试下载失败时停止提交。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good — a self-contained new benchmark recipe/sweep entry with a minimal, backward-compatible change to the shared fixed-sequence client script. Reviewed the new dsv4flash-fp4-b300-sglang-mtp recipe/sweep entry, the srt_fixed_sequence.sh allowlist addition for --dsv4, and confirmed --dsv4 is already handled downstream by benchmark_lib.sh's run_benchmark_serving (this PR only forwards an existing flag, adds no new client behavior). The model-prefix startswith-matching concern raised during triage is pre-existing sweep-filter behavior shared by other keys (e.g. dsv41flash-*), not something this PR introduces.

Extended reasoning...

Change adds a new single-node SGLang B300 TP2 8k1k MTP recipe/sweep entry for DeepSeek-V4-Flash, a one-line allowlist widening in srt_fixed_sequence.sh to forward an already-existing --dsv4 client flag, plus docs, changelog, and a new subprocess-based test. No auth, crypto, or data-exposure surface is touched; the only candidate issue (prefix-filter overlap with dsv4) is a pre-existing sweep-filter pattern also present for other keys, not new. Small, mechanical, and well-tested, so a human need not additionally review before merge.

This review covers commit 4bde51e, which is no longer the latest commit on this pull request; later commits are not covered by it.

主机下载显式使用 home 下的 HF 和 Xet 缓存,避免继承容器路径导致权限错误。
The shared Lustre copy is read through mmap page faults at ~10 MB/s, so
TP2 weight loading outlasted the 1800 s health timeout. DeepSeek-V4-Flash
is now staged under /scratch/models on every B300 node; drop the host-side
download path and list it in STAGED_MODELS. Switch MTP to 2 steps, top-k 1,
3 draft tokens.
Set cuda-graph-backend-{decode,prefill}=disabled, which also skips target
verify and draft capture, and drop the per-concurrency decode graph batch.
@functionstackx functionstackx changed the title Add B300 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B300 TP2 V4 Flash SGLang 8k1k MTP 扫描 [DNM][Experimental] Add B300 TP2 V4 Flash SGLang 8k1k MTP sweep / 新增 B300 TP2 V4 Flash SGLang 8k1k MTP 扫描 Sep 27, 2026
@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants