Skip to content

Update GB300 DeepSeek V4.1 Flash vLLM nightly / 更新 GB300 DeepSeek V4.1 Flash vLLM nightly - #3396

Merged
cquil11 merged 4 commits into
mainfrom
config/dsv41flash-gb300-vllm-nightly
Oct 2, 2026
Merged

cquil11 merged 4 commits into
mainfrom
config/dsv41flash-gb300-vllm-nightly

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Pin GB300 DeepSeek-V4.1-Flash vLLM to nightly ac68c3087215e0a4f3cdfa218508c6aada57235d on the native srt-slurm AgentX recipe. Use FlashInfer sparse attention at TP2 and TP4 with MXFP4 indexer KV, sparse logits, fp8 KV cache, and Engram CPU offload. Topology and concurrency remain unchanged. Shared TP2 and TP4 points now use B300's batching and CUDA graph tiers; GB300's extra TP2 c1 point retains its existing settings.

中文

在原生 srt-slurm AgentX 配方中,将 GB300 DeepSeek-V4.1-Flash vLLM 固定到 nightly ac68c3087215e0a4f3cdfa218508c6aada57235d。TP2 和 TP4 均使用 FlashInfer 稀疏注意力,并启用 MXFP4 索引器 KV、稀疏 logits、fp8 KV 缓存及 Engram CPU 卸载。 拓扑与并发设置不变。共有的 TP2 和 TP4 测试点现采用 B300 的批处理与 CUDA graph 档位;GB300 独有的 TP2 c1 测试点保持原设置。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@xinli-sw
xinli-sw marked this pull request as ready for review September 23, 2026 17:53
@xinli-sw
xinli-sw requested a review from a team September 23, 2026 17:53
@claude

claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @xinli-sw's task in 4m 3s —— View job


Review of PR #3396

  • Gather context (diff, changed files, repo guidelines)
  • Review benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh changes
  • Validate configs/nvidia-master.yaml image and perf-changelog.yaml entry position
  • Post review

One 🔴 BLOCKING issue (inline on benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh:80-91): the new --attention-config / --kv-cache-dtype fp8 block is gated on GPU_MEM_MIB >= 150000, but dsv41flash_fp4_b200_vllm_mtp.sh and dsv41flash_fp4_gb200_vllm_mtp.sh are symlinks to this shared script and both SKUs clear that threshold. The B200 and GB200 vLLM arms remain pinned to nightly-cd10ed6f (2026-09-15), which predates the "merged Blackwell sparse indexer settings" this PR targets, so they would likely fail at vllm serve argument parsing. Even if the older nightly accepted the flags, their recipes change without a perf-changelog.yaml entry for those config keys. Suggested fix: opt in via an env var exported from runners/launch_gb300-nv.sh (same pattern as VLLM_SAFETENSORS_LOAD_STRATEGY), or re-pin B200/GB200 and document them in the changelog. Fix this →

Other checks passed: configs/nvidia-master.yaml and perf-changelog.yaml are edited together, the new changelog entry is appended at the tail of the file, the image is a public Docker Hub tag, and the TP2/TP4 backend selection matches the PR description. Note that I could not reach the upstream vLLM source or tracking issue from this session (network fetches were not permitted), so the claim that cd10ed6f lacks the new keys is inferred from the PR description and the pin dates rather than verified against the code.
· Branch

Comment thread benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh Outdated
Comment on lines +83 to +90
ATTENTION_ARGS=()
if (( GPU_MEM_MIB >= 150000 )); then
if (( TP == 2 )); then
ATTENTION_CONFIG='{"backend":"FLASHINFER_MLA_SPARSE_DSV41","indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'
else
ATTENTION_CONFIG='{"backend":"FLASHMLA_MEGA_ATTN_DSV41","indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'
fi
ATTENTION_ARGS=(--attention-config "$ATTENTION_CONFIG" --kv-cache-dtype fp8)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 This script is symlinked to by dsv41flash_fp4_b200_vllm_mtp.sh and dsv41flash_fp4_gb200_vllm_mtp.sh too, so the new ATTENTION_ARGS block also fires for those arms, not just GB300. GPU_MEM_MIB>=150000 is true on B200 (180GB) and GB200 (256GB) as well as GB300, but only the gb300-vllm-agentic-dspark image was bumped; b200 and gb200 still pin vllm/vllm-openai:nightly-cd10ed6f9f6b37a8ace9cf380007e66fe12ec0c3, the same build documented as lacking these SM100 sparse-indexer backends on H200. Passing --attention-config with FLASHINFER_MLA_SPARSE_DSV41/FLASHMLA_MEGA_ATTN_DSV41 and --kv-cache-dtype fp8 to that older image will likely fail vllm serve startup, breaking the b200 and gb200 agentic-coding benchmark runs. …

Why this was flagged

…Fix: gate ATTENTION_ARGS on the image/nightly actually shipping this feature (e.g. per-config-key or image-tag check), not solely on GPU_MEM_MIB, and add matching perf-changelog entries if those arms' behavior does change too.

dsv41flash_fp4_b200_vllm_mtp.sh and dsv41flash_fp4_gb200_vllm_mtp.sh are symlinks to this file (benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh), confirmed via ls -la showing both -> dsv41flash_fp4_vllm_mtp.sh. Their configs (configs/nvidia-master.yaml lines ~8092 and ~8155, dsv41flash-fp4-gb200-vllm-agentic-dspark and dsv41flash-fp4-b200-vllm-agentic-dspark) still use image vllm/vllm-openai:nightly-cd10ed6f9f6b37a8ace9cf380007e66fe12ec0c3, unchanged by this diff. Both are Blackwell so GPU_MEM_MIB>=150000 at line 84 is true, so the new --attention-config/--kv-cache-dtype fp8 args at line 90 get added to their vllm serve invocation. The PR/changelog only pins and documents this feature for dsv41flash-fp4-gb300-vllm-agentic-dspark on the new nightly-0961bbae image; the older shared nightly is the same one the script's own comment says lacks SM100…

Verification: normal. The new block at benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh:83-91 is gated only on hardware memory (if (( GPU_MEM_MIB >= 150000 ))), and GPU_MEM_MIB comes from nvidia-smi --query-gpu=memory.total (line 59). The script's own comment at line 51 enumerates the Blackwell arms sharing this file — "B200 180 GB, GB200 256 GB, GB300 277 GB" — all far above 150000 MiB; the…

@xinli-sw xinli-sw added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 23, 2026
@xinli-sw
xinli-sw marked this pull request as draft September 24, 2026 15:54
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb300-vllm-nightly branch 3 times, most recently from 3494e65 to b4fa176 Compare September 25, 2026 02:03
@functionstackx

Copy link
Copy Markdown
Collaborator

InferenceX has switched away from unmaintainable bash scripts to YAML files that don't repeat the same stuff over and over again. Please merge the latest main into this PR: we have migrated single-node AgentX onto native srt-slurm (#3428), so AgentX configs are now declarative YAML recipes, not per-config 1000+ line bash slop scripts. Please also delete the old benchmarks/single_node/** scripts (see this recipe for the new format).

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb300-vllm-nightly branch 4 times, most recently from 5e51844 to 70cf20f Compare September 27, 2026 22:11
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb300-vllm-nightly branch from 70cf20f to e3dbe06 Compare September 28, 2026 12:40
@xinli-sw
xinli-sw force-pushed the config/dsv41flash-gb300-vllm-nightly branch 2 times, most recently from d62acd0 to 85c905b Compare September 30, 2026 15:26
将 GB300 vLLM 固定到 ac68c308 nightly,并保留现有批处理配置。
@cquil11

cquil11 commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Refreshed this branch with main for the GB300 rollout (#3658 and #3662). Existing runner instances now span two Slurm partitions under cluster:gb300-nv; old queued revisions lacked that routing, and inherited SBATCH_PARTITION also needed explicit enforcement. I canceled the stale sweep and the push started a fresh one: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/36955640992

This PR remains open. The rollout refresh preserves its benchmark-specific changes and appends its existing changelog entry after current main. No benchmark PR was merged as part of the infrastructure rollout.

@xinli-sw

xinli-sw commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

/use 36955640992

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@xinli-sw staged run 36955640992: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-02~r36955640992

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

@xinli-sw
xinli-sw marked this pull request as ready for review October 2, 2026 14:40
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @xinli-sw's task in 45s —— View job


LGTM - no blocking issues found
· Branch

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR; the one confirmed issue is already flagged inline. Beyond that, I checked two other candidates and ruled them out: override_tp2_c1 keeping its old cudagraph/batching settings matches the PR description's stated intent to leave GB300's TP2 c1 point unchanged, and the new kv-cache-dtype: fp8 / MXFP4 indexer KV dtype are applied uniformly to target and draft, which CONTRIBUTING.md's draft-precision rule explicitly permits.

Extended reasoning...

The diff touches only CI/benchmark config: a GB300 vLLM DeepSeek-V4.1-Flash agentic recipe, its nvidia-master.yaml entry, and the append-only perf-changelog (image bump, FlashInfer sparse-attention/KV-cache tuning, and new cudagraph compilation configs) — no auth, crypto, or runtime code touched. One confirmed finding (a stale TP2 memory-budget comment in nvidia-master.yaml) is posted inline. I separately verified override_tp2_c1's unchanged settings against the PR's own description (intentional) and the fp8/MXFP4 KV-dtype changes against CONTRIBUTING.md's draft-precision rule (compliant, since it's applied consistently to target and draft), so those are not additional bugs.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟡 inferencex-e2e/configs/nvidia-master.yaml — A maintainer reading nvidia-master.yaml's TP2 comment gets numbers that no longer match the recipe this PR just rewrote. Lines 8348-8352 still say TP2 caps batched tokens at 4096 and stops graph capture at 512, leaving ~36 GiB KV headroom, but the paired recipe (gb300-fp4-mtp/agentic.yaml) now sets max-num-batched-tokens to 8192 (or 2048) and max-cudagraph-capture-size up to 8190 (or 2046) for every TP2 point except c1. The sibling b300 entry at nvidia-master.yaml:8288-8289 was correctly updated, so this one was simply missed. Fix: update the gb300 TP2 comment to state the actual new batched-token/capture-size tiers and recompute the KV headroom math (indexer buffer now scales off 8192, not 4096).

    Why this was flagged

    The comment at inferencex-e2e/configs/nvidia-master.yaml:8348-8352 documents the pre-PR TP2 memory budget (max-num-batched-tokens 4096, graph capture stopped at 512, ~36 GiB KV headroom) for dsv41flash-fp4-gb300-vllm-agentic-dspark. The paired recipe it describes, gb300-fp4-mtp/agentic.yaml, was rewritten by this same diff so override_tp2_c8 through c64 now set max-num-batched-tokens: 8192 and max-cudagraph-capture-size: 8190, contradicting the comment's numbers. The sparse-attention indexer buffer scales with max-num-batched-tokens, so doubling it from 4096 to 8192 shrinks the documented ~36 GiB KV headroom without the comment being updated. The sibling dsv41flash-fp4-b300-vllm-agentic-dspark entry (nvidia-master.yaml:8288-8289) got its comment correctly rewritten for the same scheme change, showing this is an oversight specific to the gb300 entry. A future engineer tuning GB300 TP2 memory budgets off this comment will plan against stale figures.

    Verification: The comment at inferencex-e2e/configs/nvidia-master.yaml:8348-8352 documents the pre-PR TP2 budget: batched tokens 4096 and graph capture stopped at 512, leaving ~36 GiB of KV per GPU.

@wzhao18 wzhao18 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Assessed commit: fb2747a0c2ed70330c3499e084bb7b73f9fe4980.
  • PR validation: CI and full sweep, attempt 1 passed on the assessed commit. The sweep produced 16 successful performance points across TP2 and TP4, with concurrency 1–128. All 16 performance records have valid power data.
  • Evals: the same sweep and image passed GSM8K at TP4 c128, with em_strict = 0.9727065959059894, n_eff = 1319, and infrastructure_success: true. Evidence: eval_results_all. This is one evaluated configuration, not an eval at every performance point.
  • Reuse: authorized /use 36955640992 was accepted.
  • Recipe documentation: published DeepSeek-V4.1-Flash recipe, with the Blackwell sparse-attention settings documented in vllm-project/recipes#995, merged September 20, and GB300 TP2 Engram offload in vllm-project/recipes#1002, merged September 21.
  • Image and serving stack: upstream vllm/vllm-openai:nightly-ac68c3087215e0a4f3cdfa218508c6aada57235d, matching the recipe and master config. Performance and eval logs report vLLM 0.30.1rc1.dev396+gac68c3087. No engine patch or replacement wheel is introduced.
  • Draft precision: inspected deepseek-ai/DeepSeek-V4.1-Flash@2cba9e42aa026125f3ed06c6d98c1db82f7ca027. Safetensors headers contain 2,401 embedded mtp.* tensors across shards 44–46: packed MXFP4 expert weights, FP8 dense weights and E8M0 scales, plus BF16/F32 auxiliary tensors. The pinned upstream DSpark loader loads these embedded weights through the checkpoint’s default quantization handling. The effective recipe adds no draft checkpoint substitution, dtype override, or quantization override. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent and inapplicable to this vLLM submission.
  • AgentX acceptance and chat path: throughput uses five-token probabilistic DSpark drafting with synthetic AL 3.51, matching thinking_on K5 in the committed golden curve. Runtime artifacts confirm synthetic acceptance; eval uses real block rejection without synthetic acceptance. Replay uses /v1/chat/completions.
  • Pareto evidence: run 36955640992, attempt 1, assessed SHA above, results_bmk, AgentX P90 E2EL versus total token throughput/GPU, one GB300 vLLM FP4 visual curve and one image. There are 16 valid measured points and 8 frontier points: TP4 c8/c16/c32 and TP2 c8/c16/c32/c64/c128. The repository’s pareto_coverage helper reports PASS.
  • append-only is N/A. This PR changes existing recipes and the image and uses a regular changelog entry. All historical changelog bytes are preserved, with the new entry appended at the tail.

Signed: wzhao18

Preserve main changelog entries and append the GB300 contribution.
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

✅✅✅ Verdict: PASS ✅✅✅

Passed and not applicable checks

✅ Check 0 (CODEOWNER): PASS — @wzhao18 is a listed owner of inferencex-e2e/configs/nvidia-master.yaml and of the srt-slurm-recipes/**/gb300-*/ recipe. perf-changelog.yaml falls only under the catch-all.

✅ Check 1 (Sweep on in-PR commit): PASS — PR head is unchanged at fb2747a. That commit has 16/16 agentic / jobs and the agentic eval / job at success in run 36955640992 (attempt 1, headSha = pinned SHA). This is an AgentX-only config, so the single-node */ and eval / jobs are correctly skipped.

✅ Check 2 (Evals): PASS — In eval_results_all, GSM8K at TP4 c128 scored em_strict 0.9727 (n=1319) with infrastructure_success: true. It ran on the same image, vllm/vllm-openai:nightly-ac68c3087215e0a4f3cdfa218508c6aada57235d.

✅ Check 3 (Recipe): PASS — vllm-project/recipes#995 and #1002 are both MERGED. The published recipe matches the major args: GB300 at TP2/TP4, FLASHINFER_MLA_SPARSE_DSV41 with MXFP4 indexer KV and sparse logits, fp8 KV cache, Engram cpu_offload, and DSpark K5 probabilistic drafting. Informational only: enable_adaptive_verification:false matches the golden-AL setup; max-num-seqs, batching limits and CUDA graph sizes are scheduler tuning; VLLM_DISABLED_KERNELS works around a startup failure.

✅ Check 4 (Reuse command): PASS — /use 36955640992 was posted by xinli-sw (COLLABORATOR).

✅ Check 5 (Latest template): PASS — The sign-off contains every item of the current PR_REVIEW_CHECKLIST.md template, and all of them are checked.

✅ Check 6 (Upstream image / ordering): PASS — framework: vllm uses the upstream image vllm/vllm-openai:nightly-ac68c308…. The PR adds only open-source engine entries, so engine-first ordering is satisfied.

✅ Check 7 (Deprecated models): PASS — MODELS.md lists dsv41flash agentic coding as active, with no deprecation as of 2026-10-02.

✅ Check 8 (No arch hacks): PASS — There is no --hf-overrides, other architecture override or layer trimming. publish_events_and_metrics is a TRT-LLM setting that does not apply to this vLLM recipe. The server exposes vllm: metrics, and AIPERF_REQUIRED_SERVER_METRIC_PREFIX requires them.

✅ Check 9 (Spec-decode chat template): PASS — build_replay_cmd in benchmark_lib.sh sends replay to --endpoint /v1/chat/completions.

✅ Check 10 (No engine patches): PASS — The PR has no patch files, inline source edits, monkey-patching or wheel installs. VLLM_DISABLED_KERNELS is a supported vLLM env knob, not a patch.

✅ Check 11 (Agentic golden AL): PASS — The runtime override renders "rejection_sample_method": "synthetic", "synthetic_acceptance_length": 3.51 (seen in the TP2 c128 job log). That equals the thinking_on K5 value in dsv41flash_dspark.yaml. The eval path keeps block rejection.

➖ Check 12 (Append-only): N/A — The new changelog entry does not set append-only: true.

✅ Check 13 (Draft as shipped): PASS — The DSpark draft is the embedded mtp.* weights of deepseek-ai/DeepSeek-V4.1-Flash, loaded by the pinned loader through the target's quant_config. No draft quantization, dtype or checkpoint override is set. The fp8 KV cache applies to target and draft alike (allowed), and the draft uses only sliding-window caches, so the MXFP4 indexer setting does not reach it. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is absent.

✅ Check 14 (Pareto coverage): PASS — One curve was checked: GB300 vLLM FP4 DSpark, AgentX P90 E2EL against total_tput_tps per GPU. Source: results_bmk, run 36955640992 attempt 1, one image. It has 16 valid measured points and 8 frontier points. pareto_coverage reports PASS (8/5).

Assessed commit: fb2747a0c2ed70330c3499e084bb7b73f9fe4980.

@cquil11
cquil11 merged commit 9cae7df into main Oct 2, 2026
25 checks passed
@cquil11
cquil11 deleted the config/dsv41flash-gb300-vllm-nightly branch October 2, 2026 15:55
Juntian777 added a commit that referenced this pull request Oct 2, 2026
Resolve perf-changelog, nvidia-master and recipe conflicts with #3396 and
#3653; the GB300 recipe and master entry stay as swept in run 36955641333.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants