Skip to content

Update GB300 DSV41flash vLLM image and retune TP4/DEP2 / 更新 GB300 DSV41flash vLLM 镜像并调优 TP4/DEP2 - #3652

Merged
functionstackx merged 4 commits into
mainfrom
vllm/dsv41flash-gb300-image-bump
Oct 2, 2026
Merged

functionstackx merged 4 commits into
mainfrom
vllm/dsv41flash-gb300-image-bump

Conversation

@Juntian777

@Juntian777 Juntian777 commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Description

  • TP4 covers concurrency 1–16; DEP2 (TP1 x DP2 + EP2, MegaMoE) replaces TP2 at 8–192 behind a consistent-hash vLLM Router.
  • FlashInfer autotuning is enabled; MegaAttention is used at concurrency 128 and above.
  • vLLM recipe: vllm-project/recipes#1062 adds the vllm-router command for single-node DEP.

Validation

  • infx/tests/srt_slurm and infx/tests/matrix pass; matrix generation maps all 10 points to exactly one variant each.
  • A GPU sweep via the full-sweep-fail-fast label is pending. No performance results are claimed yet.

AI model disclosure

  • Model/version: Claude Opus 5.5 (claude-opus-5-5) in Claude Code; no delegated agents.
  • Role: drafted the recipe, master config, documentation and perf-changelog changes, ran the local validation, and wrote this description.

Related Issue

None. Supersedes #3575.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • If this PR can affect benchmark performance or adds or modifies a recipe, it carries exactly one primary sweep label (a maintainer applies it on fork PRs): full-sweep-fail-fast (recommended), full-sweep-enabled, or non-canary-full-sweep-enabled. Optional modifiers all-evals, evals-only, and agentx-fast require a primary label; the last two block reuse while applied.
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.
中文

改动说明

  • TP4 覆盖并发 1–16;DEP2(TP1 x DP2 + EP2,MegaMoE)替代 TP2,覆盖 8–192,前置一致性哈希 vLLM Router。
  • 开启 FlashInfer autotune;并发 128 及以上使用 MegaAttention。
  • vLLM recipe:vllm-project/recipes#1062 为单节点 DEP 补充 vllm-router 命令。

验证

  • infx/tests/srt_slurm 和 infx/tests/matrix 通过;矩阵生成的 10 个点均唯一匹配一个 variant。
  • 通过 full-sweep-fail-fast label 触发 GPU sweep,结果待出,暂不声明性能结果。

AI 模型使用说明

  • 模型/版本:Claude Code 中的 Claude Opus 5.5(claude-opus-5-5);没有委派其他 agent。
  • 工作内容:起草配方、master config、文档和 perf-changelog 改动,运行本地验证,并撰写本说明。

关联 issue

无。取代 #3575。

改动类型

配置变更、文档更新。

🤖 Generated with Claude Code

@Juntian777
Juntian777 requested a review from a team October 2, 2026 00:33
@Juntian777 Juntian777 added full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) agentx AgentX benchmarks, recipes, and infrastructure labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps run only when the PR appends an inferencex-e2e/perf-changelog.yaml entry and carries exactly one primary label: full-sweep-fail-fast (strongly recommended; canary plus per-matrix fail-fast), full-sweep-enabled (canary; matrix jobs continue after a failure), or non-canary-full-sweep-enabled (no canary or fail-fast). The modifiers all-evals, evals-only, and agentx-fast require a primary label. On fork PRs, a maintainer applies the label. See sweep labels and reuse.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**只有当 PR 在 inferencex-e2e/perf-changelog.yaml 末尾追加了条目,并且恰好带有一个主标签时,才会运行扫描:full-sweep-fail-fast(强烈推荐;canary 加逐矩阵 fail-fast)、full-sweep-enabled(有 canary;矩阵任务在失败后继续运行)或 non-canary-full-sweep-enabled(无 canary,也无 fail-fast)。修饰标签 all-evals、evals-only 和 agentx-fast 必须与主标签一起使用。fork PR 的标签由维护者添加。参见扫描标签与复用。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it's a large, intricate vLLM performance-tuning config change (image bump, TP4/DEP2 retune, CUDA-graph capture sizes, decoder-replay sizes, FlashInfer autotune) that the author's own description notes hasn't yet been GPU-run on GB300 with this exact recipe, a human look would still be worthwhile.

What was reviewed: the recipe YAML (agentic.yaml), the master config arm definitions, the perf-changelog entry, and the English/Chinese docs updates. Checked that the changelog entry was appended (not edited in place) and that nvidia-master.yaml's new TP4/DEP2 arms match the recipe and PR description. Also checked the engine-ready timeout vs. new FlashInfer autotuning/graph-capture startup cost (flagged as a candidate concern and ruled out as not blocking).

Extended reasoning...

The diff is config/YAML/docs only across five files: a vLLM recipe (image, attention/KV-cache/compilation/decoder-replay settings), the master config's search-space arms for this model, a perf-changelog append, and bilingual docs. No code paths, auth, or security-sensitive surface is touched. The deciding factors for defer over approve are the size and interdependency of the tuning parameters (CUDA graph capture sizes, decoder-replay sizes, FlashInfer autotune buckets, DEP2 router config) and the PR author's own admission that the recipe has not been validated on GB300 hardware with this exact image/configuration, which warrants a human performance-domain review before merge.

@Juntian777
Juntian777 force-pushed the vllm/dsv41flash-gb300-image-bump branch from d7d1c1a to c325fbe Compare October 2, 2026 01:18
Move the GB300 DeepSeek-V4.1-Flash vLLM AgentX recipe to
nightly-dev-arm64-cu130-ac9126e58aa7 and enable FlashInfer autotuning.
Run TP4 at CONC 1-16 and replace TP2 with DEP2 (TP1 x DP2 + EP2,
MegaMoE) at CONC 8-192 behind a consistent-hash vLLM Router. Append the
perf-changelog entry and document the GB300 settings.

将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 配方更新到
nightly-dev-arm64-cu130-ac9126e58aa7,并开启 FlashInfer autotune。
TP4 覆盖 CONC 1-16;用 DEP2(TP1 x DP2 + EP2,MegaMoE)替代 TP2,覆盖
CONC 8-192,前置一致性哈希 vLLM Router。追加 perf-changelog 条目,并补充
GB300 配置文档。

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@Juntian777
Juntian777 force-pushed the vllm/dsv41flash-gb300-image-bump branch from c325fbe to 32ac795 Compare October 2, 2026 01:25
@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@cquil11

cquil11 commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Refreshed this branch with main for the GB300 rollout (#3658 and #3662). Existing runner instances now span two Slurm partitions under cluster:gb300-nv; old queued revisions lacked that routing, and inherited SBATCH_PARTITION also needed explicit enforcement. I canceled the stale sweep and the push started a fresh one: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/36955641333

This PR remains open. The rollout refresh preserves its benchmark-specific changes and appends its existing changelog entry after current main. No benchmark PR was merged as part of the infrastructure rollout.

@Juntian777

Copy link
Copy Markdown
Collaborator Author

/use 36955641333

@github-actions

github-actions Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

@Juntian777 staged run 36955641333: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-02~r36955641333

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

@Juntian777
Juntian777 force-pushed the vllm/dsv41flash-gb300-image-bump branch from 7ffba3c to 6351628 Compare October 2, 2026 21:30
Resolve perf-changelog, nvidia-master and recipe conflicts with #3396 and
#3653; the GB300 recipe and master entry stay as swept in run 36955641333.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@Juntian777
Juntian777 force-pushed the vllm/dsv41flash-gb300-image-bump branch from 6351628 to 5e5c356 Compare October 2, 2026 21:31
@functionstackx
functionstackx merged commit 8988109 into main Oct 2, 2026
24 checks passed
@functionstackx
functionstackx deleted the vllm/dsv41flash-gb300-image-bump branch October 2, 2026 21:35
cursor Bot pushed a commit that referenced this pull request Oct 2, 2026
Resolve dirty mergeable_state from main changelog appends (#3402/#3652/#3653
et al.). Keep origin/main perf-changelog bytes as prefix and re-append all
#3088 tip entries including CONC 32+ max-num-seqs 1x. Recipe tip unchanged.

解决 main changelog 追加(#3402/#3652/#3653 等)导致的 dirty。保留
origin/main 的 perf-changelog 字节为前缀,并重新追加全部 #3088 tip
条目(含 CONC 32+ max-num-seqs 1×)。配方 tip 不变。

Co-authored-by: Wenyao Gao <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants