Update GB300 DSV41flash vLLM image and retune TP4/DEP2 / 更新 GB300 DSV41flash vLLM 镜像并调优 TP4/DEP2 - #3652
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it's a large, intricate vLLM performance-tuning config change (image bump, TP4/DEP2 retune, CUDA-graph capture sizes, decoder-replay sizes, FlashInfer autotune) that the author's own description notes hasn't yet been GPU-run on GB300 with this exact recipe, a human look would still be worthwhile.
What was reviewed: the recipe YAML (agentic.yaml), the master config arm definitions, the perf-changelog entry, and the English/Chinese docs updates. Checked that the changelog entry was appended (not edited in place) and that nvidia-master.yaml's new TP4/DEP2 arms match the recipe and PR description. Also checked the engine-ready timeout vs. new FlashInfer autotuning/graph-capture startup cost (flagged as a candidate concern and ruled out as not blocking).
Extended reasoning...
The diff is config/YAML/docs only across five files: a vLLM recipe (image, attention/KV-cache/compilation/decoder-replay settings), the master config's search-space arms for this model, a perf-changelog append, and bilingual docs. No code paths, auth, or security-sensitive surface is touched. The deciding factors for defer over approve are the size and interdependency of the tuning parameters (CUDA graph capture sizes, decoder-replay sizes, FlashInfer autotune buckets, DEP2 router config) and the PR author's own admission that the recipe has not been validated on GB300 hardware with this exact image/configuration, which warrants a human performance-domain review before merge.
d7d1c1a to
c325fbe
Compare
Move the GB300 DeepSeek-V4.1-Flash vLLM AgentX recipe to nightly-dev-arm64-cu130-ac9126e58aa7 and enable FlashInfer autotuning. Run TP4 at CONC 1-16 and replace TP2 with DEP2 (TP1 x DP2 + EP2, MegaMoE) at CONC 8-192 behind a consistent-hash vLLM Router. Append the perf-changelog entry and document the GB300 settings. 将 GB300 DeepSeek-V4.1-Flash vLLM AgentX 配方更新到 nightly-dev-arm64-cu130-ac9126e58aa7,并开启 FlashInfer autotune。 TP4 覆盖 CONC 1-16;用 DEP2(TP1 x DP2 + EP2,MegaMoE)替代 TP2,覆盖 CONC 8-192,前置一致性哈希 vLLM Router。追加 perf-changelog 条目,并补充 GB300 配置文档。 Co-Authored-By: Claude Opus 5.5 <[email protected]>
c325fbe to
32ac795
Compare
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36955641333 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36955641333 |
|
Refreshed this branch with main for the GB300 rollout (#3658 and #3662). Existing runner instances now span two Slurm partitions under cluster:gb300-nv; old queued revisions lacked that routing, and inherited SBATCH_PARTITION also needed explicit enforcement. I canceled the stale sweep and the push started a fresh one: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/36955641333 This PR remains open. The rollout refresh preserves its benchmark-specific changes and appends its existing changelog entry after current main. No benchmark PR was merged as part of the infrastructure rollout. |
|
/use 36955641333 |
|
@Juntian777 staged run 36955641333: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-02~r36955641333 This run remains available across future |
7ffba3c to
6351628
Compare
Resolve perf-changelog, nvidia-master and recipe conflicts with #3396 and #3653; the GB300 recipe and master entry stay as swept in run 36955641333. Co-Authored-By: Claude Opus 5.5 <[email protected]>
6351628 to
5e5c356
Compare
Resolve dirty mergeable_state from main changelog appends (#3402/#3652/#3653 et al.). Keep origin/main perf-changelog bytes as prefix and re-append all #3088 tip entries including CONC 32+ max-num-seqs 1x. Recipe tip unchanged. 解决 main changelog 追加(#3402/#3652/#3653 等)导致的 dirty。保留 origin/main 的 perf-changelog 字节为前缀,并重新追加全部 #3088 tip 条目(含 CONC 32+ max-num-seqs 1×)。配方 tip 不变。 Co-authored-by: Wenyao Gao <[email protected]>
Description
Validation
infx/tests/srt_slurmandinfx/tests/matrixpass; matrix generation maps all 10 points to exactly one variant each.full-sweep-fail-fastlabel is pending. No performance results are claimed yet.AI model disclosure
claude-opus-5-5) in Claude Code; no delegated agents.Related Issue
None. Supersedes #3575.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesfull-sweep-fail-fast(recommended),full-sweep-enabled, ornon-canary-full-sweep-enabled. Optional modifiersall-evals,evals-only, andagentx-fastrequire a primary label; the last two block reuse while applied.OWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.中文
改动说明
验证
infx/tests/srt_slurm和infx/tests/matrix通过;矩阵生成的 10 个点均唯一匹配一个 variant。full-sweep-fail-fastlabel 触发 GPU sweep,结果待出,暂不声明性能结果。AI 模型使用说明
claude-opus-5-5);没有委派其他 agent。关联 issue
无。取代 #3575。
改动类型
配置变更、文档更新。
🤖 Generated with Claude Code