[NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang aggregated and 1PxD recipes / [NVIDIA][AgentX] GB300 DeepSeek-V4.1-Flash Dynamo+SGLang 聚合与 1PxD 配方 - #3598
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it changes SLURM launcher routing (new partition/account/shared-root branch) for multi-node GB300 jobs, a human look would still be worthwhile to confirm the operational details before merge.
What was reviewed:
- Launcher branch in launch_gb300-nv.sh: check_env_vars now covers both SLURM_PARTITION and SLURM_ACCOUNT, and the new values are set via if/else rather than
:-fallback. - GB300_SHARED_ROOT is "/data/home/slurm-shared/gharunners" on the new restricted route vs "sa-shared" used everywhere else in the repo — plausible if the restricted partition has its own mount, but not verifiable from source alone.
- perf-changelog.yaml entry is appended at the file's tail as required; its pr-link is a "pull/0" placeholder rather than the documented pre-PR "TBD" (docs/configuration-procedures.md:636).
- The new agg/disagg recipe files are paired with matching master-config entries as required for multi-node srt-slurm additions.
Extended reasoning...
The diff adds a GB300 multi-node AgentX benchmark route (two srt-slurm recipe files, master-config wiring, a changelog entry, and new conditional branching in the shared launch_gb300-nv.sh script that selects a different SLURM partition/account and a new shared-storage root for this model/framework combination). It touches no auth, crypto, or data-exposure surface; the risk is purely operational (job routing and shared-cache paths for a NVIDIA benchmark CI runner). The bug hunter reported no findings, but the launcher change introduces production-affecting branching logic and an unverified shared-root path divergence from repo convention, plus a changelog pr-link that deviates from the documented pre-merge placeholder convention, which together are enough that a human familiar with the cluster layout should confirm before merge.
This review covers commit 74b1996, which is no longer the latest commit on this pull request; later commits are not covered by it.
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Port the nine exercised aggregated and disaggregated AgentX points to the pluggable launcher. Add a named restricted GB300 Slurm route for the alternate partition and shared storage.\n\n将九个已验证的聚合与分离式 AgentX 点迁移到可插拔启动器,并为备用分区和共享存储添加命名的 GB300 restricted Slurm 路由。
6ff9e47 to
7ababb3
Compare
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36658304833 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36658304833 |
Cover every vertex of the measured Dynamo+SGLang Pareto curve in the two AgentX recipe files: add the two-GPU TP2/EP2 c1 arm, the 1P2D c48 and spread c16 cells, and the HiCache prefill tier at 1P1D c160 and 2P1D c256; move the multi-decode cells (1P2D c48/c64, 1P4D c64, 2P1D c256) to the 1 s session-affinity TTL they were measured and GSM8K-gated with; refresh the measured numbers in the file headers. 在两个 AgentX 配方文件中覆盖已测得的 Dynamo+SGLang Pareto 曲线的每个顶点:新增双 GPU 的 TP2/EP2 c1 配置、1P2D c48 与跨节点 c16 单元,以及 1P1D c160 与 2P1D c256 的 HiCache prefill 层;将多 decode 单元(1P2D c48/c64、1P4D c64、2P1D c256)改为其实际测量与 GSM8K 验证所用的 1 秒会话亲和 TTL;同时更新文件头中的测量数据。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
Trim the aggregated and disaggregated AgentX recipes and the nvidia-master sweep to the eight measured Pareto vertices: agg TP4/EP1 c1 and TP2/EP2 c1; disagg 1P2D spread c16, 1P2D c48, 1P2D c64, 1P1D c96, 1P1D HiCache c160 and 2P1D HiCache c256. Drop the dominated aggregated c64 (prefill-decode-interval 8/16), 1P1D c8/c16/c64 and 1P4D c64 arms, add master entries for the new overrides, and update the perf-changelog description to the shipped set. 将聚合与分离式 AgentX 配方以及 nvidia-master sweep 精简为实测 Pareto 曲线的八个顶点:聚合 TP4/EP1 c1 与 TP2/EP2 c1;分离式 1P2D 跨节点 c16、1P2D c48、1P2D c64、1P1D c96、1P1D HiCache c160 与 2P1D HiCache c256。移除被支配的聚合 c64(prefill-decode-interval 8/16)、1P1D c8/c16/c64 与 1P4D c64 配置,为新增 override 添加 master 条目,并将 perf-changelog 描述更新为最终交付集合。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
Re-append the perf-changelog entry after main's newer entries so the changelog stays append-only. 将 upstream/main 合并进 pohanh/dsv41flash-gb300-agentx-curve,并把 perf-changelog 条目重新追加到 main 的新条目之后,保持只增不删。 Co-Authored-By: Claude Fable 5.1 <[email protected]>
Description
Add DeepSeek-V4.1-Flash FP4 AgentX recipes for GB300 with Dynamo + SGLang:
restrictedGB300 Slurm route for this workload'sbatch_2partition and shared caches, integrated with the pluggable Python launcher.The stack is digest-pinned
lmsysorg/sglang:nightly-dev-20260928-81f27fb3with ai-dynamo1.6.0.dev20260928, Dynamo KV routing with session affinity, DSpark block size 5, and Mooncake KV transfer for disaggregated points.Validation
configs/runners.yaml, including the selectedbatch_2/restrictedroute and alternate cache/squash paths;1/1/1/2/2/2/2/2/3;infx.workflows.validate_perf_changelogagainst the latestmain;AI model disclosure
main, resolved the launcher-refactor conflict, added and validated the named Slurm route, updated documentation, prepared this description, and monitored CI. Final configuration selection and submission are by Po-Han Huang.Type of Change
Checklist
inferencex-e2e/perf-changelog.yaml中文
说明
新增 GB300 上 DeepSeek-V4.1-Flash FP4 的 Dynamo + SGLang AgentX 配方:
batch_2分区和共享缓存添加命名的 GB300restrictedSlurm 路由,并接入可插拔 Python 启动器。软件栈固定为摘要锁定的
lmsysorg/sglang:nightly-dev-20260928-81f27fb3、ai-dynamo1.6.0.dev20260928、带会话亲和性的 Dynamo KV 路由、DSpark block size 5;分离式点使用 Mooncake KV 传输。验证
configs/runners.yaml,确认选择batch_2/restricted路由以及备用缓存/squash 路径;1/1/1/2/2/2/2/2/3;infx.workflows.validate_perf_changelog在最新main上通过;AI 模型披露
main,解决启动器重构冲突,添加并验证命名 Slurm 路由,更新文档,准备本说明并监控 CI。最终配置选择和提交由 Po-Han Huang 完成。