Tune MiniMax-M3 AgentX on B300 - #3547
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
50fe6d1 to
b89a085
Compare
|
Claude finished @xinli-sw's task in 1m 4s —— View job PR Review
LGTM - no blocking issues found. What I checked
|
There was a problem hiding this comment.
I reviewed this PR and didn't find any bugs. Because it raises gpu-memory-utilization alongside a vLLM image bump, changes DRAM-offload byte budgets and dram-utilization, and touches the CODEOWNER-restricted nvidia-master.yaml, a human look would still be worthwhile.
What was reviewed:
- Recipe/master-config pairing and image-tag consistency, and perf-changelog append-only compliance — both hold.
- gpu-memory-utilization raised to 0.95 together with the image bump — matches a documented CUDA-graph OOM-risk pattern in KLAUD_DEBUG.md; the PR's own sweep/eval validation is still pending.
- EAGLE3 speculative-config backend switch to FLASHINFER plus use_local_argmax_reduction — reads as a kernel/backend change, not draft-precision lowering.
- New DRAM-offload concurrency points and the tp4 dram-utilization 0.683→1.0 change — host-memory headroom under the new cap isn't yet demonstrated by a passing sweep.
Extended reasoning...
The diff tunes a single-node B300 vLLM AgentX recipe for MiniMax-M3-NVFP4 and its paired nvidia-master.yaml entry, plus an appended perf-changelog entry; it touches no auth, crypto, or data-exposure surface. nvidia-master.yaml is CODEOWNER-restricted to specific NVIDIA reviewers, and the change pairs a gpu-memory-utilization increase with an image bump — a combination KLAUD_DEBUG.md documents as an OOM risk — while the PR description itself says GPU throughput/eval results are pending and one concurrency-64 run already timed out. Those two facts (CODEOWNER-owned path plus incomplete validation on a risk-flagged tuning combination) are why a human look is still worthwhile even though no concrete bug was found.
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36567581989 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36567581989 |
292a655 to
89e79ed
Compare
1883af9 to
91138e4
Compare
更新 B300 MiniMax-M3 AgentX 的 vLLM 镜像,并为 EAGLE3 草稿启用本地 argmax 归约。
将 B300 AgentX 的每批 token 上限恢复为先前使用的 16384,以单独评估其他配置改动。
将 B300 EAGLE3 草稿注意力切换到支持 FP8 KV 与融合多步解码的 Triton 后端,目标模型注意力和 KV 精度保持不变。
恢复 B300 TP2 主机 KV 缓存预算,并将 EAGLE3 草稿注意力切回 FlashInfer。
91138e4 to
dbbc1cb
Compare
将 B300 AgentX 的最大序列数调整为并发的 1.5 倍,并使用显式 CUDA 图捕获尺寸;精简性能变更日志。
合并最新主分支并保留单条 B300 性能变更日志。
|
/use 36567581989 |
kedarpotdar-nv
left a comment
There was a problem hiding this comment.
As a PR reviewer and CODEOWNER, I have reviewed this and have:
- Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
- Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
- Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
- Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
- Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
- Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison. - For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
- Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
- Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
- If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
- If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
- Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
- I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
- Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/
<PR_NUMBER>.md— named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section. - If this PR uses
append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it. - If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
- Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.
Additional detail section:
- Validation and evals: Run Sweep 36567581989, attempt 2 completed successfully on current head
1013a47ec2811a6e9cd40c9dfb33d485f58533fb; all 13 agentic eval jobs passed with scores 0.97–0.98. Authorized reuse command:/use 36567581989. - Upstream recipe: vLLM recipes PR #1047 is merged, and the published MiniMax-M3 recipe documents the NVFP4 target and the exact GQA EAGLE3 combination with
FLASHINFERdraft attention anduse_local_argmax_reduction: trueused here. The target FlashInfer/TRT-LLM attention path and FP8 target/indexer KV cache are also documented. - Chat/AL: the AgentX client uses chat completions with
thinking_mode: enabled. Throughput runs use vLLM synthetic rejection with acceptance length2.78, matching the committed MiniMax-M3 EAGLE3-GQA thinking-on K=3 golden curve; evals use real verification. - Draft precision:
Inferact/MiniMax-M3-EAGLE3-GQArevision96692486b5fd38ebf8fd2a5f6bb53427d30819a8ships BF16 tensors and is loaded at stored BF16 by the pinned upstreamvllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92bdefault path. There is no draft quantization, dtype override, checkpoint substitution, or engine patch. FP8 KV cache is applied consistently to target and draft and does not lower draft weights/activations.SGLANG_NVFP4_CKPT_FP8_NEXTN_MOEis not present or enabled. - No architecture-FLOP hacks or serving-stack patches are introduced.
append-only: trueis not used. - Pareto coverage: the B300 MiniMax-M3 AgentX P90 curve from run 36567581989 has 13 valid measured points and 5 throughput/E2EL frontier points (5/5 recommended minimum).
Signed: kedarpotdar-nv
Description
vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92bwith FP8 KV, FlashInfer target and draft attention, and local argmax reduction.max-num-seqstoceil(1.5 × concurrency)and useFULL_AND_PIECEWISECUDA graphs at 4 × sequence counts 2 throughmax-num-seqs, plus 256, 512, 1024, and 2048.Validation
validate_recipe, andgit diff --checkpassed locally. The new sweep is running for the graph and sequence changes.AI model disclosure
Related Issue
None.
Type of Change
Checklist
perf-changelog.yamland have not edited historical entriesOWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.中文
将 MiniMax-M3 B300 AgentX 更新至
vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b,保持 FP8 KV、目标与草稿模型的 FlashInfer 注意力,并启用本地 argmax 归约。恢复 TP2 主机 KV 缓存预算。将max-num-seqs设为并发数的 1.5 倍并向上取整;使用FULL_AND_PIECEWISECUDA 图,捕获尺寸为序列数 2 至max-num-seqs的四倍,以及 256、512、1024、2048。性能变更日志仅保留一条简短记录。本 PR 仅修改 B300,未触及文档或 CI 流程。之前的完整扫描中吞吐量和准确率任务均通过,但相较已发布 B300 基线仍存在延迟回退:TP4 C20 的中位首 token 延迟为 0.547 秒,基线为 0.397 秒,吞吐量基本相同。本地已通过 YAML 解析、13 个变体配置检查、仓库的
validate_recipe及git diff --check。新扫描正在运行,以验证图捕获和序列数调整。AI 模型:GPT-6;运行环境未提供可核实的完整版本标识。用于配置修改、源码审查、本地验证及准备 PR。无关联 issue。改动类型:错误修复和配置更新。