Skip to content

Tune MiniMax-M3 AgentX on B300 - #3547

Merged
Oseltamivir merged 7 commits into
mainfrom
codex/minimaxm3-b300-agentx-split
Sep 30, 2026
Merged

Oseltamivir merged 7 commits into
mainfrom
codex/minimaxm3-b300-agentx-split

Conversation

@xinli-sw

@xinli-sw xinli-sw commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Description

  • Update MiniMax-M3 B300 AgentX to vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b with FP8 KV, FlashInfer target and draft attention, and local argmax reduction.
  • Restore the TP2 host KV cache budget. Set max-num-seqs to ceil(1.5 × concurrency) and use FULL_AND_PIECEWISE CUDA graphs at 4 × sequence counts 2 through max-num-seqs, plus 256, 512, 1024, and 2048.
  • Keep one concise performance changelog entry. This PR changes B300 only; documentation and CI procedures are untouched.

Validation

  • The previous full sweep passed throughput and accuracy jobs. It still showed higher latency than the published B300 baseline: TP4 C20 median TTFT was 0.547s versus 0.397s at nearly equal throughput.
  • YAML parsing, all 13 variant settings, repository validate_recipe, and git diff --check passed locally. The new sweep is running for the graph and sequence changes.

AI model disclosure

  • Model/version: GPT-6; the exact runtime version identifier was not exposed.
  • Role: Configuration changes, source review, local validation, and PR preparation.

Related Issue

None.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary (none needed for this configuration change)
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.
中文

将 MiniMax-M3 B300 AgentX 更新至 vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b,保持 FP8 KV、目标与草稿模型的 FlashInfer 注意力,并启用本地 argmax 归约。恢复 TP2 主机 KV 缓存预算。将 max-num-seqs 设为并发数的 1.5 倍并向上取整;使用 FULL_AND_PIECEWISE CUDA 图,捕获尺寸为序列数 2 至 max-num-seqs 的四倍,以及 256、512、1024、2048。性能变更日志仅保留一条简短记录。本 PR 仅修改 B300,未触及文档或 CI 流程。

之前的完整扫描中吞吐量和准确率任务均通过,但相较已发布 B300 基线仍存在延迟回退:TP4 C20 的中位首 token 延迟为 0.547 秒,基线为 0.397 秒,吞吐量基本相同。本地已通过 YAML 解析、13 个变体配置检查、仓库的 validate_recipe 及 git diff --check。新扫描正在运行,以验证图捕获和序列数调整。

AI 模型:GPT-6;运行环境未提供可核实的完整版本标识。用于配置修改、源码审查、本地验证及准备 PR。无关联 issue。改动类型:错误修复和配置更新。

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b300-agentx-split branch from 50fe6d1 to b89a085 Compare September 28, 2026 19:15
@xinli-sw
xinli-sw marked this pull request as ready for review September 28, 2026 19:15
@xinli-sw
xinli-sw requested a review from a team September 28, 2026 19:15
@xinli-sw xinli-sw added the full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures label Sep 28, 2026
@claude

claude Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @xinli-sw's task in 1m 4s —— View job


PR Review

  • Gather context (diff, CLAUDE.md/AGENTS.md)
  • Review recipe, master config, and perf changelog
  • Post review

LGTM - no blocking issues found.

What I checked
  • Perf changelog: The new entry is at the end of the file and not marked append-only. It covers minimaxm3-fp4-b300-vllm-agentic-mtp, and the PR link is correct.
  • Image: vllm/vllm-openai:nightly-af7f9488… is on Docker Hub, and the same image is set in both the recipe and the master config. validate_recipe requires them to match.
  • DRAM budget: agentic_dram_offload_gb limits node DRAM to 2,861,022 MiB (about 3000 GB). At dram-utilization: 1.0, a TP4 point gets 4/8 of that (1499 GB) and a TP2 point gets 2/8 (749 GB). These match the recipe's TOTAL_CPU_DRAM_GB, which validate_recipe checks against the matrix. The cpu_bytes_to_use values (1499e9 and 749e9) agree with the new "decimal GB" comment.
  • Variant selection: Each of the 13 matrix points (TP, concurrency and offload mode) matches exactly one override. TP4 c40 and TP2 c40 are told apart by gpus.
  • CUDA graph sizes: Every override sets max-cudagraph-capture-size = max-num-seqs × 4, which is the 3 speculative tokens plus 1.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it raises gpu-memory-utilization alongside a vLLM image bump, changes DRAM-offload byte budgets and dram-utilization, and touches the CODEOWNER-restricted nvidia-master.yaml, a human look would still be worthwhile.

What was reviewed:

  • Recipe/master-config pairing and image-tag consistency, and perf-changelog append-only compliance — both hold.
  • gpu-memory-utilization raised to 0.95 together with the image bump — matches a documented CUDA-graph OOM-risk pattern in KLAUD_DEBUG.md; the PR's own sweep/eval validation is still pending.
  • EAGLE3 speculative-config backend switch to FLASHINFER plus use_local_argmax_reduction — reads as a kernel/backend change, not draft-precision lowering.
  • New DRAM-offload concurrency points and the tp4 dram-utilization 0.683→1.0 change — host-memory headroom under the new cap isn't yet demonstrated by a passing sweep.
Extended reasoning...

The diff tunes a single-node B300 vLLM AgentX recipe for MiniMax-M3-NVFP4 and its paired nvidia-master.yaml entry, plus an appended perf-changelog entry; it touches no auth, crypto, or data-exposure surface. nvidia-master.yaml is CODEOWNER-restricted to specific NVIDIA reviewers, and the change pairs a gpu-memory-utilization increase with an image bump — a combination KLAUD_DEBUG.md documents as an OOM risk — while the PR description itself says GPU throughput/eval results are pending and one concurrency-64 run already timed out. Those two facts (CODEOWNER-owned path plus incomplete validation on a risk-flagged tuning combination) are why a human look is still worthwhile even though no concrete bug was found.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b300-agentx-split branch from 292a655 to 89e79ed Compare September 29, 2026 00:22
@xinli-sw xinli-sw changed the title Tune MiniMax-M3 AgentX on B300 with updated vLLM / 更新 B300 上的 MiniMax-M3 AgentX 配置 Tune MiniMax-M3 AgentX on B300 with FP8 Triton draft / 使用 FP8 Triton 草稿优化 B300 上的 MiniMax-M3 AgentX Sep 29, 2026
@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b300-agentx-split branch from 1883af9 to 91138e4 Compare September 29, 2026 01:31
@xinli-sw xinli-sw changed the title Tune MiniMax-M3 AgentX on B300 with FP8 Triton draft / 使用 FP8 Triton 草稿优化 B300 上的 MiniMax-M3 AgentX Tune MiniMax-M3 AgentX on B300 Sep 29, 2026
更新 B300 MiniMax-M3 AgentX 的 vLLM 镜像,并为 EAGLE3 草稿启用本地 argmax 归约。
将 B300 AgentX 的每批 token 上限恢复为先前使用的 16384,以单独评估其他配置改动。
将 B300 EAGLE3 草稿注意力切换到支持 FP8 KV 与融合多步解码的 Triton 后端,目标模型注意力和 KV 精度保持不变。
恢复 B300 TP2 主机 KV 缓存预算,并将 EAGLE3 草稿注意力切回 FlashInfer。
@xinli-sw
xinli-sw force-pushed the codex/minimaxm3-b300-agentx-split branch from 91138e4 to dbbc1cb Compare September 29, 2026 05:18
将 B300 AgentX 的最大序列数调整为并发的 1.5 倍,并使用显式 CUDA 图捕获尺寸;精简性能变更日志。
合并最新主分支并保留单条 B300 性能变更日志。
@xinli-sw

Copy link
Copy Markdown
Collaborator Author

/use 36567581989

@kedarpotdar-nv kedarpotdar-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that every draft model and draft head is served as it ships: the draft that ships with the served checkpoint, at its stored precision, through the pinned upstream image's default handling, with the shipped and effective draft precision recorded in the additional detail section. No submission-side quantization, dtype override, checkpoint substitution, or patch may lower draft precision below that default, regardless of eval results or AL. Explicitly verified that SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not enabled in the effective recipe, including inherited settings; enabling it is prohibited going forward, and historical runs do not grant an exception. See Draft-model precision for what counts as the default and the MLPerf comparison.
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in infx/golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; target/verifier FLOPs at lower precisions is fine, given that the config passes private evals, but this does not permit lowering draft-model or draft-head precision below what ships. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If this PR uses append-only: true, verified that it only adds generated points or recipe variants inside a selected existing config/scenario and existing same-image visual curve: every previously generated point remains present with the same recipe, no prior point is removed or rerun, and every benchmark-affecting change in the complete diff can affect only the corresponding newly appended points (never an existing point), regardless of which file contains it.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.
  • Reported measured throughput/E2EL Pareto counts and evidence per affected curve (≥5 points strongly recommended). Below 5 or unverifiable: tag a core maintainer for review; recorded admin bypass required before merge. N/A if no curves are affected. Details.

Additional detail section:

  • Validation and evals: Run Sweep 36567581989, attempt 2 completed successfully on current head 1013a47ec2811a6e9cd40c9dfb33d485f58533fb; all 13 agentic eval jobs passed with scores 0.97–0.98. Authorized reuse command: /use 36567581989.
  • Upstream recipe: vLLM recipes PR #1047 is merged, and the published MiniMax-M3 recipe documents the NVFP4 target and the exact GQA EAGLE3 combination with FLASHINFER draft attention and use_local_argmax_reduction: true used here. The target FlashInfer/TRT-LLM attention path and FP8 target/indexer KV cache are also documented.
  • Chat/AL: the AgentX client uses chat completions with thinking_mode: enabled. Throughput runs use vLLM synthetic rejection with acceptance length 2.78, matching the committed MiniMax-M3 EAGLE3-GQA thinking-on K=3 golden curve; evals use real verification.
  • Draft precision: Inferact/MiniMax-M3-EAGLE3-GQA revision 96692486b5fd38ebf8fd2a5f6bb53427d30819a8 ships BF16 tensors and is loaded at stored BF16 by the pinned upstream vllm/vllm-openai:nightly-af7f9488c2210d67e1033ecdc845b087ee7fe92b default path. There is no draft quantization, dtype override, checkpoint substitution, or engine patch. FP8 KV cache is applied consistently to target and draft and does not lower draft weights/activations. SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE is not present or enabled.
  • No architecture-FLOP hacks or serving-stack patches are introduced. append-only: true is not used.
  • Pareto coverage: the B300 MiniMax-M3 AgentX P90 curve from run 36567581989 has 13 valid measured points and 5 throughput/E2EL frontier points (5/5 recommended minimum).

Signed: kedarpotdar-nv

@Oseltamivir Oseltamivir left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@Oseltamivir
Oseltamivir merged commit 1570a54 into main Sep 30, 2026
25 checks passed
@Oseltamivir
Oseltamivir deleted the codex/minimaxm3-b300-agentx-split branch September 30, 2026 01:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

3 participants