feat(agentx): run the MiniMax-M3 MI300X LMCache point on the lmcache-server service - #3545
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
amd-master.yaml moved minimaxm3-fp8-mi300x-vllm-agentic-mtp to vllm/vllm-openai-rocm:v0.30.0 in #3361, but the recipe still named v0.29.0, so the single-node adapter rejected every point.
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also looked at the new lmcache-server service binding --host 0.0.0.0 --http-host 0.0.0.0 in the vendored srt-slurm patch (line ~600) — this exposes the RPC/HTTP ports to the whole node network rather than just localhost, but that's consistent with how srt-slurm's existing service kinds bind for multi-worker access within the job's network namespace, so it doesn't look like a new exposure introduced by this PR.
Extended reasoning...
The diff migrates LMCache DRAM cache setup from a per-TP-rank shell-scripted fleet to a vendored srt-slurm lmcache-server service; it touches a shell script, a benchmark YAML config, a large vendored patch, and a changelog, with no direct auth/crypto surface. An inline finding (pre-existing set -euo pipefail nounset violation) is already being posted, so I did not approve; I additionally checked the 0.0.0.0 bind in the new service definition and found it matches the pattern of other srt-slurm service kinds rather than being a novel exposure.
…MCache point The refreshed lmcache-server patch (SemiAnalysisAI/srt-slurm#32 at 181b2e4) keeps a role's connector on a direct vllm serve aggregate worker, so the variant names connector: lmcache-mp instead of hand-writing its kv-transfer-config. The message queue timeout falls back to LMCache's default (300 s) instead of 6000 s.
The role connector is now the lmcache-mp preset written out plus lmcache.mp.mq_timeout 6000, so the point keeps its measured timeout instead of falling back to LMCache's 300 s default.
…slurm v2.30.0 main bumped srt-slurm to v2.30.0 and dropped the 504 patch. The 507 patch is regenerated on v2.30.0 from SemiAnalysisAI/srt-slurm#32 at 181b2e4, and the #3545 perf-changelog entry is rewritten to state the benchmark change: one lmcache-server with 1296 GB L1 instead of 8 per-rank 162 GB servers, image v0.30.0, LMCacheMPConnector at localhost:8750 with mq_timeout 6000, re-measured results.
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36479049525 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36479049525 |
|
/reuse-sweep-run 36479049525 |
5a8ecf3 to
835de73
Compare
|
/stage-results 36479049525 |
…0x-lmcache-server # Conflicts: # inferencex-e2e/perf-changelog.yaml
|
@cquil11 staged run 36479049525: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-09-28~r36479049525 This run remains available across future |
Summary
Moves the MI300X MiniMax-M3 AgentX LMCache variant (
override_tp8_c16_lmcache) to srt-slurm'slmcache-serverservice. srt-slurm owns startup, readiness, and cleanup; the setup script only installs dependencies.--max-workers 2.v0.30.0.Uses the LMCache patch shared with #3543, from srt-slurm #32. Upstream: #507 and #528.
Validation
Throughput was about 14% below the earlier eight-server run (9,970 vs. 11,586 tok/s), with higher p90 TTFT. That baseline used vLLM
v0.29.0, so the comparison does not isolate the effect of changing the server layout.