Skip to content

test: validate current vLLM cache-source metrics on DeepSeek-V4-Pro - #3491

Draft
cquil11 wants to merge 6 commits into
mainfrom
codex/pr56318-bd57138-agentx
Draft

cquil11 wants to merge 6 commits into
mainfrom
codex/pr56318-bd57138-agentx

Conversation

@cquil11

@cquil11 cquil11 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Validation

Rerun vLLM #56318 at bd57138f9b98eb64c44ea6a0f10d08c8f639820e on DeepSeek-V4-Pro with hybrid KV-cache management enabled.

  • B200 TP8: native DRAM c8, Simple NVMe c14, native DRAM + NVMe c14.
  • GB300: NIXL + Mooncake c256, one DEP4 prefill worker and one DEP16 decode worker. Five nodes, 20 GPUs.
  • AIPerf AgentX, 3,600 seconds per point, MTP3, no evaluations, Python frontend.
  • Docker Hub images pinned by digest. Workers verify all PR-overlay checksums before model startup.
  • Prometheus artifacts include source counters and prefix-cache counters. GB300 scrapes all five physical DP endpoints.

Reuses the settings from successful B200 and GB300 sweeps in isolated validation recipes. Existing benchmark curves are unchanged. Adds the single-node NVMe/tier-list schema and job-owned NVMe mount needed by these points.

Local checks

  • 357 matrix, schema, and offload-gate tests passed.
  • SRT recipe resolves to five nodes and one worker per P/D role. Endpoint check includes all five metrics endpoints without changing logical routing.
  • Changed Python lint, Bash syntax, and diff whitespace checks passed.

Official full sweep is running. The B200 native DRAM canary is profiling. The other three points remain gated on the canary. CI passed.

Single-node launch migrated to srt-slurm (56136a2)

  • Replaced benchmarks/single_node/agentic/dsv4_fp4_b200_vllm_cache_sources_mtp.sh with the schema-2 recipe benchmarks/single_node/srt-slurm-recipes/dsv4/vllm/b200-fp4-mtp/cache-sources.yaml (variants override_tp8_c8_dram_native, override_tp8_c14_nvme_simple, override_tp8_c14_dramnvme_native); the three nvidia-master.yaml rows now carry srt-recipe.
  • Removed the cache-sources special case from runners/launch_b200-nscale-slurm.sh; NVMe scratch setup/teardown and the /kv-offload mount move into the recipe, and the PR56318 overlay sha256 check runs as an srt-slurm setup script (runners/srt-slurm/patches/pr56318-overlay-validation.patch).
  • Not carried over: the raw metrics-before.prom / metrics-after.prom snapshots (srtctl owns the server lifecycle); the required-server-metric check via AIPERF_REQUIRED_SERVER_METRIC_PREFIX is kept.
  • Local: new parametrized recipe-render test; full suite 1865 passed (14 known macOS bash-3 failures); ruff and changelog validation clean. B200 native DRAM runtime startup passed all 19 overlay checks. All five source counters were zero before requests, and the 87-request warmup completed without errors. The 3,600-second profile is ongoing. Full validation is not complete.

@cquil11
cquil11 force-pushed the codex/pr56318-bd57138-agentx branch from 593f473 to 1d542c7 Compare September 26, 2026 20:25
…cipe

Replace the hand-rolled bash server script
(dsv4_fp4_b200_vllm_cache_sources_mtp.sh) with a native srt-slurm
recipe at dsv4/vllm/b200-fp4-mtp/cache-sources.yaml. Three variants
select by CONC/KV_OFFLOADING: TP8 c8 native DRAM, TP8 c14 Simple NVMe,
TP8 c14 native DRAM+NVMe. NVMe variants use host_setup to
create/teardown /scratch/inferencex-kv-$SLURM_JOB_ID and
container_mounts with {job_id} templating.

Each search-space row in nvidia-master.yaml now carries srt-recipe:,
routing to the native-single-node launch path in the b200-nscale
launcher. The bash-specific routing, /ix mount override, NVMe directory
management, and GPU_MEMORY_UTILIZATION=0.85 export are removed from
launch_b200-nscale-slurm.sh (gpu-memory-utilization is now 0.85 in the
recipe).

A new srt-slurm patch (pr56318-overlay-validation.patch) adds
pr56318-overlay-check.sh to configs/patches/, referenced by the
recipe's setup_script to verify overlay checksums before engine start.

Co-Authored-By: Claude Opus 4.6 <[email protected]>
@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

@cquil11

cquil11 commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

/use 36272389390

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants