Conversation
593f473 to
1d542c7
Compare
…cipe
Replace the hand-rolled bash server script
(dsv4_fp4_b200_vllm_cache_sources_mtp.sh) with a native srt-slurm
recipe at dsv4/vllm/b200-fp4-mtp/cache-sources.yaml. Three variants
select by CONC/KV_OFFLOADING: TP8 c8 native DRAM, TP8 c14 Simple NVMe,
TP8 c14 native DRAM+NVMe. NVMe variants use host_setup to
create/teardown /scratch/inferencex-kv-$SLURM_JOB_ID and
container_mounts with {job_id} templating.
Each search-space row in nvidia-master.yaml now carries srt-recipe:,
routing to the native-single-node launch path in the b200-nscale
launcher. The bash-specific routing, /ix mount override, NVMe directory
management, and GPU_MEMORY_UTILIZATION=0.85 export are removed from
launch_b200-nscale-slurm.sh (gpu-memory-utilization is now 0.85 in the
recipe).
A new srt-slurm patch (pr56318-overlay-validation.patch) adds
pr56318-overlay-check.sh to configs/patches/, referenced by the
recipe's setup_script to verify overlay checksums before engine start.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36272389390 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36272389390 |
|
/use 36272389390 |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
Validation
Rerun vLLM #56318 at
bd57138f9b98eb64c44ea6a0f10d08c8f639820eon DeepSeek-V4-Pro with hybrid KV-cache management enabled.Reuses the settings from successful B200 and GB300 sweeps in isolated validation recipes. Existing benchmark curves are unchanged. Adds the single-node NVMe/tier-list schema and job-owned NVMe mount needed by these points.
Local checks
Official full sweep is running. The B200 native DRAM canary is profiling. The other three points remain gated on the canary. CI passed.
Single-node launch migrated to srt-slurm (56136a2)
benchmarks/single_node/agentic/dsv4_fp4_b200_vllm_cache_sources_mtp.shwith the schema-2 recipebenchmarks/single_node/srt-slurm-recipes/dsv4/vllm/b200-fp4-mtp/cache-sources.yaml(variantsoverride_tp8_c8_dram_native,override_tp8_c14_nvme_simple,override_tp8_c14_dramnvme_native); the threenvidia-master.yamlrows now carrysrt-recipe.runners/launch_b200-nscale-slurm.sh; NVMe scratch setup/teardown and the/kv-offloadmount move into the recipe, and the PR56318 overlay sha256 check runs as an srt-slurm setup script (runners/srt-slurm/patches/pr56318-overlay-validation.patch).metrics-before.prom/metrics-after.promsnapshots (srtctl owns the server lifecycle); the required-server-metric check viaAIPERF_REQUIRED_SERVER_METRIC_PREFIXis kept.