-
Notifications
You must be signed in to change notification settings - Fork 313
[Klaud Cold] Migrate all SPEED-Bench AL collectors to native srt-slurm / 将全部 SPEED-Bench AL 收集器迁移到原生 srt-slurm #3476
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
functionstackx
wants to merge
14
commits into
main
Choose a base branch
from
klaud/speedbench-srt-slurm
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
14 commits
Select commit
Hold shift + click to select a range
b81a9e6
Migrate 6 SPEED-Bench AL collectors to native srt-slurm single-node path
functionstackx 5aea3af
Fix SPEED-Bench lint and bind qwen3.8next TP4
functionstackx a43d14d
Store SPEED-Bench thinking cells as thinking_on/thinking_off
functionstackx 2bad6db
Bind SpeedBench GPUs via CUDA_VISIBLE_DEVICES
functionstackx 13afd0d
Fix srtctl validation: move THINKING/MTP from zip env lists to runtim…
functionstackx 50f9a78
Keep runtime THINKING a string through srtctl --set
functionstackx f1e8d81
Tidy SpeedBench test imports and unused variable
functionstackx df1dff9
Fix SpeedBench client workspace fallback depth
functionstackx ff8e6be
Merge remote-tracking branch 'origin/main' into klaud/speedbench-srt-…
functionstackx f481e3c
feat(speedbench): migrate remaining AL collectors to srt-slurm recipes
functionstackx 079b083
fix(tests): match speedbench recipe-render test images to actual recipes
functionstackx b3e8537
feat(speedbench): migrate remaining AL collectors to srt-slurm recipe…
functionstackx d73c4d9
Merge remote-tracking branch 'origin/klaud/speedbench-srt-slurm' into…
functionstackx 9fc4be5
fix(speedbench): fix porting regressions in 5 AL collector recipes / …
functionstackx File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
236 changes: 0 additions & 236 deletions
236
benchmarks/single_node/speedbench/dsv4_fp4_b300_vllm.sh
This file was deleted.
Oops, something went wrong.
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🔴 Running the migrated srt-slurm SPEED-Bench collector for qwen3.8next now always fails every cell, where the old script worked. The workflow sets TP='8' (line 104) and exports GPU_COUNT="$TP" (line 232) for every model prefix, but the qwen3.8next recipe (benchmarks/single_node/srt-slurm-recipes/qwen3.8next/vllm/b300-fp4-speedbench/speedbench.yaml:46,49) uses gpus: 4 / tensor-parallel-size: 4. validate_recipe in infx/srt_slurm/single_node.py (gpus check line 112, tensor x data parallel check line 51) rejects the 4 vs 8 mismatch, so select_recipe raises and every cell records N/A. Fix: give the workflow a per-prefix TP/GPU_COUNT override (as the deleted qwen3.8next_fp4_b300_vllm.sh did with its local 'TP=4' fold) so qwen3.8next resolves to TP=4 again.
Why this was flagged
Dispatching speedbench-al.yml with model-prefix=qwen3.8next runs 'export GPU_COUNT="$TP"' where TP is hardcoded '8' at .github/workflows/speedbench-al.yml:104, then loops all cells via runners/launch_b300-dsxe.sh -> launch_srt_single_node -> infx/srt_slurm/single_node.py prepare. validate_recipe compares role["gpus"]=4 and tensor-parallel-size*data_parallel=4 (qwen3.8next speedbench.yaml:38,41) against GPU_COUNT/TP=8, both checks fail, select_recipe raises 'Expected exactly one matching...'. Every cell then fails at the prepare step (no result JSON written), ALL_FAILED stays true, and the job exits 1 with 'every SPEED-Bench cell failed'. The deleted benchmarks/single_node/speedbench/qwen3.8next_fp4_b300_vllm.sh explicitly worked around this same exported TP=8 by locally reassigning 'TP=4' (its own comment: 'speedbench-al.yml exports TP=8 unconditionally; fold ... into recipe-validated TP=4'); that workaround has no equivalent in the new srt-slurm path, so the migration silently drops it for this one model.
Verification: normal. The migrated srt-slurm path always fails every cell for model-prefix=qwen3.8next, where the base branch worked. Chain: speedbench-al.yml hardcodes
TP: '8'(line 104, an env constant, not a dispatch input — the inputs block lines 12-71 has no TP override) and at line 232export GPU_COUNT="$TP"= 8 for ALL prefixes. The two sourced runtime_settings.sh files do not touch TP/GPU_COUNT…