Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
73 changes: 68 additions & 5 deletions .github/workflows/speedbench-al.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ on: # zizmor: ignore[concurrency-limits]
type: string
default: 'deepseek-ai/DeepSeek-V4-Pro'
model-prefix:
description: "Model prefix; drives launcher MODEL_PATH resolution, exp name, collector script, and artifact names"
description: "Model prefix; drives srt-slurm recipe lookup, exp name, and artifact names"
required: false
type: string
default: 'dsv4'
Expand Down Expand Up @@ -105,8 +105,6 @@ env:
EP_SIZE: '1'
DP_ATTENTION: 'false'
SPEC_DECODING: mtp
# Run the AL-matrix collector instead of the auto-selected throughput script.
BENCH_SCRIPT_OVERRIDE: benchmarks/single_node/speedbench/${{ inputs.model-prefix }}_fp4_b300_vllm.sh
SALLOC_TIME_LIMIT: ${{ inputs.salloc-time }}
# Matrix-collector tunables (propagated into the container via srun --export=ALL).
MTP_LIST: ${{ inputs.mtp-list }}
Comment on lines 105 to 110

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Running the migrated srt-slurm SPEED-Bench collector for qwen3.8next now always fails every cell, where the old script worked. The workflow sets TP='8' (line 104) and exports GPU_COUNT="$TP" (line 232) for every model prefix, but the qwen3.8next recipe (benchmarks/single_node/srt-slurm-recipes/qwen3.8next/vllm/b300-fp4-speedbench/speedbench.yaml:46,49) uses gpus: 4 / tensor-parallel-size: 4. validate_recipe in infx/srt_slurm/single_node.py (gpus check line 112, tensor x data parallel check line 51) rejects the 4 vs 8 mismatch, so select_recipe raises and every cell records N/A. Fix: give the workflow a per-prefix TP/GPU_COUNT override (as the deleted qwen3.8next_fp4_b300_vllm.sh did with its local 'TP=4' fold) so qwen3.8next resolves to TP=4 again.

Why this was flagged

Dispatching speedbench-al.yml with model-prefix=qwen3.8next runs 'export GPU_COUNT="$TP"' where TP is hardcoded '8' at .github/workflows/speedbench-al.yml:104, then loops all cells via runners/launch_b300-dsxe.sh -> launch_srt_single_node -> infx/srt_slurm/single_node.py prepare. validate_recipe compares role["gpus"]=4 and tensor-parallel-size*data_parallel=4 (qwen3.8next speedbench.yaml:38,41) against GPU_COUNT/TP=8, both checks fail, select_recipe raises 'Expected exactly one matching...'. Every cell then fails at the prepare step (no result JSON written), ALL_FAILED stays true, and the job exits 1 with 'every SPEED-Bench cell failed'. The deleted benchmarks/single_node/speedbench/qwen3.8next_fp4_b300_vllm.sh explicitly worked around this same exported TP=8 by locally reassigning 'TP=4' (its own comment: 'speedbench-al.yml exports TP=8 unconditionally; fold ... into recipe-validated TP=4'); that workaround has no equivalent in the new srt-slurm path, so the migration silently drops it for this one model.

Verification: normal. The migrated srt-slurm path always fails every cell for model-prefix=qwen3.8next, where the base branch worked. Chain: speedbench-al.yml hardcodes TP: '8' (line 104, an env constant, not a dispatch input — the inputs block lines 12-71 has no TP override) and at line 232 export GPU_COUNT="$TP" = 8 for ALL prefixes. The two sourced runtime_settings.sh files do not touch TP/GPU_COUNT…

Expand Down Expand Up @@ -239,7 +237,70 @@ jobs:
if [[ -f runners/runtime_settings.sh ]]; then
source runners/runtime_settings.sh
fi
bash ./runners/launch_${RUNNER_NAME%%_*}.sh

# All collectors use the srt-slurm single-node path with per-model
# recipes under srt-slurm-recipes/<prefix>/vllm/b300-fp4-speedbench/.
RECIPE_PATH="benchmarks/single_node/srt-slurm-recipes/${MODEL_PREFIX}/vllm/b300-fp4-speedbench/speedbench.yaml"
if [[ ! -f "$RECIPE_PATH" ]]; then
echo "ERROR: SPEED-Bench recipe not found: $RECIPE_PATH" >&2
exit 1
fi
# Plain TP8 is incompatible with the Qwen3.8-Flash-Next FP8 checkpoint's
# 128-wide quantization blocks; its recipe runs TP4 (as the legacy
# collector did), so bind the matrix TP and GPU count to match.
if [[ "$MODEL_PREFIX" == qwen3.8next ]]; then
export TP=4 GPU_COUNT=4
fi
mkdir -p speedbench_results
ALL_FAILED=true
# Harmless values for launch_srt_single_node env checks; the collector
# recipe and client do not read them.
export ISL=256
export OSL=256
export RANDOM_RANGE_RATIO=0.0
export CONC=1
export PP_SIZE=1
export DCP_SIZE=1
export PCP_SIZE=1
for mode in $THINKING_MODES; do
for mtp in $MTP_LIST; do
IDX=$((mtp - 1))
export THINKING="$mode"
export MTP="$mtp"
export SRT_RECIPE="${RECIPE_PATH}:zip_override_mtp[${IDX}]"
export RESULT_FILENAME="speedbench_${mode}_mtp${mtp}"
echo ""
echo "=========================================="
echo " srt-slurm cell: thinking=${mode} MTP=${mtp}"
echo " SRT_RECIPE=${SRT_RECIPE}"
echo "=========================================="
CELL_RC=0
bash ./runners/launch_"${RUNNER_NAME%%_*}".sh || CELL_RC=$?
# Collect per-cell artifacts into the results directory.
if [[ -f "${RESULT_FILENAME}.json" ]]; then
mv "${RESULT_FILENAME}.json" "speedbench_results/${RESULT_FILENAME}.json"
ALL_FAILED=false
fi
if [[ -f srt-single-node-logs.tar.gz ]]; then
mv srt-single-node-logs.tar.gz "speedbench_results/srt-logs_${mode}_mtp${mtp}.tar.gz"
fi
if [[ "$CELL_RC" -ne 0 ]]; then
echo " -> cell failed (rc=$CELL_RC), recording N/A"
fi
done
done
if [[ "$ALL_FAILED" == true ]]; then
echo "ERROR: every SPEED-Bench cell failed" >&2
exit 1
fi
# Aggregate per-cell result JSONs into the reference YAML.
PYTHONPATH="${GITHUB_WORKSPACE}${PYTHONPATH:+:$PYTHONPATH}" \
python3 -m infx.workflows.speedbench_matrix \
--result-dir speedbench_results \
--model-key "$(basename "$(echo "$MODEL" | tr '[:upper:]' '[:lower:]')")" \
--thinking-modes "$THINKING_MODES" \
--mtp-list "$MTP_LIST" \
> speedbench-reference-al.yaml

if [ ! -f "speedbench-reference-al.yaml" ]; then
echo "AL collection failed: speedbench-reference-al.yaml not produced." >&2
Expand Down Expand Up @@ -292,7 +353,9 @@ jobs:
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
with:
name: speedbench_server_logs-${{ inputs.model-prefix }}
path: speedbench_results/server_*.log
path: |
speedbench_results/server_*.log
speedbench_results/srt-logs_*.tar.gz
if-no-files-found: ignore

# Per-request benchmark detail (vllm bench serve --save-detailed): includes
Expand Down
236 changes: 0 additions & 236 deletions benchmarks/single_node/speedbench/dsv4_fp4_b300_vllm.sh

This file was deleted.

Loading
Loading