Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ base:
name: qwen3.5-fp4-b200-sglang-mtp-8k1k
model:
path: hf:nvidia/Qwen3.5-397B-A17B-NVFP4-V2
container: lmsysorg/sglang:v0.5.14-cu130
container: lmsysorg/sglang:v0.5.20-cu130
precision: fp4
resources:
gpu_type: b200
Expand Down Expand Up @@ -33,7 +33,7 @@ base:
mamba-ssm-dtype: bfloat16
attention-backend: trtllm_mha
moe-runner-backend: flashinfer_trtllm
cuda-graph-max-bs: 4
cuda-graph-max-bs-decode: 4
max-running-requests: 4
max-prefill-tokens: 16384
chunked-prefill-size: 16384
Expand Down Expand Up @@ -70,7 +70,7 @@ zip_override_tp2_ep1:
agg:
gpus: 2
args:
cuda-graph-max-bs: [4, 8, 16, 32, 64]
cuda-graph-max-bs-decode: [4, 8, 16, 32, 64]
max-running-requests: [4, 8, 16, 32, 64]
scheduler-recv-interval: [10, 30, 30, 30, 30]
tensor-parallel-size: 2
Expand All @@ -82,7 +82,7 @@ zip_override_tp2_ep2:
agg:
gpus: 2
args:
cuda-graph-max-bs: [16, 32, 64]
cuda-graph-max-bs-decode: [16, 32, 64]
expert-parallel-size: 2
linear-attn-backend: triton
linear-attn-decode-backend: flashinfer
Expand Down
2 changes: 1 addition & 1 deletion configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1218,7 +1218,7 @@ qwen3.5-fp4-b200-sglang:
srt-recipe: benchmarks/single_node/srt-slurm-recipes/qwen3.5/sglang/b200-fp4/8k1k.yaml

qwen3.5-fp4-b200-sglang-mtp:
image: lmsysorg/sglang:v0.5.14-cu130
image: lmsysorg/sglang:v0.5.20-cu130
model: nvidia/Qwen3.5-397B-A17B-NVFP4-V2
model-prefix: qwen3.5
runner: cluster:b200-nscale
Expand Down
6 changes: 6 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8839,3 +8839,9 @@
- "Retain 8 SWA prefix tails per concurrency at C1/C2 after rejecting the 128-tail floor in matched canonical testing; use 32 per concurrency at C4 and above, reaching 640 at C20 within the static cache budget."
- "Qualify supported DP8 attention and DP LM-head at C4/C8/C16/C20 with consistent-hash session routing and 64 SWA tails per rank; retain the complete plain TP8 curve and unchanged shipped precision, static budget and native context."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3345

- config-keys:
- qwen3.5-fp4-b200-sglang-mtp
description:
- "Update SGLang image from v0.5.14-cu130 to v0.5.20-cu130 and rename cuda-graph-max-bs to cuda-graph-max-bs-decode."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3411
Loading