Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
base:
schema: 2
name: dsv4flash-fp4-b200-sglang-mtp-8k1k
model:
path: hf:deepseek-ai/DeepSeek-V4-Flash
container: lmsysorg/sglang:nightly-dev-cu13-20260925-8ca82118@sha256:baea7ec5ea86412ce59837f17011028b940f9777fa15129254de02bd6a842559
precision: fp4
resources:
gpu_type: b200
gpus_per_node: 8
frontend:
type: sglang
enable_multiple_frontends: false
observability:
enabled: false
tachometer:
enabled: false
engine: sglang
roles:
agg:
nodes: 1
workers: 1
gpus: 2
args:
served-model-name: deepseek-ai/DeepSeek-V4-Flash
trust-remote-code: true
tensor-parallel-size: 2
data-parallel-size: 1
expert-parallel-size: 1
moe-runner-backend: flashinfer_mxfp4
disable-flashinfer-autotune: true
disable-radix-cache: true
mem-fraction-static: 0.85
swa-full-tokens-ratio: 0.1
chunked-prefill-size: 8192
context-length: 9236
reasoning-parser: deepseek-v4
tool-call-parser: deepseekv4
speculative-algorithm: EAGLE
speculative-num-steps: 2
speculative-eagle-topk: 1
speculative-num-draft-tokens: 3
# Eager decode and speculative verify/draft; no CUDA graph capture.
cuda-graph-backend-decode: disabled
cuda-graph-backend-prefill: disabled
watchdog-timeout: 3600
enable-metrics: true
env:
PYTHONNOUSERSITE: '1'
PYTHONUNBUFFERED: '1'
SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE: '0'
benchmark:
type: custom
command: bash /infmax-workspace/benchmarks/single_node/srt_fixed_sequence.sh --dsv4
env:
MODEL: deepseek-ai/DeepSeek-V4-Flash
ISL: '8192'
OSL: '1024'
RANDOM_RANGE_RATIO: '0.8'
USE_CHAT_TEMPLATE: 'true'

zip_override_tp2:
roles:
agg:
args:
max-running-requests: [1, 2, 4, 8, 16, 32, 64, 128]
benchmark:
env:
CONC: ['1', '2', '4', '8', '16', '32', '64', '128']
2 changes: 1 addition & 1 deletion benchmarks/single_node/srt_fixed_sequence.sh
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ SRT_MONITOR_INTERVAL="$GPU_MONITOR_INTERVAL"
CLIENT_ARGS=()
for argument in "$@"; do
case "$argument" in
--trust-remote-code) CLIENT_ARGS+=("$argument") ;;
--trust-remote-code|--dsv4) CLIENT_ARGS+=("$argument") ;;
*) echo "ERROR: unsupported fixed-sequence argument: $argument" >&2; exit 1 ;;
esac
done
Expand Down
15 changes: 15 additions & 0 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8928,3 +8928,18 @@ dsv41flash-fp4-gb300-sglang-agentic-dspark:
search-space:
- { tp: 2, ep: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4-mtp/agentic.yaml }
- { tp: 4, ep: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/gb300-fp4-mtp/agentic.yaml }

dsv4flash-fp4-b200-sglang-mtp:
image: lmsysorg/sglang:nightly-dev-cu13-20260925-8ca82118@sha256:baea7ec5ea86412ce59837f17011028b940f9777fa15129254de02bd6a842559
model: deepseek-ai/DeepSeek-V4-Flash
model-prefix: dsv4flash
runner: cluster:b200-nscale
precision: fp4
framework: sglang
multinode: false
scenarios:
fixed-seq-len:
- isl: 8192
osl: 1024
search-space:
- { tp: 2, ep: 1, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128], srt-recipe: benchmarks/single_node/srt-slurm-recipes/dsv4flash/sglang/b200-fp4-mtp/8k1k.yaml }
14 changes: 14 additions & 0 deletions docs/configuration-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,20 @@ The former fork's direct ATOM frontend is not required.

### Cluster profiles

The B200 Nscale fixed-sequence `deepseek-ai/DeepSeek-V4-Flash` launcher stages the
checkpoint with `hf download --local-dir` under
`/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash` before SRT submission.
Each launch resumes or validates that download; failure stops submission. This
shared writable location does not require a pre-staged `/scratch/models` copy.
The download process uses explicit writable `HF_HOME`, `HF_HUB_CACHE`, and
`HF_XET_CACHE` paths beneath the checkpoint's `.cache/huggingface` directory;
the serving process retains its container cache settings.
For an uncached digest-pinned image, this lane maps Docker's `repo:tag@sha256:...`
to Pyxis/Enroot's `repo:sha256:...` manifest reference. The recorded image and
digest stay unchanged; a valid cached squash image still takes precedence.
The recipe uses bundled MTP through `EAGLE` (2 steps, top-k 1, 3 draft tokens),
with the DeepSeek-V4 chat encoder selected by the client's `--dsv4` option.

Launchers that use srt-slurm keep their cluster configuration in
[`runners/srt-slurm/<launcher>.yaml`](../runners/srt-slurm/). The native settings
(GPU count, scheduling directives, aliases, and mounts) are separate from workload recipes.
Expand Down
12 changes: 12 additions & 0 deletions docs/configuration-procedures_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,18 @@ frontend、一个聚合 worker,并设置 `enable_multiple_frontends: false`。
镜像。TRT-LLM 配方使用原生 `engine.served_model_name`,不再通过 `roles.agg.extra_args`
重复传入该参数。不再依赖此前分叉中的 ATOM 直连 frontend。

B200 Nscale 的固定序列 `deepseek-ai/DeepSeek-V4-Flash` 启动器在提交 SRT 前,
通过 `hf download --local-dir` 将 checkpoint 下载至
`/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash`。
每次启动均续传或验证下载,失败时停止提交;此共享可写目录不依赖
`/scratch/models` 中的预置副本。下载进程显式将 `HF_HOME`、`HF_HUB_CACHE`
和 `HF_XET_CACHE` 指向 checkpoint 的 `.cache/huggingface` 下的可写路径,
服务进程仍保留容器缓存设置。配方通过 `EAGLE` 使用原生 MTP
(2 steps、top-k 1、3 draft tokens),客户端以 `--dsv4` 选择 DeepSeek-V4 chat 编码器。
对于尚未缓存的 digest 固定镜像,此路径将 Docker 的 `repo:tag@sha256:...`
转换为 Pyxis/Enroot 的 `repo:sha256:...` manifest 引用。记录的镜像和 digest
保持不变,已有有效 squash 镜像仍优先使用。

## 规程索引

1. [准备 worktree](#准备-worktree)
Expand Down
79 changes: 79 additions & 0 deletions infx/tests/srt_slurm/test_dsv4_fixed_sequence_client.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
"""Exercise the fixed-sequence client's DeepSeek-V4 encoder forwarding."""

import os
import subprocess
from pathlib import Path

ROOT = Path(__file__).resolve().parents[3]


def test_dsv4_client_forwards_encoder_and_workload(tmp_path: Path) -> None:
harness = tmp_path / "harness.sh"
harness.write_text(
"""
source() {
if [[ "$1" == */benchmark_lib.sh && "$2" != --validation-only ]]; then
start_gpu_monitor() { :; }
stop_gpu_monitor() { :; }
run_benchmark_serving() { printf '%s\\n' "$@" > "$CAPTURE"; }
else
builtin source "$@"
fi
}
pip3() { :; }
"""
)
capture = tmp_path / "arguments"
result = subprocess.run(
["bash", str(ROOT / "benchmarks/single_node/srt_fixed_sequence.sh"), "--dsv4"],
env={
**os.environ,
"BASH_ENV": str(harness),
"CAPTURE": str(capture),
"INFERENCEX_REPO_ROOT": str(ROOT),
"MODEL": "test/model",
"FRAMEWORK": "sglang",
"CONC": "2",
"ISL": "256",
"OSL": "64",
"RANDOM_RANGE_RATIO": "0.8",
"RESULT_FILENAME": "point",
"RESULT_DIR": str(tmp_path),
"SRT_FRONTEND_HOST": "127.0.0.1",
"SRT_FRONTEND_PORT": "8000",
"RUN_EVAL": "false",
"EVAL_ONLY": "false",
"GPU_MONITOR_INTERVAL": "3",
"USE_CHAT_TEMPLATE": "true",
},
capture_output=True,
text=True,
check=False,
)
assert result.returncode == 0, result.stdout + result.stderr
assert capture.read_text().splitlines() == [
"--model",
"test/model",
"--port",
"8000",
"--base-url",
"http://127.0.0.1:8000",
"--backend",
"vllm",
"--input-len",
"256",
"--output-len",
"64",
"--random-range-ratio",
"0.8",
"--num-prompts",
"20",
"--max-concurrency",
"2",
"--result-filename",
"point",
"--result-dir",
str(tmp_path),
"--dsv4",
"--use-chat-template",
]
60 changes: 56 additions & 4 deletions infx/tests/srt_slurm/test_srt_single_node.py
Original file line number Diff line number Diff line change
Expand Up @@ -259,6 +259,8 @@ def test_submission_manifest(tmp_path, record, expected):
("h200-cw", "none"), ("h100-cw", "none"), ("h100-dgxc-slurm", "none"),
("b200-cw", "none"), ("b200-nb", "none"), ("b200-nscale-slurm", "none"),
("b200-nscale-slurm", "agentic"),
("b200-nscale-slurm", "download"), ("b200-nscale-slurm", "download-failure"),
("b200-nscale-slurm", "download-cached"),
("b300-dsxe", "none"),
("mi300x-amd", "none"), ("mi325x-amds", "none"), ("mi355x-amds", "none"),
] + [(pool, "missing-recipe") for pool in (
Expand All @@ -267,7 +269,18 @@ def test_submission_manifest(tmp_path, record, expected):
"mi325x-amds", "mi355x-amds",
)])
def test_pool_launcher_stages_artifacts_and_propagates_failure(point, tmp_path, pool, failure):
path, _, point_env = point
path, recipe, point_env = point
if failure.startswith("download"):
point_env = {
**point_env, "MODEL": "deepseek-ai/DeepSeek-V4-Flash",
"IMAGE": f"example/server:nightly@sha256:{'a' * 64}",
}
recipe["model"]["path"] = "hf:deepseek-ai/DeepSeek-V4-Flash"
recipe["model"]["container"] = point_env["IMAGE"]
recipe["benchmark"]["env"]["MODEL"] = "deepseek-ai/DeepSeek-V4-Flash"
path.write_text(yaml.safe_dump({"base": recipe}))
if failure == "download-cached":
(tmp_path / f"example_server_nightly_sha256_{'a' * 64}.sqsh").write_text("cache")
binaries = tmp_path / "bin"
binaries.mkdir()
model = tmp_path / "model"
Expand All @@ -279,13 +292,23 @@ def test_pool_launcher_stages_artifacts_and_propagates_failure(point, tmp_path,
# setup/profile/acceptance helpers, binder, and artifact collection.
scripts = {
"git": 'if [[ " $* " == *" clone "* ]]; then mkdir -p "${@: -1}/configs"; else echo test-commit; fi',
"uv": 'if [[ "$1" == venv ]]; then mkdir -p .venv/bin; echo ":" > .venv/bin/activate; fi',
"uv": """
if [[ "$1" == tool ]]; then
printf '%s\\n' "$@" > "$DOWNLOAD_CAPTURE"
printf '%s\\n' "$HF_HOME" "$HF_HUB_CACHE" "$HF_XET_CACHE" > "$DOWNLOAD_CACHE_CAPTURE"
[[ "$TEST_FAILURE" != download-failure ]] || exit 17
elif [[ "$1" == venv ]]; then
mkdir -p .venv/bin
echo ":" > .venv/bin/activate
fi
""",
"make": '[[ "$TEST_FAILURE" == bootstrap ]] && exit 13; mkdir -p bin; touch bin/uv',
"squeue": '[[ "$TEST_FAILURE" == submission || "$TEST_FAILURE" == agentic ]] && echo "42"; exit 0',
"salloc": 'echo "Granted job allocation 42"',
"sacct": 'if [[ "$TEST_FAILURE" == allocation ]]; then echo "FAILED|1:0"; else echo "COMPLETED|0:0"; fi',
"scancel": 'printf "%s\\n" "$@" >> "$CANCEL_CAPTURE"',
"tail": 'exit 0',
"unsquashfs": '[[ "$TEST_FAILURE" == download-cached ]]',
}
for name, script in scripts.items():
binary = binaries / name
Expand Down Expand Up @@ -325,6 +348,9 @@ def test_pool_launcher_stages_artifacts_and_propagates_failure(point, tmp_path,
"B300_HF_CACHE_CONTAINER_DIR": "/hf", "ENROOT_IMPORT_TIME_LIMIT": "10",
"INFERENCEX_RUNTIME_ENV_VARS": "REQUIRE_POWER",
"TEST_FAILURE": failure, "CANCEL_CAPTURE": str(capture),
"DOWNLOAD_CAPTURE": str(tmp_path / "download-args"),
"DOWNLOAD_CACHE_CAPTURE": str(tmp_path / "download-cache"),
"HF_HOME": "/read-only/hf", "HF_XET_CACHE": "/read-only/xet",
"SRUN_CAPTURE": str(tmp_path / "srun.jsonl"),
"KEEP_LOGS": "0",
}
Expand All @@ -340,7 +366,21 @@ def test_pool_launcher_stages_artifacts_and_propagates_failure(point, tmp_path,
["bash", str(ROOT / f"runners/launch_{pool}.sh")], cwd=tmp_path,
env=env, capture_output=True, text=True, timeout=30,
)
assert result.returncode == {"none": 0, "allocation": 1, "submission": 7, "bootstrap": 13, "missing-recipe": 1, "agentic": 0}[failure], result.stderr
assert result.returncode == {"none": 0, "allocation": 1, "submission": 7, "bootstrap": 13, "missing-recipe": 1, "agentic": 0, "download": 0, "download-failure": 1, "download-cached": 0}[failure], result.stderr
if failure.startswith("download"):
assert Path(env["DOWNLOAD_CAPTURE"]).read_text().splitlines() == [
"tool", "run", "--from", "huggingface-hub>=0.34,<2", "hf", "download",
"deepseek-ai/DeepSeek-V4-Flash", "--local-dir",
"/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash",
]
assert Path(env["DOWNLOAD_CACHE_CAPTURE"]).read_text().splitlines() == [
"/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash/.cache/huggingface",
"/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash/.cache/huggingface/hub",
"/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash/.cache/huggingface/xet",
]
if failure == "download-failure":
assert not (tmp_path / "srt-single-node-submission.json").exists()
return
if failure == "agentic":
calls = [json.loads(line) for line in Path(env["SRUN_CAPTURE"]).read_text().splitlines()]
assert calls[-1][-2:] == ["bash", "benchmarks/single_node/agentic/fixture_fp8_b200.sh"]
Expand All @@ -362,8 +402,20 @@ def test_pool_launcher_stages_artifacts_and_propagates_failure(point, tmp_path,
assert json.loads((tmp_path / "gpu_metrics_context.json").read_text()) == {"device_count": 4}
assert (tmp_path / "srt-single-node-logs.tar.gz").stat().st_size > 0
cluster_config = yaml.safe_load(next(tmp_path.glob("srt-single.*/checkout/srtslurm.yaml")).read_text())
assert cluster_config["containers"]["test:tag"] == "test:tag"
if failure == "download":
assert cluster_config["containers"][point_env["IMAGE"]] == f"example/server:sha256:{'a' * 64}"
elif failure == "download-cached":
assert cluster_config["containers"][point_env["IMAGE"]] == str(
tmp_path / f"example_server_nightly_sha256_{'a' * 64}.sqsh"
)
else:
assert cluster_config["containers"]["test:tag"] == "test:tag"
assert cluster_config["use_exclusive_sbatch_directive"] is True
if failure in {"download", "download-cached"}:
assert cluster_config["model_paths"]["hf:deepseek-ai/DeepSeek-V4-Flash"] == (
"/data/home/sa-shared/gharunners/models/DeepSeek-V4-Flash"
)
assert cluster_config["default_mounts"]["/data/home/sa-shared/gharunners/hf-hub-cache"] == "/hf"
assert (capture.read_text() if capture.exists() else "") == ("42\n" if failure == "submission" else "")


Expand Down
18 changes: 18 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8986,3 +8986,21 @@
- "Hopper 使用 FA3,Blackwell FA4 fp8 descale 问题不适用,EAGLE3 draft 保持 attention_backend FLASH_ATTN。"
- "This change does not alter the EAGLE3 draft model data type. The draft loads unmodified from the published Inferact/MiniMax-M3-EAGLE3-GQA checkpoint via --speculative-config (method=eagle3). kv-cache-dtype fp8 sets KV-cache storage precision, not the draft weights, and no flag overrides or re-quantizes the draft weights."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3503

- config-keys:
- dsv4flash-fp4-b200-sglang-mtp
scenario-type:
- fixed-seq-len
description:
- "Add DeepSeek-V4-Flash B200 Nscale TP2 SGLang 8k1k serving at concurrency 1, 2, 4, 8, 16, 32, 64, 128 with bundled EAGLE MTP (2 steps, top-k 1, 3 draft tokens) with CUDA/HIP graphs disabled for decode, prefill, and speculative verify/draft, real verification, and GPU-resident weights/KV. Stage the unstaged checkpoint and HF/Xet download caches on writable shared storage before SRT submission; stop on download failure."
- "新增 DeepSeek-V4-Flash B200 Nscale TP2 SGLang 8k1k serving,并发为 1、2、4、8、16、32、64、128,采用原生 EAGLE MTP(2 steps、top-k 1、3 draft tokens),decode、prefill 及投机验证/draft 均关闭 CUDA/HIP graph、真实验证和 GPU 常驻权重/KV。SRT 提交前在共享可写目录下载 checkpoint 并设置 HF/Xet 下载缓存,下载失败时停止提交。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3517

- config-keys:
- dsv4flash-fp4-b200-sglang-mtp
scenario-type:
- fixed-seq-len
description:
- "Fix B200 V4 Flash single-node Pyxis image import by converting uncached Docker @sha256 references to Enroot manifest-tag syntax. Preserve the pinned image digest and prefer valid cached squash images. No serving settings change."
- "修复 B200 V4 Flash 单节点 Pyxis 镜像导入:将尚未缓存的 Docker @sha256 引用转换为 Enroot manifest-tag 语法,保留固定镜像 digest 并优先使用有效 squash 缓存,不更改 serving 设置。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3517
Loading
Loading