Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,20 @@ roles:
sbatch_directives: {cpus-per-task: "144", mem: "0"}
srun_options: {container-remap-root: ""}

telemetry:
enabled: true
collect_interval_ms: 1000
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 12
dcgm_exporter:
container_image: dcgm-exporter
port: 9401

benchmark:
concurrencies: [1]
type: custom
command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh
env:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -106,7 +106,20 @@ sbatch_directives:
srun_options:
container-remap-root: ""

telemetry:
enabled: true
collect_interval_ms: 1000
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 12
dcgm_exporter:
container_image: dcgm-exporter
port: 9401

benchmark:
concurrencies: [1]
type: custom
command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh
env:
Expand Down
2 changes: 2 additions & 0 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5702,6 +5702,8 @@ dsv4-fp4-gb200-dynamo-vllm-agentic-mtp-agg:
scenarios:
agentic-coding:
- search-space:
# Both recipes require PowerX; the launcher supplies each matrix concurrency
# to the native collector and the shared AgentX benchmark window writer.
# Match the B200 TP8 latency points for direct normalized-interactivity
# comparison, then use the public vLLM GB200 DEP8 topology for throughput.
- spec-decoding: mtp
Expand Down
27 changes: 27 additions & 0 deletions docs/configuration-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -183,6 +183,33 @@ Mapping source: [`benchmarks/multi_node/srt-slurm-recipes/RECIPES.md`](../benchm

Do not ship one side alone. `srtctl` reads the recipe, while matrix generation reads the master config. Recipe-only changes can mislabel results. Master-only changes do not alter the deployed recipe.

### GB200 AgentX measured power

GB200 SRT AgentX uses the selected recipe, after native selector expansion and
caller overrides, to enable PowerX. Use the existing `telemetry.enabled: true`,
`telemetry.required: true`, `storage_subdir: power`, and `dcgm_exporter` block
(`container_image: dcgm-exporter`); no model-specific power branch is needed.
The benchmark must use `bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh`
with `INFMAX_CONTAINER_WORKSPACE: /infmax-workspace`, `RESULT_DIR: /logs/agentic`,
and `IS_MULTINODE: "true"`. SRT validates the head-client clock, topology and
sampling settings. Model paths, serving arguments, quantization and mounts retain
their existing runtime configuration.

The launcher installs the pinned submodule runtime before inspecting the recipe,
then provisions the exporter and records the same producer SHA for collection.
Each AgentX or power-enabled matrix job must select exactly one recipe (including an indexed zip
selector); a multi-recipe selection fails before submission. Native overrides
supply matrix `CONC_LIST` to both the benchmark and `benchmark.concurrencies`,
without rewriting recipe YAML. Enabled AgentX power also requires the shared
window writer and strict result adapter; it cannot be made successful by disabling
required telemetry. Eval-only jobs retain normal eval behavior and do not publish
throughput power results.

The DSV4 GB200 vLLM TP8/DEP8 aggregate recipes demonstrate ordinary recipe-only
power enrollment. Local routing tests do not qualify live sampling: the selected
PR sweep still needs complete device/window evidence, failed-collection artifact
retention, evals, and downstream ingest/API/page verification.

## Register an llm-d recipe

Sources: [`benchmarks/llm-d/README.md`](../benchmarks/llm-d/README.md), [`benchmarks/multi_node/llm-d/README.md`](../benchmarks/multi_node/llm-d/README.md), [`llm-d-recipes/`](../benchmarks/multi_node/llm-d-recipes/), and the current [`llmd-vllm` benchmark wrapper](../benchmarks/multi_node/dsv4_fp4_gb200_llmd-vllm-disagg.sh).
Expand Down
22 changes: 22 additions & 0 deletions docs/configuration-procedures_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,28 @@ B200 Nscale 的 GLM-5.1 可用 `MODEL_PATH` 指定已有共享权重,覆盖默

不得只提交一侧:`srtctl` 读取配方,而矩阵生成读取主配置。仅改配方可能给结果贴错标签;仅改主配置不会改变实际部署的配方。

### GB200 AgentX 实测功耗

GB200 SRT AgentX 根据原生 selector 展开及调用方 overrides 后的实际配方启用
PowerX。使用现有的 `telemetry.enabled: true`、`telemetry.required: true`、
`storage_subdir: power` 及 `dcgm_exporter` 配置(`container_image: dcgm-exporter`),
无需添加模型专属功耗分支。benchmark 必须使用
`bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh`,并设置
`INFMAX_CONTAINER_WORKSPACE: /infmax-workspace`、`RESULT_DIR: /logs/agentic`
和 `IS_MULTINODE: "true"`。SRT 校验 head client 时钟、拓扑与采样设置。
模型路径、服务参数、量化和挂载保留各自现有的运行配置。

launcher 先安装固定 submodule runtime,再检查配方、准备 exporter,记录同一份
producer SHA 供采集使用。每个 AgentX 或启用功耗的矩阵 job 必须只选中一个配方(可用带索引的 zip
selector);多配方选择会在提交前失败。原生 overrides 将矩阵 `CONC_LIST` 同时
传给 benchmark 和 `benchmark.concurrencies`,不重写配方 YAML。启用 AgentX
功耗后必须使用共享窗口写入器与严格结果适配器,不能通过关闭 required telemetry
制造成功。eval-only 保留正常评估行为,不发布吞吐测量功耗结果。

DSV4 GB200 vLLM TP8/DEP8 aggregate 配方示范了普通配方的功耗接入。
本地路由测试不代表真实采样通过:所选 PR sweep 仍需完整设备/窗口证据、
采集失败后的产物保留、eval,以及下游 ingest/API/页面验收。

## 注册 llm-d 配方

来源:[`benchmarks/llm-d/README.md`](../benchmarks/llm-d/README.md)、[`benchmarks/multi_node/llm-d/README.md`](../benchmarks/multi_node/llm-d/README.md)、[`llm-d-recipes/`](../benchmarks/multi_node/llm-d-recipes/) 和当前 [`llmd-vllm` 基准 wrapper](../benchmarks/multi_node/dsv4_fp4_gb200_llmd-vllm-disagg.sh)。
Expand Down
83 changes: 76 additions & 7 deletions infx/srt_slurm/synthetic_acceptance.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
import math
import os
import re
import shlex
import subprocess
import sys
from collections.abc import Mapping
Expand Down Expand Up @@ -230,15 +231,10 @@ def plan_commands(
from srtctl.core.overrides import apply_overrides_to_recipe, parse_overrides

path, _, selector = config.partition(":")
raw = yaml.safe_load(Path(path).read_text())
if not isinstance(raw, dict):
raise ValueError("Recipe must be a mapping")
raw, caller_overrides = recipe_with_overrides(path, arguments)
parser = argparse.ArgumentParser(add_help=False, allow_abbrev=False)
parser.add_argument("--set", action="append")
parser.add_argument("--unset", action="append")
existing, _ = parser.parse_known_args(arguments)
caller_overrides = parse_overrides(existing.set, existing.unset)
apply_overrides_to_recipe(raw, caller_overrides)
commands = []
for variant, recipe in selected_recipes(raw, selector or None):
arguments_to_add = build_overrides(recipe, framework, environment, golden_dir=golden_dir)
Expand All @@ -260,17 +256,90 @@ def plan_commands(
return commands


def recipe_with_overrides(path: str, arguments: list[str]) -> tuple[dict[str, Any], list[Any]]:
"""Use the same native caller-override ordering for inspection and submission."""
from srtctl.core.overrides import apply_overrides_to_recipe, parse_overrides

raw = yaml.safe_load(Path(path).read_text())
if not isinstance(raw, dict):
raise ValueError("Recipe must be a mapping")
parser = argparse.ArgumentParser(add_help=False, allow_abbrev=False)
parser.add_argument("--set", action="append")
parser.add_argument("--unset", action="append")
existing, _ = parser.parse_known_args(arguments)
overrides = parse_overrides(existing.set, existing.unset)
apply_overrides_to_recipe(raw, overrides)
return raw, overrides


def inspect_power(config: str, arguments: list[str], environment: Mapping[str, str]) -> str:
"""Check the existing GB200 power artifact/window contract before submission."""
from marshmallow import ValidationError
from srtctl.core.config import expand_engine_config_defaults, resolve_config_with_defaults
from srtctl.core.schema import SrtConfig

path, _, selector = config.partition(":")
raw, _ = recipe_with_overrides(path, arguments)
variants = selected_recipes(raw, selector or None)
if len(variants) != 1:
# Preserve existing non-AgentX, non-power multi-variant submissions.
# Their lifecycle is separate from the single-job PowerX result contract.
if (
variants
and environment["IS_AGENTIC"] != "1"
and all(not recipe.get("telemetry", {}).get("enabled", False) for _, recipe in variants)
):
return "none"
raise ValueError("GB200 launcher requires exactly one selected recipe per job")
recipe = variants[0][1]
if not recipe.get("telemetry", {}).get("enabled", False):
return "none"
resolved = resolve_config_with_defaults(recipe, None)
expand_engine_config_defaults(resolved)
try:
typed = SrtConfig.Schema().load(resolved)
except ValidationError as error:
raise ValueError(f"Invalid power recipe: {error}") from error
if not typed.telemetry.enabled:
return "none"
if typed.telemetry.dcgm_exporter is None:
raise ValueError("GB200 PowerX requires telemetry.dcgm_exporter")
if typed.telemetry.dcgm_exporter.container_image != "dcgm-exporter":
raise ValueError("GB200 PowerX requires the dcgm-exporter container alias")
if typed.telemetry.storage_subdir != "power":
raise ValueError("GB200 PowerX requires telemetry.storage_subdir: power")
if environment["IS_AGENTIC"] != "1":
return "dcgm"
benchmark = typed.benchmark
if (
benchmark.type != "custom"
or shlex.split(benchmark.command or "")
!= ["bash", "/infmax-workspace/benchmarks/multi_node/agentic_srt.sh"]
or benchmark.env.get("RESULT_DIR") != "/logs/agentic"
or benchmark.env.get("INFMAX_CONTAINER_WORKSPACE") != "/infmax-workspace"
or benchmark.env.get("IS_MULTINODE") != "true"
):
raise ValueError("AgentX PowerX requires the shared agentic_srt.sh window/result contract")
if not typed.telemetry.required:
raise ValueError("AgentX PowerX requires telemetry.required: true")
return "agentx"


def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--inspect-power", action="store_true")
parser.add_argument("config")
parser.add_argument("framework")
parser.add_argument("arguments", nargs=argparse.REMAINDER)
args = parser.parse_args(argv)
arguments = args.arguments[1:] if args.arguments[:1] == ["--"] else args.arguments
try:
if args.inspect_power:
print(inspect_power(args.config, arguments, os.environ))
return 0
commands = plan_commands(args.config, args.framework, arguments, os.environ)
except (OSError, KeyError, ValueError, TypeError, yaml.YAMLError) as error:
print(f"ERROR: golden acceptance: {error}", file=sys.stderr)
print(f"ERROR: SRT recipe: {error}", file=sys.stderr)
return 1
for command in commands:
result = subprocess.run(command, check=False)
Expand Down
30 changes: 30 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8587,3 +8587,33 @@
- "将镜像从已失效的 nightly-eed1f3d0 重新固定到 vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984(与 MI355X 臂在 SemiAnalysisAI/InferenceX#3326 中所用相同 commit 的 ROCm 10.0 nightly 渠道构建);该镜像包含 vllm-project/vllm#57491,将两处 is_cuda() 判定放宽为 is_cuda_alike(),使 engram_config 在 gfx942 上可解析;由于 cpu_offload 现通过 VLLM_PLE_CPU_OFFLOAD 默认开启,配方按 TP 显式设置该值"
- "新增 --no-swa-bounded-replay:vllm-project/vllm#56227 在两个 pin 之间将 SWA bounded replay 默认开启,其依赖的 window clamp 在 ROCm sparse SWA 路径中缺失,曾使所有 gfx950 数据点以 HSA_STATUS_ERROR_MEMORY_FAULT 崩溃;gfx942 使用同一路径。prefix caching 保持开启"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3337

- config-keys:
- dsr1-fp4-gb200-dynamo-trt
- dsr1-fp8-gb200-dynamo-trt
- dsr1-fp8-gb200-dynamo-sglang
- dsr1-fp4-gb200-dynamo-sglang
- qwen3.5-fp8-gb200-dynamo-sglang
- qwen3.5-fp4-gb200-dynamo-sglang-agentic-mtp
- minimaxm3-fp4-gb200-dynamo-vllm-agentic-agg-mtp
- minimaxm3-fp4-gb200-dynamo-vllm-agentic-disagg-mtp
- minimaxm3-fp4-gb200-dynamo-trt-agentic-agg-mtp
- dsv4-fp4-gb200-dynamo-vllm-agentic-mtp-agg
- dsv4-fp4-gb200-dynamo-vllm-agentic-mtp-disagg
- dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-agg
- dsv4-fp4-gb200-dynamo-vllm-agentic-mtp2-disagg
- kimik3-fp4-gb200-dynamo-vllm-agentic
- kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-dcp16-agg
- kimik3-fp4-gb200-dynamo-vllm-agentic-mooncake-dcp16-agg
- kimik3-fp4-gb200-dynamo-vllm-agentic-dspark-mooncake-tp8pp2
- glm5.2-fp4-gb200-dynamo-sglang-agentic-agg
- glm5.2-fp4-gb200-dynamo-sglang-agentic-disagg
- glm5.2-fp4-gb200-dynamo-sglang-agentic-mtp
- glm5.2-fp4-gb200-dynamo-sglang-agentic-mtp-agg
- qwen3.5-fp8-gb200-dynamo-sglang-mtp
description:
- Resolve GB200 SRT power from the selected recipe and native overrides; retain
required telemetry, window identity and failure propagation.
- Enable required measured power for both DSV4 GB200 vLLM AgentX aggregate recipes
without a model-specific power branch.
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3358
85 changes: 26 additions & 59 deletions runners/launch_gb200-nv.sh
Original file line number Diff line number Diff line change
Expand Up @@ -264,52 +264,6 @@ NGINX_SQUASH_FILE="${SQUASH_DIR}/$(echo "$NGINX_IMAGE" | sed 's/[\/:@#]/_/g').sq
import_squash "$SQUASH_FILE" "$IMAGE"
import_squash "$NGINX_SQUASH_FILE" "$NGINX_IMAGE"

# The power lane is on iff the resolved recipe carries an enabled dcgm-power
# telemetry block. Read the workspace mirror; it overlays the srt-slurm clone later.
USES_DCGM_POWER=0
_RECIPE_REL="${CONFIG_FILE%%:*}"
_RECIPE_SRC="$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/${_RECIPE_REL#recipes/}"
# Scoped match: a stray "enabled: true" outside the telemetry block must not flip the lane.
if [[ -n "$CONFIG_FILE" && -f "$_RECIPE_SRC" ]] && awk '
/^telemetry:/ { t = 1; next }
t && /^[^ ]/ { t = 0 }
t && /^ dcgm_exporter:/ { p = 1 }
t && /^ enabled: true$/ { e = 1 }
END { exit !(p && e) }
' "$_RECIPE_SRC"; then
USES_DCGM_POWER=1
fi

USES_AGENTX_POWER=0
if [[ "$USES_DCGM_POWER" == "1" && "$IS_AGENTIC" == "1" ]]; then
if [[ "$MODEL_PREFIX" == "glm5.2" && "$PRECISION" == "fp4" &&
"$FRAMEWORK" == "dynamo-sglang" &&
"$_RECIPE_REL" == "recipes/glm5.2/sglang/gb200-fp4/agentx/agg.yaml" ]]; then
USES_AGENTX_POWER=1
elif [[ "$MODEL_PREFIX" == "kimik3" && "$PRECISION" == "fp4" &&
"$FRAMEWORK" == "dynamo-vllm" &&
"$_RECIPE_REL" == recipes/kimik3/vllm/gb200-fp4/agentx/* ]]; then
USES_AGENTX_POWER=1
else
echo "Error: AgentX dcgm-power requires the GLM-5.2 aggregate or supported Kimi-K3 recipe" >&2
exit 1
fi
fi
if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" != "dynamo-sglang" && "$USES_AGENTX_POWER" != "1" ]]; then
echo "Error: dcgm-power requires dynamo-sglang or the supported Kimi-K3 AgentX route" >&2
exit 1
fi

if [[ "$USES_DCGM_POWER" == "1" ]]; then
DCGM_EXPORTER_IMAGE="nvcr.io/nvidia/k8s/dcgm-exporter:4.6.0-4.8.3-distroless"
DCGM_EXPORTER_SQSH="${SQUASH_DIR}/$(echo "$DCGM_EXPORTER_IMAGE" | sed 's/[\/:@#]/_/g').sqsh"
import_squash "$DCGM_EXPORTER_SQSH" "$DCGM_EXPORTER_IMAGE"
test -r "$DCGM_EXPORTER_SQSH" || { echo "Error: DCGM exporter squash not readable: $DCGM_EXPORTER_SQSH" >&2; exit 1; }
unsquashfs -l "$DCGM_EXPORTER_SQSH" > /dev/null || { echo "Error: DCGM exporter squash invalid: $DCGM_EXPORTER_SQSH" >&2; exit 1; }
sha256sum "$DCGM_EXPORTER_SQSH" > "$GITHUB_WORKSPACE/exporter-image.sha256"
fi


export ISL="$ISL"
export OSL="$OSL"

Expand Down Expand Up @@ -379,11 +333,8 @@ if [ -d "$SRT_REPO_DIR" ]; then
rm -rf "$SRT_REPO_DIR"
fi

# This checkpoint is staged on compute-node NVMe for the power lane.
if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" == "dynamo-sglang" && "$MODEL_PREFIX" == "dsv4" && "$IS_AGENTIC" != "1" ]]; then
export MODEL_PATH="/mnt/numa1/models/DeepSeek-V4-Pro"
fi
setup_srt_slurm "$SRT_REPO_DIR" "$FRAMEWORK" "$USES_DCGM_POWER" || exit 1
# Power inspection needs this exact runtime's selector/override implementation.
setup_srt_slurm "$SRT_REPO_DIR" "$FRAMEWORK" 0 || exit 1

echo "Installing srtctl..."
curl -LsSf https://astral.sh/uv/install.sh | sh
Expand All @@ -405,6 +356,23 @@ if ! command -v srtctl &> /dev/null; then
exit 1
fi

prepare_gb200_srt_power "$CONFIG_FILE" "$FRAMEWORK" || exit 1

# This checkpoint is staged on compute-node NVMe for the non-AgentX power lane.
if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" == "dynamo-sglang" && "$MODEL_PREFIX" == "dsv4" && "$IS_AGENTIC" != "1" ]]; then
export MODEL_PATH="/mnt/numa1/models/DeepSeek-V4-Pro"
fi
if [[ "$USES_DCGM_POWER" == "1" ]]; then
cp "$GITHUB_WORKSPACE/srt-slurm-sha.txt" "$GITHUB_WORKSPACE/power-producer-sha.txt" || exit 1
DCGM_EXPORTER_IMAGE="nvcr.io/nvidia/k8s/dcgm-exporter:4.6.0-4.8.3-distroless"
DCGM_EXPORTER_SQSH="${SQUASH_DIR}/$(echo "$DCGM_EXPORTER_IMAGE" | sed 's/[\/:@#]/_/g').sqsh"
import_squash "$DCGM_EXPORTER_SQSH" "$DCGM_EXPORTER_IMAGE"
test -r "$DCGM_EXPORTER_SQSH" || { echo "Error: DCGM exporter squash not readable: $DCGM_EXPORTER_SQSH" >&2; exit 1; }
unsquashfs -l "$DCGM_EXPORTER_SQSH" > /dev/null || { echo "Error: DCGM exporter squash invalid: $DCGM_EXPORTER_SQSH" >&2; exit 1; }
sha256sum "$DCGM_EXPORTER_SQSH" > "$GITHUB_WORKSPACE/exporter-image.sha256"
fi


echo "Configs available at: $SRT_REPO_DIR/"

SRTCTL_ROOT="${GITHUB_WORKSPACE}/srt-slurm"
Expand Down Expand Up @@ -485,13 +453,7 @@ if command -v squeue >/dev/null 2>&1; then
sleep 5
done
fi
sed -i "s/^name:.*/name: \"${SRT_SLURM_JOB_NAME}\"/" "$CONFIG_PATH"

if [[ "$USES_AGENTX_POWER" == "1" ]]; then
read -r -a POWER_CONCURRENCIES <<< "$CONC_LIST"
python3 "$GITHUB_WORKSPACE/runners/inject_srt_power_concurrencies.py" \
"$CONFIG_PATH" "${POWER_CONCURRENCIES[@]}" || exit 1
fi
SRTCTL_RECIPE_ARGS+=(--set "name=$SRT_SLURM_JOB_NAME")

# sbatch's --export=ALL would carry VIRTUAL_ENV into job_script_minimal.j2,
# whose `uv run` then dies with "Broken symlink at .venv/bin/python3" because
Expand All @@ -516,8 +478,13 @@ if [[ "$FRAMEWORK" == "dynamo-sglang" ]]; then
fi
# srtctl gives RUNNER_NAME precedence over config.name; override it for the
# submission so the #SBATCH job name keeps the namespace used above.
SRTCTL_OUTPUT=$(RUNNER_NAME="$SRT_SLURM_JOB_NAME" apply_srt_recipe "$CONFIG_FILE" "$FRAMEWORK" "${SRTCTL_EVAL_ARGS[@]}" "${SRTCTL_APPLY_ARGS[@]}" 2>&1)
SRTCTL_OUTPUT=$(RUNNER_NAME="$SRT_SLURM_JOB_NAME" apply_srt_recipe "$CONFIG_FILE" "$FRAMEWORK" "${SRTCTL_RECIPE_ARGS[@]}" "${SRTCTL_APPLY_ARGS[@]}" 2>&1)
SRTCTL_RC=$?
echo "$SRTCTL_OUTPUT"
if [[ "$SRTCTL_RC" != "0" ]]; then
cancel_submitted_srt_jobs "$SRTCTL_OUTPUT"
exit "$SRTCTL_RC"
fi

JOB_ID=$(echo "$SRTCTL_OUTPUT" | grep -oP '✅ Job \K[0-9]+' || echo "$SRTCTL_OUTPUT" | grep -oP 'Job \K[0-9]+')

Expand Down
Loading
Loading