Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
41 commits
Select commit Hold shift + click to select a range
eafd4be
refactor: start native single-node SRT-Slurm migration
adibarra Sep 21, 2026
81f65a0
docs: link single-node migration changelog to draft PR
adibarra Sep 21, 2026
80a4da3
feat: wire native H200 SRT pilot into end-to-end workflow
adibarra Sep 21, 2026
42eecdb
test: cover native single-node job failures and artifacts
adibarra Sep 21, 2026
74687e3
fix: drop unused AIPerf inputs from SRT setup
adibarra Sep 21, 2026
443ff63
fix: let native SRT stage the pilot container image
adibarra Sep 21, 2026
348f773
fix: bootstrap native SRT binaries before pilot submission
adibarra Sep 21, 2026
a88a27a
feat: port H200 MTP and Qwen recipes to native SRT
adibarra Sep 22, 2026
b9757d9
docs: remove migration guide
adibarra Sep 22, 2026
8762e8c
chore: merge main into SRT migration
adibarra Sep 22, 2026
2763e8d
feat: migrate fixed-sequence SGLang recipes to native SRT
adibarra Sep 22, 2026
a67435c
feat: migrate fixed-sequence TRT recipes to native SRT
adibarra Sep 22, 2026
b4724b0
chore: merge main into SRT migration
adibarra Sep 22, 2026
83eaebe
feat: convert remaining fixed-sequence recipes to native SRT
adibarra Sep 22, 2026
a5e45fa
merge: sync main into native SRT migration
adibarra Sep 22, 2026
c40b057
fix: forward native Docker eval model identity
adibarra Sep 22, 2026
571fa51
fix: apply AMD container options as native leaf overrides
adibarra Sep 22, 2026
295d7e0
refactor: reuse native job workspace for AMD scratch files
adibarra Sep 22, 2026
0bd3bf5
fix: complete AMD native launch and terminal status handling
adibarra Sep 22, 2026
8e27918
refactor: require native recipes for the Slurm fixed-sequence cutover
adibarra Sep 22, 2026
8f3f0d5
refactor: make single-node fixed-sequence coverage SRT-only
adibarra Sep 22, 2026
c430090
fix: allow client dependencies in container virtual environments
adibarra Sep 22, 2026
6f60a8b
fix: limit native single-node steps to serving GPUs
adibarra Sep 22, 2026
b3a5a7f
chore: merge main into single-node SRT migration
adibarra Sep 22, 2026
161aa24
fix: restore repository workdir for native container steps
adibarra Sep 23, 2026
838607d
chore: sync latest main before ATOM smoke retry
adibarra Sep 23, 2026
dd95570
fix: stream AMD power samples without input buffering
adibarra Sep 23, 2026
2d532e6
chore: sync main into SRT migration
adibarra Sep 23, 2026
77961ed
fix: match registered H200 runner labels
adibarra Sep 23, 2026
814c007
chore: merge main into SRT migration
adibarra Sep 23, 2026
c6f47fa
chore: sync GB300 AgentX updates from main
adibarra Sep 23, 2026
4747349
chore: sync H200 AgentX updates from main
adibarra Sep 23, 2026
a3dc898
chore: pin srt-slurm submodule to main [skip ci]
cquil11 Sep 23, 2026
082279d
refactor(srt): return single-node migration to NVIDIA upstream
adibarra Sep 23, 2026
66781f6
merge: preserve concurrent upstream submodule update
adibarra Sep 23, 2026
2e358d1
merge: sync main result-processing Python fix
adibarra Sep 23, 2026
b0a9064
chore: bump srt-slurm submodule to v2.23.2
cquil11 Sep 23, 2026
91f040d
refactor(srt): check MI300X firmware in a setup hook
cquil11 Sep 24, 2026
5fe80b5
fix(srt): apply pending upstream srt-slurm patches at setup
cquil11 Sep 24, 2026
2016415
chore: remove perf changelog and smoke-test notes from the single-nod…
cquil11 Sep 24, 2026
f96a74e
merge: sync main into codex/single-node-srt-slurm
adibarra Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ env:
HF_HUB_CACHE: '/mnt/hf_hub_cache/'
EXP_NAME: ${{ fromJSON(inputs.config).exp-name }}
RECIPE_FINGERPRINT: ${{ fromJSON(inputs.config).recipe-fingerprint || '' }}
SRT_RECIPE: ${{ fromJSON(inputs.config).srt-recipe || '' }}
MODEL: ${{ fromJSON(inputs.config).model }}
THINKING_MODE: thinking_on
MODEL_PREFIX: ${{ fromJSON(inputs.config).model-prefix }}
Expand Down Expand Up @@ -399,6 +400,9 @@ jobs:
server.log
results/*.log
results/*_config.json
srt-single-node-logs.tar.gz
srt-single-node-submission.json
srt-slurm-sha.txt
if-no-files-found: ignore

- name: Upload GPU metrics
Expand Down
2 changes: 2 additions & 0 deletions MODELS.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,8 @@ Rationale: `dsv4` carries the largest single-turn footprint in the repository. 4

**Deprecation parity audit (2026-09-21):** Active master configs and benchmark-script locations match the enacted retirements above and in the support matrix below. GLM-5.1 B200 TileRT remains the documented exception to the earlier GLM-5/5.1 and 1k1k retirements. Conditional A/B baseline retirement remains pending; non-speculative Pareto contributors remain supported. The broader routing audit also removed stale retired-model branches from launchers/runtime settings and a GLM-5-only environment override, and corrected workflow/agent guidance that still recommended retired coverage. SPEED-Bench collectors, historical result readers, and the explicitly retained recipe YAMLs remain available. Deprecated configs are consolidated in [`configs/deprecated/amd-master.yaml`](configs/deprecated/amd-master.yaml) and [`configs/deprecated/nvidia-master.yaml`](configs/deprecated/nvidia-master.yaml).

**Single-node SRT-only cutover (2026-09-22):** Active single-node fixed-sequence recipes now use SRT-Slurm. The two Docker-only Qwen3.5 RTX PRO 6000 FP4 configs (with and without MTP) are retired, with their original settings preserved in `configs/deprecated/nvidia-master.yaml` and their scripts in `benchmarks/single_node/fixed_seq_len/deprecated/`. The unused `rtx6000pro-lat` runner mappings, launcher, and runtime settings are removed. Qwen3.5 remains active on the other supported Slurm pools; AgentX and multi-node coverage are unchanged.

## Scenarios

| Scenario | ISL/OSL | Status |
Expand Down
2 changes: 2 additions & 0 deletions MODELS_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,8 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支

**弃用状态一致性核查(2026-09-21):** 启用的主配置及基准测试脚本位置与上述已执行的退役事项和下方支持矩阵一致。GLM-5.1 B200 TileRT 仍是文档明确保留的例外,不受此前 GLM-5/5.1 和 1k1k 退役范围限制。有条件的 A/B 基线退役仍待执行;对 Pareto 前沿有贡献的非投机解码配置继续受支持。进一步的路由核查还移除了启动器和运行时设置中遗留的退役模型分支及 GLM-5 专用环境覆盖,并修正了仍推荐退役配置的工作流和智能体指南。SPEED-Bench 采集器、历史结果读取逻辑及明确保留的配方 YAML 继续保留。弃用配置现统一归档至 [`configs/deprecated/amd-master.yaml`](configs/deprecated/amd-master.yaml) 和 [`configs/deprecated/nvidia-master.yaml`](configs/deprecated/nvidia-master.yaml)。

**单节点切换为仅使用 SRT(2026-09-22):** 活跃的单节点定长配方现统一使用 SRT-Slurm。两个仅支持 Docker 的 Qwen3.5 RTX PRO 6000 FP4 配置(启用和关闭 MTP)已退役,原始设置保留在 `configs/deprecated/nvidia-master.yaml`,脚本保留在 `benchmarks/single_node/fixed_seq_len/deprecated/`。已移除不再使用的 `rtx6000pro-lat` runner 映射、启动器及运行时设置。Qwen3.5 在其他受支持的 Slurm 池上继续启用;AgentX 和多节点覆盖保持不变。

## 场景

| 场景 | ISL/OSL | 状态 |
Expand Down
34 changes: 30 additions & 4 deletions benchmarks/benchmark_lib.sh
Original file line number Diff line number Diff line change
Expand Up @@ -425,6 +425,24 @@ GPU_METRICS_CSV="${GPU_METRICS_CSV:-gpu_metrics.csv}"
NVIDIA_GPU_MONITOR_QUERY="timestamp,index,power.draw,temperature.gpu,clocks.current.sm,clocks.current.memory,utilization.gpu,utilization.memory"
export GPU_METRICS_CSV

# Keep one AMD CSV header and forward each complete row immediately. Some awk
# implementations buffer pipe input even with fflush(), losing the final ticks
# when the monitor stops.
_filter_amd_smi_metrics() {
local line header_seen=false
while IFS= read -r line; do
if [[ "$line" == timestamp,* ]]; then
if [[ "$header_seen" == true ]]; then
continue
fi
header_seen=true
fi
if [[ "$header_seen" == true ]]; then
printf '%s\n' "$line"
fi
done
}

# Background nvidia-smi/amd-smi sampler writing CSV.
# Usage: start_gpu_monitor [--output /path/to/output.csv] [--interval 1]
start_gpu_monitor() {
Expand Down Expand Up @@ -457,10 +475,9 @@ start_gpu_monitor() {
elif command -v amd-smi &>/dev/null; then
GPU_MONITOR_VENDOR="amd"
# amd-smi is Python and block-buffers stdout; without PYTHONUNBUFFERED the
# trailing ticks were lost at kill (measured on MI355X). awk keeps the first
# CSV header, drops repeated ones, and flushes every row for the same reason.
# trailing ticks were lost at kill (measured on MI355X).
PYTHONUNBUFFERED=1 amd-smi metric -p -c -t -u -w "$interval" --csv 2>/dev/null \
| awk '/^timestamp,/{if(!h){print;h=1};next} h{print;fflush()}' > "$output" &
| _filter_amd_smi_metrics > "$output" &
GPU_MONITOR_PID=$!
# Hardware energy-accumulator + identity snapshots; the end-side twin in
# stop_gpu_monitor lets auditors cross-check the integrated energy
Expand Down Expand Up @@ -800,6 +817,7 @@ run_benchmark_serving() {
local model=""
local port=""
local backend=""
local base_url=""
local endpoint=""
local input_len=""
local output_len=""
Expand Down Expand Up @@ -834,6 +852,10 @@ run_benchmark_serving() {
endpoint="$2"
shift 2
;;
--base-url)
base_url="$2"
shift 2
;;
--input-len)
input_len="$2"
shift 2
Expand Down Expand Up @@ -954,12 +976,16 @@ run_benchmark_serving() {
num_prompts="$max_concurrency"
fi

if [[ -z "$base_url" ]]; then
base_url="http://0.0.0.0:$port"
fi

local benchmark_cmd=(
env PYTHONPATH="$workspace_dir${PYTHONPATH:+:$PYTHONPATH}"
python3 -m infx.bench_serving.benchmark_serving
--model "$model"
--backend "$backend"
--base-url "http://0.0.0.0:$port"
--base-url "$base_url"
--dataset-name random
--random-input-len "$input_len"
--random-output-len "$output_len"
Expand Down
102 changes: 0 additions & 102 deletions benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh

This file was deleted.

114 changes: 0 additions & 114 deletions benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200_mtp.sh

This file was deleted.

Loading
Loading