diff --git a/.agents/skills/debug-runs/SKILL.md b/.agents/skills/debug-runs/SKILL.md index da9ed176bb..5ca087db03 100644 --- a/.agents/skills/debug-runs/SKILL.md +++ b/.agents/skills/debug-runs/SKILL.md @@ -58,7 +58,7 @@ For a **single config** (tightest CI loop, skips the rest of the matrix), dispat gh workflow run e2e-tests.yml -f generate-cli-command="test-config --config-key --config-file " -f test-name="debug " ``` -(`generate-cli-command` is the required input. Its config paths are relative to +(`generate-cli-command` carries the matrix selection; the workflow marks it `required: false` because the trusted changelog-dispatch (Klaud) mode omits it, but a manual dispatch without it fails at setup. Its config paths are relative to `inferencex-e2e/`, for example `configs/nvidia-master.yaml`. `--target` is NOT a real arg.) ### 2. Monitor continuously @@ -104,9 +104,14 @@ Steps: 1. Use the job or runner name to identify the node. Look up that cluster's access details in the canvas, then SSH in with `ssh -A` when a jumpbox or agent forwarding is involved. -2. Reproduce the exact benchmark the launcher runs. Read `inferencex-e2e/runners/launch_.sh` for - the image, container mounts, and the `inferencex-e2e/benchmarks/single_node/<...>.sh` command and env - (`IMAGE`, `TP`, `PRECISION`, `EXP_NAME`, `SPEC_DECODING`, …). On Slurm clusters, use +2. Reproduce the exact benchmark the launcher runs. Single-node jobs take the + `native-single-node` path in `inferencex-e2e/runners/launch_.sh`, which calls + `launch_srt_single_node ` in `inferencex-e2e/runners/slurm_utils.sh`: it validates the + master row's `srt-recipe:` (`inferencex-e2e/benchmarks/single_node/srt-slurm-recipes///-[-mtp]/.yaml`) + with `python3 -m infx.srt_slurm.single_node prepare`, then submits it through srtctl + with the `inferencex-e2e/runners/srt-slurm/.yaml` profile. Read the recipe for the image + (`model.container`), server args and env, and the launcher for mounts and the job env + (`IMAGE`, `TP`, `PRECISION`, `SPEC_DECODING`, `CONC`, …). On Slurm clusters, use `salloc` or `srun` with the squash image. On the **bare-metal `-tw` pools, use `docker run`** on the node directly without `srun`. 3. **Always diff against a working node or working SKU** for reference. Most node failures diff --git a/.claude/commands/find-mergeable-claude-prs.md b/.claude/commands/find-mergeable-claude-prs.md index 337aa8517b..0bfab26004 100644 --- a/.claude/commands/find-mergeable-claude-prs.md +++ b/.claude/commands/find-mergeable-claude-prs.md @@ -2,16 +2,16 @@ description: Find Claude-authored PRs with all-green full-sweep validation and confirm before merging --- -Find open PRs authored by Claude (branches starting with `claude/`) whose full-sweep validation has completed all-green, then prompt the user before merging. +Find open PRs authored by Claude (branches starting with `klaud/`, `klaud-cold/`, or the legacy `klaude/`) whose full-sweep validation has completed all-green, then prompt the user before merging. -## Step 1 — list candidate `claude/*` PRs +## Step 1 — list candidate Klaud PRs `gh pr list --json statusCheckRollup` truncates each PR's rollup, so it can't be trusted for the per-check filter. Use it only to get the candidate numbers, then re-query each PR individually. ```bash gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \ --json number,title,headRefName \ - --jq '.[] | select(.headRefName | startswith("claude/")) | .number' \ + --jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | .number' \ > /tmp/claude_pr_candidates.txt ``` diff --git a/.claude/commands/fix-klaud-cron-prs.md b/.claude/commands/fix-klaud-cron-prs.md index 3aeb86dc45..5c8e52b312 100644 --- a/.claude/commands/fix-klaud-cron-prs.md +++ b/.claude/commands/fix-klaud-cron-prs.md @@ -2,14 +2,14 @@ description: Triage failing Claude-authored PRs by reading sweep logs, debugging failures, and pushing a candidate fix per PR --- -For each open Claude-authored PR (`claude/*` branch) whose full-sweep validation produced at least one **FAILED** check, fetch the failing run's logs, diagnose the root cause, and push a candidate fix to the PR's branch. +For each open Claude-authored PR (`klaud/*`, `klaud-cold/*`, or legacy `klaude/*` branch) whose full-sweep validation produced at least one **FAILED** check, fetch the failing run's logs, diagnose the root cause, and push a candidate fix to the PR's branch. This command modifies remote PR branches. **Pause for user confirmation** after listing the candidate PRs and again before pushing each fix. -## Step 1 — find failing `claude/*` PRs whose sweep actually ran +## Step 1 — find failing Klaud PRs whose sweep actually ran A PR qualifies only if: -- `headRefName` starts with `claude/` +- `headRefName` starts with `klaud/`, `klaud-cold/`, or `klaude/` - At least one `Run Sweep` check has conclusion `SUCCESS` **or** `FAILURE` (i.e. the sweep was enabled and produced real results, rather than all checks being skipped) - At least one check has conclusion `FAILURE`, `CANCELLED`, or `TIMED_OUT` @@ -18,7 +18,7 @@ A PR qualifies only if: ```bash gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \ --json number,title,headRefName \ - --jq '.[] | select(.headRefName | startswith("claude/")) | "\(.number)\t\(.headRefName)\t\(.title)"' \ + --jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | "\(.number)\t\(.headRefName)\t\(.title)"' \ > /tmp/claude_pr_candidates.tsv : > /tmp/claude_prs_failing.tsv @@ -79,7 +79,7 @@ If the log file is very large (>2000 lines), grep it for the actual error signat ### 2c. Diagnose -Inspect the PR diff (`git -C "$WT" diff origin/main...HEAD`) and the failing-log excerpts together. Most `claude/issue-1154-*` PRs are image-bump PRs that touch a `*.yaml` recipe. Failures are usually: +Inspect the PR diff (`git -C "$WT" diff origin/main...HEAD`) and the failing-log excerpts together. Most failing Klaud PRs are image-bump PRs (e.g. `klaud/-` from `/nuke`) that touch a `*.yaml` recipe. Failures are usually: - Image tag typo / unavailable tag → fix the image reference. - Engine arg incompatibility with new image version → add/remove the affected flag in the recipe. diff --git a/.claude/commands/klaud-pr-status-html.md b/.claude/commands/klaud-pr-status-html.md index 9cdc200804..c7f070f7c4 100644 --- a/.claude/commands/klaud-pr-status-html.md +++ b/.claude/commands/klaud-pr-status-html.md @@ -2,18 +2,18 @@ description: Render an HTML dashboard of Claude/Klaud-Cold PR states (state + check breakdown per PR) and open it in the browser --- -Render an HTML dashboard for every open PR in `SemiAnalysisAI/InferenceX` that was opened by Claude (either a `claude/*` branch OR a title containing `[Klaud Cold]`). Each row shows the PR's current state, a check-status breakdown, the title, and empty "Reason"/"Suggested fix" cells you can fill in afterward by reading failed-run logs. +Render an HTML dashboard for every open PR in `SemiAnalysisAI/InferenceX` that was opened by Claude (either a `klaud/*`, `klaud-cold/*`, or legacy `klaude/*` branch OR a title containing `[Klaud Cold]`). Each row shows the PR's current state, a check-status breakdown, the title, and empty "Reason"/"Suggested fix" cells you can fill in afterward by reading failed-run logs. The dashboard lives at `/tmp/klaud_pr_status.html` and is opened with `open` (macOS) at the end. -## Step 1 — list candidate PRs (`claude/*` OR title containing `[Klaud Cold]`) +## Step 1 — list candidate PRs (`klaud/*` / `klaud-cold/*` / `klaude/*` OR title containing `[Klaud Cold]`) The title check uses `contains` (not `startswith`) so it picks up PRs whose titles embed `[Klaud Cold]` after a prefix like `[Handoff to @Oseltamivir Claude /loop]`. Handoff-style PRs from a /loop run still belong on the dashboard. ```bash gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \ --json number,title,headRefName,createdAt \ - --jq '.[] | select((.headRefName | startswith("claude/")) or (.title | contains("[Klaud Cold]"))) | "\(.number)\t\(.headRefName)\t\(.createdAt)\t\(.title)"' \ + --jq '.[] | select((.headRefName | test("^(klaud|klaud-cold|klaude)/")) or (.title | contains("[Klaud Cold]"))) | "\(.number)\t\(.headRefName)\t\(.createdAt)\t\(.title)"' \ > /tmp/klaud_pr_candidates.tsv wc -l /tmp/klaud_pr_candidates.tsv ``` @@ -65,7 +65,7 @@ while IFS=$'\t' read -r pr branch created title; do | select(.workflowName == "Run Sweep") | state as $s | select($s == "QUEUED" or $s == "IN_PROGRESS") - | ((.name | capture("(?

b200|b300|h100|h200|mi300x|mi325x|mi355x)").p) // "unknown") as $pool + | ((.name | capture("(?

gb200|gb300|b200|b300|h100|h200|mi300x|mi325x|mi355x)").p) // "unknown") as $pool | "\($pr)\t\($pool)\t\($s)" ' >> /tmp/klaud_pr_jobs.tsv done < /tmp/klaud_pr_candidates.tsv @@ -108,14 +108,16 @@ state_class = { "NO_SWEEP": "state-NOSWEEP", "NO_SUCCESS": "state-NOSWEEP", } -# Active self-hosted runner counts per pool (from GHA registrations as of the -# session this was last touched). If a pool isn't listed, falls back to 4 — a -# conservative guess. To refresh: +# Self-hosted runner counts per pool, derived from the runner-label map in +# inferencex-e2e/configs/runners.yaml (run this from the repo root). If a pool isn't listed, +# falls back to 4, a conservative guess. Online counts can be lower; to check: # gh api --paginate repos/SemiAnalysisAI/InferenceX/actions/runners \ # --jq '.runners[] | select(.status == "online") | .labels[].name' \ -# | grep -xE 'b200|b300|h200|h100|mi300x|mi325x|mi355x' | sort | uniq -c -POOL_RUNNERS = {"b200": 12, "b300": 18, "h100": 19, "h200": 18, - "mi300x": 9, "mi325x": 9, "mi355x": 9} +# | grep -xE 'gb200|gb300|b200|b300|h200|h100|mi300x|mi325x|mi355x' | sort | uniq -c +import yaml +POOL_RUNNERS = {pool: len(names) for pool, names in + yaml.safe_load(Path("inferencex-e2e/configs/runners.yaml").read_text())["labels"].items() + if not pool.startswith("cluster:")} DEFAULT_POOL_RUNNERS = 4 AVG_JOB_MIN = 7 # rough sweep-job median; eval+1k1k ~5min, 8k1k+agentic ~10-15min @@ -221,7 +223,7 @@ out = ['', ' .almost-section .eta { display:inline-block; min-width:56px; color:#3a3; font-weight:600; }', '', '

Claude / [Klaud Cold] PR status — InferenceX

', -f'
Generated {now}. Source: gh pr view --json statusCheckRollup for every claude/* or [Klaud Cold]-titled open PR. Diagnoses (if any) loaded from /tmp/klaud_pr_diag.json. ETA = pool-drain pessimistic estimate (ceil(global_pool_pending / runners) × ~{AVG_JOB_MIN}m/job); GHA dispatch isn\'t guaranteed FIFO and queue position isn\'t exposed (see docs.github.com/en/actions) so this is an upper-bound ordering hint, not a SLA.
', +f'
Generated {now}. Source: gh pr view --json statusCheckRollup for every klaud/*, klaud-cold/*, klaude/* or [Klaud Cold]-titled open PR. Diagnoses (if any) loaded from /tmp/klaud_pr_diag.json. ETA = pool-drain pessimistic estimate (ceil(global_pool_pending / runners) × ~{AVG_JOB_MIN}m/job); GHA dispatch isn\'t guaranteed FIFO and queue position isn\'t exposed (see docs.github.com/en/actions) so this is an upper-bound ordering hint, not a SLA.
', f'
Global pending across all open Klaud/claude sweeps: {total_global_pending} jobs — per-pool: {html.escape(pool_pressure_summary) if pool_pressure_summary else "—"}
'] pill_specs = [("READY", "ready"), ("RUNNING", "running"), diff --git a/.claude/commands/list-claude-pr-status.md b/.claude/commands/list-claude-pr-status.md index e5425bb69b..b8aec01313 100644 --- a/.claude/commands/list-claude-pr-status.md +++ b/.claude/commands/list-claude-pr-status.md @@ -2,16 +2,16 @@ description: List Claude-authored PRs that haven't failed (ready or still running) with their actual check state --- -List open PRs authored by Claude (branches starting with `claude/`) that have **not** had any check fail. Show each PR's actual state (READY when all checks finished green, RUNNING when sweeps are still queued/in-progress) along with a per-status check breakdown, rendered as a markdown table. +List open PRs authored by Claude (branches starting with `klaud/`, `klaud-cold/`, or the legacy `klaude/`) that have **not** had any check fail. Show each PR's actual state (READY when all checks finished green, RUNNING when sweeps are still queued/in-progress) along with a per-status check breakdown, rendered as a markdown table. -## Step 1 — list candidate `claude/*` PRs +## Step 1 — list candidate Klaud PRs `gh pr list --json statusCheckRollup` truncates each PR's rollup, so it can't be trusted for the per-check filter. Use it only to get the candidate numbers, then re-query each PR individually. ```bash gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \ --json number,title,headRefName \ - --jq '.[] | select(.headRefName | startswith("claude/")) | "\(.number)\t\(.title)"' \ + --jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | "\(.number)\t\(.title)"' \ > /tmp/claude_pr_candidates.tsv ``` @@ -64,6 +64,6 @@ Print the result directly as a markdown table. READY rows first, then RUNNING. E {printf "| [#%s](https://github.com/SemiAnalysisAI/InferenceX/pull/%s) | %s | `%s` | %s |\n", $1, $1, $2, $3, $4}' ``` -If `/tmp/claude_pr_status.tsv` is empty, print: `_No claude/* PRs are currently READY or RUNNING — all open Claude PRs have failures or no sweep results._` +If `/tmp/claude_pr_status.tsv` is empty, print: `_No Klaud PRs are currently READY or RUNNING — all open Claude PRs have failures or no sweep results._` Output the resulting markdown table to the user verbatim. This command is informational only. Do **not** auto-merge. diff --git a/.claude/commands/nuke.md b/.claude/commands/nuke.md index 1700465ae0..f8c381ba86 100644 --- a/.claude/commands/nuke.md +++ b/.claude/commands/nuke.md @@ -104,6 +104,45 @@ block += [" description:", f' - "{desc}"', " pr-link: PRLINK_PLACEHOLDER"] open(f,'w').write(content + '\n' + '\n'.join(block) + '\n') ``` +`/tmp/edit_recipe_container.py` (single-node runs go through the srt-slurm +recipe each master search-space row names in `srt-recipe:`, and +`inferencex-e2e/infx/srt_slurm/single_node.py` rejects the run unless the recipe's +`model.container` equals the master `image:` exactly): +```python +#!/usr/bin/env python3 +# Usage: edit_recipe_container.py [key2 ...] +import os, re, sys +f, new_image, keys = sys.argv[1], sys.argv[2], sys.argv[3:] +# srt-recipe paths are relative to inferencex-e2e/, the parent of configs/ +root = os.path.dirname(os.path.dirname(os.path.abspath(f))) +master = open(f).read().split('\n') +recipes = set() +for key in keys: + kre = re.compile(r'^' + re.escape(key) + r':\s*$') + start = next((i for i,l in enumerate(master) if kre.match(l)), None) + if start is None: sys.exit(f"ERROR: key not found: {key}") + for j in range(start+1, len(master)): + if re.match(r'^[A-Za-z0-9._-]+:\s*$', master[j]): break # next top-level key + recipes.update(re.findall(r'srt-recipe:\s*([^\s,}]+)', master[j])) +if not recipes: sys.exit(f"ERROR: no srt-recipe for keys {keys}") +for r in sorted(os.path.join(root, x) for x in recipes): + lines = open(r).read().split('\n') + in_model, hit = False, False + for i, l in enumerate(lines): + if re.match(r'^\s*model:\s*$', l): in_model, indent = True, len(l) - len(l.lstrip()); continue + if in_model and l.strip() and len(l) - len(l.lstrip()) <= indent: in_model = False + m = re.match(r'^(\s+)container:\s*(.+?)\s*$', l) if in_model else None + if m: + lines[i] = f"{m.group(1)}container: {new_image}"; hit = True + print(f"{r}: {m.group(2)} -> {new_image}"); break + if not hit: sys.exit(f"ERROR: no model.container in {r}") + open(r, 'w').write('\n'.join(lines)) +``` + +If `grep -rn '' inferencex-e2e/configs/*-master.yaml` shows a recipe is also +referenced by a key outside this family, stop and ask the user: bumping it would +break that other key's `model.container == image` check. + For each family, run strictly sequentially because git checkouts can't be parallel: ```bash @@ -111,6 +150,7 @@ git checkout main -q && git reset --hard origin/main -q branch="klaud/-" git checkout -b "$branch" -q python3 /tmp/edit_image.py [-mtp] +python3 /tmp/edit_recipe_container.py [-mtp] python3 /tmp/append_changelog.py inferencex-e2e/perf-changelog.yaml "" [-mtp] git add -A git commit -q -m "[Klaud Cold] Update [ (+mtp)] to " @@ -119,10 +159,14 @@ url=$(gh pr create --repo SemiAnalysisAI/InferenceX --base main --head "$branch" --title "[Klaud Cold] Update [ (+mtp)] to " \ --body "" --label full-sweep-fail-fast | grep -o 'https://github.com/[^ ]*') # patch the changelog pr-link with the real URL, then amend + force-push +# (read first, then write: open(f,'w') truncates before a same-line read runs) +before=$(wc -l < inferencex-e2e/perf-changelog.yaml) python3 - inferencex-e2e/perf-changelog.yaml "$url" <<'PY' import sys; f,u=sys.argv[1],sys.argv[2] -open(f,'w').write(open(f).read().replace("PRLINK_PLACEHOLDER",u,1)) +content = open(f).read() +open(f,'w').write(content.replace("PRLINK_PLACEHOLDER",u,1)) PY +[ "$(wc -l < inferencex-e2e/perf-changelog.yaml)" -eq "$before" ] || { echo "perf-changelog.yaml line count changed"; exit 1; } git add inferencex-e2e/perf-changelog.yaml && git commit -q --amend --no-edit && git push -q --force-with-lease ``` diff --git a/.claude/commands/recover-failed-ingest.md b/.claude/commands/recover-failed-ingest.md index 81b0a4e989..eb3f6ec180 100644 --- a/.claude/commands/recover-failed-ingest.md +++ b/.claude/commands/recover-failed-ingest.md @@ -5,7 +5,10 @@ argument-hint: [source-run-id] Recover the official database ingest for a failed or skipped InferenceX push-to-main `Run Sweep` workflow by creating a recovery PR that reuses artifacts -from an earlier PR sweep. Do not add a one-off recovery workflow. +from an earlier PR sweep. Do not add a one-off recovery workflow. The existing +`.github/workflows/recover-reused-ingest.yml` only redispatches +`ingest-agentic-results` for a failed reused **agentic** ingest (inputs +`source-run-id`, `merge-run-id`); it is not a generic fixed-sequence recovery tool. Inputs from `$ARGUMENTS`: @@ -225,9 +228,9 @@ RECOVERY_PR=$(gh pr view "$RECOVERY_PR_URL" \ --repo SemiAnalysisAI/InferenceX \ --json number --jq .number) -gh pr edit "$RECOVERY_PR" \ - --repo SemiAnalysisAI/InferenceX \ - --add-label full-sweep-fail-fast +# REST, not `gh pr edit` (projects-classic GraphQL bug) +gh api -X POST "repos/SemiAnalysisAI/InferenceX/issues/$RECOVERY_PR/labels" \ + -f "labels[]=full-sweep-fail-fast" --jq '.[].name' gh pr comment "$RECOVERY_PR" \ --repo SemiAnalysisAI/InferenceX \ --body "/reuse-sweep-run $SOURCE_RUN_ID" diff --git a/.github/workflows/README.md b/.github/workflows/README.md index a4106eebdb..9aabf8460d 100644 --- a/.github/workflows/README.md +++ b/.github/workflows/README.md @@ -16,8 +16,8 @@ positional arguments: filtering by model, precision, framework, runner type, and sequence lengths test-config Generate full sweep for specific config keys. - Supports wildcard patterns (* and ?) for matching - multiple keys at once. + Validates that all specified keys exist before + generating. options: -h, --help show this help message and exit @@ -32,12 +32,16 @@ usage: python -m infx.matrix.generate full-sweep --config-files CONFIG_FILES [CONFIG_FILES ...] [--runner-config RUNNER_CONFIG] [--no-evals | --evals-only] [--all-evals] + [--smoke] [--trim-conc] + [--runner-node-filter RUNNER_NODE_FILTER] + [--scenario-type {fixed-seq-len,agentic-coding} [{fixed-seq-len,agentic-coding} ...]] [--model-prefix MODEL_PREFIX [MODEL_PREFIX ...]] [--precision PRECISION [PRECISION ...]] [--framework FRAMEWORK [FRAMEWORK ...]] [--runner-type RUNNER_TYPE [RUNNER_TYPE ...]] [--seq-lens {1k1k,8k1k} [{1k1k,8k1k} ...]] [--step-size STEP_SIZE] + [--min-conc MIN_CONC] [--max-conc MAX_CONC] [--max-tp MAX_TP] [--max-ep MAX_EP] @@ -101,8 +105,13 @@ usage: python -m infx.matrix.generate test-config --config-files CONFIG_FILES [CONFIG_FILES ...] [--runner-config RUNNER_CONFIG] [--no-evals | --evals-only] [--all-evals] + [--smoke] [--trim-conc] + [--runner-node-filter RUNNER_NODE_FILTER] + [--scenario-type {fixed-seq-len,agentic-coding} [{fixed-seq-len,agentic-coding} ...]] --config-keys CONFIG_KEYS [CONFIG_KEYS ...] [--conc CONC [CONC ...]] + [--exp-names EXP_NAMES [EXP_NAMES ...]] + [--seq-lens {1k1k,8k1k} [{1k1k,8k1k} ...]] ``` Config keys support **wildcard patterns** using `*` (matches any characters) and `?` (matches a single character). Patterns that match no keys will raise an error. diff --git a/AGENTS.md b/AGENTS.md index 6363018c85..3997478571 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -124,7 +124,7 @@ Deleting a test that fails these questions needs no replacement. Do not preserve - Every priority-scheduled benchmark job on a self-hosted cluster must request exactly one `nodes:N` label, where `N` is the positive integer number of physical Slurm nodes required. Single-node jobs use `nodes:1`; generated multi-node jobs must forward their computed `node-count`. A queued job missing this label is ineligible for priority scheduling, and labels cannot be added retroactively, so fix the source branch and dispatch a new run. - Every change that can affect benchmark performance and every recipe addition or modification requires a new `inferencex-e2e/perf-changelog.yaml` entry. The file is append-only and byte-sensitive. Preserve all existing bytes and separator whitespace, and append only at the tail. - Multi-node srt-slurm changes update the recipe YAML and matching master config together. For image bumps, `model.container` must equal `image`. -- Every `*_mtp.sh` passes `--use-chat-template` to `run_benchmark_serving`. +- Every speculative fixed-sequence benchmark renders prompts with the chat template: single-node srt-slurm recipes that speculate set `benchmark.env.USE_CHAT_TEMPLATE: "true"` (enforced by `inferencex-e2e/infx/srt_slurm/single_node.py::validate_recipe`), which `srt_fixed_sequence.sh` turns into `--use-chat-template` for `run_benchmark_serving`. - Benchmarks create no new directories under `/workspace`. Root containers must not leave root-owned files in shared AMD runner workspaces. - Generated configuration is not runtime proof. Run the narrowest local check, then the applicable smoke, sweep, or eval procedure from [`inferencex-e2e/docs/procedures.md`](inferencex-e2e/docs/procedures.md). diff --git a/collectivex/README.md b/collectivex/README.md index bc2406557f..785f3f4d3e 100644 --- a/collectivex/README.md +++ b/collectivex/README.md @@ -38,7 +38,7 @@ in one of two modes: H100/H200 and at EP8 *and EP16* on B200 (the nscale bare-metal pool, whose gdrdrv-backed IBGDA over native IB is what a low-latency scale-out needs and no virtualized pool has) and on GB200/GB300, whose EP16 stays inside the MNNVL scale-up domain, - plus MoRI EP8 on MI300X/MI325X/MI355X and UCCL-EP EP8 on H100/H200/B200 only (UCCL's low-latency host + plus MoRI EP8 on MI355X only (the MI300X/MI325X registries carry no low-latency rows) and UCCL-EP EP8 on H100/H200/B200 only (UCCL's low-latency host assert `kNumMaxTopK + 1 <= num_warp_groups * num_warps_per_group` cannot hold on AMD, where `kNumMaxWarpGroups` is 16, since upstream raised `kNumMaxTopK` 9 -> 16 (uccl#1016, 2026-07-13) and our pin is six days later. The product is 16 for every CU count, so this is a dated regression @@ -62,7 +62,10 @@ samples per component) with 32 synchronized full roundtrip warmups before each m every trial/point. Component measurement order rotates each trial so every timed component occupies every position in the sequence, and each iteration takes the cross-rank maximum before nearest-rank p50/p90/p95/p99. A keyed BLAKE2b counter produces -byte-identical routing and gate weights on every runtime. +byte-identical routing and gate weights on every runtime. Graph-compatible backend/mode pairs (mostly decode; MoRI and prefill +stay eager) run these measurements under CUDA graph replay by default, where graphed fresh-entry +components publish only p50; `COLLX_CUDA_GRAPH=0` restores eager (see the methodology's CUDA Graph +Replay section). Those components all measure **fresh entry** (the GPU is drained around every timed window), the latency of an idle pipeline, not what a decode loop pays. So every row also carries the **chained diff --git a/collectivex/docs/methodology.md b/collectivex/docs/methodology.md index b75344bdbc..5ca666e4c3 100644 --- a/collectivex/docs/methodology.md +++ b/collectivex/docs/methodology.md @@ -126,7 +126,7 @@ and identical traffic, while its EP8 rows are correct. The deficit is confined t sustains ~34 GB/s per node against a nominal 8x400G (~4.2 GB/s per GPU-NIC pair) where bare-metal h100 reaches wire rate. Reordering the NIC-PE mapping to pair each rank with its socket-local NIC changed nothing (478µs against a 480µs baseline), which rules the selector out and points at the -GDR path being degraded wholesale inside the guest. The retired b200-nscale pool showed the same +GDR path being degraded wholesale inside the guest. The retired b200-dgxc pool showed the same shape. Treat EP16 rows from a virtualized pool as a lower bound on the hardware until the host's ACS/IOMMU configuration is confirmed. diff --git a/collectivex/docs/swap-blocks.md b/collectivex/docs/swap-blocks.md index 86f106e8e5..cfc677df4d 100644 --- a/collectivex/docs/swap-blocks.md +++ b/collectivex/docs/swap-blocks.md @@ -58,12 +58,12 @@ CPU-only machines run measurement/mapping tests and skip the real GPU test. Select `backend: swap-blocks` in **CollectiveX Sweep**, or dispatch: ```bash -gh workflow run collectivex-sweep.yml --ref codex/collectivex-swap-blocks \ +gh workflow run collectivex-sweep.yml --ref main \ -f backend=swap-blocks -f swap_profile=smoke \ -f swap_image=vllm/vllm-openai:v0.25.1 ``` -Use `--ref main` after merge. Blank `only_sku` selects the nine current Slurm GPU pools; +Blank `only_sku` selects the nine current Slurm GPU pools; set it to `h200-dgxc`, `h100-dgxc`, `b200-nscale`, `b300`, `gb200`, `gb300`, `mi300x`, `mi325x`, or `mi355x` for an isolated GPU sweep. `exclude_skus` accepts a comma-separated exclusion list. Leave EP filters blank. Each cell requests diff --git a/collectivex/docs/swap-blocks_zh.md b/collectivex/docs/swap-blocks_zh.md index f88e4909da..c5c1dd004d 100644 --- a/collectivex/docs/swap-blocks_zh.md +++ b/collectivex/docs/swap-blocks_zh.md @@ -52,12 +52,12 @@ python3 -m unittest discover collectivex/tests -p 'test_swap_blocks.py' -v 在 **CollectiveX Sweep** 中选择 `backend: swap-blocks`,或执行: ```bash -gh workflow run collectivex-sweep.yml --ref codex/collectivex-swap-blocks \ +gh workflow run collectivex-sweep.yml --ref main \ -f backend=swap-blocks -f swap_profile=smoke \ -f swap_image=vllm/vllm-openai:v0.25.1 ``` -合并后使用 `--ref main`。多平台运行和镜像选择见下方说明;`all` 仍仅运行 EP。 +多平台运行和镜像选择见下方说明;`all` 仍仅运行 EP。 `smoke` 覆盖三个方向、两种布局、257/4096/65536/262144 字节(最大 256 KiB)的块大小及 1/4/16/64/256/1024/2048 个块, diff --git a/experimental/README.md b/experimental/README.md index f39dfc4afe..62dc26b7f4 100644 --- a/experimental/README.md +++ b/experimental/README.md @@ -2,4 +2,4 @@ This folder contains experimental WIP code that is mostly Claude Code generated. -**Warning:** Code in this directory is very basic and likely contains errors or incomplete implementations. It is not intended for production use or as part of the official InferenceMAX results. +**Warning:** Code in this directory is very basic and likely contains errors or incomplete implementations. It is not intended for production use or as part of the official InferenceX results. diff --git a/experimental/token_position_decode_slo/README.md b/experimental/token_position_decode_slo/README.md index a1bc0538eb..946ed6a1ff 100644 --- a/experimental/token_position_decode_slo/README.md +++ b/experimental/token_position_decode_slo/README.md @@ -2,7 +2,7 @@ This folder contains experimental WIP code that is mostly Claude Code generated. -Warning: Code in this directory is very basic and likely contains errors or incomplete implementations. It is not intended for production use or as part of the official InferenceMAX results. +Warning: Code in this directory is very basic and likely contains errors or incomplete implementations. It is not intended for production use or as part of the official InferenceX results. image diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md index 78bce413df..371494a2ea 100644 --- a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES.md @@ -4,7 +4,7 @@ InferenceX owns the recipes in this directory. Every NVIDIA srt-slurm launcher uses `setup_srt_slurm()` in [`runners/slurm_utils.sh`](../../../runners/slurm_utils.sh), makes a job-local Git clone of the pinned submodule, and copies this entire tree into `recipes/`. The shared helper records the actual revision in `srt-slurm-sha.txt`; power lanes copy that revision into `power-producer-sha.txt` for result validation. -The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.2.1](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.2.1) (`984180e5b8755aef85e9995048b5a16cb5336bce`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers. +The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.23.2](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.23.2) (`8dace5f9596907a5075bf056251563b2e9563e7d`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers. InferenceX requires srt-slurm 2.0 or newer and `schema: 2` recipes. Legacy recipe layouts are unsupported; migrate them before adding them to this tree. @@ -25,7 +25,7 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml - Name override bundles `*-variants.yaml`. Multi-node AgentX recipes that differ only per configuration share one bundle per master-config entry, usually `agg-variants.yaml` or `disagg-variants.yaml`: `base` holds the shared settings and each former recipe becomes a named `override_` block holding only its differences (plain overrides, not `zip_override_*`). Master entries select one with `CONFIG_FILE=recipes//.yaml:override_`. Recipes read as text by a launcher, such as power recipes with top-level `telemetry:`, stay standalone. Keep distinct sweep entry files separate even when their contents match: recipe paths participate in eval grouping. The Qwen3.5 `*-stp-sweep.yaml` and `*-mtp-sweep.yaml` pair preserves that existing distinction. - Update `CONFIG_FILE` and `EVAL_CONFIG_FILE` references in active and deprecated master configs, launcher path rules, workflow filters, and local documentation together when moving a file. Preserve upstream source URLs as provenance and leave historical performance-changelog entries unchanged. No aliases for the old layout are provided. -Shared runtime assets stay under `configs/` beside the model directories; they are not standalone recipes. The four files in `configs/dsv4-moe-load-balancer-configs/` are copied verbatim from NVIDIA/srt-slurm commit `deb1dfd9934398664f92d194169c183e009da83b`, preserving the EPLB initial expert assignments used by 17 DSV4 TRT recipes. `setup_srt_slurm()` stages them into the job checkout's `configs/` directory for the recipes' bind mounts. Keeping a recipe in this tree does not activate it; the master configs determine the benchmark matrix. +Shared runtime assets stay under `configs/` beside the model directories; they are not standalone recipes. The four files in `configs/dsv4-moe-load-balancer-configs/` are copied verbatim from NVIDIA/srt-slurm commit `deb1dfd9934398664f92d194169c183e009da83b`, preserving the EPLB initial expert assignments formerly used by the DSV4 TRT recipes; no checked-in recipe currently references them. `setup_srt_slurm()` stages them into the job checkout's `configs/` directory for the recipes' bind mounts. Keeping a recipe in this tree does not activate it; the master configs determine the benchmark matrix. ## TileRT exception @@ -58,7 +58,7 @@ Install the shared pin in an isolated environment, then use its CLI: srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang # Repeat for the other model/engine directories. -# Use the pinned TileRT fork for glm5.1/tilert/. +# Use the pinned TileRT fork for tilert/ recipe directories. python -m pytest infx/tests/matrix/ -q python -m infx.matrix.generate full-sweep \ --config-files configs/nvidia-master.yaml \ @@ -69,7 +69,7 @@ Validate recipes with the exact launcher pin, including all override variants. F The initial migration also resolves compatibility issues that `srtctl migrate` cannot fix itself: -- SGLang Model Gateway recipes use `frontend.type: sglang-router`; in v2.2.1, `sglang` selects a direct worker without a router. +- SGLang Model Gateway recipes use `frontend.type: sglang-router`; in v2.23.2, `sglang` selects a direct worker without a router. - Duplicate YAML keys retain the value selected by the former PyYAML loader. - DCGM telemetry uses `collect_interval_ms: 1000` instead of `provider` and `default_frequency`. The collector derives its shutdown budget; an explicit ten-second budget is too short for the current validator. Dedicated discovery-service placement is preserved from the original recipes. The pinned upstream runtime rejects telemetry with dedicated infrastructure nodes; this remains a power compatibility blocker rather than changing the original topology to satisfy validation. H200 custom recipes declare a default concurrency that the launcher replaces before submission. - DeepSeek-V4 vLLM benchmarks use the supported `custom_tokenizer` loader. Retired `warmup_req_rate: inf` fields are removed; the current upstream client uses its fixed warmup rate of 250 requests per second. diff --git a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md index d16a4ec055..1a8c5426ec 100644 --- a/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md +++ b/inferencex-e2e/benchmarks/multi_node/srt-slurm-recipes/RECIPES_zh.md @@ -4,7 +4,7 @@ InferenceX 负责维护本目录中的配置。所有 NVIDIA srt-slurm 启动器均调用 [`runners/slurm_utils.sh`](../../../runners/slurm_utils.sh) 中的 `setup_srt_slurm()`,为作业创建固定版本子模块的本地 Git 克隆,并将整个目录复制到 `recipes/`。共享函数将实际提交记录到 `srt-slurm-sha.txt`;功耗测试路径还会将其复制到 `power-producer-sha.txt`,供结果校验使用。 -统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.2.1](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.2.1)(`984180e5b8755aef85e9995048b5a16cb5336bce`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。 +统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.23.2](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.23.2)(`8dace5f9596907a5075bf056251563b2e9563e7d`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。 InferenceX 要求 srt-slurm 2.0 或更新版本,且配置必须声明 `schema: 2`。不支持旧版配置结构;加入本目录前必须先完成迁移。 @@ -25,7 +25,7 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml - 覆盖项集合使用 `*-variants.yaml` 命名。仅在各配置间存在差异的多节点 AgentX 配置,按主配置条目合并为一个集合,通常为 `agg-variants.yaml` 或 `disagg-variants.yaml`:`base` 保存共享设置,每个原配置成为一个具名 `override_` 块,只包含其差异(使用普通覆盖项,而非 `zip_override_*`)。主配置通过 `CONFIG_FILE=recipes//.yaml:override_` 选择其一。启动器以文本方式读取的配置(例如带顶层 `telemetry:` 的功耗配置)保持独立文件。即使内容相同,也保留独立扫描入口:配置路径参与评估分组。Qwen3.5 的 `*-stp-sweep.yaml` 和 `*-mtp-sweep.yaml` 保留了这一既有区别。 - 移动文件时,同步更新当前及已弃用主配置中的 `CONFIG_FILE`、`EVAL_CONFIG_FILE`,以及启动器路径规则、工作流过滤器和本地文档。保留上游来源 URL,并保持历史性能变更日志不变。不为旧目录结构提供别名。 -共享运行时资源保留在模型目录旁的 `configs/` 中,不属于独立基准测试配置。`configs/dsv4-moe-load-balancer-configs/` 中的四个文件原样取自 NVIDIA/srt-slurm 提交 `deb1dfd9934398664f92d194169c183e009da83b`,保留了 17 个 DSV4 TRT 配置使用的 EPLB 初始专家分配。`setup_srt_slurm()` 将这些文件复制到作业仓库的 `configs/` 目录,供配置中的绑定挂载使用。将配置文件放入本目录不会启用该配置;实际基准测试矩阵由主配置决定。 +共享运行时资源保留在模型目录旁的 `configs/` 中,不属于独立基准测试配置。`configs/dsv4-moe-load-balancer-configs/` 中的四个文件原样取自 NVIDIA/srt-slurm 提交 `deb1dfd9934398664f92d194169c183e009da83b`,保留了此前 DSV4 TRT 配置使用的 EPLB 初始专家分配;目前没有已提交的配置引用这些文件。`setup_srt_slurm()` 将这些文件复制到作业仓库的 `configs/` 目录,供配置中的绑定挂载使用。将配置文件放入本目录不会启用该配置;实际基准测试矩阵由主配置决定。 ## TileRT 例外 @@ -58,7 +58,7 @@ qwen3.5/trtllm/gb300-fp4/agentx/disagg-variants.yaml srtctl migrate --verify -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang srtctl migrate --in-place -f benchmarks/multi_node/srt-slurm-recipes/dsr1/sglang # 对其他模型/引擎目录重复执行。 -# 迁移 glm5.1/tilert/ 时,使用固定提交的 TileRT 分支仓库。 +# 迁移 tilert/ 配置目录时,使用固定提交的 TileRT 分支仓库。 python -m pytest infx/tests/matrix/ -q python -m infx.matrix.generate full-sweep \ --config-files configs/nvidia-master.yaml \ @@ -69,7 +69,7 @@ python -m infx.matrix.generate full-sweep \ 本次迁移还修复了 `srtctl migrate` 无法自动处理的兼容性问题: -- SGLang Model Gateway 配置使用 `frontend.type: sglang-router`;在 v2.2.1 中,`sglang` 表示不经过路由器的独立工作进程。 +- SGLang Model Gateway 配置使用 `frontend.type: sglang-router`;在 v2.23.2 中,`sglang` 表示不经过路由器的独立工作进程。 - 对重复的 YAML 键,保留原 PyYAML 加载器实际采用的值。 - DCGM 遥测使用 `collect_interval_ms: 1000`,替代 `provider` 和 `default_frequency`。采集器自动推导退出等待时间;原先显式设置的十秒不满足当前校验要求。保留原配置中服务发现进程的专用节点部署方式。固定的上游版本不支持在专用基础设施节点上启用遥测;该功耗兼容性问题仍待解决,不通过改变原有拓扑来绕过校验。H200 自定义配置声明默认并发数,提交前由启动器替换。 - DeepSeek-V4 vLLM 基准测试使用受支持的 `custom_tokenizer` 加载器。删除已废弃的 `warmup_req_rate: inf` 字段;当前上游客户端的预热速率固定为每秒 250 个请求。 diff --git a/inferencex-e2e/configs/CONFIGS.md b/inferencex-e2e/configs/CONFIGS.md index 921eca0e00..53e6c18332 100644 --- a/inferencex-e2e/configs/CONFIGS.md +++ b/inferencex-e2e/configs/CONFIGS.md @@ -80,7 +80,7 @@ The below list describes what each field is: - `image`: The image used to serve the benchmark, e.g., `vllm/vllm-openai:v0.10.2` - `model`: The model to serve, e.g., `deepseek-ai/DeepSeek-R1-0528` -- `model-prefix`: The canonical InferenceMAX model prefix reference, i.e., `dsr1` for DeepSeek-R1 or `qwen3.5` for Qwen3.5. Consult `docs/MODELS.md` for supported model/scenario combinations. This value is used to decipher which script in `benchmarks/` should be used in order to launch the benchmark. +- `model-prefix`: The canonical InferenceX model prefix reference, i.e., `dsr1` for DeepSeek-R1 or `qwen3.5` for Qwen3.5. Consult `docs/MODELS.md` for supported model/scenario combinations. This value is used to decipher which script in `benchmarks/` should be used in order to launch the benchmark. - `runner`: This is the runner label on which to run the benchmark. This must be a valid key under `labels` in `runners.yaml`. Agentic configs must use an exact `cluster:` runner label, not a broad SKU or capacity label, so every search-space point runs on the same hardware @@ -142,7 +142,7 @@ jobs to 3600 seconds. Reusable workflow callers may override the `duration` input. Notes: -- No extra fields besides the ones listed may be specified, or else the benchmarks will fail to run. +- The fields above are the common ones, not the full schema. The Pydantic models in [`infx/matrix/validation.py`](../infx/matrix/validation.py) are the authoritative contract; they also accept fields such as `spec-decoding`, `srt-recipe`, `num-nodes`, and `require-power`, and they reject any field they do not define, which fails matrix generation. - Setting the fields above only guarantees that their values are passed as environment variables to benchmark scripts. Single-node jobs receive `PP_SIZE`, `DCP_SIZE`, and `PCP_SIZE`. Multinode jobs receive `PREFILL_PP_SIZE`, `PREFILL_DCP_SIZE`, `PREFILL_PCP_SIZE`, `DECODE_PP_SIZE`, `DECODE_DCP_SIZE`, and `DECODE_PCP_SIZE`. Actually using those variables is an implementation detail of the benchmark Bash script. ## Runners diff --git a/inferencex-e2e/docs/KLAUD_DEBUG.md b/inferencex-e2e/docs/KLAUD_DEBUG.md index be7514fe02..5b1218423b 100644 --- a/inferencex-e2e/docs/KLAUD_DEBUG.md +++ b/inferencex-e2e/docs/KLAUD_DEBUG.md @@ -23,22 +23,18 @@ Only additions to the changelog are permitted. Found deleted line: ... ``` **Root cause:** Cron-PR branches go stale. When main merges new changelog entries, the PR's local snapshot of `perf-changelog.yaml` no longer covers them, so the validator sees the missing lines as deletions. A naive rebase can also strip trailing whitespace from unrelated entries with the same effect (e.g. `pr-link: ...1311 ` → `pr-link: ...1311`). -**Fix (canonical):** +**Fix (canonical):** use the byte-preserving helper while the three conflict stages are still present (see [CI procedures](ci-procedures.md#changelog-conflict-recovery) for the full procedure): ```bash # In the PR's worktree, after `git merge origin/main` conflicts on perf-changelog.yaml: -git checkout origin/main -- perf-changelog.yaml # take main's bytes verbatim -cat >> perf-changelog.yaml < - description: - - "" - pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/ -EOF -python3 -c "import yaml; yaml.safe_load(open('perf-changelog.yaml'))" +python3 -m infx.workflows.prepare_perf_changelog_merge resolve-conflict \ + --changelog-file perf-changelog.yaml \ + --pr-number \ + --repo SemiAnalysisAI/InferenceX +git add perf-changelog.yaml +git commit --no-edit ``` -Do **not** try a 3-way merge of `perf-changelog.yaml`. Whitespace edits will silently re-trigger the deletion check. +The helper starts from main's bytes, re-appends only this PR's unique contributions and validates the result. Do **not** hand-merge or reformat `perf-changelog.yaml`, and stop if the helper refuses. Whitespace edits will silently re-trigger the deletion check. After committing and pushing the resolution, the synchronize run checks the changelog with the same matrix processor used by setup, then checks the reuse @@ -225,7 +221,7 @@ Or check whether any other recipe on main uses the proposed tag. If zero recipes git commit --allow-empty -m "Re-trigger sweep" git push ``` - The old run will be auto-cancelled by `workflow/cancel-sweep-on-merge` (provided the head SHA changed). + The old run will be auto-cancelled by the per-PR `concurrency` group in `run-sweep.yml` (`cancel-in-progress: true`). - For a `cancelled` run (not `failure`), use `gh run rerun ` without `--failed` to re-run everything. ### 7.1 Reuse after matrix-generation policy changes @@ -260,7 +256,7 @@ handoff remains untouched. ### 7.3 Final reusable sweeps stay draft until reporting finishes `run-sweep.yml` permits labeled same-repository drafts. After the smoke, append -the changelog, check the full matrix and apply `full-sweep-enabled` while DRAFT. +the changelog, check the full matrix and apply `full-sweep-fail-fast` while DRAFT (`full-sweep-enabled` only for a documented infrastructure exception). Only `finish` marks ready after full validation and final report publication. If that sweep fails, remove the label and return the PR to draft before pushing a repair, or each intermediate push starts another full sweep. The Klaud Stop @@ -290,7 +286,7 @@ are skipped; dispatch a new autosweep so recovery checks the old session first. ## 9. PR conventions for this repo - Image-bump / new-recipe PRs I open on behalf of the user (or that the user creates) get the **`[Klaud Cold]`** title prefix. -- Klaud Cold keeps targeted attempts draft and unlabeled; final validation keeps the PR draft with `full-sweep-enabled` as its sole sweep label; `finish` publishes verified results before readiness. Wait for successful completion on the exact head and reusable artifacts. See [the current Klaud guide](klaud.md); generic manual-sweep recommendations do not override this flow. +- Klaud Cold keeps targeted attempts draft and unlabeled; final validation keeps the PR draft with `full-sweep-fail-fast` as its sole sweep label; `finish` publishes verified results before readiness. Wait for successful completion on the exact head and reusable artifacts. See [the current Klaud guide](klaud.md); generic manual-sweep recommendations do not override this flow. - After any code change that shifts a PR's scope (drops a recipe, changes an image tag), **update the PR title AND body in the same step** and **verify** with `gh pr view --json title,body`. `gh pr edit` silently fails (see §8). - `uv run --extra workflows python -m infx.workflows.merge_with_reuse ` is the merge entrypoint. It handles the `perf-changelog.yaml` auto-append. diff --git a/inferencex-e2e/docs/MODELS.md b/inferencex-e2e/docs/MODELS.md index d2212e1998..cc2f74631d 100644 --- a/inferencex-e2e/docs/MODELS.md +++ b/inferencex-e2e/docs/MODELS.md @@ -12,7 +12,7 @@ InferenceX-e2e runs on a fixed, limited pool of GPUs and is maintained by a smal **Monday, August 3, 2026** is the last day for the scenarios and precisions in the first table below. They are deprecated after that date. The separate A/B baseline retirement is described below that table. -**Partially enacted on 2026-08-04** in [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493): the scenario and precision retirements in the first table were carried out. This removed 54 config keys from the active master configs and archived them under `configs/deprecated/`, with their benchmark scripts moved to the sibling `deprecated/` directories. The speculative-decoding A/B retirements in the second table are **not yet enacted**. See the note under that table. +**Partially enacted on 2026-08-04** in [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493): the scenario and precision retirements in the first table were carried out. This removed 54 config keys from the active master configs and archived them under `configs/deprecated/`, with their benchmark scripts moved to the sibling `deprecated/` directories. These archives have since been deleted (`configs/deprecated/` in [#3463](https://github.com/SemiAnalysisAI/InferenceX/pull/3463); the single-node script folders in [#3461](https://github.com/SemiAnalysisAI/InferenceX/pull/3461) and [#3464](https://github.com/SemiAnalysisAI/InferenceX/pull/3464)), so past settings live in git history and `perf-changelog.yaml`. The speculative-decoding A/B retirements in the second table are **not yet enacted**. See the note under that table. Scenario and precision retirements: @@ -42,7 +42,7 @@ The A/B retirements below concern standalone non-spec-decode baselines maintaine **Thursday, August 6, 2026** is the last day for the **Single-turn 8k1k** scenario on **Kimi-K2.5/2.6/2.7-Code** (`kimik2.5`). The scenario is deprecated for these models after that date. Rationale: Kimi-K3 launched on July 27, 2026, so GPU cluster time shifts to the newer frontier model. Combined with the Agentic coding deprecation above, this leaves `kimik2.5` with no active scenario. The model is **fully retired after August 6, 2026**. -**Enacted on 2026-08-07** in [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527): 17 `kimik2.5` config keys were removed from the active master configs and archived under `configs/deprecated/`, now consolidated in `configs/deprecated/nvidia-master.yaml` (10) and `configs/deprecated/amd-master.yaml` (7), and their 12 benchmark scripts were moved to the sibling `deprecated/` directories. `kimik2.5` now has **no active configuration in any master config** and is fully retired. The same PR archived `kimik2.5-int4-h100-vllm`, an agentic-coding key that #2493 left behind in `nvidia-master.yaml` after moving its script to `benchmarks/single_node/agentic/deprecated/`. It is now in `configs/deprecated/nvidia-master.yaml` with its siblings. The SPEED-Bench acceptance-length script `benchmarks/single_node/speedbench/kimik2.5_fp4_b300_vllm.sh` is intentionally kept. Speedbench is driven by `speedbench-al.yml`, not the master configs, matching how #2493 treated MiniMax-M3. +**Enacted on 2026-08-07** in [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527): 17 `kimik2.5` config keys were removed from the active master configs and archived under `configs/deprecated/`, later consolidated in `configs/deprecated/nvidia-master.yaml` (10) and `configs/deprecated/amd-master.yaml` (7), and their 12 benchmark scripts were moved to the sibling `deprecated/` directories. `kimik2.5` now has **no active configuration in any master config** and is fully retired. The same PR archived `kimik2.5-int4-h100-vllm`, an agentic-coding key that #2493 left behind in `nvidia-master.yaml` after moving its script to `benchmarks/single_node/agentic/deprecated/`. It was archived in `configs/deprecated/nvidia-master.yaml` with its siblings. The SPEED-Bench acceptance-length script `benchmarks/single_node/speedbench/kimik2.5_fp4_b300_vllm.sh` was kept at the time; it was later deleted in [#3471](https://github.com/SemiAnalysisAI/InferenceX/pull/3471). Speedbench is driven by `speedbench-al.yml`, not the master configs, matching how #2493 treated MiniMax-M3. ### Tuesday, September 8, 2026 @@ -54,11 +54,11 @@ The A/B retirements below concern standalone non-spec-decode baselines maintaine Rationale: `dsv4` carries the largest single-turn footprint in the repository. 45 active config keys use the 8k1k scenario, 32 in `configs/nvidia-master.yaml` and 13 in `configs/amd-master.yaml`, spanning H200, B200, B300, GB200, GB300, MI300X, MI325X, and MI355X across vLLM, SGLang, TensorRT-LLM, ATOM, Dynamo, and llm-d. That is a large share of every full sweep. AgentX trace replay is the scenario AI labs and the ML community ask about, and DeepSeek-V4-Pro's 19 agentic config keys are the part of `dsv4` that feeds the published North Star Pareto frontier. Retiring the fixed-sequence-length arm frees cluster hours for AgentX and for new frontier models such as Qwen3.8-Flash-Next without reducing what we publish for this model. Single-turn 8k1k stays active for the other models that still list it. -**Enacted on 2026-09-09** in [#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921): 46 `dsv4` 8k1k config keys were removed from the active master configs and archived under `configs/deprecated/`, now consolidated in `configs/deprecated/nvidia-master.yaml` (33) and `configs/deprecated/amd-master.yaml` (13), and their 28 benchmark scripts were moved to the sibling `deprecated/` directories (25 under `benchmarks/single_node/fixed_seq_len/`, 3 under `benchmarks/multi_node/`), matching how [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493) and [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527) were carried out. The count is 46 rather than the 45 quoted above because `dsv4-fp4-b200-dynamo-sglang` landed after this notice was written. The 19 agentic-coding keys are untouched: `dsv4` continues to run and publish with agentic coding as its only scenario. The SPEED-Bench acceptance-length scripts for `dsv4` are intentionally kept. Speedbench is driven by `speedbench-al.yml`, not the master configs. The srt-slurm and llm-d recipe YAMLs referenced by the archived multi-node keys stay in place as inert reference data, as #2493 and #2527 left theirs. +**Enacted on 2026-09-09** in [#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921): 46 `dsv4` 8k1k config keys were removed from the active master configs and archived under `configs/deprecated/`, later consolidated in `configs/deprecated/nvidia-master.yaml` (33) and `configs/deprecated/amd-master.yaml` (13), and their 28 benchmark scripts were moved to the sibling `deprecated/` directories (25 under `benchmarks/single_node/fixed_seq_len/`, 3 under `benchmarks/multi_node/`), matching how [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493) and [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527) were carried out. The count is 46 rather than the 45 quoted above because `dsv4-fp4-b200-dynamo-sglang` landed after this notice was written. The 19 agentic-coding keys are untouched: `dsv4` continues to run and publish with agentic coding as its only scenario. The SPEED-Bench acceptance-length scripts for `dsv4` are intentionally kept. Speedbench is driven by `speedbench-al.yml`, not the master configs. The srt-slurm and llm-d recipe YAMLs referenced by the archived multi-node keys stay in place as inert reference data, as #2493 and #2527 left theirs. -**Deprecation parity audit (2026-09-21):** Active master configs and benchmark-script locations match the enacted retirements above and in the support matrix below. GLM-5.1 B200 TileRT remains the documented exception to the earlier GLM-5/5.1 and 1k1k retirements. Conditional A/B baseline retirement remains pending; non-speculative Pareto contributors remain supported. The broader routing audit also removed stale retired-model branches from launchers/runtime settings and a GLM-5-only environment override, and corrected workflow/agent guidance that still recommended retired coverage. SPEED-Bench collectors, historical result readers, and the explicitly retained recipe YAMLs remain available. Deprecated configs were consolidated in `configs/deprecated/amd-master.yaml` and `configs/deprecated/nvidia-master.yaml`. +**Deprecation parity audit (2026-09-21):** Active master configs and benchmark-script locations match the enacted retirements above and in the support matrix below. GLM-5.1 B200 TileRT remains the documented exception to the earlier GLM-5/5.1 and 1k1k retirements. Conditional A/B baseline retirement remains pending; non-speculative Pareto contributors remain supported. The broader routing audit also removed stale retired-model branches from launchers/runtime settings and a GLM-5-only environment override, and corrected workflow/agent guidance that still recommended retired coverage. SPEED-Bench collectors, historical result readers, and the explicitly retained recipe YAMLs remain available. Deprecated configs were consolidated in `configs/deprecated/amd-master.yaml` and `configs/deprecated/nvidia-master.yaml`; that archive was later deleted in #3463. -**Single-node SRT-only cutover (2026-09-22):** Active single-node fixed-sequence recipes now use SRT-Slurm. The two Docker-only Qwen3.5 RTX PRO 6000 FP4 configs (with and without MTP) are retired, with their original settings preserved in `configs/deprecated/nvidia-master.yaml` and their scripts in `benchmarks/single_node/fixed_seq_len/deprecated/`. The unused `rtx6000pro-lat` runner mappings, launcher, and runtime settings are removed. Qwen3.5 remains active on the other supported Slurm pools; AgentX and multi-node coverage are unchanged. +**Single-node SRT-only cutover (2026-09-22):** Active single-node fixed-sequence recipes now use SRT-Slurm. The two Docker-only Qwen3.5 RTX PRO 6000 FP4 configs (with and without MTP) are retired, with their original settings archived in `configs/deprecated/nvidia-master.yaml` and their scripts in `benchmarks/single_node/fixed_seq_len/deprecated/` (both since deleted, in #3463 and #3464). The unused `rtx6000pro-lat` runner mappings, launcher, and runtime settings are removed. Qwen3.5 remains active on the other supported Slurm pools; AgentX and multi-node coverage are unchanged. ## Scenarios @@ -66,7 +66,7 @@ Rationale: `dsv4` carries the largest single-turn footprint in the repository. 4 |---|---|---| | Agentic coding | Long Context, Multi Turn Realistic traffic trace replay with sub agents | Active. This is the trace-replay agentic-coding benchmark (see the srt-slurm recipes under [`benchmarks/single_node/srt-slurm-recipes/`](../benchmarks/single_node/srt-slurm-recipes) and the shared client [`benchmarks/srt_agentic.sh`](../benchmarks/srt_agentic.sh)). Going forward, new models will likely be onboarded with agentic coding only. Speculative decoding may be enabled or disabled to produce the best Pareto points; a separate non-spec-decode A/B baseline is not required (see [Deprecation Notice](#deprecation-notice)). | | Single-turn 8k1k | 8192 / 1024 | Active. This is the primary fixed-sequence-length scenario. | -| Single-turn 1k1k | 1024 / 1024 | Deprecated since 2026-07-17 ([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Archived configs lived in `configs/deprecated/` until it was deleted (#3464); see git history. The GLM-5.1 B200 TileRT point added later in [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) remains active. | +| Single-turn 1k1k | 1024 / 1024 | Deprecated since 2026-07-17 ([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Archived configs lived in `configs/deprecated/` until it was deleted (#3463); see git history. The GLM-5.1 B200 TileRT point added later in [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) remains active. | | Single-turn 1k8k | 1024 / 8192 | **Deprecated for all models** since 2026-03-27 ([#911](https://github.com/SemiAnalysisAI/InferenceX/pull/911)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Configs were removed, not archived. | ## AgentX Guidelines @@ -156,7 +156,7 @@ Other offloading tiers, including NVMe KV cache offloading, are outside the init | Model architecture class | Prefix | Date added | Active scenarios | Deprecated scenarios | |---|---|---|---|---| -| DeepSeek-V4.1-Flash | `dsv41flash` | 2026-09-10 | Agentic coding (vLLM: DSpark, Engram UVA offload; SGLang: DSpark arms added per SKU from 2026-09-17; ATOM: MI355X TP2/TP4 DSpark added 2026-09-23; GPU validation pending) | | +| DeepSeek-V4.1-Flash | `dsv41flash` | 2026-09-10 | Agentic coding (vLLM: DSpark, Engram UVA offload; SGLang: DSpark arms added per SKU from 2026-09-17; ATOM: MI355X TP2/TP4 DSpark added 2026-09-23, removed in #3463; GPU validation pending) | | | GLM-5.3 | `glm5.3` | 2026-09-22 ([#3366](https://github.com/SemiAnalysisAI/InferenceX/pull/3366)) | Agentic coding (MTP only, per the Deprecation Notice) | Single-turn 1k1k (deprecated for all models before this model was added; never run) | | Qwen3.8-Flash-Next | `qwen3.8next` | 2026-08-26 ([#2742](https://github.com/SemiAnalysisAI/InferenceX/pull/2742)) | Agentic coding | | | Kimi-K3 | `kimik3` | 2026-07-27 ([#2391](https://github.com/SemiAnalysisAI/InferenceX/pull/2391)) | Agentic coding (DSpark may be disabled for better Pareto points) | Standalone non-DSpark A/B baseline (not required from day 0) | @@ -176,7 +176,7 @@ Other offloading tiers, including NVMe KV cache offloading, are outside the init ## Notes - The `Prefix` column is the canonical `model-prefix` used in `configs/*-master.yaml` and by the `full-sweep` subcommand's `--model-prefix` filter, for example `python -m infx.matrix.generate full-sweep --config-files configs/nvidia-master.yaml --model-prefix dsr1`. -- "Retired" means the model no longer has any active scenario. Retired models' configs are deleted from the master configs; the former `configs/deprecated/` archive was removed in #3464, so past settings live in git history and `perf-changelog.yaml`. +- "Retired" means the model no longer has any active scenario. Retired models' configs are deleted from the master configs; the former `configs/deprecated/` archive was removed in #3463, so past settings live in git history and `perf-changelog.yaml`. - Deprecating a precision (e.g. Qwen3.5 bf16) or one arm of an A/B pair (e.g. non-MTP) narrows a model's recipe coverage without retiring the model. The model stays listed as active as long as one scenario still runs. - `dsr1` began as the DeepSeek-V3 workflow templates in the initial repo import and was switched to DeepSeek-R1 benchmarking on 2025-08-13 (renamed `dsv3` → `dsr1` on 2025-08-20). - Adding a model? Follow [Add a model + hardware recipe](configuration-procedures.md#add-a-model--hardware-recipe) and add a row here (and in [`MODELS_zh.md`](MODELS_zh.md)) in the same PR. diff --git a/inferencex-e2e/docs/MODELS_zh.md b/inferencex-e2e/docs/MODELS_zh.md index af2f3d10e7..8a6fd5a00b 100644 --- a/inferencex-e2e/docs/MODELS_zh.md +++ b/inferencex-e2e/docs/MODELS_zh.md @@ -12,7 +12,7 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支 **2026 年 8 月 3 日(星期一)**为下方第一张表中场景与精度的最后运行日,此后即告弃用。独立 A/B 基线的下线另见该表下方说明。 -**已于 2026 年 8 月 4 日部分执行**([#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)):第一张表中的场景与精度下线已完成。此次执行从启用的 master 配置中移除 54 个配置键并归档至 `configs/deprecated/`,其基准测试脚本亦移入同级 `deprecated/` 目录。第二张表中的投机解码 A/B 下线**尚未执行**,详见该表下方说明。 +**已于 2026 年 8 月 4 日部分执行**([#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)):第一张表中的场景与精度下线已完成。此次执行从启用的 master 配置中移除 54 个配置键并归档至 `configs/deprecated/`,其基准测试脚本亦移入同级 `deprecated/` 目录。这些归档此后均已删除(`configs/deprecated/` 于 [#3463](https://github.com/SemiAnalysisAI/InferenceX/pull/3463) 删除;单节点脚本目录于 [#3461](https://github.com/SemiAnalysisAI/InferenceX/pull/3461) 和 [#3464](https://github.com/SemiAnalysisAI/InferenceX/pull/3464) 删除),历史设置保留在 Git 历史与 `perf-changelog.yaml` 中。第二张表中的投机解码 A/B 下线**尚未执行**,详见该表下方说明。 场景与精度下线: @@ -42,7 +42,7 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支 **2026 年 8 月 6 日(星期四)**为 **Kimi-K2.5/2.6/2.7-Code**(`kimik2.5`)**单轮 8k1k** 场景的最后运行日,此后该场景对这些模型弃用。原因:Kimi-K3 已于 2026 年 7 月 27 日发布,GPU 集群时间将转向更新的前沿模型。叠加上文的智能体编码弃用,`kimik2.5` 将不再有任何启用场景。该模型将于 **2026 年 8 月 6 日后完全退役**。 -**已于 2026-08-07 执行**([#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)):从启用的主配置中移除 17 个 `kimik2.5` 配置项,归档至 `configs/deprecated/`,现分别合并至 `configs/deprecated/nvidia-master.yaml`(10 个)与 `configs/deprecated/amd-master.yaml`(7 个)。对应的 12 个基准测试脚本移入同级 `deprecated/` 目录。此后 `kimik2.5` 在所有主配置中**均无启用配置**,正式完全退役。同一 PR 还归档了 `kimik2.5-int4-h100-vllm`。#2493 将其脚本移入 `benchmarks/single_node/agentic/deprecated/` 时,该智能体编码配置项被遗留在 `nvidia-master.yaml` 中,现已与同类项一并归入 `configs/deprecated/nvidia-master.yaml`。SPEED-Bench 接受长度脚本 `benchmarks/single_node/speedbench/kimik2.5_fp4_b300_vllm.sh` 予以保留。Speedbench 由 `speedbench-al.yml` 驱动,不经过主配置,与 #2493 处理 MiniMax-M3 的方式一致。 +**已于 2026-08-07 执行**([#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)):从启用的主配置中移除 17 个 `kimik2.5` 配置项,归档至 `configs/deprecated/`,后分别合并至 `configs/deprecated/nvidia-master.yaml`(10 个)与 `configs/deprecated/amd-master.yaml`(7 个)。对应的 12 个基准测试脚本移入同级 `deprecated/` 目录。此后 `kimik2.5` 在所有主配置中**均无启用配置**,正式完全退役。同一 PR 还归档了 `kimik2.5-int4-h100-vllm`。#2493 将其脚本移入 `benchmarks/single_node/agentic/deprecated/` 时,该智能体编码配置项被遗留在 `nvidia-master.yaml` 中,当时已与同类项一并归入 `configs/deprecated/nvidia-master.yaml`。SPEED-Bench 接受长度脚本 `benchmarks/single_node/speedbench/kimik2.5_fp4_b300_vllm.sh` 当时予以保留,后于 [#3471](https://github.com/SemiAnalysisAI/InferenceX/pull/3471) 删除。Speedbench 由 `speedbench-al.yml` 驱动,不经过主配置,与 #2493 处理 MiniMax-M3 的方式一致。 ### 2026 年 9 月 8 日(星期二) @@ -54,11 +54,11 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支 原因:`dsv4` 是本仓库中单轮场景占用最大的模型。当前有 45 个启用的配置项使用 8k1k 场景(`configs/nvidia-master.yaml` 32 个,`configs/amd-master.yaml` 13 个),覆盖 H200、B200、B300、GB200、GB300、MI300X、MI325X 与 MI355X,涉及 vLLM、SGLang、TensorRT-LLM、ATOM、Dynamo 与 llm-d,在每一轮完整 sweep 中占比可观。AgentX 轨迹回放才是 AI 实验室与 ML 社区真正关注的场景,而 DeepSeek-V4-Pro 的 19 个智能体编码配置项正是 `dsv4` 中支撑已发布北极星(North Star)帕累托前沿的部分。下线固定序列长度分支可为 AgentX 以及 Qwen3.8-Flash-Next 等新前沿模型腾出集群机时,同时不减少该模型对外发布的内容。对于仍列有该场景的其他模型,单轮 8k1k 保持启用。 -**已于 2026-09-09 执行**([#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921)):46 个 `dsv4` 8k1k 配置项已从启用的主配置中移除并归档至 `configs/deprecated/`,现合并至 `configs/deprecated/nvidia-master.yaml`(33 个)与 `configs/deprecated/amd-master.yaml`(13 个);对应的 28 个基准测试脚本移入同级 `deprecated/` 目录(`benchmarks/single_node/fixed_seq_len/` 下 25 个,`benchmarks/multi_node/` 下 3 个),与 [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493) 和 [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527) 的做法一致。数量为 46 而非上文所述的 45,是因为 `dsv4-fp4-b200-dynamo-sglang` 在本公告发布后才合入。19 个智能体编码配置项未做改动:`dsv4` 以智能体编码为唯一场景继续运行与发布。`dsv4` 的 SPEED-Bench 接受长度脚本予以保留。Speedbench 由 `speedbench-al.yml` 驱动,不经过主配置。已归档多节点配置项所引用的 srt-slurm 与 llm-d 配方 YAML 作为惰性参考数据原地保留,与 #2493 和 #2527 的处理一致。 +**已于 2026-09-09 执行**([#2921](https://github.com/SemiAnalysisAI/InferenceX/pull/2921)):46 个 `dsv4` 8k1k 配置项已从启用的主配置中移除并归档至 `configs/deprecated/`,后合并至 `configs/deprecated/nvidia-master.yaml`(33 个)与 `configs/deprecated/amd-master.yaml`(13 个);对应的 28 个基准测试脚本移入同级 `deprecated/` 目录(`benchmarks/single_node/fixed_seq_len/` 下 25 个,`benchmarks/multi_node/` 下 3 个),与 [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493) 和 [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527) 的做法一致。数量为 46 而非上文所述的 45,是因为 `dsv4-fp4-b200-dynamo-sglang` 在本公告发布后才合入。19 个智能体编码配置项未做改动:`dsv4` 以智能体编码为唯一场景继续运行与发布。`dsv4` 的 SPEED-Bench 接受长度脚本予以保留。Speedbench 由 `speedbench-al.yml` 驱动,不经过主配置。已归档多节点配置项所引用的 srt-slurm 与 llm-d 配方 YAML 作为惰性参考数据原地保留,与 #2493 和 #2527 的处理一致。 -**弃用状态一致性核查(2026-09-21):** 启用的主配置及基准测试脚本位置与上述已执行的退役事项和下方支持矩阵一致。GLM-5.1 B200 TileRT 仍是文档明确保留的例外,不受此前 GLM-5/5.1 和 1k1k 退役范围限制。有条件的 A/B 基线退役仍待执行;对 Pareto 前沿有贡献的非投机解码配置继续受支持。进一步的路由核查还移除了启动器和运行时设置中遗留的退役模型分支及 GLM-5 专用环境覆盖,并修正了仍推荐退役配置的工作流和智能体指南。SPEED-Bench 采集器、历史结果读取逻辑及明确保留的配方 YAML 继续保留。弃用配置曾统一归档至 `configs/deprecated/amd-master.yaml` 和 `configs/deprecated/nvidia-master.yaml`。 +**弃用状态一致性核查(2026-09-21):** 启用的主配置及基准测试脚本位置与上述已执行的退役事项和下方支持矩阵一致。GLM-5.1 B200 TileRT 仍是文档明确保留的例外,不受此前 GLM-5/5.1 和 1k1k 退役范围限制。有条件的 A/B 基线退役仍待执行;对 Pareto 前沿有贡献的非投机解码配置继续受支持。进一步的路由核查还移除了启动器和运行时设置中遗留的退役模型分支及 GLM-5 专用环境覆盖,并修正了仍推荐退役配置的工作流和智能体指南。SPEED-Bench 采集器、历史结果读取逻辑及明确保留的配方 YAML 继续保留。弃用配置曾统一归档至 `configs/deprecated/amd-master.yaml` 和 `configs/deprecated/nvidia-master.yaml`,该归档后于 #3463 删除。 -**单节点切换为仅使用 SRT(2026-09-22):** 活跃的单节点定长配方现统一使用 SRT-Slurm。两个仅支持 Docker 的 Qwen3.5 RTX PRO 6000 FP4 配置(启用和关闭 MTP)已退役,原始设置保留在 `configs/deprecated/nvidia-master.yaml`,脚本保留在 `benchmarks/single_node/fixed_seq_len/deprecated/`。已移除不再使用的 `rtx6000pro-lat` runner 映射、启动器及运行时设置。Qwen3.5 在其他受支持的 Slurm 池上继续启用;AgentX 和多节点覆盖保持不变。 +**单节点切换为仅使用 SRT(2026-09-22):** 活跃的单节点定长配方现统一使用 SRT-Slurm。两个仅支持 Docker 的 Qwen3.5 RTX PRO 6000 FP4 配置(启用和关闭 MTP)已退役,原始设置曾归档于 `configs/deprecated/nvidia-master.yaml`,脚本曾归档于 `benchmarks/single_node/fixed_seq_len/deprecated/`(二者此后分别于 #3463 和 #3464 删除)。已移除不再使用的 `rtx6000pro-lat` runner 映射、启动器及运行时设置。Qwen3.5 在其他受支持的 Slurm 池上继续启用;AgentX 和多节点覆盖保持不变。 ## 场景 @@ -66,7 +66,7 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支 |---|---|---| | 智能体编码(agentic coding) | 长上下文、多轮真实流量的轨迹回放,含子智能体(sub agents) | 启用。此场景采用基于轨迹回放的智能体编码基准测试(见 [`benchmarks/single_node/srt-slurm-recipes/`](../benchmarks/single_node/srt-slurm-recipes) 下的 srt-slurm 配方与共享客户端 [`benchmarks/srt_agentic.sh`](../benchmarks/srt_agentic.sh))。今后新模型预计将仅以智能体编码场景接入。可开启或关闭投机解码以获得最优帕累托点;不要求独立的非投机解码 A/B 基线(见[弃用公告](#弃用公告))。 | | 单轮 8k1k | 8192 / 1024 | 启用。当前主要的固定序列长度(fixed-seq-len)场景。 | -| 单轮 1k1k | 1024 / 1024 | 自 2026-07-17 起弃用([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。归档配置曾位于 `configs/deprecated/`,该目录已在 #3464 中删除,请查阅 Git 历史。后续由 [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) 加入的 GLM-5.1 B200 TileRT 测试点仍启用。 | +| 单轮 1k1k | 1024 / 1024 | 自 2026-07-17 起弃用([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。归档配置曾位于 `configs/deprecated/`,该目录已在 #3463 中删除,请查阅 Git 历史。后续由 [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) 加入的 GLM-5.1 B200 TileRT 测试点仍启用。 | | 单轮 1k8k | 1024 / 8192 | **对所有模型均已弃用**,自 2026-03-27 起([#911](https://github.com/SemiAnalysisAI/InferenceX/pull/911)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。相关配置已删除,未归档。 | ## AgentX 指南 @@ -156,7 +156,7 @@ InferenceX 支持 SGLang 和 vLLM 双方的维护者,并响应 AI 实验室和 | 模型架构类别 | 前缀 | 加入日期 | 启用场景 | 已弃用场景 | |---|---|---|---|---| -| DeepSeek-V4.1-Flash | `dsv41flash` | 2026-09-10 | 智能体编码(vLLM:DSpark、Engram UVA 卸载;SGLang:自 2026-09-17 起按 SKU 添加 DSpark 配置;ATOM:2026-09-23 添加 MI355X TP2/TP4 DSpark;GPU 待验证) | | +| DeepSeek-V4.1-Flash | `dsv41flash` | 2026-09-10 | 智能体编码(vLLM:DSpark、Engram UVA 卸载;SGLang:自 2026-09-17 起按 SKU 添加 DSpark 配置;ATOM:2026-09-23 添加 MI355X TP2/TP4 DSpark,后于 #3463 移除;GPU 待验证) | | | GLM-5.3 | `glm5.3` | 2026-09-22([#3366](https://github.com/SemiAnalysisAI/InferenceX/pull/3366)) | 智能体编码(仅 MTP,见弃用公告) | 单轮 1k1k(在本模型加入前已对所有模型弃用,从未运行) | | Qwen3.8-Flash-Next | `qwen3.8next` | 2026-08-26([#2742](https://github.com/SemiAnalysisAI/InferenceX/pull/2742)) | 智能体编码 | | | Kimi-K3 | `kimik3` | 2026-07-27 ([#2391](https://github.com/SemiAnalysisAI/InferenceX/pull/2391)) | 智能体编码(可关闭 DSpark 以获得更优帕累托点) | 独立非 DSpark A/B 基线(自第 0 天起即不要求) | @@ -176,7 +176,7 @@ InferenceX 支持 SGLang 和 vLLM 双方的维护者,并响应 AI 实验室和 ## 说明 - 「前缀」列为 `configs/*-master.yaml` 中的规范 `model-prefix`,同时用于 `full-sweep` 子命令的 `--model-prefix` 筛选参数,例如 `python -m infx.matrix.generate full-sweep --config-files configs/nvidia-master.yaml --model-prefix dsr1`。 -- 「退役」指该模型已无任何启用场景。退役模型的配置直接从主配置中删除;原 `configs/deprecated/` 归档目录已在 #3464 中删除,历史设置保留在 Git 历史与 `perf-changelog.yaml` 中。 +- 「退役」指该模型已无任何启用场景。退役模型的配置直接从主配置中删除;原 `configs/deprecated/` 归档目录已在 #3463 中删除,历史设置保留在 Git 历史与 `perf-changelog.yaml` 中。 - 弃用某一精度(如 Qwen3.5 bf16)或 A/B 对照中的某一分支(如非 MTP),只是收窄该模型的配方覆盖范围,并不等于模型退役;只要仍有一个场景在运行,该模型即继续列为启用状态。 - `dsr1` 最初以 DeepSeek-V3 workflow 模板的形式随仓库首次导入,2025-08-13 切换为 DeepSeek-R1 基准测试(2025-08-20 将 `dsv3` 重命名为 `dsr1`)。 - 新增模型时,请按[添加模型 + 硬件配方](configuration-procedures_zh.md#添加模型--硬件配方)流程操作,并在同一 PR 中同时更新本文件与 [`MODELS.md`](MODELS.md) 的表格。 diff --git a/inferencex-e2e/docs/architecture.md b/inferencex-e2e/docs/architecture.md index a0eb378cd3..45578ddee4 100644 --- a/inferencex-e2e/docs/architecture.md +++ b/inferencex-e2e/docs/architecture.md @@ -144,7 +144,7 @@ The remaining Python tools are grouped by responsibility: | `infx.datasets` | AgentX trace sampling, conversion, dataset assembly, and distribution plots | | `infx.klaud` | Klaud orchestration, lifecycle, GitHub/API adapters, and schemas | -Run commands with `python -m infx..` from `inferencex-e2e/`. Dependencies remain specific to each command; importing `infx` does not load benchmark-client or eval dependencies. Python compatibility wrappers under `utils/` have been removed. Use the canonical `infx` paths for dataset tools, AgentX aggregation and analysis, eval adapters and patches, and benchmark-client helpers. Eval documentation lives at `infx/evals/EVALS.md`, and eval tests live under `infx/tests/evals/`. Other behavioral tests, runner-provisioning shell scripts, and the external submodules remain under `utils/`. +Run commands with `python -m infx..` from `inferencex-e2e/`. Dependencies remain specific to each command; importing `infx` does not load benchmark-client or eval dependencies. Python compatibility wrappers under `utils/` have been removed. Use the canonical `infx` paths for dataset tools, AgentX aggregation and analysis, eval adapters and patches, and benchmark-client helpers. Eval documentation lives at `infx/evals/EVALS.md`, and eval tests live under `infx/tests/evals/`. Other behavioral tests live under `infx/tests/`. Only the runner-provisioning shell scripts (`utils/runner_setup/`) and the external `aiperf` and `srt-slurm` submodules remain under `utils/`. Eval adapters and patches copied into isolated environments use the actual files under `infx/evals`, so they remain standalone. Trusted workflow helpers explicitly select their tooling checkout. Fixed-sequence processing, eval-score validation, and the single-node AgentX result-validation step use the package from the workflow revision in a separate checkout, with the measured checkout as their working directory. Tooling and Python 3.12 are provisioned for eval-only jobs too. Historical measured revisions therefore do not need these helper modules. Score thresholds come from the workflow revision's packaged `infx/evals/thresholds.yaml`. @@ -412,7 +412,7 @@ Use this procedure when a row is missing, mislabeled, or unexpected. --config-keys ``` -4. **Matrix handoff:** In the `setup` job, verify the row is in the expected `single_node`, `multi_node`, `evals`, `agentic_evals`, or `multinode_evals` bucket. Confirm every required field is forwarded by the matching fan-out job. +4. **Matrix handoff:** In the `setup` job, verify the row is in the expected `single_node`, `multi_node`, `evals`, `agentic_evals`, `multinode_evals`, or `multinode_agentic_evals` bucket. Confirm every required field is forwarded by the matching fan-out job. 5. **Scheduling:** Verify the template's `runs-on` value matches the intended runner. Confirm the concrete runner name prefix resolves to an existing `runners/launch_.sh`. 6. **Runtime:** Trace the launcher branch to the exact benchmark script or external recipe. Confirm every critical matrix field reaches a consumed environment variable or command argument. 7. **Output:** Verify the workflow's required raw result exists. Then verify the expected `bmk_*`, `eval_*`, `agentic_*`, logs, or metrics artifact was uploaded. diff --git a/inferencex-e2e/docs/architecture_zh.md b/inferencex-e2e/docs/architecture_zh.md index fe579bc032..2e545dd895 100644 --- a/inferencex-e2e/docs/architecture_zh.md +++ b/inferencex-e2e/docs/architecture_zh.md @@ -144,7 +144,7 @@ flowchart LR | `infx.datasets` | AgentX 轨迹采样、转换、数据集构建及分布图 | | `infx.klaud` | Klaud 编排、生命周期、GitHub/API 适配器和模式 | -从 `inferencex-e2e/` 使用 `python -m infx..` 运行命令。依赖仍由各命令分别管理;导入 `infx` 不会加载基准测试客户端或评测依赖。`utils/` 下的 Python 兼容包装文件已删除。数据集工具、AgentX 聚合与分析、评测适配器与补丁,以及基准测试客户端辅助模块应使用规范的 `infx` 路径。评测文档位于 `infx/evals/EVALS.md`,评测测试位于 `infx/tests/evals/`。其他行为测试、运行器配置 Shell 脚本及外部子模块仍位于 `utils/`。 +从 `inferencex-e2e/`使用 `python -m infx..` 运行命令。依赖仍由各命令分别管理;导入 `infx` 不会加载基准测试客户端或评测依赖。`utils/` 下的 Python 兼容包装文件已删除。数据集工具、AgentX 聚合与分析、评测适配器与补丁,以及基准测试客户端辅助模块应使用规范的 `infx` 路径。评测文档位于 `infx/evals/EVALS.md`,评测测试位于 `infx/tests/evals/`。其他行为测试位于 `infx/tests/`。`utils/` 下仅保留运行器配置 Shell 脚本(`utils/runner_setup/`)及外部子模块 `aiperf` 和 `srt-slurm`。 复制到隔离环境中的评测适配器和补丁使用 `infx/evals` 下的实际文件,因此仍可独立运行。可信工作流辅助模块会明确选择工具代码所在的检出目录。固定序列处理、评测分数验证和单节点 AgentX 结果验证步骤使用单独检出的工作流修订版中的包,并以被测检出目录为工作目录。仅评测作业也会准备工具代码和 Python 3.12。因此,历史被测修订版无需包含这些辅助模块。分数阈值取自工作流修订版随包提供的 `infx/evals/thresholds.yaml`。 @@ -412,7 +412,7 @@ AgentX 追踪导出的体积更大,并且需要追踪发现、时间线处理 --config-keys ``` -4. **矩阵交接:** 在 `setup` 作业中,验证该行位于预期的 `single_node`、`multi_node`、`evals`、`agentic_evals` 或 `multinode_evals` 桶中。确认匹配的扇出作业转发了每个必需字段。 +4. **矩阵交接:** 在 `setup` 作业中,验证该行位于预期的 `single_node`、`multi_node`、`evals`、`agentic_evals`、`multinode_evals` 或 `multinode_agentic_evals` 桶中。确认匹配的扇出作业转发了每个必需字段。 5. **调度:** 验证模板的 `runs-on` 值与预期运行器匹配。确认具体运行器名称前缀能够解析到现有的 `runners/launch_.sh`。 6. **运行时:** 沿启动器分支追踪到确切的基准测试脚本或外部方案。确认每个关键矩阵字段均到达实际被使用的环境变量或命令参数。 7. **输出:** 验证工作流要求的原始结果存在。然后验证预期的 `bmk_*`、`eval_*`、`agentic_*`、日志或指标工件已上传。 diff --git a/inferencex-e2e/docs/ci-procedures.md b/inferencex-e2e/docs/ci-procedures.md index 45838af714..f9323105e0 100644 --- a/inferencex-e2e/docs/ci-procedures.md +++ b/inferencex-e2e/docs/ci-procedures.md @@ -50,19 +50,13 @@ These files are the contract. Follow the target ref's source rather than copying ## Source-snapshot warning -This page was authored from branch commit `0c28706b33d4a796b82f6f9c3594c19c46365575`. At that time, local `origin/main` was `de493d8597035e6692833de6189b567887968460`, and the relevant CI sources were not identical: - -- The branch-local [`e2e-tests.yml`](../../.github/workflows/e2e-tests.yml) requires `generate-cli-command` and hard-codes each matrix to `fail-fast: false`. The audited [`origin/main` version](https://github.com/SemiAnalysisAI/InferenceX/blob/de493d8597035e6692833de6189b567887968460/.github/workflows/e2e-tests.yml) makes that command conditionally optional and adds trusted-changelog dispatch, `fail-fast`, and power-validation inputs. -- The branch-local [`run-sweep.yml`](../../.github/workflows/run-sweep.yml) lacks the same-repository-head guard that `origin/main` adds before PR GPU setup. The [`trusted-external-sweep.yml` workflow](https://github.com/SemiAnalysisAI/InferenceX/blob/de493d8597035e6692833de6189b567887968460/.github/workflows/trusted-external-sweep.yml) exists on that `origin/main` snapshot but not on this branch. Do not infer external-fork secret or GPU behavior from the branch-local workflow. -- Agentic eval comments in the branch-local generator identify SWE-bench, while the audited `origin/main` generator identifies GSM8K. Inspect the target ref before describing the agentic dataset selected by `all-evals` or `evals-only`. - A `workflow_dispatch` request uses the workflow definition from its dispatch `--ref`. The separate `inputs.ref` controls what the jobs check out. Before using inputs beyond the common example below, inspect the deployed definition: ```bash gh workflow view e2e-tests.yml --repo SemiAnalysisAI/InferenceX --ref main --yaml ``` -If the target ref differs from this source snapshot, its workflow and script source wins. Do not guess that a branch-local input or fork policy is deployed. +The workflow and script source at the target ref wins. Do not guess that a branch-local input or fork policy is deployed. ## Local matrix generation @@ -578,4 +572,4 @@ use one GPU per measurement. AMD attention supports both torch and AITER. See [OperatorX GitHub Actions](../../operatorx/CI.md) for dispatch, coverage, artifacts, cancellation, and validation. -For H200 DeepSeek-V4.1 Flash SGLang AgentX performance at concurrency 64 or above, the launcher allows a 1440-minute Slurm allocation and the reusable workflow allows 1470 minutes. This accommodates normal warmup and the unchanged 3600-second profile; lower concurrencies and eval-only jobs retain the standard deadlines. Run `35775895782` exhausted the previous eight-hour allocation during progressing, error-free warmup. A failed-only retry retains the original workflow deadline, so deadline changes require a new workflow run. +For H200 DeepSeek-V4.1 Flash SGLang AgentX performance at concurrency 64 or above, the launcher allows a 1440-minute Slurm allocation and the reusable workflow allows 1470 minutes. This accommodates normal warmup and the unchanged 3600-second profile; lower concurrencies and eval-only jobs retain the standard deadlines. Run `35775895782` exhausted the previous eight-hour allocation during progressing, error-free warmup. The matching GB200 jobs get a 720-minute allocation and a 750-minute workflow deadline. A failed-only retry retains the original workflow deadline, so deadline changes require a new workflow run. diff --git a/inferencex-e2e/docs/ci-procedures_zh.md b/inferencex-e2e/docs/ci-procedures_zh.md index df64ad75ee..6023fe0f45 100644 --- a/inferencex-e2e/docs/ci-procedures_zh.md +++ b/inferencex-e2e/docs/ci-procedures_zh.md @@ -50,19 +50,13 @@ ## 源码快照警告 -本页基于分支 Commit `0c28706b33d4a796b82f6f9c3594c19c46365575` 编写。当时本地 `origin/main` 为 `de493d8597035e6692833de6189b567887968460`,相关 CI 源码并不完全相同: - -- 分支本地的 [`e2e-tests.yml`](../../.github/workflows/e2e-tests.yml) 要求提供 `generate-cli-command`,并将各矩阵硬编码为 `fail-fast: false`。审计过的 [`origin/main` 版本](https://github.com/SemiAnalysisAI/InferenceX/blob/de493d8597035e6692833de6189b567887968460/.github/workflows/e2e-tests.yml) 仅在特定条件下要求该命令,并新增受信任 Changelog 派发、`fail-fast` 与功耗验证输入。 -- 分支本地的 [`run-sweep.yml`](../../.github/workflows/run-sweep.yml) 缺少 `origin/main` 在 PR GPU Setup 之前新增的“Head 仓库必须与当前仓库相同”保护。该 `origin/main` 快照存在 [`trusted-external-sweep.yml` Workflow](https://github.com/SemiAnalysisAI/InferenceX/blob/de493d8597035e6692833de6189b567887968460/.github/workflows/trusted-external-sweep.yml),本分支则没有。不要根据分支本地 Workflow 推断外部 Fork 的 Secret 或 GPU 行为。 -- 分支本地生成器的 Agentic Eval 注释指向 SWE-bench,审计过的 `origin/main` 生成器则指向 GSM8K。在描述 `all-evals` 或 `evals-only` 选择的 Agentic 数据集之前,必须检查目标 Ref。 - `workflow_dispatch` 请求使用其派发 `--ref` 中的 Workflow 定义;单独的 `inputs.ref` 控制 Job Checkout 的内容。在使用下方公共示例以外的输入前,应检查已部署定义: ```bash gh workflow view e2e-tests.yml --repo SemiAnalysisAI/InferenceX --ref main --yaml ``` -如果目标 Ref 与本源码快照不同,以其 Workflow 和脚本源码为准。不要猜测某个分支本地输入或 Fork 策略已经部署。 +以目标 Ref 的 Workflow 和脚本源码为准。不要猜测某个分支本地输入或 Fork 策略已经部署。 ## 本地矩阵生成 @@ -560,4 +554,4 @@ attention 支持 torch 和 AITER。 触发方式、覆盖范围、产物、取消及验证说明见 [OperatorX GitHub Actions](../../operatorx/CI_zh.md)。 -H200 DeepSeek-V4.1 Flash SGLang AgentX 在并发 64 及以上的性能任务允许 1440 分钟 Slurm 分配和 1470 分钟 GitHub 任务,以容纳正常预热及保持不变的 3600 秒正式测试;更低并发和 eval-only 任务仍使用标准期限。运行 `35775895782` 在持续推进、请求无错误的预热期间耗尽了原有八小时分配。仅重试失败任务会保留原工作流期限,因此修改期限后必须启动新运行。 +H200 DeepSeek-V4.1 Flash SGLang AgentX 在并发 64 及以上的性能任务允许 1440 分钟 Slurm 分配和 1470 分钟 GitHub 任务,以容纳正常预热及保持不变的 3600 秒正式测试;更低并发和 eval-only 任务仍使用标准期限。运行 `35775895782` 在持续推进、请求无错误的预热期间耗尽了原有八小时分配。对应的 GB200 任务使用 720 分钟分配和 750 分钟 Workflow 期限。仅重试失败任务会保留原工作流期限,因此修改期限后必须启动新运行。 diff --git a/inferencex-e2e/docs/configuration-procedures.md b/inferencex-e2e/docs/configuration-procedures.md index 5f85d4c262..50986d4c72 100644 --- a/inferencex-e2e/docs/configuration-procedures.md +++ b/inferencex-e2e/docs/configuration-procedures.md @@ -140,7 +140,7 @@ STP (Single Token Prediction) is vanilla autoregressive decoding with one token 7. **Append one changelog entry** for the exact new key. See [Append the changelog safely](#append-the-changelog-safely). 8. **Validate syntax and generated output.** Inspect image, model, runner, ISL/OSL, `max-model-len`, concurrency, TP/PP/EP/DCP/PCP, and `spec-decoding`. -A `MODELS.md` row alone is not an executable recipe. The complete path is benchmark script + master entry + launcher routing + changelog trigger + generated matrix. +A `MODELS.md` row alone is not an executable recipe. The complete path is srt-slurm recipe + master entry (`srt-recipe:`) + changelog trigger + generated matrix. ## Change a master config @@ -252,19 +252,19 @@ Sources: [`AGENTS.md#non-negotiable-benchmark-invariants`](../../AGENTS.md#non-n 1. Confirm native MTP modules versus an external draft. For a draft, verify exact model ID, method (for example `eagle3`), and recommended speculative-token count from the model/upstream recipe. 2. Copy a working sibling for the same model and backend. Preserve its speculative config, attention backend, token count, model patches, and dependency setup. -3. Every `*_mtp.sh` must pass `--use-chat-template` to `run_benchmark_serving`. Raw prompts silently depress acceptance. +3. Every speculative fixed-sequence recipe variant must set `benchmark.env.USE_CHAT_TEMPLATE: "true"`; `select_recipe` rejects a speculative variant without it, and [`srt_fixed_sequence.sh`](../benchmarks/single_node/srt_fixed_sequence.sh) turns it into `--use-chat-template` for `run_benchmark_serving`. Raw prompts silently depress acceptance. 4. Size graph capture for at least `CONC * (1 + NUM_SPEC_TOKENS)`, rounded as the sibling does and capped at the framework limit (the current vLLM playbook caps at 2048). 5. Keep backend differences: do not copy CUDA-only drafter attention pins or patches into ROCm recipes. -6. Set `spec-decoding: mtp` in the relevant search-space entries and add `_mtp` launcher suffix routing. For a draft-model mode supported by the schema, use the matching generated value deliberately. Do not infer it from a filename. -7. Add script + master entry + launcher routing + changelog together. -8. Run Bash syntax and generation checks. Inspect `spec-decoding`, draft/native method, token count, chat-template use, capture range, and resolved script. +6. Set `spec-decoding: mtp` in the relevant search-space entries and point their `srt-recipe:` at the `-mtp` recipe; `select_recipe` checks the recipe's speculative config against it. For a draft-model mode supported by the schema, use the matching generated value deliberately. Do not infer it from a filename. +7. Add recipe + master entry + changelog together. +8. Run YAML and generation checks. Inspect `spec-decoding`, draft/native method, token count, chat-template use, capture range, and resolved recipe variant. ### DeepSeek-V4-Pro-0813 DSpark on MI355X ATOM -`dsv4-fp4-mi355x-atom-agentic-mtp` keeps its historical key and `_mtp.sh` -filename, while its matrix uses `spec-decoding: draft_model`. The AMD launcher -routes both speculative metadata values to that script and mounts the shared -HF cache for the 0813 checkpoint. The recipe pins revision +`dsv4-fp4-mi355x-atom-agentic-mtp` keeps its historical key, while its matrix +uses `spec-decoding: draft_model`. Every row selects +`benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml`; +recipe selection treats `draft_model` as `mtp`. The recipe pins revision `72e1d3230f6c080a530b0a1d46f8eb4602340597` and serves the resolved snapshot path; an explicit `MODEL_PATH` must pass the same checkpoint checks. Before GPU startup it verifies the config/index hashes, DSpark Markov/confidence heads, @@ -302,8 +302,8 @@ The H200 DSpark recipe uses the same minimum capture size and preserves the same B300 uses the same minimum capture size at c1/c2/c4. Its c1 CI comparison reduced request ITL P90/P99 from 38.74/41.42 ms to 2.62/3.45 ms; c2/c4 require CI confirmation. -The AgentX-only `dsv41flash-fp4--vllm-agentic-dspark` recipes use -`vllm/vllm-openai:deepseekv41-flash-0909` at TP4 on Blackwell SKUs with native five-token DSpark, +The AgentX-only `dsv41flash-fp4--vllm-agentic-dspark` recipes use the per-SKU +`image` pinned in [`nvidia-master.yaml`](../configs/nvidia-master.yaml) (originally `vllm/vllm-openai:deepseekv41-flash-0909`, which B300 still uses) at TP4 on Blackwell SKUs with native five-token DSpark, probabilistic drafting. Throughput uses the [committed golden AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml) of 3.51 for thinking on and five draft tokens, with synthetic rejection sampling and adaptive verification disabled. Accuracy evals retain real block rejection and adaptive verification. `--engram-config '{"cpu_offload":true}'` stores Engram embedding tables in pinned host DRAM accessed through UVA; `kv-offloading: none` describes the separate, GPU-resident KV cache. MXFP4 expert @@ -331,31 +331,11 @@ GPU sweep and eval evidence is required before calling any recipe validated. Source: [upstream recipe](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash). -### DeepSeek-V4.1-Flash DSpark on ATOM - -`dsv41flash-fp4-mi355x-atom-agentic-dspark` follows the -[upstream ATOM recipe](https://github.com/ROCm/ATOM/blob/53b11c9a665e786798785acbedfdfd4da3fb87c4/recipes/DeepSeek-V4.1-Flash-Agentic.md) -with `rocm/atom-dev:nightly_202609250902`. TP2 covers concurrency -`[1, 2, 8, 16, 32, 64]`; TP4 covers `[2, 8, 16, 32, 64]`, without expert -parallelism or KV offload. Every point uses BF16 KV, FP8 index cache, 128 maximum -sequences, 16K batched-token/prefill chunks, prefix caching with block size 16, -8K state checkpoints, compilation level 3 and FULL graphs. Concurrency 32 captures -every size from 1 through 32 plus 48, 64 and 128; other points use the upstream -sparse list. Five-token DSpark uses golden AL 3.51 for throughput and real -acceptance for eval, with the checkpoint's shipped draft and `dsml_v41` parser. - -The existing MI355X launcher mounts this model's shared cache and the repository -at `/ix`, preserves Slurm's GPU allocation and routes `draft_model` to the new -`dsv41flash_fp4_mi355x_atom_mtp.sh` script. Canonical AgentX runs use the uncapped -`semianalysis_cc_traces_weka_062126` corpus, 3600 seconds per point and five warmup -requests per lane. Workflow duration overrides and `agentx-fast` remain available -for diagnostics. GPU sweep and eval validation is pending. - ### DeepSeek-V4.1-Flash DSpark on H200 `dsv41flash-fp4-h200-vllm-agentic-dspark` is the H200 AgentX arm of the -DeepSeek-V4.1-Flash recipe. It shares `vllm/vllm-openai:deepseekv41-flash-0909` and the -text-only serving script with the Blackwell arms: `deepseek_v41` tokenizer and parsers, +DeepSeek-V4.1-Flash recipe. It pins `vllm/vllm-openai:nightly-cd10ed6f9f6b37a8ace9cf380007e66fe12ec0c3` (shared with B200 and GB200) and shares the +text-only serving settings with the Blackwell arms: `deepseek_v41` tokenizer and parsers, 1M context, native five-token DSpark with probabilistic drafting. Throughput uses the [committed golden AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml) of 3.51 for thinking on and five draft tokens, with synthetic rejection sampling and adaptive verification disabled. Accuracy evals retain real block rejection and adaptive verification. The arm runs **TP8**, not the upstream TP4. Upstream verifies TP4 on one GB200 NVL4 tray @@ -377,8 +357,8 @@ Trace corpus: the arm replays the uncapped `semianalysis_cc_traces_weka_062126` not the 256k-capped `..._062126_256k` variant, because the model serves 1M context. The recipe never names a corpus — `resolve_trace_source` picks the uncapped default only because its `dsv4*` case arm also matches the `dsv41flash` prefix. That is load-bearing -and invisible at the call site, so `runners/test_dsv41flash_h200.py` pins it; narrowing -the arm would silently downgrade this recipe's traces. +and invisible at the call site, and no test pins it (the former `runners/test_dsv41flash_h200.py` +was removed in #3141); narrowing the arm would silently downgrade this recipe's traces. **The H100 arm is separate.** H100 is not in the upstream hardware table, and the blocker is not the weights. At 1M context the sparse attention indexer allocates a @@ -386,8 +366,8 @@ blocker is not the weights. At 1M context the sparse attention indexer allocates `fp8_fp4_paged_mqa_logits`, which at the default 8192 batched tokens is exactly 16 GiB. That is a fixed startup cost paid during memory profiling, independent of concurrency, so it fails at concurrency 1 on an 80 GB card even though the resident weights fit. The H100 -arm therefore ships its own script with capped batched tokens instead of the shared -symlink; see the H100 section below. +arm therefore ships its own recipe settings with capped batched tokens; see the H100 +section below. The launcher mounts the repository at `/ix` for this recipe so AgentX runtime directories are not created under `/workspace`, and it already mounts the shared HF cache, so the @@ -404,18 +384,19 @@ DeepSeek-V4.1-Flash recipe, added after the H200 arm and deliberately separate f it. H100 is **not** in the upstream hardware table, which lists h200, gb200, gb300, and mi350x. -Unlike the other SKUs, H100 does not use the shared `dsv41flash_fp4_vllm_mtp.sh`. It has -its own copy, because the shared flags cannot serve 1M context on an 80 GB card. At 1M +Unlike the other SKUs, H100 did not inherit the shared serving flags. Its recipe, +`benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/h100-fp4-mtp/agentic.yaml`, carries its own, +because the shared flags cannot serve 1M context on an 80 GB card. At 1M context the sparse attention indexer allocates a `[max-num-batched-tokens, max-model-len]` logits buffer in `fp8_fp4_paged_mqa_logits`: -at the shared script's effective 8192 batched tokens that is 8192 x 1048576 x 2 bytes, +at the shared flags' effective 8192 batched tokens that is 8192 x 1048576 x 2 bytes, exactly 16.00 GiB. It is a fixed cost paid during startup memory profiling, independent of concurrency, so it OOMed at concurrency 1 in run [34467029236](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34467029236) next to roughly 35.9 GiB per GPU of resident weights — trimming the concurrency list cannot help. -The H100 script therefore caps `--max-num-batched-tokens` at 4096, putting the indexer +The H100 recipe therefore caps `--max-num-batched-tokens` at 4096, putting the indexer buffer at 8 GiB. Capping `--max-model-len` instead would shrink it just as well, but a context cap forces the 256k-capped trace corpus onto a model that serves 1M, so batched tokens is the right lever. The script also sets `--max-num-seqs` to twice the trajectory @@ -458,14 +439,15 @@ The H100 SGLang candidate sweeps DSpark at concurrency 1/2/4/8/16/20. It retains The same sweep also qualifies supported TP8/EP8/DP8 attention at C4/C8/C16/C20. DP uses a stock consistent-hash router with stable session keys, DP LM-head execution, and 64 SWA prefix tails per rank. Full C16 GSM8K passed on all 1,319 examples; its performance contribution remains under measurement. The native 1M context and the AgentX subagent/session semantics are preserved. -The nightly candidate uses `nightly-dev-cu13-20260922-582389ce`, native MXFP4 Marlin MoE, and `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`. A resolved local snapshot lets the upstream allocator evict checkpoint file cache before allocating anonymous host tables. Draft precision follows the pinned image's default handling, including its WO_A FP8-to-BF16 conversion. GPU-specific block32 FP8 launch configurations use the upstream kernel and supported `SPLIT_K`/`SWAP_AB` options; checkpoint data, scales, output dtype and context limits remain unchanged. Small-batch configurations require FP32-reference and CUDA-graph validation at selection boundaries, followed by full-model accuracy and serving measurements. Larger batches retain their previous configurations. Full canonical qualification is still required. +The nightly candidate uses `nightly-dev-cu13-20260922-582389ce`, native MXFP4 Marlin MoE, and `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`. A resolved local snapshot lets the upstream allocator evict checkpoint file cache before allocating anonymous host tables. Draft precision follows the pinned image's default handling, including its WO_A FP8-to-BF16 conversion. Block32 FP8 GEMMs use SGLang's default tilings; the custom block32 launch configurations were removed in [#3463](https://github.com/SemiAnalysisAI/InferenceX/pull/3463). Full canonical qualification is still required. `dsv41flash-fp4--sglang-agentic-dspark` are the SGLang counterparts of the vLLM arms, one PR per SKU across h100, h200, b200, b300, gb200, gb300 and mi355x. They follow the [SGLang cookbook](https://lmsysorg.mintlify.app/cookbook/autoregressive/DeepSeek/DeepSeek-V4_1), -which has no released SGLang version for this model yet. B200 pins the CUDA 13 nightly -`lmsysorg/sglang:nightly-dev-cu13-20260922-4cbf290f` by digest; the other NVIDIA arms use -`lmsysorg/sglang:dev-dsv41` and MI355X uses `lmsysorg/sglang:dev-dsv41-mi35x`. +which has no released SGLang version for this model yet. B200, B300, GB300 and H100 pin the CUDA 13 nightly +`lmsysorg/sglang:nightly-dev-cu13-20260922-582389ce` by digest; GB200 and H200 use +`lmsysorg/sglang:nightly-dev-cu13-20260923-06008c17` (GB200 by digest), and MI355X pins +`lmsysorg/sglang:dev-dsv41-mi35x` by digest. Each master entry's `image` is authoritative. B200 uses shipped-default DSpark across TP4/EP4 C1–128 and TP2/EP2 C1–8. Engram stays in host DRAM with `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`. @@ -542,12 +524,12 @@ so the Hopper arms do the same. `--mem-fraction-static 0.8` is the cookbook's lo setting. `--max-running-requests` is `2 * CONC` for AgentX subagent fan-out and the decode graph batch covers it, floored at the cookbook's 64 and capped at 128. -Each SKU ships its own `dsv41flash_fp4__sglang_mtp.sh` in its own PR. H100 is not in -the cookbook's hardware table, so its script differs: the Engram tables move to a single shared host copy +Each SKU ships its own recipe, `benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/-fp4-mtp/agentic.yaml`, in its own PR. H100 is not in +the cookbook's hardware table, so its recipe differs: the Engram tables move to a single shared host copy (`SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1`, the SGLang analogue of the vLLM arm's Engram CPU offload) and the prefill chunk is capped at 4096, the same batched-token cap the vLLM H100 arm needed for the sparse-attention indexer buffer on an 80 GB card. Concurrency -stops at 8 there until the KV ceiling is measured. MI355X has its own script with the +stops at 8 there until the KV ceiling is measured. MI355X has its own recipe with the cookbook's ROCm environment (`SGLANG_USE_AITER=1`, `SGLANG_MOE_PADDING=1`, `AITER_FLYDSL_FORCE_REDUCE=1`, `ROCM_QUICK_REDUCE_QUANTIZATION=NONE`), `--disable-radix-cache`, and breakable prefill graphs capped at 4096 tokens. @@ -664,7 +646,7 @@ Stop before dispatching GPU work or claiming the configuration complete when any - Calculated topology exceeds the fleet, DCP does not divide TP, heterogeneous hardware metadata is one-sided, or generated topology differs from the intended recipe. - An srt-slurm recipe and master entry disagree, `model.container != image`, or upstream recipe validation has not run. - An llm-d recipe is missing and would fall back unintentionally, allocation counts disagree, or endpoint discovery cannot satisfy literal-IPv4/unique-name/valid-port rules. -- An MTP script lacks chat-template benchmarking, the speculative method/token count is unverified, or graph capture exceeds the backend limit. +- An MTP recipe lacks chat-template benchmarking, the speculative method/token count is unverified, or graph capture exceeds the backend limit. - The changelog change would modify historical bytes, is not at EOF, has a conflict, or still has `TBD` when the PR is otherwise ready for sweep. - YAML, Bash, strict schema, exact-key generation, launcher simulation, or recipe validation fails. @@ -676,13 +658,13 @@ The `dsv41flash-fp4-mi355x-vllm-agentic-dspark` recipe extends [#2958](https://g Follow the AMD overrides in the merged [upstream recipe #968](https://github.com/vllm-project/recipes/pull/968): `VLLM_ROCM_USE_AITER=1`, `VLLM_ROCM_USE_AITER_MOE=1`, and `--moe-backend aiter`. The generic AITER selector lets vLLM pick the CK a8w4 experts, matching the DSV4-Pro MI355X recipe. The recipe pins `semianalysis_cc_traces_weka_062126` (the unfiltered corpus) via `WEKA_LOADER_OVERRIDE`. KV stays GPU-resident. Engram stayed on GPU under the upstream AMD defaults until [vllm-project/vllm#57491](https://github.com/vllm-project/vllm/pull/57491) widened the two `is_cuda()` gates to `is_cuda_alike()`. From that commit on, ROCm resolves an `EngramConfig` and `cpu_offload` defaults to on through `VLLM_PLE_CPU_OFFLOAD`, so the recipe sets `--engram-config` explicitly rather than leaning on that default. TP=2 always offloads, since the tables need 94.4 GiB per rank there; TP=4 keeps them resident through concurrency 64, where the KV pool is not the constraint, and offloads only at 128. The recipe likewise trims `--max-num-batched-tokens` only above concurrency 32, to 8192 at TP=2 c64 and TP=4 c128 and to 4096 at TP=2 c128, because the sparse-attention indexer and its companion per-rank buffers grow at roughly 4.4 MiB per batched token. Where that chunk falls below six times the API-server default of 1024 sequences, `--max-num-seqs` is capped at the graph-capture shape: DSpark verifies 1+5 tokens per sequence, and at 4096 against 1024 sequences the engram projection faults during profiling. The rule in every case is to spend device memory on KV only at the concurrencies that ran short of it, leaving the validated low-concurrency settings alone. Images built before that merge still reject the option on ROCm. The MI355X launcher uses the shared HF cache and mounts this model's repository at `/ix`, and exports `INFMAX_CONTAINER_WORKSPACE=/ix` so AgentX dependencies and outputs resolve inside that mount. -**GPU validation:** The recipe uses `vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984` (digest `sha256:eccb72b7…`, published 2026-09-21) on the ROCm 10.0 nightly channel that `kimik3-fp4-mi355x-vllm-agentic-mtp` already runs on this cluster. The sweep in [#3326](https://github.com/SemiAnalysisAI/InferenceX/pull/3326) qualifies that pin across TP4 and TP2 at concurrency 1–128, and is the only evidence for it: [run 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) covered TP4 concurrency 1–32 plus eval-only concurrency 32 on the superseded `nightly-eed1f3d0c6043bd494424a22443ee198dd56f657`, so its points do not carry onto this image. The merged [upstream recipe #1006](https://github.com/vllm-project/recipes/pull/1006) documents the MI355X TP2 Engram offload and `--no-swa-bounded-replay`, and the merged [#968](https://github.com/vllm-project/recipes/pull/968) records the original AMD overrides and the complete InferenceX command. Follow the [AgentX procedure](eval-agentx-procedures.md#7-run-agentx-fast-feedback-versus-canonical-evidence) for future runtime evidence; local generation and registry metadata alone are not GPU proof. +**GPU validation:** The recipe uses `vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098`, the first ROCm 10.0 nightly carrying vllm#58510, re-swept in [#3420](https://github.com/SemiAnalysisAI/InferenceX/pull/3420). The sweep in [#3326](https://github.com/SemiAnalysisAI/InferenceX/pull/3326) qualified the earlier `nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984` pin (digest `sha256:eccb72b7…`, published 2026-09-21) across TP4 and TP2 at concurrency 1–128, and was the only evidence for it: [run 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) covered TP4 concurrency 1–32 plus eval-only concurrency 32 on the superseded `nightly-eed1f3d0c6043bd494424a22443ee198dd56f657`, so its points do not carry onto this image. The merged [upstream recipe #1006](https://github.com/vllm-project/recipes/pull/1006) documents the MI355X TP2 Engram offload and `--no-swa-bounded-replay`, and the merged [#968](https://github.com/vllm-project/recipes/pull/968) records the original AMD overrides and the complete InferenceX command. Follow the [AgentX procedure](./eval-agentx-procedures.md#7-run-agentx-fast-feedback-versus-canonical-evidence) for future runtime evidence; local generation and registry metadata alone are not GPU proof. ## DeepSeek-V4.1-Flash on MI300X and MI325X `dsv41flash-fp4-mi300x-vllm-agentic-dspark` and `dsv41flash-fp4-mi325x-vllm-agentic-dspark` -copy the validated MI355X vLLM arm onto gfx942, on the `nightly-eed1f3d0` ROCm nightly the -MI355X arm used before it moved to the `nightly-rocm100` channel, and with the same +copy the validated MI355X vLLM arm onto gfx942, on the ROCm 10.0 nightly +`nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984`, and with the same AMD overrides (`VLLM_ROCM_USE_AITER=1`, `VLLM_ROCM_USE_AITER_MOE=1`, `VLLM_USE_BREAKABLE_CUDAGRAPH=1`, `--moe-backend aiter`, adaptive verification off). gfx942 is not in the upstream hardware table, and it has no FP4 MFMA: the plain `aiter` MoE @@ -695,7 +677,9 @@ hold its share of the 511 GB checkpoint plus the GPU-resident Engram tables (ups defaults; no CPU offload) and still leave a 1M-context KV pool. MI300X additionally caps `--max-num-batched-tokens` at 8192 because the sparse-attention indexer allocates a `[batched-tokens, max-model-len]` fp8 logits buffer at startup (16 GiB at 8192, 32 GiB at -the MI355X arm's 16384). Concurrency is 1–32 on both. +the MI355X arm's 16384). Concurrency is 1–32 on both. Both entries also carry smaller +layouts (MI300X TP4; MI325X TP2 and TP4) that move the Engram tables to host memory with +`--engram-config '{"cpu_offload":true}'`. `runners/launch_mi300x-amd.sh` and `runners/launch_mi325x-amds.sh` mount the checkout at `/ix` for this checkpoint and rewrite `RESULT_DIR`, as the MI355X launcher does, so AgentX diff --git a/inferencex-e2e/docs/configuration-procedures_zh.md b/inferencex-e2e/docs/configuration-procedures_zh.md index cc7abc13a9..089763c4e0 100644 --- a/inferencex-e2e/docs/configuration-procedures_zh.md +++ b/inferencex-e2e/docs/configuration-procedures_zh.md @@ -91,7 +91,7 @@ STP(Single Token Prediction,单 Token 预测)是每次前向传播生成 7. **追加一条 changelog**,精确选择新 key;参见[安全追加 changelog](#安全追加-changelog)。 8. **验证语法和生成结果。**检查 image、model、runner、ISL/OSL、`max-model-len`、并发、TP/PP/EP/DCP/PCP 和 `spec-decoding`。 -仅添加 `MODELS.md` 行并不会产生可执行配方。完整路径是:基准脚本 + 主配置条目 + launcher 路由 + changelog 触发项 + 生成的矩阵。 +仅添加 `MODELS.md` 行并不会产生可执行配方。完整路径是:srt-slurm 配方 + 主配置条目(`srt-recipe:`)+ changelog 触发项 + 生成的矩阵。 ## 修改主配置 @@ -201,18 +201,19 @@ llm-d 不是 srt-slurm 路径:InferenceX 自己持有 Slurm allocation,并 1. 确认使用原生 MTP 模块还是外部 draft。使用 draft 时,从模型/上游配方验证精确模型 ID、方法(例如 `eagle3`)和建议 speculative token 数。 2. 复制相同模型和 backend 的可工作同类项。保留其 speculative config、attention backend、token 数、模型补丁和依赖设置。 -3. 每个 `*_mtp.sh` 都必须向 `run_benchmark_serving` 传入 `--use-chat-template`;原始 prompt 会静默降低 acceptance。 +3. 每个投机解码的定长配方变体都必须设置 `benchmark.env.USE_CHAT_TEMPLATE: "true"`;`select_recipe` 会拒绝缺少该设置的投机解码变体,[`srt_fixed_sequence.sh`](../benchmarks/single_node/srt_fixed_sequence.sh) 会将其转换为传给 `run_benchmark_serving` 的 `--use-chat-template`。原始 prompt 会静默降低 acceptance。 4. graph capture 至少按 `CONC * (1 + NUM_SPEC_TOKENS)` 确定规模,采用同类项的取整方式,并限制在框架上限内(当前 vLLM playbook 上限为 2048)。 5. 保留 backend 差异:不要把 CUDA 专用 drafter attention pin 或补丁复制到 ROCm 配方。 -6. 在相应搜索空间条目设置 `spec-decoding: mtp`,并添加 `_mtp` launcher 后缀路由。若使用 schema 支持的 draft-model 模式,要有意设置匹配的生成值;不要根据文件名推断。 -7. 同时添加脚本 + 主配置条目 + launcher 路由 + changelog。 -8. 运行 Bash 语法和生成检查;检查 `spec-decoding`、draft/native 方法、token 数、chat-template 使用、capture 范围和解析出的脚本。 +6. 在相应搜索空间条目设置 `spec-decoding: mtp`,并将其 `srt-recipe:` 指向 `-mtp` 配方;`select_recipe` 会据此校验配方的 speculative 配置。若使用 schema 支持的 draft-model 模式,要有意设置匹配的生成值;不要根据文件名推断。 +7. 同时添加配方 + 主配置条目 + changelog。 +8. 运行 YAML 和生成检查;检查 `spec-decoding`、draft/native 方法、token 数、chat-template 使用、capture 范围和解析出的配方变体。 ### MI355X ATOM 上的 DeepSeek-V4-Pro-0813 DSpark -`dsv4-fp4-mi355x-atom-agentic-mtp` 保留历史配置 key 和 `_mtp.sh` 文件名, -矩阵元数据改为 `spec-decoding: draft_model`。AMD launcher 将这两种投机解码 -元数据都路由到该脚本,并为 0813 checkpoint 挂载共享 HF 缓存。配方固定 revision +`dsv4-fp4-mi355x-atom-agentic-mtp` 保留历史配置 key, +矩阵元数据改为 `spec-decoding: draft_model`。所有行均选择 +`benchmarks/single_node/srt-slurm-recipes/dsv4/atom/mi355x-fp4-mtp/agentic.yaml`; +配方选择时将 `draft_model` 视为 `mtp`。配方固定 revision `72e1d3230f6c080a530b0a1d46f8eb4602340597`,以实际 snapshot 路径启动服务; 显式传入的 `MODEL_PATH` 也必须通过相同检查。GPU 启动前核对 config/index 哈希、 DSpark Markov/confidence head、全部 66 个分片的 header 与 payload 边界, @@ -247,7 +248,7 @@ H200 的 DSpark 配方使用相同的最小捕获范围,并保持相同的工 B300 在 c1/c2/c4 使用相同的最小捕获范围。其 c1 CI 对比中,请求 ITL P90/P99 从 38.74/41.42 ms 降至 2.62/3.45 ms;c2/c4 仍需 CI 验证。 仅运行 AgentX 的 `dsv41flash-fp4--vllm-agentic-dspark` 配方使用 -`vllm/vllm-openai:deepseekv41-flash-0909`,在 Blackwell SKU 上采用 TP4、原生五 token DSpark、 +[`nvidia-master.yaml`](../configs/nvidia-master.yaml) 中按 SKU 固定的 `image`(最初为 `vllm/vllm-openai:deepseekv41-flash-0909`,B300 仍在使用),在 Blackwell SKU 上采用 TP4、原生五 token DSpark、 概率采样草稿。吞吐测试使用[已提交的黄金 AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml):thinking 开启、五个草稿 token 对应 3.51,采用合成拒绝采样并关闭自适应验证。准确率 eval 保留真实块拒绝采样和自适应验证。 `--engram-config '{"cpu_offload":true}'` 将 Engram 嵌入表放在固定页主机 DRAM 中,通过 UVA 访问;`kv-offloading: none` 描述的是另行保留在 GPU 上的 KV cache。 @@ -270,29 +271,11 @@ GB300 launcher 将引擎就绪等待时间设为 7200 秒。在[运行 345049691 来源:[上游配方](https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4.1-Flash)。 -### ATOM 上的 DeepSeek-V4.1-Flash DSpark - -`dsv41flash-fp4-mi355x-atom-agentic-dspark` 按照 -[ATOM 上游配方](https://github.com/ROCm/ATOM/blob/53b11c9a665e786798785acbedfdfd4da3fb87c4/recipes/DeepSeek-V4.1-Flash-Agentic.md) -使用 `rocm/atom-dev:nightly_202609250902`。TP2 覆盖并发 -`[1, 2, 8, 16, 32, 64]`,TP4 覆盖 `[2, 8, 16, 32, 64]`,不启用专家并行或 -KV 卸载。所有点均使用 BF16 KV、FP8 index cache、128 个最大序列、16K -批处理 token/prefill chunk、block size 16 的前缀缓存、8K 状态检查点、 -编译 level 3 和 FULL graphs。并发 32 捕获 1 到 32 的全部尺寸以及 48、64、128, -其他点使用上游稀疏列表。5-token DSpark 在吞吐测试中使用 golden AL 3.51, -eval 使用真实 acceptance,并保留检查点随附的 draft 和 `dsml_v41` parser。 - -现有 MI355X launcher 挂载该模型的共享缓存,并将仓库挂载到 `/ix`,保留 Slurm -分配的 GPU,将 `draft_model` 路由到新增的 `dsv41flash_fp4_mi355x_atom_mtp.sh`。 -标准 AgentX 运行使用未截断的 `semianalysis_cc_traces_weka_062126` 数据集, -每个点测量 3600 秒,每条 lane 预热 5 个请求。诊断时仍可使用 workflow 的时长覆盖 -和 `agentx-fast`。GPU sweep 和 eval 尚待验证。 - ### H200 上的 DeepSeek-V4.1-Flash DSpark `dsv41flash-fp4-h200-vllm-agentic-dspark` 是 DeepSeek-V4.1-Flash 配方的 H200 AgentX -分支。它与 Blackwell 分支共用 `vllm/vllm-openai:deepseekv41-flash-0909` 和纯文本服务 -脚本:`deepseek_v41` tokenizer 和解析器、1M 上下文、原生五 token DSpark(概率采样草稿)。吞吐测试使用[已提交的黄金 AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml):thinking 开启、五个草稿 token 对应 3.51,采用合成拒绝采样并关闭自适应验证。准确率 eval 保留真实块拒绝采样和自适应验证。 +分支。它固定使用 `vllm/vllm-openai:nightly-cd10ed6f9f6b37a8ace9cf380007e66fe12ec0c3`(与 B200、GB200 相同),并与 Blackwell 分支共用纯文本服务 +设置:`deepseek_v41` tokenizer 和解析器、1M 上下文、原生五 token DSpark(概率采样草稿)。吞吐测试使用[已提交的黄金 AL](../infx/golden_al_distribution/dsv41flash_dspark.yaml):thinking 开启、五个草稿 token 对应 3.51,采用合成拒绝采样并关闭自适应验证。准确率 eval 保留真实块拒绝采样和自适应验证。 该分支使用 **TP8**,而非上游的 TP4。上游在一个 GB200 NVL4 tray 上验证 TP4,并说明在 8 GPU 节点上同一布局每个角色变为 TP8,而 H200 DGXC 节点正是 8 GPU 节点。 @@ -309,15 +292,15 @@ MI355X 分支保持一致。Hopper 没有 FP4 tensor core,因此这些权重 轨迹语料:该分支回放未截断的 `semianalysis_cc_traces_weka_062126` 语料,而不是 256k 截断的 `..._062126_256k` 变体,因为该模型服务 1M 上下文。配方本身并未指定语料 —— `resolve_trace_source` 选中未截断的默认值,仅仅是因为其 `dsv4*` 分支同时匹配了 -`dsv41flash` 前缀。这一依赖在调用处并不可见却至关重要,因此由 -`runners/test_dsv41flash_h200.py` 固定;收窄该分支会静默地降级本配方的轨迹。 +`dsv41flash` 前缀。这一依赖在调用处并不可见却至关重要,且目前没有测试固定它(原 +`runners/test_dsv41flash_h200.py` 已在 #3141 中删除);收窄该分支会静默地降级本配方的轨迹。 **H100 分支单独实现。** H100 不在上游硬件表中,且瓶颈不在权重。在 1M 上下文下,稀疏 注意力 indexer 会在 `fp8_fp4_paged_mqa_logits` 中分配一个 `[max-num-batched-tokens, max-model-len]` 的 logits 缓冲区,在默认 8192 batched tokens 下恰好为 16 GiB。这是显存 profiling 阶段固定支付的启动开销,与并发无关,因此即使驻留 -权重放得下,在 80 GB 卡上并发 1 也会失败。因此 H100 分支使用独立脚本并收窄 batched -tokens,而非共享符号链接,详见下文 H100 小节。 +权重放得下,在 80 GB 卡上并发 1 也会失败。因此 H100 分支使用独立的配方设置并收窄 batched +tokens,详见下文 H100 小节。 launcher 为该配方将仓库挂载到 `/ix`,避免在 `/workspace` 下创建 AgentX 运行目录;它本来 就挂载了共享 HF 缓存,因此脚本通过 `HF_HUB_CACHE` 解析模型,而不依赖各节点的独立路径。 @@ -332,15 +315,16 @@ launcher 为该配方将仓库挂载到 `/ix`,避免在 `/workspace` 下创建 分支,在 H200 分支之后加入,并有意与其分开。H100 **不在**上游硬件表中(该表列出 h200、gb200、gb300、mi350x)。 -与其他 SKU 不同,H100 不使用共享的 `dsv41flash_fp4_vllm_mtp.sh`,而是拥有独立副本, +与其他 SKU 不同,H100 不沿用共享服务参数,其配方 +`benchmarks/single_node/srt-slurm-recipes/dsv41flash/vllm/h100-fp4-mtp/agentic.yaml` 使用独立参数, 因为共享参数无法在 80 GB 卡上服务 1M 上下文。在 1M 上下文下,稀疏注意力 indexer 会在 `fp8_fp4_paged_mqa_logits` 中分配 `[max-num-batched-tokens, max-model-len]` 的 logits -缓冲区:按共享脚本实际生效的 8192 batched tokens 计算,即 8192 x 1048576 x 2 字节, +缓冲区:按共享参数实际生效的 8192 batched tokens 计算,即 8192 x 1048576 x 2 字节, 恰好 16.00 GiB。这是启动阶段显存 profiling 固定支付的开销,与并发无关,因此在 [34467029236](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34467029236) 中于并发 1 即 OOM(此时每 GPU 驻留权重约 35.9 GiB)—— 收窄并发列表无济于事。 -因此 H100 脚本将 `--max-num-batched-tokens` 限制为 4096,使 indexer 缓冲区降至 8 GiB。 +因此 H100 配方将 `--max-num-batched-tokens` 限制为 4096,使 indexer 缓冲区降至 8 GiB。 改为收窄 `--max-model-len` 同样有效,但上下文上限会迫使一个服务 1M 上下文的模型使用 256k 截断语料,因此 batched tokens 才是正确的调节点。脚本还将 `--max-num-seqs` 设为 轨迹并发的两倍(而非沿用 vLLM 默认的 1024)、设置 `--gpu-memory-utilization 0.92`, @@ -376,15 +360,16 @@ H100 SGLang 候选配方在并发 1/2/4/8/16/20 下测试 DSpark。C1/C2 按并 同一 sweep 还会在 C4/C8/C16/C20 下验证受支持的 TP8/EP8/DP8 attention。DP 使用原生一致性哈希路由器与稳定会话键、DP LM-head,以及每 rank 64 个 SWA 前缀尾部。C16 的完整 GSM8K 已通过全部 1,319 个样本;其性能贡献仍在测量中。原生 1M 上下文与 AgentX 子代理/会话语义保持不变。 -nightly 候选配方使用 `nightly-dev-cu13-20260922-582389ce`、原生 MXFP4 Marlin MoE,以及 `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`。解析后的本地快照路径让上游分配器能够在分配匿名主机表之前清理检查点文件缓存。draft 精度遵循固定镜像的默认处理,包括 WO_A 从 FP8 到 BF16 的转换。针对 GPU 的 block32 FP8 启动配置使用上游内核及其支持的 `SPLIT_K`/`SWAP_AB` 选项;检查点数据、scale、输出 dtype 和上下文限制均不变。小批次配置必须在选择边界通过 FP32 参考值和 CUDA graph 验证,再进行完整模型准确率评测及服务性能测试。较大批次保留原有配置。仍需完成规范全量 sweep 验收。 +nightly 候选配方使用 `nightly-dev-cu13-20260922-582389ce`、原生 MXFP4 Marlin MoE,以及 `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`。解析后的本地快照路径让上游分配器能够在分配匿名主机表之前清理检查点文件缓存。draft 精度遵循固定镜像的默认处理,包括 WO_A 从 FP8 到 BF16 的转换。block32 FP8 GEMM 使用 SGLang 的默认 tiling;自定义 block32 启动配置已在 [#3463](https://github.com/SemiAnalysisAI/InferenceX/pull/3463) 中移除。仍需完成规范全量 sweep 验收。 `dsv41flash-fp4--sglang-agentic-dspark` 是 vLLM 配方在 h100、h200、b200、b300、gb200、gb300 与 mi355x 上的 SGLang 对应版本(每个 SKU 一个 PR),遵循 [SGLang cookbook](https://lmsysorg.mintlify.app/cookbook/autoregressive/DeepSeek/DeepSeek-V4_1)。 -该模型尚无正式发布的 SGLang 版本。B200 通过 digest 固定 CUDA 13 nightly 镜像 -`lmsysorg/sglang:nightly-dev-cu13-20260922-4cbf290f`;其他 NVIDIA 配方使用 -`lmsysorg/sglang:dev-dsv41`,MI355X 使用 `lmsysorg/sglang:dev-dsv41-mi35x`。 +该模型尚无正式发布的 SGLang 版本。B200、B300、GB300 与 H100 通过 digest 固定 CUDA 13 nightly 镜像 +`lmsysorg/sglang:nightly-dev-cu13-20260922-582389ce`;GB200 与 H200 使用 +`lmsysorg/sglang:nightly-dev-cu13-20260923-06008c17`(GB200 通过 digest 固定),MI355X 通过 digest 固定 +`lmsysorg/sglang:dev-dsv41-mi35x`。以各主配置条目的 `image` 为准。 B200 在 TP4/EP4 C1–128 与 TP2/EP2 C1–8 全部使用上游默认 DSpark。 Engram 保留在主机 DRAM,设置 `SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT=per_rank`。 @@ -449,12 +434,12 @@ cookbook 会自动解析 attention、MoE 与 FP8 GEMM 后端,并警告手动 为 cookbook 的低延迟设置。`--max-running-requests` 为 `2 * CONC` 以容纳 AgentX 子代理扇出, decode 图的 batch 覆盖该值,下限为 cookbook 的 64,上限为 128。 -每个 SKU 在各自的 PR 中提供独立的 `dsv41flash_fp4__sglang_mtp.sh`。H100 不在 cookbook 的 -硬件表中,其脚本有所不同:Engram 表移至 +每个 SKU 在各自的 PR 中提供独立配方 `benchmarks/single_node/srt-slurm-recipes/dsv41flash/sglang/-fp4-mtp/agentic.yaml`。H100 不在 cookbook 的 +硬件表中,其配方有所不同:Engram 表移至 单一共享主机副本(`SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1`,对应 vLLM 配方的 Engram CPU offload),prefill 分块上限设为 4096,即 vLLM H100 配方在 80 GB 显卡上为稀疏注意力 indexer 缓冲区所需的相同批处理 token 上限。在测得 KV 上限之前,该配方并发止于 8。MI355X 也有独立 -脚本,包含 cookbook 的 ROCm 环境变量(`SGLANG_USE_AITER=1`、`SGLANG_MOE_PADDING=1`、 +配方,包含 cookbook 的 ROCm 环境变量(`SGLANG_USE_AITER=1`、`SGLANG_MOE_PADDING=1`、 `AITER_FLYDSL_FORCE_REDUCE=1`、`ROCM_QUICK_REDUCE_QUANTIZATION=NONE`)、 `--disable-radix-cache`,以及上限 4096 token 的 breakable prefill 图。 @@ -568,7 +553,7 @@ python -m pytest infx/tests/matrix/ -v - 计算出的拓扑超过 fleet、DCP 不能整除 TP、异构 hardware 元数据只写一侧,或生成拓扑与目标配方不一致。 - srt-slurm 配方与主条目不一致、`model.container != image`,或尚未运行上游配方验证。 - llm-d 配方缺失并会意外 fallback、allocation 数不一致,或 endpoint discovery 无法满足 IPv4 字面量/唯一名称/有效端口规则。 -- MTP 脚本缺少 chat-template 基准、speculative 方法/token 数未验证,或 graph capture 超过 backend 上限。 +- MTP 配方缺少 chat-template 基准、speculative 方法/token 数未验证,或 graph capture 超过 backend 上限。 - changelog 变更会修改历史字节、没有位于 EOF、存在冲突,或 PR 已准备请求 sweep 但仍保留 `TBD`。 - YAML、Bash、严格 schema、精确 key 生成、launcher 模拟或配方验证失败。 @@ -580,13 +565,13 @@ python -m pytest infx/tests/matrix/ -v 遵循已合并的[上游配方 #968](https://github.com/vllm-project/recipes/pull/968) 中的 AMD 设置:`VLLM_ROCM_USE_AITER=1`、`VLLM_ROCM_USE_AITER_MOE=1` 和 `--moe-backend aiter`。通用 AITER 选择器允许 vLLM 选择 CK a8w4 专家内核,与 DSV4-Pro MI355X 配方一致。配方通过 `WEKA_LOADER_OVERRIDE` 固定使用完整语料 `semianalysis_cc_traces_weka_062126`。KV 驻留 GPU。在 [vllm-project/vllm#57491](https://github.com/vllm-project/vllm/pull/57491) 将两处 `is_cuda()` 判断放宽为 `is_cuda_alike()` 之前,Engram 按上游 AMD 默认设置常驻 GPU。自该提交起,ROCm 会解析 `EngramConfig`,且 `cpu_offload` 经由 `VLLM_PLE_CPU_OFFLOAD` 默认开启,因此配方显式设置 `--engram-config`,而不依赖该默认值。TP=2 始终下放,因为此时表每 rank 需 94.4 GiB;TP=4 在并发 64 及以下保持常驻(此时 KV 池并非瓶颈),仅在并发 128 时下放。同样地,配方仅在并发高于 32 时调低 `--max-num-batched-tokens`:TP=2 c64 与 TP=4 c128 为 8192,TP=2 c128 为 4096,因为稀疏注意力 indexer 及其配套的每 rank 缓冲区按每个批量 token 约 4.4 MiB 增长。当该分块低于 API server 默认 1024 序列所需的六倍时,`--max-num-seqs` 会被限制为 CUDA graph 捕获的规模:DSpark 每序列验证 1+5 个 token,4096 对 1024 序列会使 engram 投影在 profiling 阶段崩溃。所有情况下的原则一致:只在确实出现 KV 不足的并发点上把设备内存让给 KV,保持低并发处已验证的设置不变。早于该合并的镜像在 ROCm 上仍会拒绝该选项。MI355X launcher 使用共享 HF 缓存,并将此模型的仓库挂载至 `/ix`,同时导出 `INFMAX_CONTAINER_WORKSPACE=/ix`,确保 AgentX 依赖与输出路径位于该挂载中。 -**GPU 验证:** 配方使用 `vllm/vllm-openai-rocm:nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984`(摘要 `sha256:eccb72b7…`,发布于 2026-09-21),位于 ROCm 10.0 nightly 渠道,本集群上的 `kimik3-fp4-mi355x-vllm-agentic-mtp` 已在使用该渠道。[#3326](https://github.com/SemiAnalysisAI/InferenceX/pull/3326) 的 sweep 在 TP4 与 TP2、并发 1–128 下验证该镜像,且是其唯一证据:[运行 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) 覆盖的是已被取代的 `nightly-eed1f3d0c6043bd494424a22443ee198dd56f657` 上 TP4 并发 1–32 的吞吐测试与仅评测并发 32,其数据点不能沿用到本镜像。已合并的[上游配方 #1006](https://github.com/vllm-project/recipes/pull/1006) 记录了 MI355X 的 TP2 Engram 卸载与 `--no-swa-bounded-replay`,已合并的 [#968](https://github.com/vllm-project/recipes/pull/968) 则记录了最初的 AMD 设置与完整的 InferenceX 命令。后续运行时证据请遵循 [AgentX 流程](eval-agentx-procedures_zh.md);仅有本地矩阵生成和镜像元数据不能证明 GPU 验证完成。 +**GPU 验证:** 配方使用 `vllm/vllm-openai-rocm:nightly-rocm100-29468dde8b515031dc6d4d9d06bf0a2fa0442098`,即首个包含 vllm#58510 的 ROCm 10.0 nightly,已在 [#3420](https://github.com/SemiAnalysisAI/InferenceX/pull/3420) 中重新扫描。[#3326](https://github.com/SemiAnalysisAI/InferenceX/pull/3326) 的 sweep 在 TP4 与 TP2、并发 1–128 下验证的是此前的 `nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984`(摘要 `sha256:eccb72b7…`,发布于 2026-09-21),且是该镜像的唯一证据:[运行 34710937012](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34710937012) 覆盖的是已被取代的 `nightly-eed1f3d0c6043bd494424a22443ee198dd56f657` 上 TP4 并发 1–32 的吞吐测试与仅评测并发 32,其数据点不能沿用到本镜像。已合并的[上游配方 #1006](https://github.com/vllm-project/recipes/pull/1006) 记录了 MI355X 的 TP2 Engram 卸载与 `--no-swa-bounded-replay`,已合并的 [#968](https://github.com/vllm-project/recipes/pull/968) 则记录了最初的 AMD 设置与完整的 InferenceX 命令。后续运行时证据请遵循 [AgentX 流程](./eval-agentx-procedures_zh.md);仅有本地矩阵生成和镜像元数据不能证明 GPU 验证完成。 ## MI300X 与 MI325X 上的 DeepSeek-V4.1-Flash `dsv41flash-fp4-mi300x-vllm-agentic-dspark` 与 `dsv41flash-fp4-mi325x-vllm-agentic-dspark` -将已验证的 MI355X vLLM 配方复制到 gfx942,使用 MI355X 迁移到 `nightly-rocm100` 通道之前所用的 -`nightly-eed1f3d0` ROCm nightly,以及相同的 AMD 设置 +将已验证的 MI355X vLLM 配方复制到 gfx942,使用 ROCm 10.0 nightly +`nightly-rocm100-3df4ae153eb385e27b52f26c81f8edb9e20b9984`,以及相同的 AMD 设置 (`VLLM_ROCM_USE_AITER=1`、`VLLM_ROCM_USE_AITER_MOE=1`、`VLLM_USE_BREAKABLE_CUDAGRAPH=1`、 `--moe-backend aiter`、关闭自适应验证)。gfx942 不在上游硬件表中,且没有 FP4 MFMA:通用的 `aiter` MoE 后端允许 vLLM 选择器跳过仅 gfx950 可用的 CK a8w4 专家内核;若启动时所有候选均被 @@ -596,7 +581,8 @@ python -m pytest infx/tests/matrix/ -v 511 GB 检查点的分片以及常驻 GPU 的 Engram 表(上游 AMD 默认,不做 CPU offload),并仍留出 1M 上下文 KV 池。MI300X 另将 `--max-num-batched-tokens` 上限设为 8192:稀疏注意力 indexer 在启动时分配 `[batched-tokens, max-model-len]` 的 fp8 logits 缓冲区(8192 时 16 GiB,MI355X -配方的 16384 时 32 GiB)。两者并发均为 1–32。 +配方的 16384 时 32 GiB)。两者并发均为 1–32。两个条目还包含更小的布局(MI300X 的 TP4;MI325X 的 TP2 与 TP4), +通过 `--engram-config '{"cpu_offload":true}'` 将 Engram 表移至主机内存。 `runners/launch_mi300x-amd.sh` 与 `runners/launch_mi325x-amds.sh` 与 MI355X launcher 一样,为该 检查点将仓库挂载到 `/ix` 并重写 `RESULT_DIR`,使 AgentX 运行目录不落在 `/workspace` 下。MI300X diff --git a/inferencex-e2e/docs/eval-agentx-procedures.md b/inferencex-e2e/docs/eval-agentx-procedures.md index d373942495..d29e0a0a7b 100644 --- a/inferencex-e2e/docs/eval-agentx-procedures.md +++ b/inferencex-e2e/docs/eval-agentx-procedures.md @@ -64,7 +64,7 @@ uv run --no-project --exclude-newer PT12H --python 3.12 --with pydantic --with p --config-files configs/nvidia-master.yaml | jq . ``` -A correct AgentX eval row contains `"scenario-type": "agentic-coding"`, `"run-eval": true`, and `"eval-only": true`. The workflow splits generated rows into throughput, fixed-sequence eval, and agentic eval jobs in [`.github/workflows/e2e-tests.yml`](../../.github/workflows/e2e-tests.yml#L278-L293). +A correct AgentX eval row contains `"scenario-type": "agentic-coding"`, `"run-eval": true`, and `"eval-only": true`. The workflow splits generated rows into throughput, fixed-sequence eval, and agentic eval jobs in [`.github/workflows/e2e-tests.yml`](../../.github/workflows/e2e-tests.yml#L351-L358). ## 2. Add a graded eval @@ -101,7 +101,7 @@ append_lm_eval_summary python3 -m infx.evals.validate_scores --model-prefix "$MODEL_PREFIX" ``` -`run_lm_eval` passes concurrency through `num_concurrent` in `--model_args`. It is deliberately an environment variable, not a `run_eval` CLI option. The exact invocation is in [`run_lm_eval()`](../benchmarks/benchmark_lib.sh#L1080-L1162). +`run_lm_eval` passes concurrency through `num_concurrent` in `--model_args`. It is deliberately an environment variable, not a `run_eval` CLI option. The exact invocation is in [`run_lm_eval()`](../benchmarks/benchmark_lib.sh#L2044-L2128). ## 3. `EVAL_ONLY` is a launcher contract @@ -113,9 +113,9 @@ Set `EVAL_ONLY=true` **before server launch**. It is not merely a switch inside 4. Throughput returns immediately or is skipped. 5. `run_eval` and artifact staging run. -Relevant implementation: [context setup](../benchmarks/benchmark_lib.sh#L1049-L1078), [eval dispatch and failure policy](../benchmarks/benchmark_lib.sh#L1789-L1923), and [workflow inputs](../../.github/workflows/benchmark-tmpl.yml#L79-L97). +Relevant implementation: [context setup](../benchmarks/benchmark_lib.sh#L2016-L2042), [eval dispatch and failure policy](../benchmarks/benchmark_lib.sh#L2893-L3073), and [workflow inputs](../../.github/workflows/benchmark-tmpl.yml#L40-L57). -Do not toggle `EVAL_ONLY` after a throughput-sized server is already running and assume the context changed. Restart through the recipe. In eval-only mode an eval failure is returned after available artifacts are staged. In a workflow, upload happens with `always()` before score validation so failed evidence survives ([single-node upload and gate](../../.github/workflows/benchmark-tmpl.yml#L399-L417), [multi-node upload and gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L466-L488)). +Do not toggle `EVAL_ONLY` after a throughput-sized server is already running and assume the context changed. Restart through the recipe. In eval-only mode an eval failure is returned after available artifacts are staged. In a workflow, upload happens with `always()` before score validation so failed evidence survives ([single-node upload and gate](../../.github/workflows/benchmark-tmpl.yml#L467-L494), [multi-node upload and gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L487-L518)). ## 4. Batched eval concurrency @@ -137,9 +137,9 @@ The batch runner creates a fresh temporary output directory per point, stages fi - `completed_eval_concs`: eval and staging both succeeded. - `failed_eval_concs`: either eval or staging failed. -A failed point is deferred so artifacts from every attempted point can upload. The post-upload validator then fails the job. Batched mode accepts positive integers and supports only `lm-eval`. See [`run_eval` batching](../benchmarks/benchmark_lib.sh#L1839-L1900), [artifact suffixing](../benchmarks/benchmark_lib.sh#L1163-L1222), and [manifest validation](../infx/evals/validate_scores.py#L72-L171). +A failed point is deferred so artifacts from every attempted point can upload. The post-upload validator then fails the job. Batched mode accepts positive integers and supports only `lm-eval`. See [`run_eval` batching](../benchmarks/benchmark_lib.sh#L2980-L3031), [artifact suffixing](../benchmarks/benchmark_lib.sh#L2130-L2188), and [manifest validation](../infx/evals/validate_scores.py#L119-L216). -For multi-node `all-evals`, the workflow constructs `EVAL_CONC` by joining the topology's concurrency list ([dispatch](../../.github/workflows/e2e-tests.yml#L397-L400)). Never compare a point if its `_conc` result or completed-manifest entry is missing. +For multi-node `all-evals`, the workflow constructs `EVAL_CONC` by joining the topology's concurrency list ([dispatch](../../.github/workflows/e2e-tests.yml#L417-L419)). Never compare a point if its `_conc` result or completed-manifest entry is missing. ## 5. Validate scores, not file existence @@ -160,7 +160,7 @@ python3 -m infx.evals.validate_scores \ --thresholds infx/evals/thresholds.yaml ``` -Validation resolves the threshold in this order: `models..`, `default.`, then `--min-score` (default `0.85`). By default it checks numeric, non-stderr metrics beginning with `exact_match,`. It fails when a score is below threshold, no metric matches, a requested concurrency is absent, metadata has duplicates/invalid values, any point is marked failed, or result suffixes do not match the manifest. Current floors are authoritative in [`thresholds.yaml`](../infx/evals/thresholds.yaml). See [threshold resolution](../infx/evals/validate_scores.py#L61-L69) and the [validation flow](../infx/evals/validate_scores.py#L174-L302). +Validation resolves the threshold in this order: `models..`, `default.`, then `--min-score` (default `0.85`). By default it checks numeric, non-stderr metrics beginning with `exact_match,`. It fails when a score is below threshold, no metric matches, a requested concurrency is absent, metadata has duplicates/invalid values, any point is marked failed, or result suffixes do not match the manifest. Current floors are authoritative in [`thresholds.yaml`](../infx/evals/thresholds.yaml). See [threshold resolution](../infx/evals/validate_scores.py#L63-L73) and the [validation flow](../infx/evals/validate_scores.py#L219-L373). A manual combined throughput+eval recipe uploads eval output but the template's automatic score gate is specific to eval-only jobs. Run the validator explicitly for manual or combined runs. @@ -191,7 +191,7 @@ Retain `meta_env.json`, `results*.json`, and `sample*.jsonl`. Agentic SWE-bench [`install_agentic_deps()`](../benchmarks/benchmark_lib.sh) declares the AgentX client dependencies directly alongside the editable `utils/aiperf` install. It installs them into the isolated `AIPERF_RUNTIME_DIR` environment with the caller-supplied `AIPERF_PYTHON_VERSION`. -AgentX is AIPerf `inferencex-agentx-mvp` trace replay, not a fixed-token synthetic benchmark. The checked-in default uses ten additional warmup requests per trajectory lane and the recipe's configured profile duration. `agentx-fast` forces one warmup request per lane and a 1,200-second profile. It affects single- and multi-node AgentX throughput only. Fixed-sequence throughput and evals remain canonical. Fast runs are not eligible for artifact reuse ([workflow policy](../../.github/workflows/README.md#agentx-fast-mode), [fast replay settings](../benchmarks/benchmark_lib.sh#L2104-L2128)). +AgentX is AIPerf `inferencex-agentx-mvp` trace replay, not a fixed-token synthetic benchmark. The checked-in default uses ten additional warmup requests per trajectory lane and the recipe's configured profile duration. `agentx-fast` forces one warmup request per lane and a 1,200-second profile. It affects single- and multi-node AgentX throughput only. Fixed-sequence throughput and evals remain canonical. Fast runs are not eligible for artifact reuse ([workflow policy](../../.github/workflows/README.md#agentx-fast-mode), [fast replay settings](../benchmarks/benchmark_lib.sh#L3255-L3259)). For multi-node srt-slurm jobs, the benchmark client may run on a different host from the frontend. `srt_agentic.sh` uses an explicit `AIPERF_SERVER_URL` when supplied, otherwise derives it from `SRT_FRONTEND_HOST` and `SRT_FRONTEND_PORT`, and falls back to `localhost:$PORT` only when no remote endpoint is available. Trace replay and inter-point drain checks must use that same resolved endpoint. @@ -228,11 +228,11 @@ gh workflow run e2e-tests.yml --repo SemiAnalysisAI/InferenceX --ref "$REF" \ For a publishable SWE-bench score, omit `eval-limit`. Do not use `single-shot`, which is only a debugging escape hatch. SWE-bench generation/scoring controls and its `0.50` full-split threshold are documented next to the implementation in [`infx/evals/EVALS.md`](../infx/evals/EVALS.md#swe-bench-lite---framework-swebench). -Treat fast results as bring-up evidence, never as a replacement for the canonical candidate. A duration below 900 seconds or `AIPERF_UNSAFE_OVERRIDE=true` adds AIPerf's `--unsafe-override` and flags the submission invalid. Use it only for smoke diagnosis ([source](../benchmarks/benchmark_lib.sh#L2266-L2268)). After a fast run is healthy, run the exact candidate canonically before claiming benchmark success. +Treat fast results as bring-up evidence, never as a replacement for the canonical candidate. A duration below 900 seconds or `AIPERF_UNSAFE_OVERRIDE=true` adds AIPerf's `--unsafe-override` and flags the submission invalid. Use it only for smoke diagnosis ([source](../benchmarks/benchmark_lib.sh#L3362-L3364)). After a fast run is healthy, run the exact candidate canonically before claiming benchmark success. ## 8. Preserve trace and run provenance -AgentX defaults to recorded assistant-response replay. Live server outputs are measured but discarded when constructing later turns. Set `AIPERF_DATASET_WEKA_LIVE_ASSISTANT_RESPONSES=1` only for an explicitly different live-assistant experiment. The selected trace corpus is model-family dependent unless `WEKA_LOADER_OVERRIDE` pins it. The resolver logs both loader and Hugging Face dataset ([trace resolution](../benchmarks/benchmark_lib.sh#L2023-L2102), [replay semantics](../benchmarks/benchmark_lib.sh#L2104-L2270)). +AgentX defaults to recorded assistant-response replay. Live server outputs are measured but discarded when constructing later turns. Set `AIPERF_DATASET_WEKA_LIVE_ASSISTANT_RESPONSES=1` only for an explicitly different live-assistant experiment. The selected trace corpus is model-family dependent unless `WEKA_LOADER_OVERRIDE` pins it. The resolver logs both loader and Hugging Face dataset ([trace resolution](../benchmarks/benchmark_lib.sh#L3165-L3234), [replay semantics](../benchmarks/benchmark_lib.sh#L3236-L3366)). Capture orchestration provenance immediately: @@ -264,7 +264,7 @@ For each concurrency retain: - server/frontend logs and every metrics endpoint represented. - run URL/ID, attempt, head SHA, recipe/config identity, image, topology, fast flag, and any override. -The runner writes the command before replay and validates raw results after aggregation ([execution path](../benchmarks/benchmark_lib.sh#L2320-L2360)). Aggregation preserves dataset provenance and hardware/model/topology fields ([aggregate construction](../infx/results/agentic/__init__.py)). Raw workflow uploads intentionally omit very large `inputs.json` and `profile_export_raw.jsonl`. If those are required for an investigation, preserve them from the live allocation before cleanup ([single-node artifact contract](../../.github/workflows/benchmark-tmpl.yml#L349-L358), [multi-node contract](../../.github/workflows/benchmark-multinode-tmpl.yml#L455-L464)). +The runner writes the command before replay and validates raw results after aggregation ([execution path](../benchmarks/benchmark_lib.sh#L3412-L3566)). Aggregation preserves dataset provenance and hardware/model/topology fields ([aggregate construction](../infx/results/agentic/__init__.py)). Raw workflow uploads intentionally omit very large `inputs.json` and `profile_export_raw.jsonl`. If those are required for an investigation, preserve them from the live allocation before cleanup ([single-node artifact contract](../../.github/workflows/benchmark-tmpl.yml#L400-L409), [multi-node contract](../../.github/workflows/benchmark-multinode-tmpl.yml#L476-L485)). ## 9. Debug long AgentX runs from live evidence @@ -313,7 +313,7 @@ curl -fsS '' | \ rg -i 'request|queue|cache|token|prefill|decode|error|fail' ``` -Track trends over repeated samples: running/waiting requests, KV usage, prefix hits, input/output token rates, completed/cancelled/errored requests, frontend routing balance, and disaggregated KV transfer. AIPerf records endpoint identity for every server series ([metrics wiring](../benchmarks/benchmark_lib.sh#L2236-L2260)). +Track trends over repeated samples: running/waiting requests, KV usage, prefix hits, input/output token rates, completed/cancelled/errored requests, frontend routing balance, and disaggregated KV transfer. AIPerf records endpoint identity for every server series ([metrics wiring](../benchmarks/benchmark_lib.sh#L3344-L3359)). Use phase markers, not total Slurm age: @@ -334,7 +334,7 @@ Recommend stopping early when direct evidence is already disqualifying: - persistent near-100% KV usage plus a growing queue and unusable latency. - throughput has plateaued while more concurrency only worsens TTFT/TPOT. - any disaggregated pool or required metrics source never registers. -- AIPerf validation shows zero completed requests or error rate above the configured `0.10` limit ([validator](../infx/results/agentic/validate_agentic_result.py#L49-L87)). +- AIPerf validation shows zero completed requests or error rate above the configured `0.10` limit ([validator](../infx/results/agentic/validate_agentic_result.py#L48-L88)). Do **not** stop merely because model loading, dataset configuration, warmup, cutoff drain, or profiling is slow while completions advance and queues remain stable. Before any cancellation, capture timestamps, exact topology, relevant log lines, at least two metric samples showing the trend, current phase, and diagnosis. diff --git a/inferencex-e2e/docs/eval-agentx-procedures_zh.md b/inferencex-e2e/docs/eval-agentx-procedures_zh.md index 3774d8c94c..8f2cd5fec2 100644 --- a/inferencex-e2e/docs/eval-agentx-procedures_zh.md +++ b/inferencex-e2e/docs/eval-agentx-procedures_zh.md @@ -62,7 +62,7 @@ uv run --no-project --exclude-newer PT12H --python 3.12 --with pydantic --with p --config-files configs/nvidia-master.yaml | jq . ``` -正确的 AgentX eval 行包含 `"scenario-type": "agentic-coding"`、`"run-eval": true` 和 `"eval-only": true`。工作流会在 [`.github/workflows/e2e-tests.yml`](../../.github/workflows/e2e-tests.yml#L278-L293) 中将生成的行拆分到吞吐量、定长序列 eval 和 agentic eval 作业。 +正确的 AgentX eval 行包含 `"scenario-type": "agentic-coding"`、`"run-eval": true` 和 `"eval-only": true`。工作流会在 [`.github/workflows/e2e-tests.yml`](../../.github/workflows/e2e-tests.yml#L351-L358) 中将生成的行拆分到吞吐量、定长序列 eval 和 agentic eval 作业。 ## 2. 添加评分 eval @@ -99,7 +99,7 @@ append_lm_eval_summary python3 -m infx.evals.validate_scores --model-prefix "$MODEL_PREFIX" ``` -`run_lm_eval` 通过 `--model_args` 中的 `num_concurrent` 传递并发;它刻意采用环境变量,而不是 `run_eval` CLI 选项。准确调用见 [`run_lm_eval()`](../benchmarks/benchmark_lib.sh#L1080-L1162)。 +`run_lm_eval` 通过 `--model_args` 中的 `num_concurrent` 传递并发;它刻意采用环境变量,而不是 `run_eval` CLI 选项。准确调用见 [`run_lm_eval()`](../benchmarks/benchmark_lib.sh#L2044-L2128)。 ## 3. `EVAL_ONLY` 是 launcher 约定 @@ -111,9 +111,9 @@ python3 -m infx.evals.validate_scores --model-prefix "$MODEL_PREFIX" 4. 吞吐量路径立即返回或被跳过。 5. 运行 `run_eval` 和 artifact staging。 -相关实现:[context 设置](../benchmarks/benchmark_lib.sh#L1049-L1078)、[eval 分派与失败策略](../benchmarks/benchmark_lib.sh#L1789-L1923) 和[工作流输入](../../.github/workflows/benchmark-tmpl.yml#L79-L97)。 +相关实现:[context 设置](../benchmarks/benchmark_lib.sh#L2016-L2042)、[eval 分派与失败策略](../benchmarks/benchmark_lib.sh#L2893-L3073) 和[工作流输入](../../.github/workflows/benchmark-tmpl.yml#L40-L57)。 -不要在吞吐量规格的服务已经运行后才切换 `EVAL_ONLY`,并假定 context 会随之变化。应通过 recipe 重启。Eval-only 模式会在暂存已有 artifact 后返回 eval 失败;在工作流中,上传步骤使用 `always()`,并位于分数校验前,因此失败证据仍会保留([单节点上传与 gate](../../.github/workflows/benchmark-tmpl.yml#L399-L417)、[多节点上传与 gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L466-L488))。 +不要在吞吐量规格的服务已经运行后才切换 `EVAL_ONLY`,并假定 context 会随之变化。应通过 recipe 重启。Eval-only 模式会在暂存已有 artifact 后返回 eval 失败;在工作流中,上传步骤使用 `always()`,并位于分数校验前,因此失败证据仍会保留([单节点上传与 gate](../../.github/workflows/benchmark-tmpl.yml#L467-L494)、[多节点上传与 gate](../../.github/workflows/benchmark-multinode-tmpl.yml#L487-L518))。 ## 4. 批量 eval 并发 @@ -135,9 +135,9 @@ python3 -m infx.evals.validate_scores --expected-concs '16 32 64' - `completed_eval_concs`:eval 与 staging 均成功的点; - `failed_eval_concs`:eval 或 staging 失败的点。 -失败点会延迟报错,使所有已尝试点的 artifact 都能上传;随后 post-upload validator 会使作业失败。批量模式只接受正整数,且仅支持 `lm-eval`。参见 [`run_eval` batching](../benchmarks/benchmark_lib.sh#L1839-L1900)、[artifact 后缀处理](../benchmarks/benchmark_lib.sh#L1163-L1222) 和[manifest 校验](../infx/evals/validate_scores.py#L72-L171)。 +失败点会延迟报错,使所有已尝试点的 artifact 都能上传;随后 post-upload validator 会使作业失败。批量模式只接受正整数,且仅支持 `lm-eval`。参见 [`run_eval` batching](../benchmarks/benchmark_lib.sh#L2980-L3031)、[artifact 后缀处理](../benchmarks/benchmark_lib.sh#L2130-L2188) 和[manifest 校验](../infx/evals/validate_scores.py#L119-L216)。 -对于多节点 `all-evals`,工作流通过连接拓扑的并发列表构造 `EVAL_CONC`([分派](../../.github/workflows/e2e-tests.yml#L397-L400))。如果缺少某点的 `_conc` 结果或 completed manifest 条目,绝不能比较该点。 +对于多节点 `all-evals`,工作流通过连接拓扑的并发列表构造 `EVAL_CONC`([分派](../../.github/workflows/e2e-tests.yml#L417-L419))。如果缺少某点的 `_conc` 结果或 completed manifest 条目,绝不能比较该点。 ## 5. 校验分数,而不只是检查文件存在 @@ -158,7 +158,7 @@ python3 -m infx.evals.validate_scores \ --thresholds infx/evals/thresholds.yaml ``` -阈值按以下顺序解析:`models..`、`default.`,最后是 `--min-score`(默认 `0.85`)。默认检查名称以 `exact_match,` 开头、数值类型且非 stderr 的指标。当分数低于阈值、没有匹配指标、缺少请求的并发、metadata 有重复或无效值、任何点被标记为失败,或结果后缀与 manifest 不一致时,校验都会失败。当前权威下限位于 [`thresholds.yaml`](../infx/evals/thresholds.yaml)。参见[阈值解析](../infx/evals/validate_scores.py#L61-L69)与[校验流程](../infx/evals/validate_scores.py#L174-L302)。 +阈值按以下顺序解析:`models..`、`default.`,最后是 `--min-score`(默认 `0.85`)。默认检查名称以 `exact_match,` 开头、数值类型且非 stderr 的指标。当分数低于阈值、没有匹配指标、缺少请求的并发、metadata 有重复或无效值、任何点被标记为失败,或结果后缀与 manifest 不一致时,校验都会失败。当前权威下限位于 [`thresholds.yaml`](../infx/evals/thresholds.yaml)。参见[阈值解析](../infx/evals/validate_scores.py#L63-L73)与[校验流程](../infx/evals/validate_scores.py#L219-L373)。 手动的吞吐量+eval 组合 recipe 会上传 eval 输出,但模板的自动分数 gate 专用于 eval-only 作业。对手动或组合运行必须显式执行 validator。 @@ -189,7 +189,7 @@ gh run download "$RUN_ID" --repo SemiAnalysisAI/InferenceX \ [`install_agentic_deps()`](../benchmarks/benchmark_lib.sh) 在安装可编辑模式的 `utils/aiperf` 时直接声明 AgentX client 所需的依赖,并使用调用方提供的 `AIPERF_PYTHON_VERSION` 将它们安装到隔离的 `AIPERF_RUNTIME_DIR` 环境中。 -AgentX 是 AIPerf `inferencex-agentx-mvp` trace replay,不是固定 token 的合成 benchmark。仓库默认设置对每条 trajectory lane 额外执行十个 warmup 请求,并使用 recipe 配置的 profile 时长。`agentx-fast` 强制每条 lane 只运行一个 warmup 请求,并将 profile 设为 1,200 秒。它只影响单节点和多节点 AgentX 吞吐量;定长序列吞吐量与 eval 保持 canonical。Fast 运行不符合 artifact reuse 条件([工作流策略](../../.github/workflows/README.md#agentx-fast-mode)、[fast replay 设置](../benchmarks/benchmark_lib.sh#L2104-L2128))。 +AgentX 是 AIPerf `inferencex-agentx-mvp` trace replay,不是固定 token 的合成 benchmark。仓库默认设置对每条 trajectory lane 额外执行十个 warmup 请求,并使用 recipe 配置的 profile 时长。`agentx-fast` 强制每条 lane 只运行一个 warmup 请求,并将 profile 设为 1,200 秒。它只影响单节点和多节点 AgentX 吞吐量;定长序列吞吐量与 eval 保持 canonical。Fast 运行不符合 artifact reuse 条件([工作流策略](../../.github/workflows/README.md#agentx-fast-mode)、[fast replay 设置](../benchmarks/benchmark_lib.sh#L3255-L3259))。 对于多节点 srt-slurm 作业,benchmark client 与 frontend 可能运行在不同主机上。`srt_agentic.sh` 会优先使用显式提供的 `AIPERF_SERVER_URL`;否则从 `SRT_FRONTEND_HOST` 和 `SRT_FRONTEND_PORT` 推导地址;仅在没有远端 endpoint 时回退到 `localhost:$PORT`。Trace replay 和并发点之间的 drain 检查必须使用同一个解析后的 endpoint。 @@ -226,11 +226,11 @@ gh workflow run e2e-tests.yml --repo SemiAnalysisAI/InferenceX --ref "$REF" \ 要得到可发布的 SWE-bench 分数,省略 `eval-limit`;不要使用 `single-shot`,它只是诊断逃生选项。SWE-bench generation/scoring 控制项以及完整 split 的 `0.50` 阈值在实现旁的 [`infx/evals/EVALS.md`](../infx/evals/EVALS.md#swe-bench-lite---framework-swebench) 中说明。 -Fast 结果只能作为 bring-up 证据,绝不能替代 canonical candidate。小于 900 秒的 duration 或 `AIPERF_UNSAFE_OVERRIDE=true` 会添加 AIPerf 的 `--unsafe-override` 并将 submission 标记为无效;只能用于 smoke 诊断([源码](../benchmarks/benchmark_lib.sh#L2266-L2268))。Fast 运行健康后,必须对完全相同的 candidate 进行 canonical 运行,才能宣称 benchmark 成功。 +Fast 结果只能作为 bring-up 证据,绝不能替代 canonical candidate。小于 900 秒的 duration 或 `AIPERF_UNSAFE_OVERRIDE=true` 会添加 AIPerf 的 `--unsafe-override` 并将 submission 标记为无效;只能用于 smoke 诊断([源码](../benchmarks/benchmark_lib.sh#L3362-L3364))。Fast 运行健康后,必须对完全相同的 candidate 进行 canonical 运行,才能宣称 benchmark 成功。 ## 8. 保留 trace 与运行 provenance -AgentX 默认 replay 已记录的 assistant response。实时服务输出会被测量,但构造后续 turn 时会丢弃。只有在明确要进行不同的 live-assistant 实验时,才设置 `AIPERF_DATASET_WEKA_LIVE_ASSISTANT_RESPONSES=1`。除非用 `WEKA_LOADER_OVERRIDE` 固定,否则所选 trace corpus 依赖模型 family;resolver 会同时记录 loader 与 Hugging Face dataset([trace 解析](../benchmarks/benchmark_lib.sh#L2023-L2102)、[replay 语义](../benchmarks/benchmark_lib.sh#L2104-L2270))。 +AgentX 默认 replay 已记录的 assistant response。实时服务输出会被测量,但构造后续 turn 时会丢弃。只有在明确要进行不同的 live-assistant 实验时,才设置 `AIPERF_DATASET_WEKA_LIVE_ASSISTANT_RESPONSES=1`。除非用 `WEKA_LOADER_OVERRIDE` 固定,否则所选 trace corpus 依赖模型 family;resolver 会同时记录 loader 与 Hugging Face dataset([trace 解析](../benchmarks/benchmark_lib.sh#L3165-L3234)、[replay 语义](../benchmarks/benchmark_lib.sh#L3236-L3366))。 立即记录 orchestration provenance: @@ -262,7 +262,7 @@ gh run download "$RUN_ID" --repo SemiAnalysisAI/InferenceX \ - server/frontend 日志以及所代表的每个 metrics endpoint; - run URL/ID、attempt、head SHA、recipe/config 标识、image、topology、fast 标志和所有 override。 -Runner 会在 replay 前写入命令,并在聚合后校验原始结果([执行路径](../benchmarks/benchmark_lib.sh#L2320-L2360))。聚合会保留 dataset provenance 以及硬件/模型/拓扑字段([aggregate 构造](../infx/results/agentic/__init__.py))。工作流的 raw upload 会有意排除体积很大的 `inputs.json` 和 `profile_export_raw.jsonl`;如果调查需要这些文件,应在清理前从实时 allocation 保存([单节点 artifact 约定](../../.github/workflows/benchmark-tmpl.yml#L349-L358)、[多节点约定](../../.github/workflows/benchmark-multinode-tmpl.yml#L455-L464))。 +Runner 会在 replay 前写入命令,并在聚合后校验原始结果([执行路径](../benchmarks/benchmark_lib.sh#L3412-L3566))。聚合会保留 dataset provenance 以及硬件/模型/拓扑字段([aggregate 构造](../infx/results/agentic/__init__.py))。工作流的 raw upload 会有意排除体积很大的 `inputs.json` 和 `profile_export_raw.jsonl`;如果调查需要这些文件,应在清理前从实时 allocation 保存([单节点 artifact 约定](../../.github/workflows/benchmark-tmpl.yml#L400-L409)、[多节点约定](../../.github/workflows/benchmark-multinode-tmpl.yml#L476-L485))。 ## 9. 用实时证据调试长时间 AgentX 运行 @@ -311,7 +311,7 @@ curl -fsS '' | \ rg -i 'request|queue|cache|token|prefill|decode|error|fail' ``` -通过重复 sample 跟踪趋势:running/waiting request、KV usage、prefix hit、input/output token rate、completed/cancelled/errored request、frontend routing balance,以及 disaggregated KV transfer。AIPerf 会为每条 server series 记录 endpoint identity([metrics 接线](../benchmarks/benchmark_lib.sh#L2236-L2260))。 +通过重复 sample 跟踪趋势:running/waiting request、KV usage、prefix hit、input/output token rate、completed/cancelled/errored request、frontend routing balance,以及 disaggregated KV transfer。AIPerf 会为每条 server series 记录 endpoint identity([metrics 接线](../benchmarks/benchmark_lib.sh#L3344-L3359))。 应使用 phase marker,而不是 Slurm 总运行时间: @@ -332,7 +332,7 @@ date -u - KV usage 长期接近 100%,queue 持续增长且 latency 已不可用; - 吞吐量已经平台化,而更高并发只会恶化 TTFT/TPOT; - 任意 disaggregated pool 或必需 metrics source 始终未注册; -- AIPerf 校验显示 completed request 为零,或错误率超过配置的 `0.10` 上限([validator](../infx/results/agentic/validate_agentic_result.py#L49-L87))。 +- AIPerf 校验显示 completed request 为零,或错误率超过配置的 `0.10` 上限([validator](../infx/results/agentic/validate_agentic_result.py#L48-L88))。 如果 completion 持续增加且 queue 稳定,不要仅因模型加载、dataset 配置、warmup、cutoff drain 或 profiling 较慢而停止。任何取消前,都要捕获时间戳、准确拓扑、相关日志行、至少两个体现趋势的 metric sample、当前 phase 和诊断结论。 diff --git a/inferencex-e2e/docs/klaud-reporting.md b/inferencex-e2e/docs/klaud-reporting.md index 8bdfc2b095..21579bbba9 100644 --- a/inferencex-e2e/docs/klaud-reporting.md +++ b/inferencex-e2e/docs/klaud-reporting.md @@ -106,7 +106,7 @@ Before adding `full-sweep-fail-fast`, validate the exact pushed head's full matr ```bash head_sha=$(git rev-parse HEAD) -uv run --no-project --python 3.12 --with 'pydantic>=2.10,<3' --with pyyaml \ +uv run --no-project --exclude-newer PT12H --python 3.12 --with 'pydantic>=2.10,<3' --with pyyaml \ python -m infx.matrix.plan --base-ref origin/main --head-ref "$head_sha" \ --changelog-file perf-changelog.yaml > "$KLAUD_EVIDENCE/final-matrix.json" "${KLAUD[@]}" check-final --matrix-file "$KLAUD_EVIDENCE/final-matrix.json" diff --git a/inferencex-e2e/docs/klaud-reporting_zh.md b/inferencex-e2e/docs/klaud-reporting_zh.md index c3ce58dc81..ee01e8b9db 100644 --- a/inferencex-e2e/docs/klaud-reporting_zh.md +++ b/inferencex-e2e/docs/klaud-reporting_zh.md @@ -106,7 +106,7 @@ KLAUD=(uv run --no-project --exclude-newer PT12H --python 3.12 \ ```bash head_sha=$(git rev-parse HEAD) -uv run --no-project --python 3.12 --with 'pydantic>=2.10,<3' --with pyyaml \ +uv run --no-project --exclude-newer PT12H --python 3.12 --with 'pydantic>=2.10,<3' --with pyyaml \ python -m infx.matrix.plan --base-ref origin/main --head-ref "$head_sha" \ --changelog-file perf-changelog.yaml > "$KLAUD_EVIDENCE/final-matrix.json" "${KLAUD[@]}" check-final --matrix-file "$KLAUD_EVIDENCE/final-matrix.json" diff --git a/inferencex-e2e/docs/recovery-results-procedures.md b/inferencex-e2e/docs/recovery-results-procedures.md index a966aac81a..fb30d86522 100644 --- a/inferencex-e2e/docs/recovery-results-procedures.md +++ b/inferencex-e2e/docs/recovery-results-procedures.md @@ -323,7 +323,7 @@ Canonical source: [MI355X root-owned file recovery](https://github.com/SemiAnaly ## MI300X cluster debugging: enroot/pyxis user-namespace failures -Canonical signature on `mi300x-amds_*` / `chi-mi300x-*`: +Canonical signature on `mi300x-amd_*` / `chi-mi300x-*`: ```text error: pyxis: enroot-nsenter: failed to create user namespace: Permission denied diff --git a/inferencex-e2e/docs/recovery-results-procedures_zh.md b/inferencex-e2e/docs/recovery-results-procedures_zh.md index 49cbdb9d24..fea02892eb 100644 --- a/inferencex-e2e/docs/recovery-results-procedures_zh.md +++ b/inferencex-e2e/docs/recovery-results-procedures_zh.md @@ -322,7 +322,7 @@ jumpbox 没有 sudo;需要使用 agent forwarding 连接到在 `/it-share` 上 ## MI300X 集群调试:enroot/pyxis 用户命名空间故障 -`mi300x-amds_*` / `chi-mi300x-*` 上的典型特征: +`mi300x-amd_*` / `chi-mi300x-*` 上的典型特征: ```text error: pyxis: enroot-nsenter: failed to create user namespace: Permission denied diff --git a/inferencex-e2e/infx/KNOWN_LIMITATION.md b/inferencex-e2e/infx/KNOWN_LIMITATION.md index 342de691f5..6dde266f2e 100644 --- a/inferencex-e2e/infx/KNOWN_LIMITATION.md +++ b/inferencex-e2e/infx/KNOWN_LIMITATION.md @@ -1,7 +1,7 @@ # KNOWN_LIMITATION -When benchmarking small models (e.g., Gemma 1B or Llama 8B) and at ultra-high querys per second (QPS) and with short input and output lengths, the current InferenceX bench_serving client becomes client-bound. +When benchmarking small models (e.g., Gemma 1B or Llama 8B) and at ultra-high queries per second (QPS) and with short input and output lengths, the current InferenceX bench_serving client becomes client-bound. InferenceX does not currently benchmark in this regime and has no plans to. Our roadmap skews the opposite direction, larger models, longer ISL/OSL (i.e., agentic workloads), and interactive TTFT & tok/s/user scenarios rather than ultra-high QPS. -Should the need arise, we shall fix this known limitation & migrate to a multi-process benchmark client. If your needs are benchmarking small models at high QPS at small ISL/OSL lengths, we recommend xyz benchmark instead. +Should the need arise, we shall fix this known limitation & migrate to a multi-process benchmark client. diff --git a/inferencex-e2e/infx/golden_al_distribution/README.md b/inferencex-e2e/infx/golden_al_distribution/README.md index 18cd652434..1cb14ec9fd 100644 --- a/inferencex-e2e/infx/golden_al_distribution/README.md +++ b/inferencex-e2e/infx/golden_al_distribution/README.md @@ -134,6 +134,7 @@ Before accepting an updated curve, reviewers should verify: | MiniMax-M3 | EAGLE3 | [`minimaxm3_eagle3.yaml`](minimaxm3_eagle3.yaml) | [28061204145](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28061204145) | | MiniMax-M3 | EAGLE3 (GQA) | [`minimaxm3_eagle3_gqa.yaml`](minimaxm3_eagle3_gqa.yaml) | [29784780049](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/29784780049) | | GLM-5.2 | MTP | [`glm5.2_mtp.yaml`](glm5.2_mtp.yaml) | [28058352479](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28058352479) | +| GLM-5.3 | MTP (provisional, K=3 copied from GLM-5.2) | [`glm5.3_mtp.yaml`](glm5.3_mtp.yaml) | [28058352479](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28058352479) (GLM-5.2) | | Qwen3.8-Flash-Next | MTP (native) | [`qwen3.8next_mtp.yaml`](qwen3.8next_mtp.yaml) | [33034290269](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/33034290269) | ## Querying the curves diff --git a/inferencex-e2e/infx/golden_al_distribution/README_zh.md b/inferencex-e2e/infx/golden_al_distribution/README_zh.md index 793e936a92..e17338c529 100644 --- a/inferencex-e2e/infx/golden_al_distribution/README_zh.md +++ b/inferencex-e2e/infx/golden_al_distribution/README_zh.md @@ -134,6 +134,7 @@ gh workflow run speedbench-al.yml \ | MiniMax-M3 | EAGLE3 | [`minimaxm3_eagle3.yaml`](minimaxm3_eagle3.yaml) | [28061204145](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28061204145) | | MiniMax-M3 | EAGLE3(GQA) | [`minimaxm3_eagle3_gqa.yaml`](minimaxm3_eagle3_gqa.yaml) | [29784780049](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/29784780049) | | GLM-5.2 | MTP | [`glm5.2_mtp.yaml`](glm5.2_mtp.yaml) | [28058352479](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28058352479) | +| GLM-5.3 | MTP(临时,K=3 取自 GLM-5.2) | [`glm5.3_mtp.yaml`](glm5.3_mtp.yaml) | [28058352479](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/28058352479)(GLM-5.2) | | Qwen3.8-Flash-Next | MTP (native) | [`qwen3.8next_mtp.yaml`](qwen3.8next_mtp.yaml) | [33034290269](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/33034290269) | ## 查询曲线 diff --git a/inferencex-e2e/utils/runner_setup/RUNNER_SETUP.md b/inferencex-e2e/utils/runner_setup/RUNNER_SETUP.md index a8fa43fcb5..9b1f69929e 100644 --- a/inferencex-e2e/utils/runner_setup/RUNNER_SETUP.md +++ b/inferencex-e2e/utils/runner_setup/RUNNER_SETUP.md @@ -119,7 +119,7 @@ Required permissions (all of these endpoints require **admin access to the repos 5. Start the runners: ```bash - ./InferenceX/utils/runner_setup/start_runners.sh 0 13 ~/gharunners + ./InferenceX/utils/runner_setup/start_runners.sh 0 17 ~/gharunners ``` This (re)creates a tmux session (default name: `github-actions`) with one tiled pane @@ -151,7 +151,7 @@ key off that name: [`configs/runners.yaml`](../../configs/runners.yaml). New runners do **not** receive sweep jobs until they are added there, and the entries must match the registered names exactly, including zero-padding. Some older fleets predate the - padded convention, such as `h200-dgxc-slurm_0`. Because `setup.sh` always zero-pads, new + padded convention, such as `gb200-nv_0`. Because `setup.sh` always zero-pads, new entries should use the padded form. ## Labels / `ADDITIONAL_RUNNER_TAGS` @@ -251,11 +251,11 @@ The host side of each is defined in that cluster's `runners/launch_.sh` each node downloads its own copy, so prefer shared storage where available). 3. **Pre-staged model weights.** Large models are not downloaded from HF in CI. The launch scripts override `MODEL_PATH` to per-cluster staging directories - (e.g. `/lustre/fsw/models/...` on b200-nscale, `/data/models/...` on b300, + (e.g. `/scratch/models/...` on b200-nscale, `/data/models/...` on b300, read-only `/scratch/models/` on b300 multinode). Bringing up a new model on a cluster means staging the weights there first. 4. **Squash images.** Launch scripts `enroot import` each Docker image once into a - `.sqsh` file under a shared `SQUASH_DIR` (e.g. `/home/sa-shared/containers` on + `.sqsh` file under a shared `SQUASH_DIR` (e.g. `/data/home/sa-shared/containers` on b200-nscale, `/mnt/lustre01/users-public/sa-shared` on gb200), then launch with `--container-image=.sqsh`. This must be on shared storage because pyxis reads the file on the **compute** node, and it lets concurrent diff --git a/operatorx/CI.md b/operatorx/CI.md index 8c0e1381f4..d62fe4b60b 100644 --- a/operatorx/CI.md +++ b/operatorx/CI.md @@ -25,15 +25,14 @@ from CollectiveX's platform registry, and both planning and execution validate t ## Dispatch -Once GitHub has registered the workflow, select **OperatorX Sweep → Run workflow**, +Select **OperatorX Sweep → Run workflow**, choose the source branch, and keep the initial defaults: `pool=h100-dgxc`, `backends=torch`, `testlists=gemm`, `world_sizes=1`, `chunk_size=500`. This schedules the complete checked-in GEMM catalog in bounded shards (currently 7,212 cases in 15 shards). The catalog includes formats unsupported by a selected backend and shapes that can exceed device memory. Unsupported rows remain visible; actual kernel and allocation errors fail CI. A full catalog run is not a promise -that every case fits or is supported on H100. A newly added workflow -may need to reach the default branch before GitHub accepts manual dispatch. +that every case fits or is supported on H100. ```bash gh workflow run operatorx-sweep.yml --repo SemiAnalysisAI/InferenceX \ diff --git a/operatorx/CI_zh.md b/operatorx/CI_zh.md index ecbdc79bca..e9d379fa74 100644 --- a/operatorx/CI_zh.md +++ b/operatorx/CI_zh.md @@ -25,13 +25,12 @@ GB200/GB300 每次使用一个四卡计算托盘,不会占用整个 NVL72 机 ## 触发运行 -GitHub 注册该工作流后,选择 **OperatorX Sweep → Run workflow**,指定源码分支, +选择 **OperatorX Sweep → Run workflow**,指定源码分支, 并保留初始默认值:`pool=h100-dgxc`、`backends=torch`、`testlists=gemm`、 `world_sizes=1`、`chunk_size=500`。这会将完整的 GEMM 测试列表拆分为有界分片(目前为 7,212 个测试、15 个分片)。 列表中包含所选后端不支持的精度,以及可能超出设备显存的形状。不支持的测试会保留 在结果中;实际的内核和显存分配错误仍会使 CI 失败。运行完整列表不代表其中每个 测试都能在 H100 上执行。 -新工作流可能需要先进入默认分支,GitHub 才允许手动触发。 ```bash gh workflow run operatorx-sweep.yml --repo SemiAnalysisAI/InferenceX \ diff --git a/operatorx/CLUSTERS.md b/operatorx/CLUSTERS.md index 7491e7248f..046c132cef 100644 --- a/operatorx/CLUSTERS.md +++ b/operatorx/CLUSTERS.md @@ -145,16 +145,25 @@ one LNC=2 unit (1 HBM bank + 2 paired NC-v4 cores). ```python CLUSTER_PLATFORMS = { - "b200_dgx_8x": "nvidia", - "b300_hgx_8x": "nvidia", - "b200_nvl72": "nvidia", - "mi355x_8x": "amd", - "v6e_1x": "tpu", - "v6e_4x": "tpu", - "v6e_pod": "tpu", - "trn3_1x": "trainium", - "trn3_8x": "trainium", - "trn3_16x": "trainium", + "h100_dgxc_8x": "nvidia", + "h200_dgxc_8x": "nvidia", + "b200_dgx_8x": "nvidia", + "b200_nscale_8x": "nvidia", + "b300_dsxe_8x": "nvidia", + "gb200_nvl72_4x": "nvidia", + "gb300_nvl72_4x": "nvidia", + "b300_hgx_8x": "nvidia", + "b200_nvl72": "nvidia", + "mi355x_8x": "amd", + "mi300x_amds_8x": "amd", + "mi325x_amds_8x": "amd", + "v6e_1x": "tpu", + "v6e_4x": "tpu", + "v6e_pod": "tpu", + "v7x_4x": "tpu", + "trn3_1x": "trainium", + "trn3_8x": "trainium", + "trn3_16x": "trainium", } ``` diff --git a/operatorx/README.md b/operatorx/README.md index 6d547e77eb..0251028c46 100644 --- a/operatorx/README.md +++ b/operatorx/README.md @@ -37,7 +37,6 @@ cd "$HOME/inferencex/operatorx" OPERATORX_CLUSTER=b300_hgx_8x \ OPERATORX_PARTITION=batch_1 \ OPERATORX_ACCOUNT=benchmark \ -OPERATORX_QOS=batch_1_qos \ OPERATORX_SQUASH_DIR="$HOME/containers" \ OPERATORX_BACKENDS=torch,deepgemm,flashinfer,sglang \ python3 scripts/submit_run.py nvidia @@ -74,7 +73,7 @@ rather than running our old single-device dense fallback. | `OPERATORX_QOS` | (omitted) | SLURM `--qos`. | | `OPERATORX_SQUASH_DIR` | `/home/sa-shared/containers` | Where `.sqsh` lives. | | `OPERATORX_BACKENDS` | all backends for the platform | CSV allowlist. | -| `OPERATORX_JOB_NAME` | `benchmark` | SLURM job name. Use `h-benchmark` for benchmark runs (see `CLUSTERS.md`). | +| `OPERATORX_JOB_NAME` | `benchmark` | SLURM job name. Set a recognizable name for your runs (see `CLUSTERS.md`). | `WORLD_SIZES` in the script is `[1, 2, 4, 8]` and supports single-node runs only. Values above 8 are disabled because multi-node NCCL IB bring-up currently hangs on b200/b300.