Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 9 additions & 4 deletions .agents/skills/debug-runs/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,7 +58,7 @@ For a **single config** (tightest CI loop, skips the rest of the matrix), dispat
gh workflow run e2e-tests.yml -f generate-cli-command="test-config --config-key <KEY> --config-file <PATH/to/master.yaml>" -f test-name="debug <KEY>"
```

(`generate-cli-command` is the required input. Its config paths are relative to
(`generate-cli-command` carries the matrix selection; the workflow marks it `required: false` because the trusted changelog-dispatch (Klaud) mode omits it, but a manual dispatch without it fails at setup. Its config paths are relative to
`inferencex-e2e/`, for example `configs/nvidia-master.yaml`. `--target` is NOT a real arg.)

### 2. Monitor continuously
Expand Down Expand Up @@ -104,9 +104,14 @@ Steps:

1. Use the job or runner name to identify the node. Look up that cluster's access details in
the canvas, then SSH in with `ssh -A` when a jumpbox or agent forwarding is involved.
2. Reproduce the exact benchmark the launcher runs. Read `inferencex-e2e/runners/launch_<cluster>.sh` for
the image, container mounts, and the `inferencex-e2e/benchmarks/single_node/<...>.sh` command and env
(`IMAGE`, `TP`, `PRECISION`, `EXP_NAME`, `SPEC_DECODING`, …). On Slurm clusters, use
2. Reproduce the exact benchmark the launcher runs. Single-node jobs take the
`native-single-node` path in `inferencex-e2e/runners/launch_<cluster>.sh`, which calls
`launch_srt_single_node <cluster>` in `inferencex-e2e/runners/slurm_utils.sh`: it validates the
master row's `srt-recipe:` (`inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/<model>/<engine>/<sku>-<precision>[-mtp]/<scenario>.yaml`)
with `python3 -m infx.srt_slurm.single_node prepare`, then submits it through srtctl
with the `inferencex-e2e/runners/srt-slurm/<cluster>.yaml` profile. Read the recipe for the image
(`model.container`), server args and env, and the launcher for mounts and the job env
(`IMAGE`, `TP`, `PRECISION`, `SPEC_DECODING`, `CONC`, …). On Slurm clusters, use
`salloc` or `srun` with the squash image. On the **bare-metal `-tw` pools, use `docker run`**
on the node directly without `srun`.
3. **Always diff against a working node or working SKU** for reference. Most node failures
Expand Down
6 changes: 3 additions & 3 deletions .claude/commands/find-mergeable-claude-prs.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,16 +2,16 @@
description: Find Claude-authored PRs with all-green full-sweep validation and confirm before merging
---

Find open PRs authored by Claude (branches starting with `claude/`) whose full-sweep validation has completed all-green, then prompt the user before merging.
Find open PRs authored by Claude (branches starting with `klaud/`, `klaud-cold/`, or the legacy `klaude/`) whose full-sweep validation has completed all-green, then prompt the user before merging.

## Step 1 — list candidate `claude/*` PRs
## Step 1 — list candidate Klaud PRs

`gh pr list --json statusCheckRollup` truncates each PR's rollup, so it can't be trusted for the per-check filter. Use it only to get the candidate numbers, then re-query each PR individually.

```bash
gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \
--json number,title,headRefName \
--jq '.[] | select(.headRefName | startswith("claude/")) | .number' \
--jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | .number' \
> /tmp/claude_pr_candidates.txt
```

Expand Down
10 changes: 5 additions & 5 deletions .claude/commands/fix-klaud-cron-prs.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,14 +2,14 @@
description: Triage failing Claude-authored PRs by reading sweep logs, debugging failures, and pushing a candidate fix per PR
---

For each open Claude-authored PR (`claude/*` branch) whose full-sweep validation produced at least one **FAILED** check, fetch the failing run's logs, diagnose the root cause, and push a candidate fix to the PR's branch.
For each open Claude-authored PR (`klaud/*`, `klaud-cold/*`, or legacy `klaude/*` branch) whose full-sweep validation produced at least one **FAILED** check, fetch the failing run's logs, diagnose the root cause, and push a candidate fix to the PR's branch.

This command modifies remote PR branches. **Pause for user confirmation** after listing the candidate PRs and again before pushing each fix.

## Step 1 — find failing `claude/*` PRs whose sweep actually ran
## Step 1 — find failing Klaud PRs whose sweep actually ran

A PR qualifies only if:
- `headRefName` starts with `claude/`
- `headRefName` starts with `klaud/`, `klaud-cold/`, or `klaude/`
- At least one `Run Sweep` check has conclusion `SUCCESS` **or** `FAILURE` (i.e. the sweep was enabled and produced real results, rather than all checks being skipped)
- At least one check has conclusion `FAILURE`, `CANCELLED`, or `TIMED_OUT`

Expand All @@ -18,7 +18,7 @@ A PR qualifies only if:
```bash
gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \
--json number,title,headRefName \
--jq '.[] | select(.headRefName | startswith("claude/")) | "\(.number)\t\(.headRefName)\t\(.title)"' \
--jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | "\(.number)\t\(.headRefName)\t\(.title)"' \
> /tmp/claude_pr_candidates.tsv

: > /tmp/claude_prs_failing.tsv
Expand Down Expand Up @@ -79,7 +79,7 @@ If the log file is very large (>2000 lines), grep it for the actual error signat

### 2c. Diagnose

Inspect the PR diff (`git -C "$WT" diff origin/main...HEAD`) and the failing-log excerpts together. Most `claude/issue-1154-*` PRs are image-bump PRs that touch a `*.yaml` recipe. Failures are usually:
Inspect the PR diff (`git -C "$WT" diff origin/main...HEAD`) and the failing-log excerpts together. Most failing Klaud PRs are image-bump PRs (e.g. `klaud/<basekey>-<TAG>` from `/nuke`) that touch a `*.yaml` recipe. Failures are usually:

- Image tag typo / unavailable tag → fix the image reference.
- Engine arg incompatibility with new image version → add/remove the affected flag in the recipe.
Expand Down
24 changes: 13 additions & 11 deletions .claude/commands/klaud-pr-status-html.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,18 +2,18 @@
description: Render an HTML dashboard of Claude/Klaud-Cold PR states (state + check breakdown per PR) and open it in the browser
---

Render an HTML dashboard for every open PR in `SemiAnalysisAI/InferenceX` that was opened by Claude (either a `claude/*` branch OR a title containing `[Klaud Cold]`). Each row shows the PR's current state, a check-status breakdown, the title, and empty "Reason"/"Suggested fix" cells you can fill in afterward by reading failed-run logs.
Render an HTML dashboard for every open PR in `SemiAnalysisAI/InferenceX` that was opened by Claude (either a `klaud/*`, `klaud-cold/*`, or legacy `klaude/*` branch OR a title containing `[Klaud Cold]`). Each row shows the PR's current state, a check-status breakdown, the title, and empty "Reason"/"Suggested fix" cells you can fill in afterward by reading failed-run logs.

The dashboard lives at `/tmp/klaud_pr_status.html` and is opened with `open` (macOS) at the end.

## Step 1 — list candidate PRs (`claude/*` OR title containing `[Klaud Cold]`)
## Step 1 — list candidate PRs (`klaud/*` / `klaud-cold/*` / `klaude/*` OR title containing `[Klaud Cold]`)

The title check uses `contains` (not `startswith`) so it picks up PRs whose titles embed `[Klaud Cold]` after a prefix like `[Handoff to @Oseltamivir Claude /loop]`. Handoff-style PRs from a /loop run still belong on the dashboard.

```bash
gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \
--json number,title,headRefName,createdAt \
--jq '.[] | select((.headRefName | startswith("claude/")) or (.title | contains("[Klaud Cold]"))) | "\(.number)\t\(.headRefName)\t\(.createdAt)\t\(.title)"' \
--jq '.[] | select((.headRefName | test("^(klaud|klaud-cold|klaude)/")) or (.title | contains("[Klaud Cold]"))) | "\(.number)\t\(.headRefName)\t\(.createdAt)\t\(.title)"' \
> /tmp/klaud_pr_candidates.tsv
wc -l /tmp/klaud_pr_candidates.tsv
```
Expand Down Expand Up @@ -65,7 +65,7 @@ while IFS=$'\t' read -r pr branch created title; do
| select(.workflowName == "Run Sweep")
| state as $s
| select($s == "QUEUED" or $s == "IN_PROGRESS")
| ((.name | capture("(?<p>b200|b300|h100|h200|mi300x|mi325x|mi355x)").p) // "unknown") as $pool
| ((.name | capture("(?<p>gb200|gb300|b200|b300|h100|h200|mi300x|mi325x|mi355x)").p) // "unknown") as $pool
| "\($pr)\t\($pool)\t\($s)"
' >> /tmp/klaud_pr_jobs.tsv
done < /tmp/klaud_pr_candidates.tsv
Expand Down Expand Up @@ -108,14 +108,16 @@ state_class = {
"NO_SWEEP": "state-NOSWEEP", "NO_SUCCESS": "state-NOSWEEP",
}

# Active self-hosted runner counts per pool (from GHA registrations as of the
# session this was last touched). If a pool isn't listed, falls back to 4 — a
# conservative guess. To refresh:
# Self-hosted runner counts per pool, derived from the runner-label map in
# inferencex-e2e/configs/runners.yaml (run this from the repo root). If a pool isn't listed,
# falls back to 4, a conservative guess. Online counts can be lower; to check:
# gh api --paginate repos/SemiAnalysisAI/InferenceX/actions/runners \
# --jq '.runners[] | select(.status == "online") | .labels[].name' \
# | grep -xE 'b200|b300|h200|h100|mi300x|mi325x|mi355x' | sort | uniq -c
POOL_RUNNERS = {"b200": 12, "b300": 18, "h100": 19, "h200": 18,
"mi300x": 9, "mi325x": 9, "mi355x": 9}
# | grep -xE 'gb200|gb300|b200|b300|h200|h100|mi300x|mi325x|mi355x' | sort | uniq -c
import yaml
POOL_RUNNERS = {pool: len(names) for pool, names in
yaml.safe_load(Path("inferencex-e2e/configs/runners.yaml").read_text())["labels"].items()
if not pool.startswith("cluster:")}
DEFAULT_POOL_RUNNERS = 4
AVG_JOB_MIN = 7 # rough sweep-job median; eval+1k1k ~5min, 8k1k+agentic ~10-15min

Expand Down Expand Up @@ -221,7 +223,7 @@ out = ['<!doctype html>',
' .almost-section .eta { display:inline-block; min-width:56px; color:#3a3; font-weight:600; }',
'</style></head><body>',
'<h1>Claude / [Klaud Cold] PR status &mdash; InferenceX</h1>',
f'<div class="meta">Generated {now}. Source: <code>gh pr view --json statusCheckRollup</code> for every <code>claude/*</code> or <code>[Klaud Cold]</code>-titled open PR. Diagnoses (if any) loaded from <code>/tmp/klaud_pr_diag.json</code>. <strong>ETA</strong> = pool-drain pessimistic estimate (<code>ceil(global_pool_pending / runners) × ~{AVG_JOB_MIN}m/job</code>); GHA dispatch isn\'t guaranteed FIFO and queue position isn\'t exposed (see <a href="https://docs.github.com/en/actions">docs.github.com/en/actions</a>) so this is an upper-bound ordering hint, not a SLA.</div>',
f'<div class="meta">Generated {now}. Source: <code>gh pr view --json statusCheckRollup</code> for every <code>klaud/*</code>, <code>klaud-cold/*</code>, <code>klaude/*</code> or <code>[Klaud Cold]</code>-titled open PR. Diagnoses (if any) loaded from <code>/tmp/klaud_pr_diag.json</code>. <strong>ETA</strong> = pool-drain pessimistic estimate (<code>ceil(global_pool_pending / runners) × ~{AVG_JOB_MIN}m/job</code>); GHA dispatch isn\'t guaranteed FIFO and queue position isn\'t exposed (see <a href="https://docs.github.com/en/actions">docs.github.com/en/actions</a>) so this is an upper-bound ordering hint, not a SLA.</div>',
f'<div class="meta">Global pending across all open Klaud/claude sweeps: <strong>{total_global_pending}</strong> jobs &mdash; per-pool: {html.escape(pool_pressure_summary) if pool_pressure_summary else "—"}</div>']

pill_specs = [("READY", "ready"), ("RUNNING", "running"),
Expand Down
8 changes: 4 additions & 4 deletions .claude/commands/list-claude-pr-status.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,16 +2,16 @@
description: List Claude-authored PRs that haven't failed (ready or still running) with their actual check state
---

List open PRs authored by Claude (branches starting with `claude/`) that have **not** had any check fail. Show each PR's actual state (READY when all checks finished green, RUNNING when sweeps are still queued/in-progress) along with a per-status check breakdown, rendered as a markdown table.
List open PRs authored by Claude (branches starting with `klaud/`, `klaud-cold/`, or the legacy `klaude/`) that have **not** had any check fail. Show each PR's actual state (READY when all checks finished green, RUNNING when sweeps are still queued/in-progress) along with a per-status check breakdown, rendered as a markdown table.

## Step 1 — list candidate `claude/*` PRs
## Step 1 — list candidate Klaud PRs

`gh pr list --json statusCheckRollup` truncates each PR's rollup, so it can't be trusted for the per-check filter. Use it only to get the candidate numbers, then re-query each PR individually.

```bash
gh pr list --repo SemiAnalysisAI/InferenceX --state open --limit 200 \
--json number,title,headRefName \
--jq '.[] | select(.headRefName | startswith("claude/")) | "\(.number)\t\(.title)"' \
--jq '.[] | select(.headRefName | test("^(klaud|klaud-cold|klaude)/")) | "\(.number)\t\(.title)"' \
> /tmp/claude_pr_candidates.tsv
```

Expand Down Expand Up @@ -64,6 +64,6 @@ Print the result directly as a markdown table. READY rows first, then RUNNING. E
{printf "| [#%s](https://github.com/SemiAnalysisAI/InferenceX/pull/%s) | %s | `%s` | %s |\n", $1, $1, $2, $3, $4}'
```

If `/tmp/claude_pr_status.tsv` is empty, print: `_No claude/* PRs are currently READY or RUNNING — all open Claude PRs have failures or no sweep results._`
If `/tmp/claude_pr_status.tsv` is empty, print: `_No Klaud PRs are currently READY or RUNNING — all open Claude PRs have failures or no sweep results._`

Output the resulting markdown table to the user verbatim. This command is informational only. Do **not** auto-merge.
46 changes: 45 additions & 1 deletion .claude/commands/nuke.md
Original file line number Diff line number Diff line change
Expand Up @@ -104,13 +104,53 @@ block += [" description:", f' - "{desc}"', " pr-link: PRLINK_PLACEHOLDER"]
open(f,'w').write(content + '\n' + '\n'.join(block) + '\n')
```

`/tmp/edit_recipe_container.py` (single-node runs go through the srt-slurm
recipe each master search-space row names in `srt-recipe:`, and
`inferencex-e2e/infx/srt_slurm/single_node.py` rejects the run unless the recipe's
`model.container` equals the master `image:` exactly):
```python
#!/usr/bin/env python3
# Usage: edit_recipe_container.py <master_yaml> <new_image> <key1> [key2 ...]
import os, re, sys
f, new_image, keys = sys.argv[1], sys.argv[2], sys.argv[3:]
# srt-recipe paths are relative to inferencex-e2e/, the parent of configs/
root = os.path.dirname(os.path.dirname(os.path.abspath(f)))
master = open(f).read().split('\n')
recipes = set()
for key in keys:
kre = re.compile(r'^' + re.escape(key) + r':\s*$')
start = next((i for i,l in enumerate(master) if kre.match(l)), None)
if start is None: sys.exit(f"ERROR: key not found: {key}")
for j in range(start+1, len(master)):
if re.match(r'^[A-Za-z0-9._-]+:\s*$', master[j]): break # next top-level key
recipes.update(re.findall(r'srt-recipe:\s*([^\s,}]+)', master[j]))
if not recipes: sys.exit(f"ERROR: no srt-recipe for keys {keys}")
for r in sorted(os.path.join(root, x) for x in recipes):
lines = open(r).read().split('\n')
in_model, hit = False, False
for i, l in enumerate(lines):
if re.match(r'^\s*model:\s*$', l): in_model, indent = True, len(l) - len(l.lstrip()); continue
if in_model and l.strip() and len(l) - len(l.lstrip()) <= indent: in_model = False
m = re.match(r'^(\s+)container:\s*(.+?)\s*$', l) if in_model else None
if m:
lines[i] = f"{m.group(1)}container: {new_image}"; hit = True
print(f"{r}: {m.group(2)} -> {new_image}"); break
if not hit: sys.exit(f"ERROR: no model.container in {r}")
open(r, 'w').write('\n'.join(lines))
```

If `grep -rn '<recipe path>' inferencex-e2e/configs/*-master.yaml` shows a recipe is also
referenced by a key outside this family, stop and ask the user: bumping it would
break that other key's `model.container == image` check.

For each family, run strictly sequentially because git checkouts can't be parallel:

```bash
git checkout main -q && git reset --hard origin/main -q
branch="klaud/<basekey>-<TAG>"
git checkout -b "$branch" -q
python3 /tmp/edit_image.py <master.yaml> <NEW_IMAGE> <key> [<key>-mtp]
python3 /tmp/edit_recipe_container.py <master.yaml> <NEW_IMAGE> <key> [<key>-mtp]
python3 /tmp/append_changelog.py inferencex-e2e/perf-changelog.yaml "<DESC>" <key> [<key>-mtp]
git add -A
git commit -q -m "[Klaud Cold] Update <basekey>[ (+mtp)] <PHRASE> to <TAG>"
Expand All @@ -119,10 +159,14 @@ url=$(gh pr create --repo SemiAnalysisAI/InferenceX --base main --head "$branch"
--title "[Klaud Cold] Update <basekey>[ (+mtp)] <PHRASE> to <TAG>" \
--body "<BODY>" --label full-sweep-fail-fast | grep -o 'https://github.com/[^ ]*')
# patch the changelog pr-link with the real URL, then amend + force-push
# (read first, then write: open(f,'w') truncates before a same-line read runs)
before=$(wc -l < inferencex-e2e/perf-changelog.yaml)
python3 - inferencex-e2e/perf-changelog.yaml "$url" <<'PY'
import sys; f,u=sys.argv[1],sys.argv[2]
open(f,'w').write(open(f).read().replace("PRLINK_PLACEHOLDER",u,1))
content = open(f).read()
open(f,'w').write(content.replace("PRLINK_PLACEHOLDER",u,1))
PY
[ "$(wc -l < inferencex-e2e/perf-changelog.yaml)" -eq "$before" ] || { echo "perf-changelog.yaml line count changed"; exit 1; }
git add inferencex-e2e/perf-changelog.yaml && git commit -q --amend --no-edit && git push -q --force-with-lease
```

Expand Down
11 changes: 7 additions & 4 deletions .claude/commands/recover-failed-ingest.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,10 @@ argument-hint: <failed-run-or-job-url | pr-number> [source-run-id]

Recover the official database ingest for a failed or skipped InferenceX
push-to-main `Run Sweep` workflow by creating a recovery PR that reuses artifacts
from an earlier PR sweep. Do not add a one-off recovery workflow.
from an earlier PR sweep. Do not add a one-off recovery workflow. The existing
`.github/workflows/recover-reused-ingest.yml` only redispatches
`ingest-agentic-results` for a failed reused **agentic** ingest (inputs
`source-run-id`, `merge-run-id`); it is not a generic fixed-sequence recovery tool.

Inputs from `$ARGUMENTS`:

Expand Down Expand Up @@ -225,9 +228,9 @@ RECOVERY_PR=$(gh pr view "$RECOVERY_PR_URL" \
--repo SemiAnalysisAI/InferenceX \
--json number --jq .number)

gh pr edit "$RECOVERY_PR" \
--repo SemiAnalysisAI/InferenceX \
--add-label full-sweep-fail-fast
# REST, not `gh pr edit` (projects-classic GraphQL bug)
gh api -X POST "repos/SemiAnalysisAI/InferenceX/issues/$RECOVERY_PR/labels" \
-f "labels[]=full-sweep-fail-fast" --jq '.[].name'
gh pr comment "$RECOVERY_PR" \
--repo SemiAnalysisAI/InferenceX \
--body "/reuse-sweep-run $SOURCE_RUN_ID"
Expand Down
13 changes: 11 additions & 2 deletions .github/workflows/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,8 +16,8 @@ positional arguments:
filtering by model, precision, framework, runner type,
and sequence lengths
test-config Generate full sweep for specific config keys.
Supports wildcard patterns (* and ?) for matching
multiple keys at once.
Validates that all specified keys exist before
generating.

options:
-h, --help show this help message and exit
Expand All @@ -32,12 +32,16 @@ usage: python -m infx.matrix.generate full-sweep
--config-files CONFIG_FILES [CONFIG_FILES ...]
[--runner-config RUNNER_CONFIG]
[--no-evals | --evals-only] [--all-evals]
[--smoke] [--trim-conc]
[--runner-node-filter RUNNER_NODE_FILTER]
[--scenario-type {fixed-seq-len,agentic-coding} [{fixed-seq-len,agentic-coding} ...]]
[--model-prefix MODEL_PREFIX [MODEL_PREFIX ...]]
[--precision PRECISION [PRECISION ...]]
[--framework FRAMEWORK [FRAMEWORK ...]]
[--runner-type RUNNER_TYPE [RUNNER_TYPE ...]]
[--seq-lens {1k1k,8k1k} [{1k1k,8k1k} ...]]
[--step-size STEP_SIZE]
[--min-conc MIN_CONC]
[--max-conc MAX_CONC]
[--max-tp MAX_TP]
[--max-ep MAX_EP]
Expand Down Expand Up @@ -101,8 +105,13 @@ usage: python -m infx.matrix.generate test-config
--config-files CONFIG_FILES [CONFIG_FILES ...]
[--runner-config RUNNER_CONFIG]
[--no-evals | --evals-only] [--all-evals]
[--smoke] [--trim-conc]
[--runner-node-filter RUNNER_NODE_FILTER]
[--scenario-type {fixed-seq-len,agentic-coding} [{fixed-seq-len,agentic-coding} ...]]
--config-keys CONFIG_KEYS [CONFIG_KEYS ...]
[--conc CONC [CONC ...]]
[--exp-names EXP_NAMES [EXP_NAMES ...]]
[--seq-lens {1k1k,8k1k} [{1k1k,8k1k} ...]]
```

Config keys support **wildcard patterns** using `*` (matches any characters) and `?` (matches a single character). Patterns that match no keys will raise an error.
Expand Down
Loading
Loading