Skip to content

srt-slurm: consolidate sibling recipes into named override variants - #3391

Open
cquil11 wants to merge 1 commit into
ix/srt-recipe-selector-toolingfrom
ix/srt-recipe-zip-consolidation
Open

cquil11 wants to merge 1 commit into
ix/srt-recipe-selector-toolingfrom
ix/srt-recipe-zip-consolidation

Conversation

@cquil11

@cquil11 cquil11 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #3390 (base branch ix/srt-recipe-selector-tooling). Review and merge that first; this PR retargets to main afterwards.

What

Replaces 357 flat srt-slurm recipes referenced by the active configs/nvidia-master.yaml with one variants.yaml per directory (40 directories). Each file has a shared base plus one named override block per benchmark configuration, separated by blank lines:

base:
  # everything the configurations share
  ...

override_disagg_1p1d_dep8_b8_eplb0_mtp3:
  name: ctx1_gen1_dep8_batch8_eplb0_mtp3
  roles:
    decode:
      nodes: 1
      args: { max_batch_size: 8, ... }
  benchmark:
    concurrencies: '90'
  • Override names come from the replaced filename, which already follows the RECIPES.md convention (topology plus distinguishing settings), lowercased with underscores. Each block keeps its original recipe name and lists only what differs from base.
  • Master entries now use CONFIG_FILE=recipes/<dir>/variants.yaml:override_<name>. The 357 old→new pairs are recorded in benchmarks/multi_node/srt-slurm-recipe-identities.yaml, so fingerprints and curve identity are unchanged.
  • Named overrides were chosen over zip_override_* lists: each configuration reads on its own, and selectors are names rather than positions, so adding or removing a configuration never shifts others. srt-slurm: resolve CONFIG_FILE recipe selectors in launchers and the matrix #3390's resolver handles both forms.
  • RECIPES.md (+ _zh) documents the layout.

Net diff: 401 files, +25.2k / −54.6k lines.

Not changed

  • Deprecated recipes and configs. A recipe was only a candidate if the active master config references it and no deprecated config does. The one file shared with a launcher path rule (glm5.2/sglang/gb200-fp4/agentx/agg.yaml) was also left alone.
  • Three directories stay flat:
    • dsr1/sglang/b200-fp4/8k1k and dsr1/trtllm/b200-fp8/8k1k: siblings differ by an explicit null, and a null in an override deletes the key instead of setting it.
    • glm5.2/sglang/h200-fp8/agentx: one sibling is not schema: 2.
  • Existing override files, single-recipe directories, and upstream-only recipes.

Equivalence evidence

The scripts were run locally; they are one-off and not committed.

  1. Recipe content. For all 357 configurations, yaml.safe_load(old file) equals srtctl's generate_override_configs(new, "override_<name>") (pinned submodule), and also equals the InferenceX resolver's output.
  2. Matrix. full-sweep --multi-node against srt-slurm: resolve CONFIG_FILE recipe selectors in launchers and the matrix #3390, in default (518 rows) and --all-evals (459 rows) modes, is identical row for row once the new selectors are mapped back to their old paths. That covers every node count and every recipe fingerprint, including the 454 / 411 rows that now use selectors. Eval grouping is unchanged, because old paths map one-to-one onto distinct selectors.
  3. Launcher text steps. On each materialized variant versus its old file, I ran the power-lane awk and GNU-sed-equivalent name / max_attempts / dist-timeout patches. The power lane is identical for all 357, including the 51 power recipes. The only difference is max_attempts on 15 GB200 AgentX recipes: the originals use inline health_check: {…}, which the b200-nscale sed never matched. Those recipes run through launch_gb200-nv.sh, which has no such patch, so there is no runtime change.
  4. Comments. Every original comment line is present, placed above the same key path. Comments shared by all configurations sit in base; configuration-specific ones sit inside that configuration's override.

perf-changelog

This PR has deliberately no perf-changelog.yaml entry. No resolved recipe, node count or fingerprint changes. An entry covering these config keys would re-run essentially every NVIDIA multi-node benchmark on merge to reproduce identical configurations. Reviewers can overrule this if the repository rule should apply regardless.

Testing

  • pytest utils/ runners/: 1711 passed locally. Skipped only suites needing deps not installed locally.
  • No GPU run. Recommended before merge: one sweep-enabled smoke per affected launcher (b200-nscale, b300-dsxe, gb200-nv, gb300-nv, h100-dgxc, h200-dgxc), to exercise materialize_srt_configs on the real login nodes, including uv run --with pyyaml.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it's a very large, mechanical restructuring (402 files, +16.8k/-54.6k lines) that changes how ~357 recipes are addressed (moving to positional zip_override_variants[i] indices) and is stacked on an unmerged base branch, a human look would still be worthwhile before merge.

What was reviewed: the variants.yaml consolidation pattern (shared base + zip_override_variants, header mapping old files to indices) in several directories, the corresponding configs/nvidia-master.yaml selector rewrites (spot-checked several CONFIG_FILE=...:zip_override_variants[i] mappings against the variant headers), and the three directories intentionally left flat. The perf-changelog omission and positional-index fragility were already flagged and reasoned through in the ruled-out list.

Extended reasoning...

Large mechanical PR (402 files) consolidating srt-slurm benchmark recipe YAMLs into per-directory variants.yaml files with zip_override_variants, plus corresponding configs/nvidia-master.yaml selector updates; no application/security-sensitive code is touched, only benchmark config data. Spot-checked one variants.yaml (dsr1/sglang/b200-fp8/8k1k) and matching nvidia-master.yaml selector edits (dsr1/trtllm/b200-fp4/8k1k) and both look internally consistent with the stated header-to-index mapping. Decided defer over approve because of sheer scale, the new positional-index addressing scheme (a design choice a human should weigh in on), and that the PR is stacked on an unmerged base branch.

…variants

Replace 357 flat recipes referenced by the active NVIDIA master config with
one variants.yaml per directory (40 directories): a shared base plus one
override_<name> block per benchmark configuration, named after the replaced
file (disagg-1p1d-dep8-b8-eplb0-mtp3.yaml -> override_disagg_1p1d_dep8_b8_eplb0_mtp3)
and keeping its original recipe name. Master entries select
variants.yaml:override_<name>; srt-slurm-recipe-identities.yaml maps each
selector back to the file it replaced.

No resolved recipe changes:
- every override expands (srtctl generate_override_configs and infx) to exactly
  the flat recipe it replaces;
- the full multi-node matrix, including node counts and recipe fingerprints,
  is identical in default and all-evals modes;
- launcher text steps (power-lane detection, name/max_attempts/dist-timeout
  patches) give the same result on the materialized variant, except
  max_attempts on 15 GB200 recipes whose launcher does not apply that patch.

Every original comment is kept beside the same key; comments specific to one
configuration sit inside its override. Skipped: dsr1/sglang/b200-fp4/8k1k and
dsr1/trtllm/b200-fp8/8k1k (a differing explicit null deletes rather than sets)
and glm5.2/sglang/h200-fp8/agentx (a non-schema-2 sibling). Recipes used by
deprecated configs are untouched.
@cquil11
cquil11 force-pushed the ix/srt-recipe-zip-consolidation branch from aec6c5b to 85c369c Compare September 23, 2026 17:01
@cquil11 cquil11 changed the title srt-slurm: consolidate sibling recipes into zip_override variants srt-slurm: consolidate sibling recipes into named override variants Sep 23, 2026
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants