Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

InferenceX owns the recipes in this directory. Every NVIDIA srt-slurm launcher uses `setup_srt_slurm()` in [`runners/slurm_utils.sh`](../../../runners/slurm_utils.sh), makes a job-local Git clone of the pinned submodule, and copies this entire tree into `recipes/`. The shared helper records the actual revision in `srt-slurm-sha.txt`; power lanes copy that revision into `power-producer-sha.txt` for result validation.

The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.30.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.30.0) (`0b37c791fc95a7cb42e8d2b281d44b642cda75d3`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers.
The shared version is the Git submodule pointer at [`utils/srt-slurm`](../../../utils/srt-slurm), currently [v2.36.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.36.0) (`7b5863a7837673d81403b076be219bbf18a7700f`). Update that submodule pointer when upgrading, then run the recipe and integration checks. Do not add model-specific checkout branches to launchers.

InferenceX requires srt-slurm 2.0 or newer and `schema: 2` recipes. Legacy recipe layouts are unsupported; migrate them before adding them to this tree.

Expand Down Expand Up @@ -69,7 +69,7 @@ Validate recipes with the exact launcher pin, including all override variants. F

The initial migration also resolves compatibility issues that `srtctl migrate` cannot fix itself:

- SGLang Model Gateway recipes use `frontend.type: sglang-router`; in v2.30.0, `sglang` selects a direct worker without a router.
- SGLang Model Gateway recipes use `frontend.type: sglang-router`; in v2.36.0, `sglang` selects a direct worker without a router.
- Duplicate YAML keys retain the value selected by the former PyYAML loader.
- DCGM telemetry uses `collect_interval_ms: 1000` instead of `provider` and `default_frequency`. The collector derives its shutdown budget; an explicit ten-second budget is too short for the current validator. Dedicated discovery-service placement is preserved from the original recipes. The pinned upstream runtime rejects telemetry with dedicated infrastructure nodes; this remains a power compatibility blocker rather than changing the original topology to satisfy validation. H200 custom recipes declare a default concurrency that the launcher replaces before submission.
- DeepSeek-V4 vLLM benchmarks use the supported `custom_tokenizer` loader. Retired `warmup_req_rate: inf` fields are removed; the current upstream client uses its fixed warmup rate of 250 requests per second.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@

InferenceX 负责维护本目录中的配置。所有 NVIDIA srt-slurm 启动器均调用 [`runners/slurm_utils.sh`](../../../runners/slurm_utils.sh) 中的 `setup_srt_slurm()`,为作业创建固定版本子模块的本地 Git 克隆,并将整个目录复制到 `recipes/`。共享函数将实际提交记录到 `srt-slurm-sha.txt`;功耗测试路径还会将其复制到 `power-producer-sha.txt`,供结果校验使用。

统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.30.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.30.0)(`0b37c791fc95a7cb42e8d2b281d44b642cda75d3`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。
统一版本由 [`utils/srt-slurm`](../../../utils/srt-slurm) 的 Git 子模块指针指定,目前为 [v2.36.0](https://github.com/NVIDIA/srt-slurm/releases/tag/v2.36.0)(`7b5863a7837673d81403b076be219bbf18a7700f`)。升级时更新该子模块指针,然后运行配置和集成检查。不要在启动器中新增按模型选择检出版本的分支。

InferenceX 要求 srt-slurm 2.0 或更新版本,且配置必须声明 `schema: 2`。不支持旧版配置结构;加入本目录前必须先完成迁移。

Expand Down Expand Up @@ -69,7 +69,7 @@ python -m infx.matrix.generate full-sweep \

本次迁移还修复了 `srtctl migrate` 无法自动处理的兼容性问题:

- SGLang Model Gateway 配置使用 `frontend.type: sglang-router`;在 v2.30.0 中,`sglang` 表示不经过路由器的独立工作进程。
- SGLang Model Gateway 配置使用 `frontend.type: sglang-router`;在 v2.36.0 中,`sglang` 表示不经过路由器的独立工作进程。
- 对重复的 YAML 键,保留原 PyYAML 加载器实际采用的值。
- DCGM 遥测使用 `collect_interval_ms: 1000`,替代 `provider` 和 `default_frequency`。采集器自动推导退出等待时间;原先显式设置的十秒不满足当前校验要求。保留原配置中服务发现进程的专用节点部署方式。固定的上游版本不支持在专用基础设施节点上启用遥测;该功耗兼容性问题仍待解决,不通过改变原有拓扑来绕过校验。H200 自定义配置声明默认并发数,提交前由启动器替换。
- DeepSeek-V4 vLLM 基准测试使用受支持的 `custom_tokenizer` 加载器。删除已废弃的 `warmup_req_rate: inf` 字段;当前上游客户端的预热速率固定为每秒 250 个请求。
Expand Down
Loading
Loading