Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
ec73d1c
refactor(launch): port hardware launchers to a pluggable Python launcher
adibarra Sep 29, 2026
a2c073b
Merge remote-tracking branch 'origin/main' into feat/python-launchers
adibarra Sep 29, 2026
d18ce8b
test(operatorx): use the clusters runner schema in the generator fixture
adibarra Sep 29, 2026
ffc07e1
refactor(launch): resolve models from recipe aliases and normalize sr…
adibarra Sep 29, 2026
33ad5ab
refactor(launch): trim srt driver helpers, comments and tests
adibarra Sep 29, 2026
82d3ae0
fix(launch): default the srt account to the user's Slurm default; iso…
adibarra Sep 29, 2026
2d930e6
fix(launch): request typed GRES on CW clusters and skip --segment for…
adibarra Sep 29, 2026
1069244
fix(launch): let clusters opt single-node srt jobs out of --exclusive
adibarra Sep 29, 2026
6280dc6
fix(launch): give single-node srt jobs the same 2 h health-check budg…
adibarra Sep 29, 2026
1e66256
fix(launch): cluster runtime settings override the runner host's envi…
adibarra Sep 29, 2026
c3ce5ae
fix(launch): namespace srtctl jobs, restore pristine model defaults, …
adibarra Sep 29, 2026
52802e5
fix(clusters): mi300x-amd partition is now MI300X-UBUNTU
adibarra Sep 29, 2026
4ce5c26
fix(launch): drop the whole-node CPU request for TRT-LLM srt jobs
adibarra Sep 29, 2026
de5fb0e
fix(launch): split the node's CPU budget across TRT-LLM's per-GPU tasks
adibarra Sep 29, 2026
fc6a81e
fix(clusters): request b300 CPUs per GPU, not per task
adibarra Sep 29, 2026
df1a780
fix(clusters): list the gb200-nv runners that exist
adibarra Sep 29, 2026
95f13d4
Merge origin/main into feat/python-launchers
adibarra Sep 29, 2026
5cccb7c
fix(bench): run the multi-node client with PYTHONSAFEPATH instead of -P
adibarra Sep 29, 2026
bede049
chore: drop inline comments from the launcher port
adibarra Sep 29, 2026
34f50c2
refactor(launch): move the srtctl job tag and MI355X RDMA TC into run…
adibarra Sep 29, 2026
aeee6a9
Merge remote-tracking branch 'origin/main' into feat/python-launchers
adibarra Sep 29, 2026
1fa6e2f
revert(bench): restore main's comments in files with no code change
adibarra Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 8 additions & 9 deletions .agents/skills/debug-runs/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,9 +24,9 @@ paths live in an access-controlled **InferenceX Clusters** Slack canvas, NOT in

Before SSHing to a cluster, look up that cluster's row in the canvas for: **login address**,
**GHA runner user**, **runner directory**, any **jumpbox / ProxyJump**, whether it's
**Slurm or bare-metal**, and the **per-node host RAM**. The matching
`inferencex-e2e/runners/launch_<cluster>.sh` is the source of truth for the exact container image mounts
and the benchmark command.
**Slurm or bare-metal**, and the **per-node host RAM**. The cluster's `clusters.<id>` record in
`inferencex-e2e/configs/runners.yaml` and the launch drivers under `inferencex-e2e/infx/launch/`
are the source of truth for the exact container image, mounts and benchmark command.

- If you **can't read the canvas** (no Slack access, or unsure), **ask the user** for the
cluster's SSH target + runner user rather than guessing or pasting infra into the repo.
Expand Down Expand Up @@ -104,13 +104,12 @@ Steps:

1. Use the job or runner name to identify the node. Look up that cluster's access details in
the canvas, then SSH in with `ssh -A` when a jumpbox or agent forwarding is involved.
2. Reproduce the exact benchmark the launcher runs. Single-node jobs take the
`native-single-node` path in `inferencex-e2e/runners/launch_<cluster>.sh`, which calls
`launch_srt_single_node <cluster>` in `inferencex-e2e/runners/slurm_utils.sh`: it validates the
2. Reproduce the exact benchmark the launcher runs. Single-node jobs go through the srt driver
of `python -m infx.launch run` (`inferencex-e2e/infx/launch/drivers/srt/`): it binds the
master row's `srt-recipe:` (`inferencex-e2e/benchmarks/single_node/srt-slurm-recipes/<model>/<engine>/<sku>-<precision>[-mtp]/<scenario>.yaml`)
with `python3 -m infx.srt_slurm.single_node prepare`, then submits it through srtctl
with the `inferencex-e2e/runners/srt-slurm/<cluster>.yaml` profile. Read the recipe for the image
(`model.container`), server args and env, and the launcher for mounts and the job env
with `python -m infx.srt_slurm.single_node prepare`, renders the job-local `srtslurm.yaml`
from the cluster's record, then submits it through srtctl. Read the recipe for the image
(`model.container`), server args and env, and the cluster record for mounts and the job env
(`IMAGE`, `TP`, `PRECISION`, `SPEC_DECODING`, `CONC`, …). On Slurm clusters, use
`salloc` or `srun` with the squash image. On the **bare-metal `-tw` pools, use `docker run`**
on the node directly without `srun`.
Expand Down
2 changes: 1 addition & 1 deletion .claude/commands/add-model-hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,7 +130,7 @@ Confirm which master file by SKU: `mi*` → `amd-master.yaml`, everything else

## Step 4 — no launcher routing

Single-node points with an `srt-recipe:` go through `launch_srt_single_node`, which picks the
Single-node points with an `srt-recipe:` go through the srt driver of `python -m infx.launch`, which picks the
one recipe variant whose TP/GPU count, `CONC`, `KV_OFFLOADING` and image match the matrix
point (`inferencex-e2e/infx/srt_slurm/single_node.py::select_recipe`). No per-script launcher routing is
needed; if a point matches zero or several variants, fix the recipe, not the launcher.
Expand Down
4 changes: 2 additions & 2 deletions .github/klaud-candidate-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,8 +61,8 @@ candidate.json provides the planner-verified exact `baseline-model`; use that va
The planner supplies `baseline-preflight.json` beside candidate.json. It contains the verified
benchmark roster bound to the selected candidate, base SHA, source observation and model.
After resolving the exact old/new image goal, prepare-baseline checks this binding and uses
that roster without refetching it. If candidate.json requires the preflight and it is absent
or invalid, stop with `baseline-preflight-mismatch`; only legacy candidates may reconstruct.
that roster without refetching it. If the preflight is absent or invalid, stop with
`baseline-preflight-mismatch`.
The preflight is not a published or final baseline. Supplement verified public
eval/dataset evidence before freezing; never replace a failed lookup with a partial roster.
Never reduce the baseline to overlapping points, displayed rows or a smaller current family. Never dispatch the old
Expand Down
49 changes: 29 additions & 20 deletions .github/workflows/benchmark-multinode-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -312,18 +312,36 @@ jobs:
path: .result-tooling
persist-credentials: false

- name: Set up uv for result processing
- name: Set up uv
uses: astral-sh/setup-uv@20cfd1bf945f4377ade1205e4dbc17946fc9a30d # v10.0.1
with:
enable-cache: false

- name: Prepare result-processing Python
env: &uv-setup-env
UV_CACHE_DIR: ${{ github.workspace }}/../.infx-uv-cache
UV_PYTHON_INSTALL_DIR: ${{ github.workspace }}/../.infx-uv-python
run: |
uv python install --quiet --no-bin 3.12
result_python=$(uv python find --system 3.12)
"$result_python" --version
echo "INFERENCEX_RESULTS_PYTHON=$result_python" >> "$GITHUB_ENV"

- name: Prepare launcher Python
env: *uv-setup-env
run: |
launch_venv="$RUNNER_TEMP/infx-launch-venv"
uv venv --quiet --python 3.12 "$launch_venv"
uv pip install --quiet --exclude-newer PT12H --python "$launch_venv/bin/python" \
-r "$GITHUB_WORKSPACE/.result-tooling/inferencex-e2e/pyproject.toml"
echo "INFERENCEX_LAUNCH_PYTHON=$launch_venv/bin/python" >> "$GITHUB_ENV"

- name: Launcher cleanup (pre-run)
env:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: &launcher-cleanup |
"$INFERENCEX_LAUNCH_PYTHON" -P -m infx.launch cleanup

- name: Launch multi-node job script
env:
VALIDATION_BENCHMARK_LIB: ${{ github.workspace }}/.result-tooling/inferencex-e2e/benchmarks/benchmark_lib.sh
Expand All @@ -337,31 +355,16 @@ jobs:
export GITHUB_WORKSPACE="$PWD"
echo "INFERENCEX_E2E_ROOT=$PWD" >> "$GITHUB_ENV"
set -x
if [[ -f infx/results/result_filename.py ]]; then
RESULT_FILENAME=$(python3 -m infx.results.result_filename)
else
# Older measured commits predate the helper; preserve safe identity.
RESULT_FILENAME=$(printf '%s\0' "$RESULT_FILENAME_BASE" "$RECIPE_FINGERPRINT" | sha256sum)
RESULT_FILENAME="${RESULT_FILENAME%% *}"
fi
RESULT_FILENAME=$(python3 -m infx.results.result_filename)
export RESULT_FILENAME
# Export RESULT_FILENAME early so it's available for artifact uploads even if cancelled
echo "RESULT_FILENAME=${RESULT_FILENAME}" >> "$GITHUB_ENV"
echo "EVAL_ARTIFACT_RECIPE=${RECIPE_FINGERPRINT:0:16}" >> "$GITHUB_ENV"
eval_artifact_conc="$(python3 -c 'import hashlib,os; print(hashlib.sha256(os.environ["CONC_LIST"].encode()).hexdigest()[:12])')"
echo "EVAL_ARTIFACT_CONC=${eval_artifact_conc}" >> "$GITHUB_ENV"

# Historical measured revisions own their configuration in the launchers.
if [[ -f benchmarks/runtime_settings.sh ]]; then
source benchmarks/runtime_settings.sh
fi
if [[ -f runners/runtime_settings.sh ]]; then
source runners/runtime_settings.sh
fi
# Historical measured revisions own their configuration in the launchers.
if [[ -f benchmarks/multi_node/runtime_settings.sh ]]; then
source benchmarks/multi_node/runtime_settings.sh
fi
source benchmarks/runtime_settings.sh
source benchmarks/multi_node/runtime_settings.sh
# Assign data literally; recipe values must never become shell code.
settings_json=$(jq -cen \
--argjson prefill "$PREFILL_ADDITIONAL_SETTINGS" \
Expand Down Expand Up @@ -390,7 +393,7 @@ jobs:
fi
export EVAL_CONC
export IS_MULTINODE=true
bash "./runners/launch_${RUNNER_NAME%%_*}.sh"
"$INFERENCEX_LAUNCH_PYTHON" -m infx.launch run
if [ "${EVAL_ONLY}" = "true" ]; then
echo "Eval-only mode: skipping benchmark result file check"
# Verify eval produced results
Expand Down Expand Up @@ -545,3 +548,9 @@ jobs:
- name: Slurm cleanup (post-run)
if: always()
run: *slurm-cleanup

- name: Launcher cleanup (post-run)
if: ${{ always() && env.INFERENCEX_LAUNCH_PYTHON != '' }}
env:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: *launcher-cleanup
78 changes: 40 additions & 38 deletions .github/workflows/benchmark-tmpl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -206,7 +206,7 @@ jobs:
) ||
format('[{0}]', toJSON(inputs.runner))
) }}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Single-node CI jobs on gb200-nv running dsv41flash sglang agentic-coding at conc>=64 now get only the 500-minute default job timeout instead of 750, so they get killed by GitHub Actions before finishing where the base branch let them run to completion. The refactor's new timeout-minutes expression keeps the h200-dgxc 1470-minute bump but drops the inputs.runner == 'cluster:gb200-nv' && ... && 750 branch, falling straight to || 500. Fix: restore a gb200-nv (or equivalent) timeout bump for this workload so its long points keep the time they had on main, alongside the surviving h200-dgxc bump.

Why this was flagged

On the base branch (e90870b), benchmark-tmpl.yml:208 sets timeout-minutes to 750 when inputs.runner == 'cluster:gb200-nv' and the config is model-prefix dsv41flash, framework sglang, scenario-type agentic-coding, not eval-only, and conc >= 64. The diff's version of that line keeps only the h200-dgxc 1470-minute clause and falls through to the 500-minute default for every other runner, silently removing the gb200-nv clause. This combination is a real, active config: inferencex-e2e/configs/nvidia-master.yaml defines dsv41flash-fp4-gb200-sglang-agentic-dspark with conc-list up to 128 for gb200-nv sglang. Any dispatched job matching that config on gb200-nv now runs into the GitHub Actions 500-minute step timeout and is cancelled mid-run instead of completing, a regression versus main's 750-minute allowance; no other new timeout mechanism in the diff covers this case.

Verification: normal. The base benchmark-tmpl.yml:208 timeout-minutes expression had two bump clauses: `inputs.runner == 'cluster:h200-dgxc' && ... && 1470 || inputs.runner == 'cluster:gb200-nv' && fromJSON(inputs.config).model-prefix == 'dsv41flash' && fromJSON(inputs.config).framework == 'sglang' && inputs.scenario-type == 'agentic-coding' && !inputs.eval-only && fromJSON(inputs.config).conc >= 64 && 750…

timeout-minutes: ${{ inputs.runner == 'cluster:h200-dgxc' && fromJSON(inputs.config).model-prefix == 'dsv41flash' && fromJSON(inputs.config).framework == 'sglang' && inputs.scenario-type == 'agentic-coding' && !inputs.eval-only && fromJSON(inputs.config).conc >= 64 && 1470 || inputs.runner == 'cluster:gb200-nv' && fromJSON(inputs.config).model-prefix == 'dsv41flash' && fromJSON(inputs.config).framework == 'sglang' && inputs.scenario-type == 'agentic-coding' && !inputs.eval-only && fromJSON(inputs.config).conc >= 64 && 750 || 500 }}
timeout-minutes: ${{ inputs.runner == 'cluster:h200-dgxc' && fromJSON(inputs.config).model-prefix == 'dsv41flash' && fromJSON(inputs.config).framework == 'sglang' && inputs.scenario-type == 'agentic-coding' && !inputs.eval-only && fromJSON(inputs.config).conc >= 64 && 1470 || 500 }}
name: >-
${{ inputs.klaud-run && 'klaud | ' || '' }}p${{ inputs.priority }} | ${{ fromJSON(inputs.config).model-prefix }} ${{ fromJSON(inputs.config).precision }} ${{ inputs.runner }} ${{ fromJSON(inputs.config).framework == 'sglang' && 'sgl' || fromJSON(inputs.config).framework == 'dynamo-sglang' && 'dyn-sgl' || fromJSON(inputs.config).framework == 'sglang-disagg' && 'sgl-disagg' || fromJSON(inputs.config).framework }}
TP${{ fromJSON(inputs.config).tp }}${{ format('{0}', fromJSON(inputs.config).pp) != '' && format('{0}', fromJSON(inputs.config).pp) != '1' && format('/PP{0}', fromJSON(inputs.config).pp) || '' }}${{ format('{0}', fromJSON(inputs.config).dcp-size) != '' && format('{0}', fromJSON(inputs.config).dcp-size) != '1' && format('/DCP{0}', fromJSON(inputs.config).dcp-size) || '' }}${{ format('{0}', fromJSON(inputs.config).pcp-size) != '' && format('{0}', fromJSON(inputs.config).pcp-size) != '1' && format('/PCP{0}', fromJSON(inputs.config).pcp-size) || '' }}${{ format('{0}', fromJSON(inputs.config).ep) != '' && format('{0}', fromJSON(inputs.config).ep) != '1' && format('/EP{0}', fromJSON(inputs.config).ep) || '' }}${{ inputs.dp-attn && '/DPA' || '' }}
Expand All @@ -222,26 +222,16 @@ jobs:
sudo chown -R "$(id -u):$(id -g)" "$GITHUB_WORKSPACE"
fi

- name: Resource cleanup (pre-run)
run: &resource-cleanup |
# Cleanup Docker resources
if command -v docker >/dev/null 2>&1 && docker info >/dev/null 2>&1; then
echo "[Docker] Cleaning up resources ..."
docker ps -aq | xargs -r docker rm -f
docker network prune -f
while [ -n "$(docker ps -aq)" ]; do
docker ps -a
sleep 5
done
fi

# Cleanup SLURM resources
- name: Slurm cleanup (pre-run)
run: &slurm-cleanup |
if command -v squeue >/dev/null 2>&1; then
echo "[Slurm] Cleaning up jobs with name: ${RUNNER_NAME} ..."
scancel --name="${RUNNER_NAME}" || true
while [ -n "$(squeue --name="${RUNNER_NAME}" --noheader --format='%i')" ]; do
squeue --name="${RUNNER_NAME}"
sleep 5
for job_name in "${RUNNER_NAME}" "inferencex-${RUNNER_NAME}"; do
echo "[Slurm] Cleaning up jobs with name: ${job_name} ..."
scancel --name="${job_name}" || true
while [ -n "$(squeue --name="${job_name}" --noheader --format='%i')" ]; do
squeue --name="${job_name}"
sleep 5
done
done
fi

Expand Down Expand Up @@ -285,18 +275,36 @@ jobs:
path: .result-tooling
persist-credentials: false

- name: Set up uv for result processing
- name: Set up uv
uses: astral-sh/setup-uv@20cfd1bf945f4377ade1205e4dbc17946fc9a30d # v10.0.1
with:
enable-cache: false

- name: Prepare result-processing Python
env: &uv-setup-env
UV_CACHE_DIR: ${{ github.workspace }}/../.infx-uv-cache
UV_PYTHON_INSTALL_DIR: ${{ github.workspace }}/../.infx-uv-python
run: |
uv python install --quiet --no-bin 3.12
result_python=$(uv python find --system 3.12)
"$result_python" --version
echo "INFERENCEX_RESULTS_PYTHON=$result_python" >> "$GITHUB_ENV"

- name: Prepare launcher Python
env: *uv-setup-env
run: |
launch_venv="$RUNNER_TEMP/infx-launch-venv"
uv venv --quiet --python 3.12 "$launch_venv"
uv pip install --quiet --exclude-newer PT12H --python "$launch_venv/bin/python" \
-r "$GITHUB_WORKSPACE/.result-tooling/inferencex-e2e/pyproject.toml"
echo "INFERENCEX_LAUNCH_PYTHON=$launch_venv/bin/python" >> "$GITHUB_ENV"

- name: Launcher cleanup (pre-run)
env:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: &launcher-cleanup |
"$INFERENCEX_LAUNCH_PYTHON" -P -m infx.launch cleanup

- name: Launch job script
env:
RUNNER_NAME: ${{ runner.name }}
Expand All @@ -309,28 +317,16 @@ jobs:
if [[ -f inferencex-e2e/configs/runners.yaml ]]; then cd inferencex-e2e; fi
export GITHUB_WORKSPACE="$PWD"
echo "INFERENCEX_E2E_ROOT=$PWD" >> "$GITHUB_ENV"
if [[ -f infx/results/result_filename.py ]]; then
RESULT_FILENAME=$(python3 -m infx.results.result_filename)
else
# Older measured commits predate the helper; preserve safe identity.
RESULT_FILENAME=$(printf '%s\0' "$RESULT_FILENAME_BASE" "$RECIPE_FINGERPRINT" | sha256sum)
RESULT_FILENAME="${RESULT_FILENAME%% *}"
fi
RESULT_FILENAME=$(python3 -m infx.results.result_filename)
export RESULT_FILENAME
export GPU_COUNT=$((TP * PP_SIZE * PCP_SIZE))
echo "GPU_COUNT=${GPU_COUNT}" >> "$GITHUB_ENV"

# Export RESULT_FILENAME early so it's available for artifact uploads even if cancelled
echo "RESULT_FILENAME=${RESULT_FILENAME}" >> "$GITHUB_ENV"

# Historical measured revisions own their configuration in the launchers.
if [[ -f benchmarks/runtime_settings.sh ]]; then
source benchmarks/runtime_settings.sh
fi
if [[ -f runners/runtime_settings.sh ]]; then
source runners/runtime_settings.sh
fi
bash "./runners/launch_${RUNNER_NAME%%_*}.sh"
source benchmarks/runtime_settings.sh
"$INFERENCEX_LAUNCH_PYTHON" -m infx.launch run

if [ "${EVAL_ONLY}" = "true" ]; then
echo "Eval-only mode: skipping benchmark result file check"
Expand Down Expand Up @@ -513,6 +509,12 @@ jobs:
sudo chown -R "$(id -u):$(id -g)" "$GITHUB_WORKSPACE"
fi

- name: Resource cleanup (post-run)
- name: Slurm cleanup (post-run)
if: always()
run: *resource-cleanup
run: *slurm-cleanup

- name: Launcher cleanup (post-run)
if: ${{ always() && env.INFERENCEX_LAUNCH_PYTHON != '' }}
env:
PYTHONPATH: ${{ github.workspace }}/.result-tooling/inferencex-e2e
run: *launcher-cleanup
1 change: 1 addition & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@ on:
- 'inferencex-e2e/.python-version'
- 'inferencex-e2e/infx/ruff.toml'
- '**/pytest.ini'
- 'inferencex-e2e/configs/**'
- 'inferencex-e2e/utils/srt-slurm'
push:
branches: [main]
Expand Down
25 changes: 9 additions & 16 deletions .github/workflows/claude.yml
Original file line number Diff line number Diff line change
Expand Up @@ -466,32 +466,25 @@ jobs:
- Comment: "Image must be publicly accessible on NGC, Docker Hub, or another public registry. Local paths like `/scratch/...` or `.sqsh` files are generally not accepted. Please push the container to a public registry (e.g., `nvcr.io/nvidia/...` for NGC) and update the config with the public image reference."
- Link to the specific line with the invalid image path

## Enroot Import Validation for Launch Scripts:
When reviewing changes to `inferencex-e2e/runners/launch_*.sh` files, verify that the script properly transforms public Docker images to enroot local images for reproducibility.
## Enroot Import Validation for Launch Code:
When reviewing changes to `inferencex-e2e/infx/launch/` (including `infx/launch/backends/`) or a cluster's `slurm.squash` record in `inferencex-e2e/configs/runners.yaml`, verify that public Docker images still reach containers through the cluster's backend for reproducibility, and that no driver hand-rolls its own `enroot import`.

**Expected pattern:**
The script should include an enroot import command like:
```bash
srun --jobid=$JOB_ID bash -c "enroot import -o $SQUASH_FILE docker://$IMAGE"
```
or similar variations such as:
```bash
srun -N 1 -A $SLURM_ACCOUNT -p $SLURM_PARTITION bash -c "enroot import -o $SQUASH_FILE docker://$IMAGE"
```
Drivers get images from the backend: `backend.prepare_image(image)` for containers the backend runs, or, in the Slurm-only srt driver, `SlurmBackend.stage_image(...)`. On Slurm these call `ensure_image` in `infx/launch/backends/slurm/squash.py`, which imports `docker://` references into `.sqsh` files according to the cluster's `slurm.squash.import` mode (`submit-host`, `compute`, `all-nodes`, `pre-staged`). A cluster without `slurm.squash` hands Pyxis the registry reference to import inside the job.

**Why this matters:**
- Ensures the exact same public NGC/Docker Hub image is used
- Makes benchmarks reproducible by anyone with access to the public image
- Prevents reliance on pre-existing local container images that others cannot access

**Validation Steps:**
1. Look for `enroot import` commands that convert `docker://` images to local `.sqsh` files
2. The image source should be a public registry (NGC, Docker Hub, etc.), not a local path
3. If the script uses containers but does NOT have an `enroot import docker://` pattern:
1. Look for container images started without going through the backend's image preparation.
2. The image source should be a public registry (NGC, Docker Hub, etc.), not a local path.
3. If a driver starts a container without the backend, or a cluster switches to `pre-staged` without explanation:
- This is a 🟡 **WARNING** issue
- Comment: "This launch script uses container images but does not appear to transform a public Docker image to an enroot local image using `enroot import -o $SQUASH_FILE docker://$IMAGE`. For reproducibility, please either:
1. Add the enroot import pattern to pull from a public registry, OR
2. Explain why this script has a different workflow (e.g., uses a different container runtime, pre-built images are acceptable for this use case, etc.)"
- Comment: "This launch code starts a container image without staging it through the cluster backend (`prepare_image` / `stage_image`). For reproducibility, please either:
1. Stage the public registry image through the backend, OR
2. Explain why this path has a different workflow (e.g., pre-staged images are acceptable for this cluster)"
- Ask the developer to provide a reasonable explanation if the pattern is intentionally omitted

## Line Count Report for the matrix generator:
Expand Down
Loading
Loading