Skip to content

CollectiveX: point mi300x at the compute-0 partition - #3538

Merged
Oseltamivir merged 2 commits into
mainfrom
cx-mi300x-partition
Sep 28, 2026
Merged

Oseltamivir merged 2 commits into
mainfrom
cx-mi300x-partition

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #3537.

platform_config.json sent mi300x shards to Slurm partition compute, which does not exist on the mi300x-amd cluster. salloc failed with "invalid partition specified: compute", as the mi300x swap-blocks shard did in 36334394602.

That cluster's partition is compute-0, which runners/launch_mi300x-amd.sh and runners/srt-slurm/mi300x-amd.yaml already use. The squash directory (/raid/inferencex/squash) already matched.

Validation

Swap-blocks on mi300x (36378673274) is green: salloc now lands on compute-0, and all 126 cases (63 random + 63 contiguous layout) pass correctness.

@Oseltamivir
Oseltamivir requested a review from a team as a code owner September 28, 2026 04:40
@Oseltamivir
Oseltamivir added this pull request to stack #3539 September 28, 2026 04:45

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, straightforward config fix.

What was reviewed: the single-line change in experimental/CollectiveX/configs/platform_config.json, which renames the mi300x entry's operator.partition value from "compute" to "compute-0". Confirmed this is a pure JSON data value change with no code, schema, or test modifications, and no other platform entries are affected.

Extended reasoning...

The diff touches exactly one line in a JSON config file for an experimental benchmarking tool, changing a Slurm partition name string; there is no code, auth, or security-sensitive surface involved. The change is mechanical and self-contained, matching the approval criteria for simple config tweaks, and no CODEOWNER or outstanding reviewer objections were found.

The default graph usage mode (2, mixing) wraps every captured collective in an external
event wait and record. Graphed HT decode captures one host collective per pair, the routing
ncclAllGather, and the h100 EP8 trace showed ~14us idle before it, making the graphed pair
period 1.14x eager. Mode 1 (one graph at a time, never concurrent with uncaptured work on the
communicator) removes it. Graphed HT decode rows carry the -gum1 generation suffix.
The mi300x-amd runners submit to a cluster whose partition is compute-0 (as
runners/launch_mi300x-amd.sh and runners/srt-slurm/mi300x-amd.yaml already use); the
registry's "compute" failed every mi300x shard at salloc.
Base automatically changed from cx-nccl-graph-mode to main September 28, 2026 10:10
@Oseltamivir
Oseltamivir merged commit a4c497c into main Sep 28, 2026
1 check passed
@Oseltamivir
Oseltamivir deleted the cx-mi300x-partition branch September 28, 2026 10:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant