Skip to content

CollectiveX: run nccl-ep's communicator in NCCL graph usage mode 1 - #3537

Merged
Oseltamivir merged 1 commit into
mainfrom
cx-nccl-graph-mode
Sep 28, 2026
Merged

Oseltamivir merged 1 commit into
mainfrom
cx-nccl-graph-mode

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Graphed nccl-ep HT decode was the one graphed row family that ran slower than eager in the full sweep 36334392363: pair period 1.03× the eager Sep 25 baseline on average, and up to 1.30×.

Cause

An h100 EP8 trace (torch.profiler/CUPTI, HT decode, graphed vs eager) showed that every kernel's duration is identical in both regimes. The whole graphed penalty is idle gaps between kernels, dominated by about 14 µs just before the routing ncclAllGather (13.7 µs graphed vs 1.5 µs eager at T=1).

That all-gather is the only host NCCL collective in any capture. NCCL's default graph usage mode is 2 ("mixing"): it wraps every captured collective in an external event wait and an event record, so graph and uncaptured work can interleave on one communicator (NCCL 2.30.7 src/misc/strongstream.cc:241-247, :347).

Fix

nccl-ep's own communicator is created with NCCLConfig(graph_usage_mode=1): one graph at a time, never concurrent with uncaptured work on that communicator. That is how both the harness and a captured decode step use it. In NCCL 2.30.7, modes 0 and 1 behave the same; only == 2 enables the serialization.

Graphed HT decode rows get the -gum1 generation suffix. LL graphs capture no host collective and HT prefill stays eager, so neither changes.

Trace result, h100 EP8 HT decode:

T eager graphed, mode 2 graphed, mode 0/1
1 127.2 µs 144.5 µs 131.0 µs
256 267.8 µs 283.6 µs 266.6 µs

Ruled out with no effect: NCCL_GRAPH_STREAM_ORDERING=0, NCCL_MEM_SYNC_DOMAIN=0, NCCL_CGA_CLUSTER_SIZE=0, and a symmetric routing window (which switches NCCL to its NVLS all-gather kernel but leaves the gap).

At EP16 the captured all-gather also keeps NCCL's per-replay proxy host callback (enqueue.cc:1703-1722). This change does not touch that.

Validation

nccl-ep normal mode, h100 and h200, EP8 and EP16 (36378559608, 36378561991): all 4 shards green, 56/56 rows pass correctness, and graphed HT decode rows carry nccl-ep-v02-ht-routed-zc-static-gum1-cudagraph.

Graphed HT decode pair period, geomean over T=1..512 (range in brackets):

vs main's graphed rows (36334392363) vs eager (Sep 25 baseline)
h100 EP8 0.886× (0.76–0.94) 0.976× (max 1.00)
h100 EP16 0.923× (0.83–0.96) 1.004× (max 1.01)
h200 EP8 0.885× (0.78–0.93) 0.989× (max 1.01)
h200 EP16 0.910× (0.84–0.95) 0.940× (max 0.97)

Graphed HT decode goes from 1.03× eager on average (up to 1.30×) to parity or better on every cell. The largest gains are at T=256–512, where main's graphed rows ran furthest behind (h100 EP8 T=512: 418.4 → 317.0 µs). The eager baseline predates the full-plane combine (-zc rather than -zc-static), so treat those ratios as approximate.

HT prefill (eager, unchanged generation) stays at 0.93–1.02× main, which is run-to-run spread.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment on lines +159 to +162
elif self.cuda_graph_supported:
# Only graphed HT captures a host NCCL collective (the routing ncclAllGather), so
# only its rows change with the communicator's graph usage mode.
self.kernel_generation = f"{type(self).kernel_generation}-gum1"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) When COLLX_CUDA_GRAPH=0 forces HT decode eager (methodology.md: "restores the eager pipeline"), rows now get the -gum1 kernel_generation suffix even though no graph is ever captured. self.kernel_generation is fixed at construction from the static cuda_graph_supported property (line 159), not from cuda_graph_enabled, which is what ep_harness.kernel_generation() (ep_harness.py:195-197) checks at record time to append -cudagraph. So an eager-forced decode row is stored as nccl-ep-v02-ht-routed-zc-static-gum1 with no -cudagraph suffix, a new family that no longer pools with the pre-PR eager rows (nccl-ep-v02-ht-routed-zc-static) even though its measured behavior is unchanged (mode 0/1 only diverges from mode 2 when a collective is actually captured). …

Why this was flagged

…Fix: gate the -gum1 suffix on the same dynamic cuda_graph_enabled condition the harness uses for -cudagraph, or apply it in kernel_generation() alongside that suffix, so eager-forced rows keep the pre-existing family name.

Trigger: run nccl-ep HT decode with COLLX_CUDA_GRAPH=0 (documented in methodology.md as the eager override). NCCLEPBackend.init (ep_nccl.py:159-162) sets kernel_generation using self.cuda_graph_supported, a static capability check, before COLLX_CUDA_GRAPH is consulted. ep_harness.kernel_generation() (ep_harness.py:195-197) only appends -cudagraph when backend.cuda_graph_enabled is true, so the stored family becomes ...-gum1 with no cudagraph marker, distinct from the base branch's plain nccl-ep-v02-ht-routed-zc-static for the same eager run. This fragments the durable store's history/trend for that eager configuration with no corresponding behavior change, defeating the suffix mechanism methodology.md describes for preventing incomparable rows from being silently pooled or split.

Verification: nit. The mechanism is real and reachable. ep_nccl.py:159-162 sets self.kernel_generation = f"{type(self).kernel_generation}-gum1" under elif self.cuda_graph_supported:, keyed on the static capability property. ep_nccl.py:112-118 shows cuda_graph_supported only checks mode=="normal" and phase=="decode"; it never reads COLLX_CUDA_GRAPH. The eager override is instead handled by… | nit.…

@Oseltamivir
Oseltamivir force-pushed the cx-nccl-graph-mode branch 2 times, most recently from a62d581 to 3474bf6 Compare September 28, 2026 10:03
Base automatically changed from cx-bench-refactor to main September 28, 2026 10:05
The default graph usage mode (2, mixing) wraps every captured collective in an external
event wait and record. Graphed HT decode captures one host collective per pair, the routing
ncclAllGather, and the h100 EP8 trace showed ~14us idle before it, making the graphed pair
period 1.14x eager. Mode 1 (one graph at a time, never concurrent with uncaptured work on the
communicator) removes it. Graphed HT decode rows carry the -gum1 generation suffix.
@Oseltamivir
Oseltamivir merged commit 5568982 into main Sep 28, 2026
3 checks passed
@Oseltamivir
Oseltamivir deleted the cx-nccl-graph-mode branch September 28, 2026 10:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant