CollectiveX: run nccl-ep's communicator in NCCL graph usage mode 1 - #3537
Conversation
4895387 to
fbbc876
Compare
| elif self.cuda_graph_supported: | ||
| # Only graphed HT captures a host NCCL collective (the routing ncclAllGather), so | ||
| # only its rows change with the communicator's graph usage mode. | ||
| self.kernel_generation = f"{type(self).kernel_generation}-gum1" |
There was a problem hiding this comment.
🟡 (optional) When COLLX_CUDA_GRAPH=0 forces HT decode eager (methodology.md: "restores the eager pipeline"), rows now get the -gum1 kernel_generation suffix even though no graph is ever captured. self.kernel_generation is fixed at construction from the static cuda_graph_supported property (line 159), not from cuda_graph_enabled, which is what ep_harness.kernel_generation() (ep_harness.py:195-197) checks at record time to append -cudagraph. So an eager-forced decode row is stored as nccl-ep-v02-ht-routed-zc-static-gum1 with no -cudagraph suffix, a new family that no longer pools with the pre-PR eager rows (nccl-ep-v02-ht-routed-zc-static) even though its measured behavior is unchanged (mode 0/1 only diverges from mode 2 when a collective is actually captured). …
Why this was flagged
…Fix: gate the -gum1 suffix on the same dynamic cuda_graph_enabled condition the harness uses for -cudagraph, or apply it in kernel_generation() alongside that suffix, so eager-forced rows keep the pre-existing family name.
Trigger: run nccl-ep HT decode with COLLX_CUDA_GRAPH=0 (documented in methodology.md as the eager override). NCCLEPBackend.init (ep_nccl.py:159-162) sets kernel_generation using self.cuda_graph_supported, a static capability check, before COLLX_CUDA_GRAPH is consulted. ep_harness.kernel_generation() (ep_harness.py:195-197) only appends -cudagraph when backend.cuda_graph_enabled is true, so the stored family becomes ...-gum1 with no cudagraph marker, distinct from the base branch's plain nccl-ep-v02-ht-routed-zc-static for the same eager run. This fragments the durable store's history/trend for that eager configuration with no corresponding behavior change, defeating the suffix mechanism methodology.md describes for preventing incomparable rows from being silently pooled or split.
Verification: nit. The mechanism is real and reachable. ep_nccl.py:159-162 sets self.kernel_generation = f"{type(self).kernel_generation}-gum1" under elif self.cuda_graph_supported:, keyed on the static capability property. ep_nccl.py:112-118 shows cuda_graph_supported only checks mode=="normal" and phase=="decode"; it never reads COLLX_CUDA_GRAPH. The eager override is instead handled by… | nit.…
a62d581 to
3474bf6
Compare
The default graph usage mode (2, mixing) wraps every captured collective in an external event wait and record. Graphed HT decode captures one host collective per pair, the routing ncclAllGather, and the h100 EP8 trace showed ~14us idle before it, making the graphed pair period 1.14x eager. Mode 1 (one graph at a time, never concurrent with uncaptured work on the communicator) removes it. Graphed HT decode rows carry the -gum1 generation suffix.
3474bf6 to
bd5c440
Compare
Graphed nccl-ep HT decode was the one graphed row family that ran slower than eager in the full sweep 36334392363: pair period 1.03× the eager Sep 25 baseline on average, and up to 1.30×.
Cause
An h100 EP8 trace (torch.profiler/CUPTI, HT decode, graphed vs eager) showed that every kernel's duration is identical in both regimes. The whole graphed penalty is idle gaps between kernels, dominated by about 14 µs just before the routing
ncclAllGather(13.7 µs graphed vs 1.5 µs eager at T=1).That all-gather is the only host NCCL collective in any capture. NCCL's default graph usage mode is 2 ("mixing"): it wraps every captured collective in an external event wait and an event record, so graph and uncaptured work can interleave on one communicator (NCCL 2.30.7
src/misc/strongstream.cc:241-247,:347).Fix
nccl-ep's own communicator is created with
NCCLConfig(graph_usage_mode=1): one graph at a time, never concurrent with uncaptured work on that communicator. That is how both the harness and a captured decode step use it. In NCCL 2.30.7, modes 0 and 1 behave the same; only== 2enables the serialization.Graphed HT decode rows get the
-gum1generation suffix. LL graphs capture no host collective and HT prefill stays eager, so neither changes.Trace result, h100 EP8 HT decode:
Ruled out with no effect:
NCCL_GRAPH_STREAM_ORDERING=0,NCCL_MEM_SYNC_DOMAIN=0,NCCL_CGA_CLUSTER_SIZE=0, and a symmetric routing window (which switches NCCL to its NVLS all-gather kernel but leaves the gap).At EP16 the captured all-gather also keeps NCCL's per-replay proxy host callback (
enqueue.cc:1703-1722). This change does not touch that.Validation
nccl-ep normal mode, h100 and h200, EP8 and EP16 (36378559608, 36378561991): all 4 shards green, 56/56 rows pass correctness, and graphed HT decode rows carry
nccl-ep-v02-ht-routed-zc-static-gum1-cudagraph.Graphed HT decode pair period, geomean over T=1..512 (range in brackets):
Graphed HT decode goes from 1.03× eager on average (up to 1.30×) to parity or better on every cell. The largest gains are at T=256–512, where main's graphed rows ran furthest behind (h100 EP8 T=512: 418.4 → 317.0 µs). The eager baseline predates the full-plane combine (
-zcrather than-zc-static), so treat those ratios as approximate.HT prefill (eager, unchanged generation) stays at 0.93–1.02× main, which is run-to-run spread.