CollectiveX: let h200 deepep-v2 keep PCIe relaxed ordering on its GIN window / CollectiveX:允许 h200 deepep-v2 的 GIN 窗口保留 PCIe 宽松排序 - #3568
Conversation
There was a problem hiding this comment.
I reviewed this PR and did not find any bugs. Because it patches third-party DeepEP build source, changes PCIe/RDMA memory-ordering semantics, and touches collectivex/ paths that are CODEOWNER-owned, a human look is still worthwhile.
What was reviewed: the fail-closed deepep_patch_window_ordering source patch and build-gen cache bump; collx_apply_network_profile's new COLLX_RDMA_RELAXED_ORDERING validation/export and the kernel_generation suffixing in ep_deepep_v2.py; whether the inferencex-e2e/perf-changelog.yaml invariant applies here — ruled out, since the prior analogous collectivex perf change (#3537, the -gum1 discriminator) also skipped it; and the single-node/mnnvl early-return path in collx_apply_network_profile, where a stale EP_WIN_RELAXED_ORDERING from an earlier scale-out call in the same shell isn't cleared like the other network vars are (already investigated by the automated hunt and not flagged as a bug).
Extended reasoning...
The diff (8 files under collectivex/) adds an opt-in relaxed-ordering patch to DeepEP's GIN window registration, wires a new validated RDMA env var through shell config and a JSON pool profile, adds a benchmark row discriminator, and rewrites methodology docs; it touches no auth/crypto surface but does change memory-ordering semantics for RDMA that affect correctness of collective communication. Test coverage is solid (stubs/subprocess driving real code, not source inspection) and the build patch fails closed, but the change is non-mechanical, alters production pool behavior (h200-dgxc), and collectivex/ is CODEOWNER-owned, which argues for a human look rather than automated approval; I independently confirmed the perf-changelog invariant doesn't apply here by checking prior precedent, and examined but did not further pursue the already-ruled-out stale-env-var gap in the single-node/mnnvl path.
|
Note: Upstream DeepEP has correctness issues deepseek-ai/DeepEP#661 and fixed by deepseek-ai/DeepEP#674 Unfortunately on H200 VM strict ordering causes this slowdown |
… window DeepEP's ElasticBuffer registers its GIN window with NCCL_WIN_STRICT_ORDERING, which drops IBV_ACCESS_RELAXED_ORDERING from the memory region regardless of NCCL_IB_PCI_RELAXED_ORDERING. In h200's KVM guests that capped the EP16 scale-out hop near 10 GB/s per GPU, 3-8x slower than h100/b200. The DeepEP build seds the flag into a switch on EP_WIN_RELAXED_ORDERING (fails if the pin moves it; BUILD_GEN bumped), and a pool opts in with network.rdma_relaxed_ordering. Only h200-dgxc does. Its rows carry a "-relaxed-ordering" kernel_generation suffix. Upstream made the window strict on purpose (DeepEP #661/#674): relaxed ordering could let a GIN signal overtake its data. Unreproduced, and oracle-green does not rule it out. Measured on h200: EP16 bf16 decode T=512 3088us -> 620us; oracle green in bf16 and fp8. 中文:DeepEP 的 ElasticBuffer 以 NCCL_WIN_STRICT_ORDERING 注册 GIN 窗口, 使内存区域失去 PCIe 宽松排序。在 h200 的 KVM 虚拟机中,这将 EP16 跨节点 带宽限制在每 GPU 约 10 GB/s,比 h100/b200 慢 3 到 8 倍。DeepEP 构建时用 sed 将该标志改为由 EP_WIN_RELAXED_ORDERING 控制的开关(锚点变化即失败,并更新 BUILD_GEN),资源池通过 network.rdma_relaxed_ordering 开启,目前仅 h200-dgxc。 其结果行的 kernel_generation 带 "-relaxed-ordering" 后缀。上游有意使用严格 排序(DeepEP #661/#674):宽松排序下 GIN 信号理论上可能先于数据到达;该竞争 尚未复现,正确性校验通过也不能排除它。h200 实测:EP16 bf16 解码 T=512 往返 从 3088us 降至 620us;bf16 与 fp8 校验均通过。
9c80091 to
8cc4db2
Compare
Summary
h200-dgxc deepep-v2 EP16 has paid 3–8× the cross-node cost of h100/b200, while nccl-ep on the same nodes runs at h100 rates. The cause is one flag in DeepEP:
NCCL_WIN_STRICT_ORDERINGsetsNCCL_NET_MR_FLAG_FORCE_SO, which stripsIBV_ACCESS_RELAXED_ORDERINGfrom the memory registration (net_ib/reg.cc,gdaki/gin_host_gdaki.cc). This overridesNCCL_IB_PCI_RELAXED_ORDERING, which is why earlier environment A/Bs showed nothing. nccl-ep registers its windowNCCL_WIN_COLL_SYMMETRIC, so it keeps relaxed ordering.In h200's KVM guests every GPU-NIC write crosses the host root complex, so strict-ordered writes cap the scale-out hop near 10 GB/s per GPU.
Change
sedindeepep_installturns that flag intoget_env("EP_WIN_RELAXED_ORDERING", 0) ? NCCL_WIN_DEFAULT : NCCL_WIN_STRICT_ORDERING.nccl.cudoes not contain the flag exactly once.COLLX_DEEPEP_V2_BUILD_GENis bumped, so every pool rebuilds its DeepEP venv once.network.rdma_relaxed_orderingregistry field;collx_apply_network_profileexportsEP_WIN_RELAXED_ORDERING=1when it is1. Only h200-dgxc sets it.-relaxed-orderingsuffix onkernel_generation, so they never pool with the strict-ordered series (no version bump, per the methodology's discriminator precedent).methodology.mdreplaces the old "GDR path degraded wholesale" explanation with the measured cause and the caveat below.Evidence
All numbers are from hand runs of
run_ep.pyon one h200 node pair, bf16 EP16 normal mode.hybrid_{dispatch,combine}_impl(~1.5–1.7 ms vs nccl-ep 0.26–0.40 ms), not on the CPU proxynvidia-smi dmon)Relaxed ordering at CI trial counts (256 trials, 32 warmup) passed the oracle on all four cases:
Risk
Testing
python3 -m unittest discover -s tests: OK (8 skipped).GinWindowOrdering: only relaxed, scale-out, normal-mode rows carry the suffix.RelaxedOrderingProfileTests: flag1exports the switch;0or absent clears a stale export.sedoutput on DeepEP's pinnednccl.cumatches the build measured above (it only adds parentheses around the ternary).sweep_matrix.pyoutput is byte-identical to main.