Skip to content

CollectiveX: let h200 deepep-v2 keep PCIe relaxed ordering on its GIN window / CollectiveX:允许 h200 deepep-v2 的 GIN 窗口保留 PCIe 宽松排序 - #3568

Merged
Oseltamivir merged 1 commit into
mainfrom
cx-deepep-v2-relaxed-ordering
Sep 29, 2026
Merged

Oseltamivir merged 1 commit into
mainfrom
cx-deepep-v2-relaxed-ordering

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

h200-dgxc deepep-v2 EP16 has paid 3–8× the cross-node cost of h100/b200, while nccl-ep on the same nodes runs at h100 rates. The cause is one flag in DeepEP:

// DeepEP csrc/kernels/backend/nccl.cu:140
ncclCommWindowRegister(comm, ptr, num_bytes, &window, NCCL_WIN_STRICT_ORDERING);

NCCL_WIN_STRICT_ORDERING sets NCCL_NET_MR_FLAG_FORCE_SO, which strips IBV_ACCESS_RELAXED_ORDERING from the memory registration (net_ib/reg.cc, gdaki/gin_host_gdaki.cc). This overrides NCCL_IB_PCI_RELAXED_ORDERING, which is why earlier environment A/Bs showed nothing. nccl-ep registers its window NCCL_WIN_COLL_SYMMETRIC, so it keeps relaxed ordering.

In h200's KVM guests every GPU-NIC write crosses the host root complex, so strict-ordered writes cap the scale-out hop near 10 GB/s per GPU.

Change

  • Build: one sed in deepep_install turns that flag into get_env("EP_WIN_RELAXED_ORDERING", 0) ? NCCL_WIN_DEFAULT : NCCL_WIN_STRICT_ORDERING.
    • The build fails if the pinned nccl.cu does not contain the flag exactly once.
    • With the variable unset, the window registers exactly as upstream does.
    • COLLX_DEEPEP_V2_BUILD_GEN is bumped, so every pool rebuilds its DeepEP venv once.
  • Opt-in per pool: a new network.rdma_relaxed_ordering registry field; collx_apply_network_profile exports EP_WIN_RELAXED_ORDERING=1 when it is 1. Only h200-dgxc sets it.
  • Row discriminator: relaxed-ordering scale-out rows get a -relaxed-ordering suffix on kernel_generation, so they never pool with the strict-ordered series (no version bump, per the methodology's discriminator precedent).
  • Docs: methodology.md replaces the old "GDR path degraded wholesale" explanation with the measured cause and the caveat below.

Evidence

All numbers are from hand runs of run_ep.py on one h200 node pair, bf16 EP16 normal mode.

Test Result
NCCL 2.30.4 → 2.30.7 (DeepEP rebuilt) no change (3096 → 3097 µs at T=512)
RDMA service level 3 → 0 no change
NIC mapping NCCL topology already gives each GPU its PIX NIC
Nsight Systems time is inside hybrid_{dispatch,combine}_impl (~1.5–1.7 ms vs nccl-ep 0.26–0.40 ms), not on the CPU proxy
SM count 10 / 20 / 32 flat (3090 / 3094 / 3048 µs)
Hidden 3584 / 7168 / 14336 1593 / 3090 / 5981 µs, so the cost scales with bytes
Per-GPU PCIe (nvidia-smi dmon) even across all 8 GPUs, ~7 GB/s vs nccl-ep ~17 GB/s
Window flag, same build strict 3088 µs → relaxed 620 µs at T=512

Relaxed ordering at CI trial counts (256 trials, 32 warmup) passed the oracle on all four cases:

Case T Round trip
decode bf16 1 / 64 / 512 96 / 150 / 617 µs
decode fp8 512 514 µs
prefill bf16 8192 6.04 ms (nccl-ep / h100 ≈ 5.7 ms)
prefill fp8 8192 4.58 ms

Risk

  • Upstream made the window strict on purpose (DeepEP #661, fixed by [AMD] bump MI355X MORI FP8 image to latest Feb 10 #674): with relaxed ordering, a GIN signal atomic can in principle overtake the relaxed-ordered data writes it announces. The maintainer confirmed the hazard, but it was never reproduced. A green oracle on this pool is evidence, not proof.
  • The switch is off everywhere except h200, and the rows are discriminated, so the strict series is untouched.
  • h100 and b200 have not been measured with it, so they keep upstream's registration.
  • An upstream request (relaxed data, strict signals) has not been filed.

Testing

  • python3 -m unittest discover -s tests: OK (8 skipped).
  • New GinWindowOrdering: only relaxed, scale-out, normal-mode rows carry the suffix.
  • New RelaxedOrderingProfileTests: flag 1 exports the switch; 0 or absent clears a stale export.
  • Both new tests fail against the pre-change code.
  • The sed output on DeepEP's pinned nccl.cu matches the build measured above (it only adds parentheses around the ternary).
  • sweep_matrix.py output is byte-identical to main.
  • No CI hardware run yet.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and did not find any bugs. Because it patches third-party DeepEP build source, changes PCIe/RDMA memory-ordering semantics, and touches collectivex/ paths that are CODEOWNER-owned, a human look is still worthwhile.

What was reviewed: the fail-closed deepep_patch_window_ordering source patch and build-gen cache bump; collx_apply_network_profile's new COLLX_RDMA_RELAXED_ORDERING validation/export and the kernel_generation suffixing in ep_deepep_v2.py; whether the inferencex-e2e/perf-changelog.yaml invariant applies here — ruled out, since the prior analogous collectivex perf change (#3537, the -gum1 discriminator) also skipped it; and the single-node/mnnvl early-return path in collx_apply_network_profile, where a stale EP_WIN_RELAXED_ORDERING from an earlier scale-out call in the same shell isn't cleared like the other network vars are (already investigated by the automated hunt and not flagged as a bug).

Extended reasoning...

The diff (8 files under collectivex/) adds an opt-in relaxed-ordering patch to DeepEP's GIN window registration, wires a new validated RDMA env var through shell config and a JSON pool profile, adds a benchmark row discriminator, and rewrites methodology docs; it touches no auth/crypto surface but does change memory-ordering semantics for RDMA that affect correctness of collective communication. Test coverage is solid (stubs/subprocess driving real code, not source inspection) and the build patch fails closed, but the change is non-mechanical, alters production pool behavior (h200-dgxc), and collectivex/ is CODEOWNER-owned, which argues for a human look rather than automated approval; I independently confirmed the perf-changelog invariant doesn't apply here by checking prior precedent, and examined but did not further pursue the already-ruled-out stale-env-var gap in the single-node/mnnvl path.

@Oseltamivir

Copy link
Copy Markdown
Collaborator Author

Note: Upstream DeepEP has correctness issues deepseek-ai/DeepEP#661 and fixed by deepseek-ai/DeepEP#674

Unfortunately on H200 VM strict ordering causes this slowdown

… window

DeepEP's ElasticBuffer registers its GIN window with NCCL_WIN_STRICT_ORDERING,
which drops IBV_ACCESS_RELAXED_ORDERING from the memory region regardless of
NCCL_IB_PCI_RELAXED_ORDERING. In h200's KVM guests that capped the EP16
scale-out hop near 10 GB/s per GPU, 3-8x slower than h100/b200.

The DeepEP build seds the flag into a switch on EP_WIN_RELAXED_ORDERING
(fails if the pin moves it; BUILD_GEN bumped), and a pool opts in with
network.rdma_relaxed_ordering. Only h200-dgxc does. Its rows carry a
"-relaxed-ordering" kernel_generation suffix. Upstream made the window strict
on purpose (DeepEP #661/#674): relaxed ordering could let a GIN signal
overtake its data. Unreproduced, and oracle-green does not rule it out.
Measured on h200: EP16 bf16 decode T=512 3088us -> 620us; oracle green in
bf16 and fp8.

中文:DeepEP 的 ElasticBuffer 以 NCCL_WIN_STRICT_ORDERING 注册 GIN 窗口,
使内存区域失去 PCIe 宽松排序。在 h200 的 KVM 虚拟机中,这将 EP16 跨节点
带宽限制在每 GPU 约 10 GB/s,比 h100/b200 慢 3 到 8 倍。DeepEP 构建时用 sed
将该标志改为由 EP_WIN_RELAXED_ORDERING 控制的开关(锚点变化即失败,并更新
BUILD_GEN),资源池通过 network.rdma_relaxed_ordering 开启,目前仅 h200-dgxc。
其结果行的 kernel_generation 带 "-relaxed-ordering" 后缀。上游有意使用严格
排序(DeepEP #661/#674):宽松排序下 GIN 信号理论上可能先于数据到达;该竞争
尚未复现,正确性校验通过也不能排除它。h200 实测:EP16 bf16 解码 T=512 往返
从 3088us 降至 620us;bf16 与 fp8 校验均通过。
@Oseltamivir
Oseltamivir force-pushed the cx-deepep-v2-relaxed-ordering branch from 9c80091 to 8cc4db2 Compare September 29, 2026 03:52
@Oseltamivir
Oseltamivir merged commit dc911f6 into main Sep 29, 2026
3 checks passed
@Oseltamivir
Oseltamivir deleted the cx-deepep-v2-relaxed-ordering branch September 29, 2026 03:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant