Skip to content

CollectiveX: kv-transfer suite — NIXL, Mooncake, MoRI-IO KV-cache handoff - #2510

Open
Oseltamivir wants to merge 4 commits into
mainfrom
collectivex-kv-transfer
Open

Oseltamivir wants to merge 4 commits into
mainfrom
collectivex-kv-transfer

Conversation

@Oseltamivir

@Oseltamivir Oseltamivir commented Aug 6, 2026 •

Copy link
Copy Markdown
Collaborator

This adds a kv-transfer suite: the prefill-to-decode KV handoff of disaggregated serving, measured with the transfer libraries engines ship (NIXL, Mooncake, MoRI-IO) on real fabrics.

  • Shape. A leg is 2 nodes x 1 GPU. It moves bursts of 1 to 32 concurrent requests' paged KV as vLLM's packed block-major descriptor lists over seed-keyed random block tables, pull and push, against a single-descriptor bulk baseline.
  • Verification. Every request in a burst is pattern-verified against its own block tables. A failed verify makes the artifact invalid and the leg red.
  • Workload. kv-dsv4 is DeepSeek-V4-Pro as vLLM allocates it, at ISL 2k to 512k with the 256-token block. Geometry and the measurement model are in docs/methodology.md.

Ported onto the suite structure

This branch was rebuilt on current main (with #3541's suite structure) rather than rebased. The old branch (51 commits, from before the collectivex/ move) is kept locally as kv-transfer-pre-port at dc68d14cd.

The benchmark code (bench/kv_*.py, run_kv.py) and its tests carry over with three changes:

  • The pool budget is now a case argument (below).
  • The result document is built by the case-attempt envelope EP already uses, so identity and provenance can't drift apart. kv now validates COLLX_ATTEMPT_ID the way EP does.
  • The never-set --kv-mori-qp and --kv-mori-chunking flags are gone. Library defaults (1 QP, no chunking) were already the only values used, after 4 QPs plus chunking hung transfers on hardware.

The integration layer is new:

  • Matrix. sweep_matrix.py --suites kv-transfer resolves shards from configs/kv_sweep.json crossed with the registry's kv_backends. One shard per (backend, fabric), each on the pool's own launcher. EP output is unchanged.
  • Codec. config.py case-args encodes kv cases through the suite codec, and the rank wrapper execs run_kv. The old if suite == kv-transfer branch inside the EP codec is gone.
  • Scheduling is data, not launcher branches. The per-pool allocation and hang guard now live in kv_sweep.json scheduling: gb200 460/420 min, gb300 690/660, others 210/190. Before, they were hardcoded in case $COLLX_BENCH in nixl|mooncake|mori-io) blocks in three launchers. Each shard also carries its own GitHub job ceiling, so EP jobs keep 350 minutes instead of everything moving to 720.
  • Pool budget is a case argument. The mi355x mooncake 20 GiB cap is a registry pool_budget that reaches run_kv --pool-budget, instead of a launcher exporting COLLX_KV_POOL_BUDGET.
  • Image pin. mi355x mooncake's atom-dev image rides the shard image override.
  • GB rdma legs set COLLX_FABRIC=rdma, so collx_set_placement labels them mnnvl-rdma and gb-nv validates their network profile. gb200 and gb300 gain the network selectors those legs need; EP stays MNNVL, and the profile is a no-op there. Mooncake opts out of MC_FORCE_MNNVL.
  • Backend preparation installs the pinned nixl-cu13==1.3.2 and mooncake-transfer-engine==0.3.12.post1 wheels, preferring an image-provided build, and asserts mori.io.
  • Summary. summarize.py renders a KV table next to the EP one, and bandwidth.py skips kv rows.

Scope changes from the old branch

  • No b300 rows. main's b300 is now the AWS EFA pool, and the old rows were one-rail measurements on the previous RoCE b300 cluster. The b300-specific launcher budget and retry go with them.
  • kv is opt-in. The default suites is ep, so kv legs run only on dispatches that name kv-transfer. Before, a full --backend all dispatch pulled in hours-long kv legs.
  • Dropped. The skip_queue_pr workflow input, and the test-job numpy install (the CI test group already has numpy).

Enabled rows

pool backends fabrics
h200-dgxc, b200-nscale nixl, mooncake rdma
gb200, gb300 nixl rdma, mnnvl
gb200, gb300 mooncake rdma
mi355x mori-io rdma
mi355x mooncake rdma, push-only

That is 12 shards.

Validation

  • Unit tests: 180 pass locally. That covers matrix, codec round-trip through run_kv's own parser, identity recompute, scheduling invariants, the launcher identity gate, grid budgets, registration chunking and the summary.
  • On hardware: not yet run on this wiring. The last all-green nine-shard reference (31600727215) was on the pre-port branch and grid.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional findings (outside current diff — PR may have been updated during review):

  • 🟡 experimental/CollectiveX/configs/platform_config.json:105-112 — Enabling kv_backends on gb200 routes its KV leg through launch_gb-nv.sh, which unconditionally exports COLLX_TRANSPORT=mnnvl for every shard and never calls collx_validate_network_profile_on_job — unlike launch_single-slurm.sh/launch_mi-amds.sh, which branch to an -rdma transport for scale-out and validate the fabric. Because transport stays mnnvl, collx_apply_network_profile, the rank wrapper's network branch, and prepare_backend.sh's validate_container_network all skip the fail-closed HCA/interface checks for this leg, even though this PR's own fabric note calls the gb200 KV leg real cross-node InfiniBand. The transfer itself likely still works (run_kv pins UCX/gloo selectors from the operator config independently of transport), so the concrete loss is the missing pre-flight fabric validation, not a guaranteed break.

    Extended reasoning...

    What's happening: configs/platform_config.json now sets kv_backends: {nixl: [rdma]} on gb200 (line 108) alongside a new network.rdma_devices/socket_ifname block and a fabric note that explicitly calls the KV leg "4x ConnectX-7 NDR400 InfiniBand (KV scale-out; EP stays MNNVL)". That's a genuine cross-node RDMA transfer.

    sweep_matrix._kv_cases() schedules this as a 2-node x 1-GPU shard using PLATFORMS["gb200"]["launcher"], which is gb-nv. launchers/launch_gb-nv.sh (untouched by this PR) unconditionally does export COLLX_TRANSPORT=mnnvl for every shard it runs and never calls collx_validate_network_profile_on_job. Compare that to launch_single-slurm.sh and launch_mi-amds.sh, which branch COLLX_TRANSPORT to an -rdma variant when NODES>1 and then validate the fabric on the job before proceeding.

    Why this was fine before, and why it isn't now: gb200/gb300 previously only ran the EP suite, whose EP16 always stays inside the 72-GPU MNNVL scale-up domain, so mnnvl was always the correct transport label for anything gb-nv launched. This PR is the first thing that schedules a real scale-out RDMA leg (KV transfer) on a gb-nv-launched SKU, and the launcher has no branch to distinguish that case from the EP/MNNVL case.

    Concrete effect: because COLLX_TRANSPORT stays mnnvl for the KV shard, three fail-closed validation paths all skip:

    • collx_apply_network_profile (runtime/common.sh) early-returns on its nodes>1 && transport!=mnnvl gate, so it never validates or exports NCCL_IB_HCA/GLOO_SOCKET_IFNAME for this leg.
    • The rank wrapper's own network branch in common.sh is gated the same way and is skipped.
    • prepare_backend.sh's validate_container_network is likewise gated on transport != mnnvl and returns early.

    So the gb200 KV leg is the only scale-out RDMA row in the registry that never gets the "prove the configured socket interface and RDMA HCA actually exist on every allocated node" check every other scale-out fabric (b200-nscale, mi355x, and the EP16 x86 rows) gets.

    Step-by-step to see it:

    1. platform_config.json:105-112 — gb200 has launcher: gb-nv and kv_backends: {nixl: [rdma]}.
    2. sweep_matrix._kv_cases() builds a case with nodes=2, gpus_per_node=1 for this SKU/backend.
    3. That case dispatches through launch_gb-nv.sh, which does export COLLX_TRANSPORT=mnnvl with no NODES-based branch (contrast launch_single-slurm.sh's scale-out branch).
    4. collx_apply_network_profile 2 mnnvl is called somewhere downstream; its gate [ "$nodes" -gt 1 ] && [ "$transport" != mnnvl ] evaluates 2 -gt 1 && mnnvl != mnnvl → false, so it returns immediately without validating anything.
    5. Same story for validate_container_network and the rank wrapper's branch — both keyed off the same transport != mnnvl test.
    6. Net effect: the KV shard runs without ever confirming the InfiniBand interfaces/HCAs named in the new network block actually exist on the allocated nodes.

    Why it's not a hard break: run_kv.py's export_ucx_selectors() pins UCX_NET_DEVICES/UCX_IB_GID_INDEX directly from COLLX_RDMA_DEVICES/COLLX_IB_GID_INDEX, independent of the mnnvl branch, and it derives GLOO_SOCKET_IFNAME similarly. So the actual UCX/gloo transfer likely still selects the right devices and runs correctly on a healthy node — the loss is specifically the pre-flight, fail-closed proof (that methodology.md documents as required for every non-MNNVL scale-out node) that those devices exist and are up, not a guaranteed crash or silently wrong measurement.

    How to fix: give launch_gb-nv.sh the same NODES>1 branch the other two launchers have — select an -rdma transport variant for the KV shard and call collx_validate_network_profile_on_job on it — so a bad or missing IB config on a gb200 node fails the shard early instead of running the transfer unvalidated. Note this is independent of the separately-reported COLLX_BENCH allowlist issue on the same launcher family: that gate is a backend-name check that happens after the COLLX_TRANSPORT=mnnvl assignment, so fixing it alone would still leave this transport hardcoding in place.

Comment thread experimental/CollectiveX/runtime/common.sh Outdated
Comment thread experimental/CollectiveX/bench/run_kv.py Outdated
Oseltamivir added a commit that referenced this pull request Aug 7, 2026
…date gb-nv rdma legs

Review findings on #2510, both real. The launchers collx_die on COLLX_BENCH
values outside their EP enum, so every kv shard died at the identity stage;
nixl/mooncake/mori-io are now accepted where the registry schedules them,
pinned by a test that greps each SKU's launcher for its kv backends.

launch_gb-nv.sh also exported COLLX_TRANSPORT=mnnvl unconditionally, which
made a gb200 kv rdma leg the only scale-out fabric that skipped
collx_apply_network_profile, the rank wrapper's network branch, and
validate_container_network. kv rdma shards now carry mnnvl-rdma (the
workflow exports the shard mode) and the launcher proves the pinned socket
interface and HCAs on the allocation before running, like every other
scale-out launcher; mnnvl shards keep skipping, as elsewhere.
@Oseltamivir

Copy link
Copy Markdown
Collaborator Author

On the gb-nv transport finding from the review: fixed in dac114e. kv rdma shards on gb-nv now carry COLLX_TRANSPORT=mnnvl-rdma (the workflow exports the shard mode), which re-enables collx_apply_network_profile, the rank wrapper's network branch, and validate_container_network for those legs, and the launcher gained the same on-allocation collx_validate_network_profile_on_job check the other scale-out launchers run. mnnvl shards keep the mnnvl label and skip, as before.

Base automatically changed from collectivex-fixes to main August 10, 2026 06:19
@Oseltamivir
Oseltamivir requested a review from a team August 10, 2026 06:19
Oseltamivir added a commit that referenced this pull request Aug 10, 2026
…date gb-nv rdma legs

Review findings on #2510, both real. The launchers collx_die on COLLX_BENCH
values outside their EP enum, so every kv shard died at the identity stage;
nixl/mooncake/mori-io are now accepted where the registry schedules them,
pinned by a test that greps each SKU's launcher for its kv backends.

launch_gb-nv.sh also exported COLLX_TRANSPORT=mnnvl unconditionally, which
made a gb200 kv rdma leg the only scale-out fabric that skipped
collx_apply_network_profile, the rank wrapper's network branch, and
validate_container_network. kv rdma shards now carry mnnvl-rdma (the
workflow exports the shard mode) and the launcher proves the pinned socket
interface and HCAs on the allocation before running, like every other
scale-out launcher; mnnvl shards keep skipping, as elsewhere.
@Oseltamivir
Oseltamivir force-pushed the collectivex-kv-transfer branch from 72c1c3b to f2f36c3 Compare August 10, 2026 06:37
Comment thread experimental/CollectiveX/runtime/common.sh Outdated
Comment thread .github/workflows/collectivex-sweep.yml Outdated
matrix.queue-token
)),
toJSON(format('ci-attempt-{0}', github.run_attempt))
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Skip-queue ignored without node slots

Medium Severity

skip_queue_pr is nested under NODE_SLOT_SCHEDULER_ENABLED, so the ci-skip-queue-pr-* label is only requested when the node-slot flag is on. With that flag unset or false, a filled skip_queue_pr falls through to the three-label runs-on path and the job queues normally. skip_queue_pr and node-slot matching are independent; the other sweep templates attach the skip-queue label whenever the priority scheduler is on.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit c6e2d26. Configure here.

Comment thread experimental/CollectiveX/launchers/launch_gb-nv.sh Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

There are 2 total unresolved issues (including 1 from previous review).

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 93feafc. Configure here.

Comment thread experimental/CollectiveX/launchers/launch_single-slurm.sh Outdated
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@Oseltamivir
Oseltamivir force-pushed the collectivex-kv-transfer branch from dc68d14 to 249ee19 Compare September 28, 2026 14:58
@Oseltamivir Oseltamivir changed the title CollectiveX: kv-transfer suite — NIXL + MoRI-IO KV-cache handoff benchmark (stacked on #2489) CollectiveX: kv-transfer suite — NIXL, Mooncake, MoRI-IO KV-cache handoff (stacked on #3541) Sep 28, 2026
@Oseltamivir
Oseltamivir changed the base branch from main to cx-swap-suite September 28, 2026 14:58
@Oseltamivir
Oseltamivir force-pushed the collectivex-kv-transfer branch from 249ee19 to 60dbf2b Compare September 28, 2026 15:20
Base automatically changed from cx-swap-suite to main September 28, 2026 15:37
…che handoff)

A leg is 2 nodes x 1 GPU moving bursts of concurrent requests' paged KV
(vLLM's packed block-major DSV4 descriptors over random block tables), pull
and push, against a single-descriptor bulk baseline, with every request
pattern-verified. It resolves in sweep_matrix.py (--suites kv-transfer) from
configs/kv_sweep.json and the registry's kv_backends map, runs on each pool's
own launcher, and reaches run_kv through the suite codec. Per-pool allocation
and hang-guard budgets live in kv_sweep.json and per-backend pool budgets in
the registry, not in launcher branches. No b300 rows: the current b300 pool
is the EFA cluster.
run_ep and run_kv build their identity, provenance and outcome through one
ep_harness.case_attempt (kv now validates COLLX_ATTEMPT_ID as EP does). The kv
codec and shard builder use the shared helpers, run_kv folds its row tail and
drops the never-set MoRI QP/chunking flags (library defaults were already the
only value), and the UCX selector and pool-budget tests are table-driven. EP
documents, emitted argv and resolved matrices are unchanged.
@Oseltamivir
Oseltamivir force-pushed the collectivex-kv-transfer branch 2 times, most recently from 60dbf2b to 3b42529 Compare September 28, 2026 16:03
@Oseltamivir Oseltamivir changed the title CollectiveX: kv-transfer suite — NIXL, Mooncake, MoRI-IO KV-cache handoff (stacked on #3541) CollectiveX: kv-transfer suite — NIXL, Mooncake, MoRI-IO KV-cache handoff Sep 28, 2026
Correctness:
- Wire the sweep seed into the block tables and record it.
- Salt each rank's pool pattern so a loopback transfer fails verification.
- Paint the fabric pool on-device instead of from a pool-sized host array.
- Fail closed on an empty grid and on a multi-page-family pool over budget.
- Release NIXL transfer handles after each row.
- Keep hostname rendezvous on MNNVL: the GB network block is kv-rdma only.
- Never set a kv job timeout below the fleet-wide 350 minutes.

Trims: shared library_version/spans/offset_lists helpers, the one-call
registry-spec helper inlined, the kv precision filter and dead
MC_FORCE_MNNVL export dropped, pip_install reused, one-pass summarize
split, table-driven and registry-independent kv tests (the launcher-grep
test removed, the argv codec moved onto the real case-args seam), and
condensed kv docs with the stale b300 and 576 B claims fixed.
…FABRIC

b300 is now the AWS p6-b300 EFA pool. EFA is not a verbs HCA and UCX has
no transport for it, so the NIXL adapter selects the wheel's LIBFABRIC
plugin when the network profile marks the pool rdma_fabric=efa, and rows
record the plugin in implementation.transport. A two-node hand probe on
the pool moved 94 GB/s READ and 97 GB/s WRITE per GPU at 1 GiB, verified.
No mooncake leg: the PyPI wheel's transport is verbs RC only.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants