Skip to content

[TE][EFA] GPU buffers are registered on all NICs, costing 7x vs per-GPU rail selection #3217

Description

@whn09

Summary

EfaTransport registers every single-chunk buffer on all EFA NICs. On p5.48xlarge (32 NICs, 8 GPUs) that is 8x more fi_mr_reg calls than NIXL's libfabric plugin issues for the same GPU buffer, and it costs 7.0x in registration time. This is the largest single term in Mooncake's EFA registration cost — larger than either #3210 (concurrency cap, 2.6x) or #3216 (ascending sort, 3.1x worst case).

Filing this for discussion rather than as a PR, because reducing the NIC set is a throughput trade-off, not a free win, and the right policy needs bandwidth numbers that I have not measured yet.

Current behavior

EfaTransport::registerLocalMemoryInternal() (mooncake-transfer-engine/src/transport/efa_transport/efa_transport.cpp):

if (chunks.size() <= 1) {
    // Single chunk: all NICs
    for (size_t n = 0; n < num_nics; ++n) {
        nic_assignments[0].push_back(n);
    }
}

Every KV-cache buffer takes this branch, so each one is registered on all 32 domains. And since use_parallel_reg is 0 for device memory — the NIC-parallel path requires pre-touch, which is skipped for VRAM because a CPU store into a cudaMalloc pointer segfaults — those 32 registrations happen serially per buffer.

What NIXL does

nixlLibfabricRailManager::selectRailsForMemory() (src/utils/libfabric/libfabric_rail_manager.cpp:735) branches on memory type:

  • VRAM_SEG: queries the buffer's PCI bus ID via cudaQueryAddr(), then returns only that GPU's topology-local rails (getEfaDevicesForPci(), built from hwloc). On p5.48xlarge that is 4 of 32.
  • DRAM_SEG: all rails.

So NIXL's fan-out for GPU memory is NICs / GPUs, and it falls back to all rails only when the PCI mapping lookup fails.

Measurement

P5-1 (p5.48xlarge, 32 EFA NICs, 8x H100), 48 x 391 MB GPU buffers, one buffer per batch_register_memory() call, strictly serial. Only the engine's device list varies — no code change, the 4th argument of the Python initialize() is the device filter:

NICs per buffer total 1st reg 48th reg avg
32 (Mooncake today) 123.7 s 128 ms 5014 ms 2578 ms
4 (NIXL's VRAM choice) 17.6 s 18 ms 595 ms 367 ms
1 4.4 s 6 ms 150 ms 91 ms

7.0x from 32 → 4, and near-linear in the NIC count (123.7 / 4.4 = 28x for 32x the NICs).

It also multiplies a second effect. Device-memory registration on EFA costs roughly k x (device bytes already registered on that domain), k ~= 260 ms/GiB (measured in #3216) — and that is per libfabric domain, so 32 domains pay it 32 times. At 4 NICs the same curve is still present (18 ms first vs 595 ms last) but every term is 8x cheaper. The fan-out and the accumulation are not independent problems; the fan-out amplifies the accumulation.

Why this is not a straightforward fix

Registering a VRAM buffer on 4 of 32 NICs means only those 4 NICs can serve a transfer touching that buffer. That is a deliberate exchange of registration time for potential per-buffer bandwidth, and whether it is a good trade depends on:

  • whether the transfer path already prefers topology-local NICs, in which case the other 28 registrations were mostly unused and this is close to free;
  • how transfers striped across a buffer set are distributed across NICs at the deployed process count — this must be measured with all ranks running, not single-process;
  • what happens on failover when a local NIC is down and the buffer is not registered anywhere else.

There is also a middle option: overlapping rather than disjoint windows (see the sliding window MR idea), giving each buffer more than NICs/GPUs coverage without going to all 32.

Questions for maintainers

  1. Is registering GPU memory on all NICs intentional, or is it the incidental result of the chunks.size() <= 1 branch? My assumption had been that operators would narrow the device list by hand when they wanted this, but NIXL makes topology-local selection the default for VRAM, which suggests the opposite default is defensible.
  2. Should the selection be topology-derived (hwloc/PCI, like NIXL) or should it reuse the existing Topology machinery? Mooncake already has Topology::selectDevice() with location-string affinity; the information needed may already be present.
  3. Disjoint (NICs/GPUs) or overlapping windows? Disjoint matches NIXL and is simplest; overlapping trades some registration time back for redundancy and per-buffer bandwidth.
  4. Opt-in or default? Given the throughput implications I would expect an initial landing behind a flag, with the default flipped once bandwidth numbers exist.

I am happy to do the bandwidth measurements and write the PR — mainly looking for agreement on the policy before implementing, since the choice among 2–4 above determines the shape of the change.

Related

Environment: p5.48xlarge (32 EFA NICs, 8x H100, 192 cores), USE_EFA=ON USE_CUDA=ON, libfabric with the efa provider.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions