Summary
EfaTransport registers every single-chunk buffer on all EFA NICs. On p5.48xlarge (32 NICs, 8 GPUs) that is 8x more fi_mr_reg calls than NIXL's libfabric plugin issues for the same GPU buffer, and it costs 7.0x in registration time. This is the largest single term in Mooncake's EFA registration cost — larger than either #3210 (concurrency cap, 2.6x) or #3216 (ascending sort, 3.1x worst case).
Filing this for discussion rather than as a PR, because reducing the NIC set is a throughput trade-off, not a free win, and the right policy needs bandwidth numbers that I have not measured yet.
Current behavior
EfaTransport::registerLocalMemoryInternal() (mooncake-transfer-engine/src/transport/efa_transport/efa_transport.cpp):
if (chunks.size() <= 1) {
// Single chunk: all NICs
for (size_t n = 0; n < num_nics; ++n) {
nic_assignments[0].push_back(n);
}
}
Every KV-cache buffer takes this branch, so each one is registered on all 32 domains. And since use_parallel_reg is 0 for device memory — the NIC-parallel path requires pre-touch, which is skipped for VRAM because a CPU store into a cudaMalloc pointer segfaults — those 32 registrations happen serially per buffer.
What NIXL does
nixlLibfabricRailManager::selectRailsForMemory() (src/utils/libfabric/libfabric_rail_manager.cpp:735) branches on memory type:
VRAM_SEG: queries the buffer's PCI bus ID via cudaQueryAddr(), then returns only that GPU's topology-local rails (getEfaDevicesForPci(), built from hwloc). On p5.48xlarge that is 4 of 32.
DRAM_SEG: all rails.
So NIXL's fan-out for GPU memory is NICs / GPUs, and it falls back to all rails only when the PCI mapping lookup fails.
Measurement
P5-1 (p5.48xlarge, 32 EFA NICs, 8x H100), 48 x 391 MB GPU buffers, one buffer per batch_register_memory() call, strictly serial. Only the engine's device list varies — no code change, the 4th argument of the Python initialize() is the device filter:
| NICs per buffer |
total |
1st reg |
48th reg |
avg |
| 32 (Mooncake today) |
123.7 s |
128 ms |
5014 ms |
2578 ms |
| 4 (NIXL's VRAM choice) |
17.6 s |
18 ms |
595 ms |
367 ms |
| 1 |
4.4 s |
6 ms |
150 ms |
91 ms |
7.0x from 32 → 4, and near-linear in the NIC count (123.7 / 4.4 = 28x for 32x the NICs).
It also multiplies a second effect. Device-memory registration on EFA costs roughly k x (device bytes already registered on that domain), k ~= 260 ms/GiB (measured in #3216) — and that is per libfabric domain, so 32 domains pay it 32 times. At 4 NICs the same curve is still present (18 ms first vs 595 ms last) but every term is 8x cheaper. The fan-out and the accumulation are not independent problems; the fan-out amplifies the accumulation.
Why this is not a straightforward fix
Registering a VRAM buffer on 4 of 32 NICs means only those 4 NICs can serve a transfer touching that buffer. That is a deliberate exchange of registration time for potential per-buffer bandwidth, and whether it is a good trade depends on:
- whether the transfer path already prefers topology-local NICs, in which case the other 28 registrations were mostly unused and this is close to free;
- how transfers striped across a buffer set are distributed across NICs at the deployed process count — this must be measured with all ranks running, not single-process;
- what happens on failover when a local NIC is down and the buffer is not registered anywhere else.
There is also a middle option: overlapping rather than disjoint windows (see the sliding window MR idea), giving each buffer more than NICs/GPUs coverage without going to all 32.
Questions for maintainers
- Is registering GPU memory on all NICs intentional, or is it the incidental result of the
chunks.size() <= 1 branch? My assumption had been that operators would narrow the device list by hand when they wanted this, but NIXL makes topology-local selection the default for VRAM, which suggests the opposite default is defensible.
- Should the selection be topology-derived (hwloc/PCI, like NIXL) or should it reuse the existing
Topology machinery? Mooncake already has Topology::selectDevice() with location-string affinity; the information needed may already be present.
- Disjoint (
NICs/GPUs) or overlapping windows? Disjoint matches NIXL and is simplest; overlapping trades some registration time back for redundancy and per-buffer bandwidth.
- Opt-in or default? Given the throughput implications I would expect an initial landing behind a flag, with the default flipped once bandwidth numbers exist.
I am happy to do the bandwidth measurements and write the PR — mainly looking for agreement on the policy before implementing, since the choice among 2–4 above determines the shape of the change.
Related
Environment: p5.48xlarge (32 EFA NICs, 8x H100, 192 cores), USE_EFA=ON USE_CUDA=ON, libfabric with the efa provider.
Summary
EfaTransportregisters every single-chunk buffer on all EFA NICs. On p5.48xlarge (32 NICs, 8 GPUs) that is 8x morefi_mr_regcalls than NIXL's libfabric plugin issues for the same GPU buffer, and it costs 7.0x in registration time. This is the largest single term in Mooncake's EFA registration cost — larger than either #3210 (concurrency cap, 2.6x) or #3216 (ascending sort, 3.1x worst case).Filing this for discussion rather than as a PR, because reducing the NIC set is a throughput trade-off, not a free win, and the right policy needs bandwidth numbers that I have not measured yet.
Current behavior
EfaTransport::registerLocalMemoryInternal()(mooncake-transfer-engine/src/transport/efa_transport/efa_transport.cpp):Every KV-cache buffer takes this branch, so each one is registered on all 32 domains. And since
use_parallel_regis 0 for device memory — the NIC-parallel path requires pre-touch, which is skipped for VRAM because a CPU store into acudaMallocpointer segfaults — those 32 registrations happen serially per buffer.What NIXL does
nixlLibfabricRailManager::selectRailsForMemory()(src/utils/libfabric/libfabric_rail_manager.cpp:735) branches on memory type:VRAM_SEG: queries the buffer's PCI bus ID viacudaQueryAddr(), then returns only that GPU's topology-local rails (getEfaDevicesForPci(), built from hwloc). On p5.48xlarge that is 4 of 32.DRAM_SEG: all rails.So NIXL's fan-out for GPU memory is
NICs / GPUs, and it falls back to all rails only when the PCI mapping lookup fails.Measurement
P5-1 (p5.48xlarge, 32 EFA NICs, 8x H100), 48 x 391 MB GPU buffers, one buffer per
batch_register_memory()call, strictly serial. Only the engine's device list varies — no code change, the 4th argument of the Pythoninitialize()is the device filter:7.0x from 32 → 4, and near-linear in the NIC count (123.7 / 4.4 = 28x for 32x the NICs).
It also multiplies a second effect. Device-memory registration on EFA costs roughly
k x (device bytes already registered on that domain), k ~= 260 ms/GiB (measured in #3216) — and that is per libfabric domain, so 32 domains pay it 32 times. At 4 NICs the same curve is still present (18 ms first vs 595 ms last) but every term is 8x cheaper. The fan-out and the accumulation are not independent problems; the fan-out amplifies the accumulation.Why this is not a straightforward fix
Registering a VRAM buffer on 4 of 32 NICs means only those 4 NICs can serve a transfer touching that buffer. That is a deliberate exchange of registration time for potential per-buffer bandwidth, and whether it is a good trade depends on:
There is also a middle option: overlapping rather than disjoint windows (see the
sliding window MRidea), giving each buffer more thanNICs/GPUscoverage without going to all 32.Questions for maintainers
chunks.size() <= 1branch? My assumption had been that operators would narrow the device list by hand when they wanted this, but NIXL makes topology-local selection the default for VRAM, which suggests the opposite default is defensible.Topologymachinery? Mooncake already hasTopology::selectDevice()with location-string affinity; the information needed may already be present.NICs/GPUs) or overlapping windows? Disjoint matches NIXL and is simplest; overlapping trades some registration time back for redundancy and per-buffer bandwidth.I am happy to do the bandwidth measurements and write the PR — mainly looking for agreement on the policy before implementing, since the choice among 2–4 above determines the shape of the change.
Related
MC_MAX_CONCURRENT_REG_MR), 2.6xEnvironment: p5.48xlarge (32 EFA NICs, 8x H100, 192 cores),
USE_EFA=ON USE_CUDA=ON, libfabric with theefaprovider.