cluster/drill

NCCL · intermediate

On a routed leaf-spine RoCEv2 fabric, NCCL jobs run fine when all nodes share a rack but hang at init when they span racks. show_gids on a worker prints the table below. What is the fix?

DEV     PORT  INDEX  GID                                      IPV4          VER   DEV
mlx5_0  1     0      fe80:0000:0000:0000:0e42:a1ff:fe5b:10a4                v1    eth2
mlx5_0  1     1      fe80:0000:0000:0000:0e42:a1ff:fe5b:10a4                v2    eth2
mlx5_0  1     2      0000:0000:0000:0000:0000:ffff:0a14:0305  10.20.3.5     v1    eth2
mlx5_0  1     3      0000:0000:0000:0000:0000:ffff:0a14:0305  10.20.3.5     v2    eth2

The options

The answer

D. Set NCCL_IB_GID_INDEX=3: index 3 is the IPv4-based RoCEv2 GID, which is routable across subnets; the lower indexes are RoCEv1 or link-local entries that only work inside a single L2 domain

Why

Working in-rack but hanging across racks is the signature of a wrong GID: RoCEv1 GIDs and the fe80:: link-local entries are L2-only, while RoCEv2 encapsulates in UDP/IP and routes across subnets. In this table only index 3 combines v2 with a real IPv4-mapped GID, so NCCL_IB_GID_INDEX=3 is the standard setting on such fabrics. Index 1 is v2 but link-local, so it still cannot cross a router - that is the tempting near-miss. NCCL_SOCKET_IFNAME only steers bootstrap/socket traffic, not the RDMA GID, and disabling IB abandons the fabric instead of fixing one environment variable.

More NCCL questions

This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes


← Back to ClusterDrill