RoCE · basic
ib_write_bw -R from a node in another subnet cannot connect to this host on mlx5_2. What does the show_gids output reveal about the cause?
$ show_gids mlx5_2
DEV PORT INDEX GID IPv4 VER DEV
--- ---- ----- --- ------------ --- ---
mlx5_2 1 0 fe80:0000:0000:0000:1270:fdff:fe2c:9f31 v1 eth2
mlx5_2 1 1 fe80:0000:0000:0000:1270:fdff:fe2c:9f31 v2 eth2
n_gids_found=2
The options
- AThe firmware on this NIC only supports a two-entry GID table, so the routable v2 entry has nowhere to live until the card is flashed with a newer release
- Beth2 has no IP address assigned, so only link-local GIDs exist and there is no routable IPv4-based RoCE v2 GID; assign an address to the interface
- CGID index 1 is already a v2 entry, so RoCE v2 is fully functional on this host and the connection failure has to be a filtering problem on the switches
- DThe fe80 prefix on both GIDs shows the port came up in InfiniBand mode, so the link type must be switched to Ethernet before usable RoCE v2 GIDs appear
The answer
B. eth2 has no IP address assigned, so only link-local GIDs exist and there is no routable IPv4-based RoCE v2 GID; assign an address to the interface
Why
RoCE GIDs are derived from the netdev's addresses: the fe80:: entries come from the link-local IPv6 address every interface has, and an IPv4-mapped GID (::ffff:a.b.c.d) only appears once the interface has an IPv4 address. With no IP on eth2, rdma_cm has nothing to bind or route with across subnets. The two-entry table is a symptom of the missing address, not a firmware limit. Index 1 is indeed v2, but a link-local GID is not routable, so it cannot serve cross-subnet traffic. fe80 GIDs are normal on RoCE ports and say nothing about IB versus Ethernet mode.
This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes
← Back to ClusterDrill