InfiniBand · intermediate
During a bandwidth regression investigation on an HDR cluster, ibstatus on the slow node prints the following. How do you read this?
Infiniband device 'mlx5_0' port 1 status:
default gid: fe80:0000:0000:0000:1070:fd03:0045:d1a2
base lid: 0x4e
sm lid: 0x1
state: 4: ACTIVE
phys state: 5: LinkUp
rate: 50 Gb/sec (1X HDR)
link_layer: InfiniBand
The options
- AThe link trained at 1X width instead of 4X — lanes are failing, so reseat or replace the cable and retrain
- BHDR ports report the per-lane rate rather than the aggregate, so 50 Gb/sec is what a healthy 4X link shows here
- CThe SM restricted this port to 1X width to save power; raise the minimum allowed width in opensm.conf and resweep
- DThe card sits in a PCIe x4 slot, and that narrow bus caps the negotiated IB link width at 1X here
The answer
A. The link trained at 1X width instead of 4X — lanes are failing, so reseat or replace the cable and retrain
Why
A healthy HDR port shows 'rate: 200 Gb/sec (4X HDR)'; '(1X HDR)' means the link negotiated full HDR lane speed but only one of the four lanes survived training — a classic symptom of a damaged cable, a marginal transceiver, or dirty/bent connector pins. Width is negotiated between the two PHYs at link training time; the SM activates ports but does not decide the trained width, and PCIe lane count is a completely separate bus from the IB link (a narrow PCIe slot would cap throughput but ibstatus would still show 4X). Reseat first, then swap the cable if 1X returns.
More InfiniBand questions
This is 1 of 10 free questions. The full bank is 150 questions and 10 incident labs against a simulated 4-node HGX cluster you can break and repair — €7.99. All free questions · Field notes
← Back to ClusterDrill