AI Fabric Operations (InfiniBand / RoCE)
This course covers the scale-out network fabric that connects GPU nodes into a training or inference cluster, covering both InfiniBand and RDMA over Converged Ethernet (RoCE) fabric technologies as alternative implementations of the same lossless, low-latency transport requirement. Students configure and query an InfiniBand subnet manager, use NVIDIA UFM (Unified Fabric Manager) to monitor fabric health and congestion, and diagnose link and port errors. On the Ethernet side, students configure Priority Flow Control (IEEE 802.1Qbb) and Explicit Congestion Notification/DCQCN to build a lossless RoCE fabric, and verify GPUDirect RDMA operation between GPUs across the fabric without CPU-memory bounce buffers. The course covers non-blocking leaf-spine fabric topology at teaching scale, cable and transceiver verification, and NCCL-based fabric bandwidth testing across multiple nodes as the acceptance criterion for a commissioned fabric segment. Students troubleshoot induced congestion and packet-loss scenarios and document root cause using fabric telemetry.
What you'll be able to do
- Configure an InfiniBand subnet manager to discover and initialize a multi-switch fabric topology.
- Use NVIDIA UFM to monitor fabric health and identify a degraded or congested link.
- Configure Priority Flow Control (IEEE 802.1Qbb) on a lossless RoCE fabric to prevent buffer-overflow packet loss.
- Configure DCQCN/ECN marking thresholds on switch buffers to provide early congestion signaling for RoCE traffic.
- Verify GPUDirect RDMA operation between GPUs on different nodes across the fabric without a CPU bounce buffer.
- Run a multi-node NCCL bandwidth test across the commissioned fabric segment as a fabric acceptance criterion.
- Verify cable and transceiver integrity for a fabric link using vendor diagnostic counters.
- Diagnose an induced packet-loss or congestion event on the lab fabric and document root cause using fabric telemetry.
- Design a non-blocking leaf-spine fabric topology for a specified number of GPU nodes at teaching scale.
- Distinguish InfiniBand and RoCE fabric architectures by mechanism and select the appropriate technology for a stated site constraint.
- Document a fabric commissioning report including topology diagram, subnet manager or PFC/ECN configuration, and NCCL acceptance results.
- Escalate a fabric hardware failure requiring vendor replacement, citing the specific failed component and diagnostic evidence.
Train the team that runs the factory.
The Institute travels with every SAVRN campus.