AI Fabric Engineering
This course develops the specialized networking skills required to build and operate the InfiniBand and RDMA over Converged Ethernet (RoCE) fabrics that interconnect GPU clusters in an AI factory. Students cable and verify a rail-optimized fat-tree fabric spanning leaf and spine switches, configure Quantum-X800/Spectrum-X class switch ports, and validate GPU-to-GPU communication performance using NVIDIA Collective Communications Library (NCCL) bandwidth and latency tests. Topics include InfiniBand subnet management and topology verification, rail-optimized versus non-rail-optimized GPU-NIC mapping, RoCEv2 congestion control (DCQCN, PFC) as an Ethernet-fabric alternative to native InfiniBand, and troubleshooting NCCL AllReduce performance degradation caused by fabric misconfiguration. Laboratory sessions use the Atom Lab Pod's 4-leaf/2-spine lab fabric to cable a representative rail-optimized topology, run subnet manager topology verification, and execute NCCL benchmark tests against a reference bandwidth target. This course maps to NVIDIA Deep Learning Institute (DLI) AI infrastructure and networking modules and requires access to a live multi-node GPU/fabric lab environment.
What you'll be able to do
- Cable a rail-optimized fat-tree fabric connecting GPU nodes to leaf and spine switches per a reference topology diagram.
- Configure InfiniBand subnet manager settings and verify fabric topology against the reference design.
- Execute an NCCL AllReduce bandwidth and latency benchmark across a multi-GPU fabric and evaluate results against a reference target.
- Diagnose a degraded NCCL AllReduce benchmark result to a specific fabric misconfiguration.
- Configure RoCEv2 congestion-control parameters (DCQCN, Priority Flow Control) on an Ethernet-fabric alternative to native InfiniBand.
- Distinguish rail-optimized from non-rail-optimized GPU-to-NIC mapping and explain the resulting impact on NCCL collective performance.
- Interpret switch port error counters and physical-layer diagnostics to identify a marginal fabric link before it causes a job failure.
- Document a fabric commissioning report including topology verification, cabling map, and benchmark results for turnover to operations.
- Explain the power and thermal interdependency between the AI fabric and a liquid-cooled GPU rack, referencing the reference rack's rated power envelope.
- Complete NVIDIA Deep Learning Institute practice modules covering InfiniBand/RoCE fabric fundamentals and NCCL performance concepts.
- Perform a fabric leak-of-isolation safety check before working on fabric cabling adjacent to an energized liquid-cooling manifold. Safety-critical
Train the team that runs the factory.
The Institute travels with every SAVRN campus.