GPU Cluster Operations
This course develops the operational skills required to commission, monitor, and maintain a liquid-cooled GPU cluster in an AI factory, using GB300 NVL72 and VR200 NVL72-class reference hardware. Students perform direct-to-chip liquid-cooling commissioning steps including pressure leak testing, coolant quality verification, and cold-plate flow-rate confirmation, and operate an in-row or sidecar coolant distribution unit (CDU) under normal and failover conditions. Topics include GPU thermal management and junction-temperature monitoring, NVIDIA driver and container-runtime installation for GPU workloads, GPU health monitoring using NVIDIA System Management Interface (nvidia-smi) and Data Center GPU Manager (DCGM), quick-disconnect (QDC) fitting torque and leak-check procedure, and escalation practice for a coolant-loop anomaly. Laboratory sessions use the Atom Lab Pod's NVL72-form mechanical rack, cold-plate thermal emulator array, in-row CDU N+1 pumps, and L2A sidecar CDU to simulate commissioning and operating tasks on a non-production rack. This course requires access to a live multi-rack liquid-cooling and GPU environment and maps to NVIDIA AI Operations Academy operational practice.
What you'll be able to do
- Perform a pressure leak test on a direct-to-chip liquid-cooling loop before energizing a GPU rack. Safety-critical
- Verify coolant quality (conductivity and pH) against the manufacturer's specification before returning a loop to service. Safety-critical
- Verify per-GPU cold-plate flow rate meets the manufacturer's minimum specification during rack commissioning. Safety-critical
- Torque a quick-disconnect (QDC) coolant fitting to the manufacturer's specification and verify zero-leak performance. Safety-critical
- Operate an in-row CDU through a simulated N+1 pump failover and verify continued cooling delivery. Safety-critical
- Monitor GPU health, utilization, and thermal status using nvidia-smi and NVIDIA Data Center GPU Manager (DCGM).
- Install and validate the NVIDIA driver and container runtime stack on a GPU node for containerized workload execution.
- Diagnose and correctly escalate a coolant-loop anomaly (pressure drop, elevated supply temperature, or leak indication) per the facility's escalation matrix. Safety-critical
- Interpret the power and cooling envelope of a reference NVL72-class rack to plan a safe rack-commissioning sequence. Safety-critical
- Document a GPU cluster commissioning and health-monitoring report for turnover to the operations team.
- Verify air-side environmental compliance (ASHRAE A2 allowable envelope) for a liquid-cooled rack's residual air-cooled components.
- Complete NVIDIA AI Operations Academy practice modules covering GPU cluster operations and liquid-cooling operational practice.
Train the team that runs the factory.
The Institute travels with every SAVRN campus.