GPU Cluster Operations

Advanced · HIT
128Total hours
32Lecture
96Hands-on lab
HIT 1302; NET 2202Prerequisite
NVIDIA AI Operations AcademyCredential
96 lab hours 32 lecture hours

This course develops the operational skills required to commission, monitor, and maintain a liquid-cooled GPU cluster in an AI factory, using GB300 NVL72 and VR200 NVL72-class reference hardware. Students perform direct-to-chip liquid-cooling commissioning steps including pressure leak testing, coolant quality verification, and cold-plate flow-rate confirmation, and operate an in-row or sidecar coolant distribution unit (CDU) under normal and failover conditions. Topics include GPU thermal management and junction-temperature monitoring, NVIDIA driver and container-runtime installation for GPU workloads, GPU health monitoring using NVIDIA System Management Interface (nvidia-smi) and Data Center GPU Manager (DCGM), quick-disconnect (QDC) fitting torque and leak-check procedure, and escalation practice for a coolant-loop anomaly. Laboratory sessions use the Atom Lab Pod's NVL72-form mechanical rack, cold-plate thermal emulator array, in-row CDU N+1 pumps, and L2A sidecar CDU to simulate commissioning and operating tasks on a non-production rack. This course requires access to a live multi-rack liquid-cooling and GPU environment and maps to NVIDIA AI Operations Academy operational practice.

What you'll be able to do

  1. Perform a pressure leak test on a direct-to-chip liquid-cooling loop before energizing a GPU rack. Safety-critical
  2. Verify coolant quality (conductivity and pH) against the manufacturer's specification before returning a loop to service. Safety-critical
  3. Verify per-GPU cold-plate flow rate meets the manufacturer's minimum specification during rack commissioning. Safety-critical
  4. Torque a quick-disconnect (QDC) coolant fitting to the manufacturer's specification and verify zero-leak performance. Safety-critical
  5. Operate an in-row CDU through a simulated N+1 pump failover and verify continued cooling delivery. Safety-critical
  6. Monitor GPU health, utilization, and thermal status using nvidia-smi and NVIDIA Data Center GPU Manager (DCGM).
  7. Install and validate the NVIDIA driver and container runtime stack on a GPU node for containerized workload execution.
  8. Diagnose and correctly escalate a coolant-loop anomaly (pressure drop, elevated supply temperature, or leak indication) per the facility's escalation matrix. Safety-critical
  9. Interpret the power and cooling envelope of a reference NVL72-class rack to plan a safe rack-commissioning sequence. Safety-critical
  10. Document a GPU cluster commissioning and health-monitoring report for turnover to the operations team.
  11. Verify air-side environmental compliance (ASHRAE A2 allowable envelope) for a liquid-cooled rack's residual air-cooled components.
  12. Complete NVIDIA AI Operations Academy practice modules covering GPU cluster operations and liquid-cooling operational practice.

Train the team that runs the factory.

The Institute travels with every SAVRN campus.

Engage SAVRN → Open the catalog