GPU Cluster Bring-Up & Commissioning
This course covers the bring-up, burn-in, and acceptance testing of GPU compute nodes and racks in an AI-factory data hall, using NVIDIA VR200 NVL72 (190-230 kW/rack, fully liquid-cooled) and GB300 NVL72 (132-142 kW/rack) as reference platforms. Students perform node-level hardware verification, execute burn-in test suites under sustained synthetic load, and validate NVLink and NVSwitch connectivity using NCCL collective-communication tests (all-reduce, all-gather) to confirm expected bandwidth and latency before a node is accepted into production. Coverage includes DCGM (Data Center GPU Manager) diagnostics and health checks, GPU Xid error interpretation, direct-to-chip liquid cooling loop verification (flow rate, differential pressure, leak detection) at the rack manifold level, and rack power-on sequencing for a liquid-cooled, no-air-cooled-option platform. Students complete a documented acceptance-test package for a rack before sign-off, mirroring the vendor and site acceptance-testing process used before a rack is released to production tenancy.
What you'll be able to do
- Verify GPU node hardware inventory against the bill of materials before power-on.
- Verify direct-to-chip liquid cooling loop flow rate and differential pressure at the rack manifold before energizing compute. Safety-critical
- Execute a rack power-on sequencing procedure for a liquid-cooled, no-air-cooled-option NVL72-class rack. Safety-critical
- Run an NCCL all-reduce bandwidth test across multiple GPUs and evaluate results against expected NVLink bandwidth.
- Interpret DCGM diagnostic output to classify a GPU as healthy, degraded, or failed.
- Execute a multi-hour GPU burn-in test under sustained synthetic load and log thermal and power stability. Safety-critical
- Diagnose a failed NVLink or PCIe link identified during bring-up and isolate the faulty component.
- Assemble a documented rack acceptance-test package prior to production sign-off.
- Calculate expected versus measured power draw for a populated rack against the platform's rated power envelope.
- Verify leak-detection sensor placement and function at rack manifold quick-disconnect points before production release. Safety-critical
- Document and escalate a bring-up failure that exceeds the technician's authorized corrective-action scope.
- Compare GB300 NVL72 and VR200 NVL72 platform specifications to determine site infrastructure readiness requirements.
Train the team that runs the factory.
The Institute travels with every SAVRN campus.