GPU Cluster Bring-Up & Commissioning

Advanced · AIO
120Total hours
32Lecture
88Hands-on lab
AIO 1902Prerequisite
Credential
88 lab hours 32 lecture hours

This course covers the bring-up, burn-in, and acceptance testing of GPU compute nodes and racks in an AI-factory data hall, using NVIDIA VR200 NVL72 (190-230 kW/rack, fully liquid-cooled) and GB300 NVL72 (132-142 kW/rack) as reference platforms. Students perform node-level hardware verification, execute burn-in test suites under sustained synthetic load, and validate NVLink and NVSwitch connectivity using NCCL collective-communication tests (all-reduce, all-gather) to confirm expected bandwidth and latency before a node is accepted into production. Coverage includes DCGM (Data Center GPU Manager) diagnostics and health checks, GPU Xid error interpretation, direct-to-chip liquid cooling loop verification (flow rate, differential pressure, leak detection) at the rack manifold level, and rack power-on sequencing for a liquid-cooled, no-air-cooled-option platform. Students complete a documented acceptance-test package for a rack before sign-off, mirroring the vendor and site acceptance-testing process used before a rack is released to production tenancy.

What you'll be able to do

  1. Verify GPU node hardware inventory against the bill of materials before power-on.
  2. Verify direct-to-chip liquid cooling loop flow rate and differential pressure at the rack manifold before energizing compute. Safety-critical
  3. Execute a rack power-on sequencing procedure for a liquid-cooled, no-air-cooled-option NVL72-class rack. Safety-critical
  4. Run an NCCL all-reduce bandwidth test across multiple GPUs and evaluate results against expected NVLink bandwidth.
  5. Interpret DCGM diagnostic output to classify a GPU as healthy, degraded, or failed.
  6. Execute a multi-hour GPU burn-in test under sustained synthetic load and log thermal and power stability. Safety-critical
  7. Diagnose a failed NVLink or PCIe link identified during bring-up and isolate the faulty component.
  8. Assemble a documented rack acceptance-test package prior to production sign-off.
  9. Calculate expected versus measured power draw for a populated rack against the platform's rated power envelope.
  10. Verify leak-detection sensor placement and function at rack manifold quick-disconnect points before production release. Safety-critical
  11. Document and escalate a bring-up failure that exceeds the technician's authorized corrective-action scope.
  12. Compare GB300 NVL72 and VR200 NVL72 platform specifications to determine site infrastructure readiness requirements.

Train the team that runs the factory.

The Institute travels with every SAVRN campus.

Engage SAVRN → Open the catalog