AI Factory Capstone — Live Site Practicum

Capstone · AIO
240Total hours
8Lecture
232Hands-on lab
AIO 2901; AIO 2902; AIO 2903; AIO 2904; AIO 2905; AIO 2906Prerequisite
Credential
232 lab hours 8 lecture hours

This capstone is a supervised production practicum conducted on live AI-factory infrastructure rather than a classroom course. Students rotate through operational shifts on a real four-rack NVIDIA VR200 NVL72 live block (760-920 kW of compute) carrying real tenant workloads, under the direct supervision of employer mentors and site operations staff. Rotations integrate the technical domains covered across the AIO sequence: cluster health monitoring and DCGM-based triage, fabric and scheduler operations on production tenancy, inference-service observability and incident response, and PUE/WUE and utilization reporting against live facility data. Students carry a documented apprenticeship-style log of shift activities, incidents handled, and escalations made, and complete a capstone practicum portfolio combining that log with employer-mentor evaluations across each rotation. Because this course requires real production hardware and live tenant workloads, access to an AL-3 live block is the binding constraint on offering this course: it cannot be delivered on emulated or bench-scale equipment, and a partner site without an operating AL-3 live block cannot offer this capstone to its students regardless of classroom capacity.

What you'll be able to do

  1. Conduct a live-shift cluster health check across production VR200 racks using DCGM and site monitoring tools, escalating any anomaly per site procedure. Safety-critical
  2. Respond to a live production fabric or scheduler alert during a supervised shift and execute the site's documented first-response procedure. Safety-critical
  3. Perform a supervised rack-level physical inspection on a live, energized VR200 rack following lockout/tagout and PPE procedures. Safety-critical
  4. Monitor and report live tenant workload utilization, PUE, and WUE metrics for an assigned shift against site targets.
  5. Triage a live inference-service performance degradation using the production observability stack and recommend a corrective action.
  6. Execute an assigned MIG or Kubernetes tenant-provisioning change on production infrastructure under mentor approval and change-control procedure. Safety-critical
  7. Maintain a continuous apprenticeship-style shift log documenting activities performed, incidents handled, and escalations made.
  8. Participate in a live cross-functional incident bridge (facilities, network, and compute operations) during a multi-domain production event. Safety-critical
  9. Compile a capstone practicum portfolio synthesizing shift logs, mentor evaluations, and at least one significant incident write-up across the rotation.
  10. Receive and act on real-time employer-mentor feedback to correct a procedural deviation during a live shift.
  11. Deliver an oral defense of the capstone portfolio to a panel including the employer mentor and program faculty of record.
  12. Demonstrate sustained safe work practice across the full live-site rotation with zero reportable safety incidents attributable to technician error. Safety-critical

Train the team that runs the factory.

The Institute travels with every SAVRN campus.

Engage SAVRN → Open the catalog