Scheduler & Tenancy Operations (Slurm, Kubernetes, MIG)
This course covers workload scheduling and multi-tenant GPU allocation on a commissioned AI-factory cluster, using Slurm and Kubernetes as the two dominant scheduling frameworks. Students configure Slurm partitions, QOS, and GRES (generic resource) GPU allocation; submit and manage batch and interactive GPU jobs; and use SchedMD Slinky components (slurm-bridge and slurm-operator) to run Slurm workloads under Kubernetes orchestration. The course covers the NVIDIA GPU Operator for automated driver installation, container runtime configuration, and device-plugin deployment, and configures DCGM Exporter to expose per-job GPU telemetry labeled with Slurm job IDs for tenant chargeback and health monitoring. Students partition GPUs using Multi-Instance GPU (MIG) to isolate tenant workloads with guaranteed compute and memory slices, and configure resource quotas and namespaces to enforce multi-tenant isolation on Kubernetes. The course closes with a tenant-onboarding exercise establishing a new project's quota, priority, and monitoring dashboard on a shared cluster.
What you'll be able to do
- Configure Slurm partitions, QOS, and GRES GPU resource definitions for a multi-tenant cluster.
- Submit and troubleshoot a batch GPU job that fails due to resource contention or misconfigured GRES request.
- Deploy the NVIDIA GPU Operator on a Kubernetes cluster to automate driver, container runtime, and device-plugin configuration.
- Partition a GPU using Multi-Instance GPU (MIG) to provide isolated compute and memory slices to separate tenant workloads.
- Configure DCGM Exporter to expose per-job GPU telemetry labeled with Slurm job IDs for tenant chargeback.
- Configure Kubernetes resource quotas and namespaces to enforce multi-tenant GPU isolation.
- Run Slurm workloads under Kubernetes orchestration using the Slinky slurm-operator or slurm-bridge components.
- Onboard a new tenant project onto a shared cluster with defined quota, scheduling priority, and monitoring dashboard.
- Compare Slurm-only, Kubernetes-only, and Slurm-on-Kubernetes scheduling architectures for a given tenancy and workload scenario.
Train the team that runs the factory.
The Institute travels with every SAVRN campus.