Course Introduction
This course builds an enterprise AI infrastructure and large-model production line from the ground up, covering networking, containers, Kubernetes, GPU clusters, image distribution, data storage, model training, evaluation, distributed inference, performance optimization, and observability. It combines NVIDIA official solutions with open-source approaches to develop reproducible full-stack AI implementation skills.
CURRICULUM
Course Curriculum
14
-
01
Accounts, DNS, NTP, CA, .env, and Domain Names
-
02
BCM HeadNode GUI Installation and Basic Tools
-
03
cnode Image, Harbor Trust, NVIDIA User Space, and Pyxis/Enroot
-
04
DOCA-OFED, PXE Provisioning, and Hardware Acceptance
-
05
Docker, Kubernetes, GPU/Network/NIM Operators, BCM Gateway, and MIG
-
06
Harbor Installation, Projects, Image Preloading, and Acceptance
-
07
MinIO Installation, Buckets, and Object Paths
-
08
Lustre Server/Client and Kubernetes PV/PVC
-
09
GPU/MIG Labels, Run:AI Control Plane/Cluster, Projects, and Lab Spaces
-
10
NeMo Platform, Asset Import, NCCL, Customizer, MIG Evaluator, and NIM
-
11
Dynamo TensorRT-LLM Production Path, Four KV Scenarios, and Ideal TP Rank/GPU/NIC/Rail Mapping
-
12
BCM Prometheus/Grafana Monitoring for Live Dynamo
-
13
Slurm Fundamentals, CPU Examples, Pyxis/Enroot, and Dual-RTX Training
-
14
SGLang PD Disaggregation, Four KV-aware Routing Scenarios, OpenAI API, and Grafana Dashboard
COURSE RECORDINGS
Course Video
Watch the public preview here or open it on the original video platform.
BILIBILI01
PUBLIC PREVIEW
