Course Introduction
This episode accelerates large-language-model inference with NVIDIA TensorRT-LLM. It covers model conversion, precision and quantization choices, optimized kernels and execution graphs, KV cache, batching, and parallelism, then evaluates the tradeoffs among throughput, time to first token, memory use, and model quality.
Total recording duration: 24 min 51 sec.
Companion Git repository: https://qytgit.qytang.com/qytadmin/2026-10-TensorRT-LLM
COURSE RECORDINGS
Course Video
Watch the full recording here or open it on the original video platform.
BILIBILI01
FULL SESSION
2026 Evolution Ep. 10: NVIDIA TensorRT-LLM Inference Acceleration
25 MIN
Watch on the Original Platform