NVIDIA's unified library implements quantization, distillation, pruning, and speculative decoding. It is designed to optimize models for deployment on TensorRT-LLM, TensorRT, and vLLM.
HOW THIS AFFECTS YOU
●
builderYou can use these SOTA techniques to significantly reduce latency and memory footprints in production.