TensorRT
NVIDIA TensorRT is an ecosystem of APIs for high-performance deep learning inference. TensorRT includes an inference runtime and model optimizations that deliver low latency and high throughput for production applications. The TensorRT ecosystem includes TensorRT, TensorRT-LLM, TensorRT Model Optimizer, and TensorRT Cloud. TensorRT Advantage:
- Speed Up Inference by 36X
- Optimize Inference Performance
- Accelerate Every Workload
- Deploy, Run, and Scale With Triton

