Are there benchmarks on how much faster TensorRT vs native torch/cuda?
It was one of the fastest backends last time I checked (with vLLM and lmdeploy being comparable), but the space moves fast. It uses cuda under the hood, torch is not relevant in this context.