TensorRT offers significant advantages wrt inference and it takes ONNX files. Best I can tell this does not have a TensorRT backend (https://github.com/uber/neuropod/search?q=nvinfer.h&unscoped...). Why not?
We actually do use TensorRT with several of our models, but our approach is generally to do all TRT related processing before the Neuropod export step. For example, we might do something like
TF model -> TF-TRT optimization -> Neuropod export
or PyTorch model
-> (convert subset of model to a torchscript engine)
-> PyTorch model + custom op to run TRT engine
-> TorchScript model + custom op to run TRT engine
-> Neuropod export
Since Neuropod wraps the underlying model (including custom ops), this approach works well for us.That said, I'm getting ridiculously good performance with it, even without using the TensorCores.