Certain inference serving solutions like Nvidia Triton Inference Server will even take an ONNX model and then do TRT compilation (with cache!) on the actual inference hardware dynamically at model load time. This is really nice because you can deploy a standard ONNX model across instances and varying GPU hardware and always get TRT optimized and compatible with Compute Capability, etc. Really handy and basically comes down to a few lines of config in the model configuration.
I'm not terribly familiar with JAX but I have to imagine there's ONNX export or straight to TRT export somewhere.
There's some effort going into systems for saving and restoring the computation graph for Jax programs, which will help a lot. I'm surprised it didn't happen sooner, as it seems like quite a natural fit with the jax architecture.