I'm currently doing inference on GPUs with libtorch and have a few concerns: (1) It seems like libtorch/torchscript are on a path to getting deprecated and (2) libtorch/torchscript pull in enormously bloated libraries. Should I be looking at executorch? I currently don't see an nvidia backend / integration with tensor rt in https://github.com/pytorch/executorch/tree/main/backends , but seems like it might be possible. Is this something you are thinking about?
OTOH, writing platform specific backends is a huge undertaking. Do you think backends such as MPS may become a shared effort?
[1] https://github.com/huggingface/candle [2] https://github.com/huggingface/candle/issues/313
As an aside, the Vulkan backend is tied to TorchScript at the moment, so it is not yet compatible with ExecuTorch. However, we are also planning to introduce a Vulkan delegate for ExecuTorch which will enable GPU delegation through ExecuTorch.