ExecuTorch: Run PyTorch programs on mobile and edge devices
pytorch.org
pytorch.org
[Edit] And here is the answer [2]
> PyTorch Mobile uses TorchScript to allow PyTorch models to run on devices with limited resources. ExecuTorch has a significantly smaller memory size and a dynamic memory footprint resulting in superior performance and portability compared to PyTorch Mobile. Also, ExecuTorch does not rely on TorchScript, and instead leverages PyTorch 2 compiler and export functionality for on-device execution of PyTorch models.
[1] https://pytorch.org/mobile/home/ [2] https://pytorch.org/executorch/stable/intro-overview.html#ho...
> Also, ExecuTorch does not rely on TorchScript, and instead leverages PyTorch 2 compiler and export functionality for on-device execution of PyTorch models.
I think this is the major change: AFAIU you could consider ExecuTorch as ""just a rewrite"" of PyTorch mobile using PyTorch 2 compiler.
BUT this is should still be pretty major. It was pretty hard to retarget models to be compatible with TorchScript. But most recent models will want to target PyTorch 2 compiler for the performance boost. Which means they'll automatically get ExecuTorch support.
As someone who got burned several times in trying to port models to TorchScript and run them on Androids, I'm happy they did that.
I'm currently doing inference on GPUs with libtorch and have a few concerns: (1) It seems like libtorch/torchscript are on a path to getting deprecated and (2) libtorch/torchscript pull in enormously bloated libraries. Should I be looking at executorch? I currently don't see an nvidia backend / integration with tensor rt in https://github.com/pytorch/executorch/tree/main/backends , but seems like it might be possible. Is this something you are thinking about?
OTOH, writing platform specific backends is a huge undertaking. Do you think backends such as MPS may become a shared effort?
[1] https://github.com/huggingface/candle [2] https://github.com/huggingface/candle/issues/313
As an aside, the Vulkan backend is tied to TorchScript at the moment, so it is not yet compatible with ExecuTorch. However, we are also planning to introduce a Vulkan delegate for ExecuTorch which will enable GPU delegation through ExecuTorch.
1. MPS backend uses MPSGraph exclusively, might hit some performance ceilings limited by MPSGraph. s4nnc moved more and more ops from MPSGraph to Metal directly to have better control on both allocation and some erratic behaviors from MPSGraph.
2. CoreML backend uses coremltools, thus it carries all the baggage of that: requiring to generate CoreML model AOT because Python dependency, have no control over memory planning, weight quantization scheme, or where to put the weights. Dynamic shape might further making memory planning worse as static shape is the main use-case of CoreML so far and far better tested.
I would love to see an updated port of coremltools either in Swift or C++ (ONNX's coremltools implementation is in v3 I believe, and coremltools moved to v7 spec?).
(submitted title was "ExecuTorch: Enabling On-Device interference for embedded devices")
Edit: I switched "anywhere" to "mobile and edge devices" per https://pytorch.org/blog/pytorch-edge/. One does not get away with saying things like "anywhere" in HN titles...
It wouldn't be the first time autocorrect has messed up a headline on HN.
EDIT: dang fixed the headline. Thanks dang!