> Hopefully Apple optimizes Core ML to map transformer workloads to the ANE.
If you want to convert models to run on the ANE there are tools provided:
> Convert models from TensorFlow, PyTorch, and other libraries to Core ML.
If you want to convert models to run on the ANE there are tools provided:
> Convert models from TensorFlow, PyTorch, and other libraries to Core ML.
That’s just an issue with stale and incorrect information. Here are the docs https://opensource.apple.com/projects/mlx/
The issue is in targeting specific hardware blocks. When you convert with coremltools, Core ML takes over and doesn't provide fine-grained control - run on GPU, CPU or ANE. Also, ANE isn't really designed with transformers in mind, so most LLM inference defaults to GPU.
Look for Apple to add matmul acceleration into the GPU instead. Thats how to truly speed up local LLMs.