Yep, but isn’t there an integrated ML chip that makes it faster than cpu? Or does llama.cpp not use that?
unfortunately that chip is proprietary and undocumented, it's very difficult for open source programs to make use of. I think there is some reverse engineering work being done but it's not complete.