Apple has been stingy with RAM in consumer hardware. RAM prices will continue escalating, some analysts say into 2030, and this will make it more difficult to build next generation phones with sufficient memory for meaningful ML workloads.
I hope there's some kind of inversion in the current chip economics, because I love distributed/democratized/private compute, but currently cloud based LLM inference seems to be much more viable. I don't see local llms meaningfully viable for the general usecase in the near future.