1) They think the existing hardware is powerful enough.
2) The utilization of the ANE doesn't justify increased resources.
3) They plan to re-generalize AI compute through things like vector operations.
4) (the most pessimistic) They are saving large increases for future releases when they need to force upgrades.
I could see the math for on-device just not working for things like massive LLMs. The amount of silicon you'd need to make it possible would be large, and the frequency of use low. The same silicon for one person's phone could likely support dozens if it were in a datacenter instead.