If the rumors about LLM stuff baked into Siri on the next iPhone are true, it will be interesting see how they make the most use of the hardware with that.
It's rumored the difference is to aid in photography and possibly Face ID, but of course this will be helpful for any onboard inference.
You can also think about it as compounding errors - at any one weight index the bit values might not be too meaningful, but cascaded over a lot of tensor multiplications they will be.
Siri deserves more slander than you delivered.