Naw man, you crazy. If you tell me that in 5 years, consumer chips will be so good that I can run GPT-5.4-level AI on my phone, I'd find that plausible (I buy cheap phones). If you're telling me that in 5 years we won't need _servers_ because our _phones and/or desktops_ will be powerful enough to run the biggest newest LLMs in existence, I question your judgment, I think that prediction shows a deep uncreativity about how massively compute-hungry SOTA models will get.
The valuable things to do with inference will keep being a server niche because they'll keep being 1-2 OOM more compute-hungry than whatever consumer hardware can handle. Like gaming: my laptop can run games from 2015 at max settings no problem but the games actually worth getting excited about in 2026 still melt a $2k GPU, because whatever headroom the hardware gains, developers immediately spend on ray tracing and Nanite and modelling individual skin cells or whatever. I don't see any plausible reason to expect that the ceiling on "valuable server-side compute" or "inference capacity" will rise any more slowly than the on-device capability is rising.
My assumption is that in 2031, SOTA top-intelligence AI will be hosted on cloud servers like it is today, offering dirt-cheap access to capabilities we can't even dream of today, while your Android will be running some open-source GPT-5+ equivalent.