"probably still requires a lot of GPU compute" is something that would be very desirable to undercut.
There's a galaxy of potential AI markets that are local-only for compliance/privacy/deployment environment reasons, and an equally large and overlapping set of use cases where you won't be able to bolt a rack of thousand-watt GPUs on the side of the device.
What if the next "DeepSeek shock" is something we can run for non-toy use cases on our box of old Android phones? The GPU data centres could be expensive albatrosses very quickly.
There's probably a case that right now, the bigger-is-better paradigm protects incumbents-- a smaller model will always have a FOMO factor unless it can be proven competitive, so that leaves the playing field to those who can afford to train and deploy a new Fable or Sol every few months.