I'd disagree here. I see two avenues for an efficiency multiple, albeit a single-digit multiple:
* Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute.
* Dynamic batching allows typical requests to run in batches of more-than-1 and/or overlap, offering better internal compute utilization ratios (e.g. interleaving output and input streams). The small limit of on-device LLMs will run with batch sizes of one with strong memory bandwidth bottlenecks.
For an example of these factors in action, see the API cost differential between batch, standard, and 'fast' processing. OpenAI prices these tiers at a 1:2:4 ratio.
Go get a job a hyperscaler, they want to smoke what you are smoking.
I am not talking about diurnal cloud workloads, I am talking about the native efficiencies of hyperscalers vs on-prem. They have no magic and they all think they are going to make up their business overheads in exorbitant saas pricing.
- buy in bulk, for lower prices and access to better hardware through big contracts
- build in bulk (i.e. spread out software improvements over a lot of data centres and customers)
- offer additional services such as edge caching and multi-region data redundancy
That doesn't mean that hyperscalers don't also do things wrong, but it seems very odd to pretend there's nothing to them.