I think we’ll start seeing people focus on optimisation, we already see companies like Apple focus on it.
LLM are still to new, and still advancing to quickly for optimisation to take place. It’s like we’re back in the MHz wars of old between CPU manufacturers. The goal is just more performance, regardless of cost, because it was clear that even in the consumer space, people wanted more performance.
Then we hit a kind of plateau in last 10 years, where basic compute is so powerful that your average consumer is not longer upgrading every year for better performance. A 5 year old machine has enough performance for most people. Then the focus on energy efficiency kicked in, because people didn’t want faster computers, they wanted battery life and cheaper computers.
No doubt we’ll see the same with LLM, possibly quite soon. Claude Sonnet 4 and similar class models have enough reasoning performance, that agentic systems can be quite reliable. Which means we hit the base level of “reasoning” performance needed, and we can extend that “performance” in domain specific ways by lightly customising the agentic framework, with no need to fine tuning. The elimination of fine tuning to build domain specific agents is a huge game changer. But it also means that putting together a 10x or 100x efficient model, with “reasoning” performance equivalent to current gen LLM would also be a huge game changer. It opens up the possibility to apply this tech into spaces that currently require either lots of specialists knowledge to fine tune an LLM, or a huge amount of on tap compute to allow the agents to take enough turns to slowly “reason” they’re way through problems.
But a Claude Sonnet 4 that runs on a iPhone for example. That would make Apple’s complete failure to improve Siri look like a genius level move. Why bother with small incremental improvements using current tech, when waiting a few years, and just stuffing a full fat LLM and agent system into an iPhone will basically give you the ultimate Siri.