What is interesting is that it seems like the ever larger sums of money sloshing around are resulting in bigger, faster hype cycles. We are already seeing some companies face issues after blowback from adopting AI too fast.
(It might be too expensive to pay for LLM subscriptions when every device in your house is "thinking" all day long. A 3-5k Computer for a local llm might pay itself off after a year or two. )
The next frontier would be training directly with block floating point, where you have a shared exponent plus the two remaining bits. It's getting tight.
Maybe it is possible to have mini LoRA blocks where an n times n block is approximated by the outer product of two n sized vectors. For n = 4 the savings would be 50% less FLOPs and for n=8 the savings would be 75% less FLOPs.