Today the only way to scale compute is to throw more power at it or settle for the 5% per year real single core performance improvement.
- Running up single-core performance through increasingly sophisticated core design and clock speed (which is now at the 5% per year point mentioned)
- Going wider by throwing more SMT, more cores, and larger caches at the problem.
Assuming here that x86 was the last major architecture that was going for high single-thread performance at all costs, the first phase lasted us a good 30 years-- from the 4004 to the flameout of Netburst.
We could consider the second phase starting when they started delivering the P4 with Hyperthreading, and its true-dual-core predecessors shortly thereafter, so we're now about 20 years into that era.
Do we have another 10 left in it?
The argument is something like that is not really possible anymore given the absurd upfront investments we're seeing existing AI companies need in order to further their offerings.
But yes, there was a window of opportunity when it was possible to do cutting-edge work without billions of investment. That window of opportunity is now past, at least for LLMs. Many new technologies follow a similar pattern.
– Hank Rutherford Hill